Pith. sign in

REVIEW 2 cited by

One-shot Voice Conversion by Separating Speaker and Content Representations with Instance Normalization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.05742 v4 pith:CM2P436J submitted 2019-04-10 cs.LG cs.SDeess.ASstat.ML

classification cs.LGcs.SDeess.ASstat.ML
keywords speakervoicemodelablerepresentationstargetcontentconversion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, voice conversion (VC) without parallel data has been successfully adapted to multi-target scenario in which a single model is trained to convert the input voice to many different speakers. However, such model suffers from the limitation that it can only convert the voice to the speakers in the training data, which narrows down the applicable scenario of VC. In this paper, we proposed a novel one-shot VC approach which is able to perform VC by only an example utterance from source and target speaker respectively, and the source and target speaker do not even need to be seen during training. This is achieved by disentangling speaker and content representations with instance normalization (IN). Objective and subjective evaluation shows that our model is able to generate the voice similar to target speaker. In addition to the performance measurement, we also demonstrate that this model is able to learn meaningful speaker representations without any supervision.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding

    eess.AS 2025-06 conditional novelty 5.0 of 10

    RT-VC converts a speaker's voice to a new target voice in real time on a CPU with 61.4ms latency, matching the quality of the current SOTA StreamVC.

  2. Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training

    cs.SD 2025-06 reject novelty 5.0 of 10

    Pureformer-VC is a transformer-based encoder-decoder for non-parallel voice conversion that reports competitive, but not state-of-the-art, results on VCTK and AISHELL-3.

Pith tools