Pith. sign in

REVIEW 3 cited by

QR-VC: Leveraging Quantization Residuals for Linear Disentanglement in Zero-Shot Voice Conversion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.16147 v2 pith:ZY5CCQVY submitted 2024-11-25 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords quantizationspeakerconversiondisentanglementidentitylinearresidualsvoice
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Zero-shot voice conversion is a technique that alters the speaker identity of an input speech to match a target speaker using only a single reference utterance, without requiring additional training. Recent approaches extensively utilize self-supervised learning features with K-means quantization to extract high-quality content representations while removing speaker identity. However, this quantization process also eliminates fine-grained phonetic and prosodic variations, degrading intelligibility and prosody preservation. While prior works have primarily focused on quantized representations, quantization residuals remain underutilized and deserve further exploration. In this paper, we introduce a novel approach that fully utilizes quantization residuals by leveraging temporal properties of speech components. This facilitates the disentanglement of speaker identity and the recovery of phonetic and prosodic details lost during quantization. By applying only K-means quantization and linear projections, our method achieves simple yet effective disentanglement, without requiring complex architectures or explicit supervision. This allows for high-fidelity voice conversion trained solely with reconstruction losses. Experiments show that the proposed model outperforms existing methods across both subjective and objective metrics. It achieves superior intelligibility and speaker similarity, along with improved prosody preservation, highlighting the impact of our Linear Disentangler module.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Private kNN-VC: Interpretable Anonymization of Converted Speech

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Phone duration prediction and per-phone k-means quantization raise kNN-VC's privacy EER from 10% to nearly 50%, but target-selection changes can erase most of that gain.

  2. REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers

    eess.AS 2025-08 unverdicted novelty 5.0 of 10

    REF-VC is a zero-shot voice conversion system that random-erases redundant parts of speech-embedding features to stay robust to noise, and uses shortcut-distilled flow matching to convert speech in only four steps.

  3. HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement

    eess.AS 2025-06 conditional novelty 5.0 of 10

    HASRD factorizes SSL speech representations into a first semantic codebook and residual acoustic codebooks, reporting improved ASR and reconstruction at 3.1 kbps versus SpeechTokenizer's 6.0 kbps.

Pith tools