Pith. sign in

REVIEW 14 cited by

Zero-shot Voice Conversion with Diffusion Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.09943 v1 pith:2Q6L6V7L submitted 2024-11-15 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords timbreconversionvoicespeechzero-shotseed-vctrainingdiffusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Zero-shot voice conversion aims to transform a source speech utterance to match the timbre of a reference speech from an unseen speaker. Traditional approaches struggle with timbre leakage, insufficient timbre representation, and mismatches between training and inference tasks. We propose Seed-VC, a novel framework that addresses these issues by introducing an external timbre shifter during training to perturb the source speech timbre, mitigating leakage and aligning training with inference. Additionally, we employ a diffusion transformer that leverages the entire reference speech context, capturing fine-grained timbre features through in-context learning. Experiments demonstrate that Seed-VC outperforms strong baselines like OpenVoice and CosyVoice, achieving higher speaker similarity and lower word error rates in zero-shot voice conversion tasks. We further extend our approach to zero-shot singing voice conversion by incorporating fundamental frequency (F0) conditioning, resulting in comparative performance to current state-of-the-art methods. Our findings highlight the effectiveness of Seed-VC in overcoming core challenges, paving the way for more accurate and versatile voice conversion systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors

    eess.AS 2026-07 conditional novelty 6.0 of 10

    On 1,168 professional voice actors, a misidentification floor in speaker embeddings survives calibration, normalization, and discriminative re-ranking, and the same floor makes fixed-threshold voice-clone attribution ...

  2. VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Pretrained spoofing detectors reach at best 28.98% EER on a new English-Spanish benchmark of 10 LLM-era TTS/VC systems under 10 post-processing conditions, with most near chance.

  3. GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech

    cs.LG 2026-07 accept novelty 6.0 of 10

    GRAFT splices a short spoken word sample into a neural codec TTS prompt and uses voice-conversion training so the model copies that pronunciation into any target voice, cutting target-word phoneme error 22-39%.

  4. TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    TRACE detects synthetic emotional entrainment disruption in dyadic speech at up to 93.47% accuracy when conditioned on relationship, using windowed emotion-Whisper sequences on the new DyadEE dataset.

  5. Entropy-based Coarse and Compressed Semantic Speech Representation Learning

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Predictive entropy from a token-level speech language model finds merge boundaries, producing compressed semantic tokens that keep ASR and translation accuracy at 15 Hz while lowering latency.

  6. SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods

    cs.SD 2025-07 conditional novelty 6.0 of 10

    SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.

  7. The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents

    cs.CR 2025-07 unverdicted novelty 6.0 of 10

    A multi-agent audio-language model framework can automatically profile private attributes, such as age, health, and income, directly from general audio recordings.

  8. Universal Speech Content Factorization

    eess.AS 2026-03 conditional novelty 5.0 of 10

    A universal least-squares speech-to-content map plus few-second speaker transforms yields open-set, low-rank, timbre-suppressed features competitive for zero-shot VC and TTS.

  9. QASA: Quality-Aware Semantic Augmentation for Robust Multimodal Sentiment Analysis

    cs.LG 2026-01 reject novelty 5.0 of 10

    Diffusion-generated video/audio samples weighted by a learned quality scorer are claimed to improve multimodal sentiment analysis on CH-SIMS, CMU-MOSI, and MUStARD.

  10. REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers

    eess.AS 2025-08 unverdicted novelty 5.0 of 10

    REF-VC is a zero-shot voice conversion system that random-erases redundant parts of speech-embedding features to stay robust to noise, and uses shortcut-distilled flow matching to convert speech in only four steps.

  11. De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks

    cs.SD 2025-07 conditional novelty 5.0 of 10

    Existing voice-protection perturbations succeed only against naive attackers; a phoneme-guided purification-refinement pipeline restores cloneability of protected speech for most VC models.

  12. IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

    cs.CL 2025-06 conditional novelty 5.0 of 10

    IndexTTS2 achieves precise token-count-based duration control and emotion/speaker disentanglement in an autoregressive zero-shot TTS, reporting SOTA WER, speaker similarity, and emotional fidelity.

  13. Semantic-Aware Ship Detection with Vision-Language Integration

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    Abstract claims a VLM-plus-adaptive-window framework and a new semantic ship dataset, but the manuscript body is a different voice-timbre paper.

  14. EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

    cs.SD 2025-05 reject novelty 4.0 of 10

    EZ-VC combines discrete units from a multilingual self-supervised encoder (Xeus) with an F5-TTS flow-matching decoder to achieve zero-shot any-to-any voice conversion, without text labels or multiple disentangling encoders.

Pith tools