REVIEW 14 cited by
Zero-shot Voice Conversion with Diffusion Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Zero-shot voice conversion aims to transform a source speech utterance to match the timbre of a reference speech from an unseen speaker. Traditional approaches struggle with timbre leakage, insufficient timbre representation, and mismatches between training and inference tasks. We propose Seed-VC, a novel framework that addresses these issues by introducing an external timbre shifter during training to perturb the source speech timbre, mitigating leakage and aligning training with inference. Additionally, we employ a diffusion transformer that leverages the entire reference speech context, capturing fine-grained timbre features through in-context learning. Experiments demonstrate that Seed-VC outperforms strong baselines like OpenVoice and CosyVoice, achieving higher speaker similarity and lower word error rates in zero-shot voice conversion tasks. We further extend our approach to zero-shot singing voice conversion by incorporating fundamental frequency (F0) conditioning, resulting in comparative performance to current state-of-the-art methods. Our findings highlight the effectiveness of Seed-VC in overcoming core challenges, paving the way for more accurate and versatile voice conversion systems.
Forward citations
Cited by 14 Pith papers
-
A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors
On 1,168 professional voice actors, a misidentification floor in speaker embeddings survives calibration, normalization, and discriminative re-ranking, and the same floor makes fixed-threshold voice-clone attribution ...
-
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Pretrained spoofing detectors reach at best 28.98% EER on a new English-Spanish benchmark of 10 LLM-era TTS/VC systems under 10 post-processing conditions, with most near chance.
-
GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech
GRAFT splices a short spoken word sample into a neural codec TTS prompt and uses voice-conversion training so the model copies that pronunciation into any target voice, cutting target-word phoneme error 22-39%.
-
TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech
TRACE detects synthetic emotional entrainment disruption in dyadic speech at up to 93.47% accuracy when conditioned on relationship, using windowed emotion-Whisper sequences on the new DyadEE dataset.
-
Entropy-based Coarse and Compressed Semantic Speech Representation Learning
Predictive entropy from a token-level speech language model finds merge boundaries, producing compressed semantic tokens that keep ASR and translation accuracy at 15 Hz while lowering latency.
-
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.
-
The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents
A multi-agent audio-language model framework can automatically profile private attributes, such as age, health, and income, directly from general audio recordings.
-
Universal Speech Content Factorization
A universal least-squares speech-to-content map plus few-second speaker transforms yields open-set, low-rank, timbre-suppressed features competitive for zero-shot VC and TTS.
-
QASA: Quality-Aware Semantic Augmentation for Robust Multimodal Sentiment Analysis
Diffusion-generated video/audio samples weighted by a learned quality scorer are claimed to improve multimodal sentiment analysis on CH-SIMS, CMU-MOSI, and MUStARD.
-
REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers
REF-VC is a zero-shot voice conversion system that random-erases redundant parts of speech-embedding features to stay robust to noise, and uses shortcut-distilled flow matching to convert speech in only four steps.
-
De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks
Existing voice-protection perturbations succeed only against naive attackers; a phoneme-guided purification-refinement pipeline restores cloneability of protected speech for most VC models.
-
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
IndexTTS2 achieves precise token-count-based duration control and emotion/speaker disentanglement in an autoregressive zero-shot TTS, reporting SOTA WER, speaker similarity, and emotional fidelity.
-
Semantic-Aware Ship Detection with Vision-Language Integration
Abstract claims a VLM-plus-adaptive-window framework and a new semantic ship dataset, but the manuscript body is a different voice-timbre paper.
-
EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion
EZ-VC combines discrete units from a multilingual self-supervised encoder (Xeus) with an F5-TTS flow-matching decoder to achieve zero-shot any-to-any voice conversion, without text labels or multiple disentangling encoders.
Discussion (0). Sign in to comment.