REVIEW 31 cited by
NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall short in speech quality, similarity, and prosody. Considering speech intricately encompasses various attributes (e.g., content, prosody, timbre, and acoustic details) that pose significant challenges for generation, a natural idea is to factorize speech into individual subspaces representing different attributes and generate them individually. Motivated by it, we propose NaturalSpeech 3, a TTS system with novel factorized diffusion models to generate natural speech in a zero-shot way. Specifically, 1) we design a neural codec with factorized vector quantization (FVQ) to disentangle speech waveform into subspaces of content, prosody, timbre, and acoustic details; 2) we propose a factorized diffusion model to generate attributes in each subspace following its corresponding prompt. With this factorization design, NaturalSpeech 3 can effectively and efficiently model intricate speech with disentangled subspaces in a divide-and-conquer way. Experiments show that NaturalSpeech 3 outperforms the state-of-the-art TTS systems on quality, similarity, prosody, and intelligibility, and achieves on-par quality with human recordings. Furthermore, we achieve better performance by scaling to 1B parameters and 200K hours of training data.
Forward citations
Cited by 31 Pith papers
-
MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control
MAVIN proposes boundary-aware attention, ID-aware propagation, a multi-agent scripting pipeline, and the MAVINSet dataset as the first framework for multi-shot audio-visual generation with narrative control, claiming ...
-
DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech
A LoRA-plus-convolution adaptation converts an AR TTS backbone into a confidence-ordered discrete diffusion model that improves WER and speed on limited data.
-
MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation
MeloCodec quantizes chromagram-derived melody tokens and fuses them with acoustic tokens using a two-stage training scheme, improving pitch consistency and enabling controllable pitch shifting in a singing-voice codec.
-
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
Autoregressive TTS from 8-Hz, 768-dimensional continuous tokens works when the tokenizer shapes its latent space with a low-dimensional core and an energy hierarchy, and the generator separates guidance into local, se...
-
StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis
A mobile-oriented 83M-parameter masked transformer with sparse phone-anchored temporal embeddings achieves RTF 0.08 and lower WER than MaskGCT/F5-TTS on Seed-TTS test sets.
-
AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
AuEmoChat's learned 1,000-code emotion token space, combined with emotion-guided token merging and classifier-guided flow matching, yields higher naturalness and emotion scores than four CSS baselines on NCSSD-EmCap.
-
TokAN: Accent Normalization Using Self-Supervised Speech Tokens
TokAN maps L2 speech tokens to L1-like tokens via an autoregressive converter plus GRPO rewards, cutting WER to 9.23% on seven English accents without natural parallel L1-L2 recordings.
-
Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models
Synthetic data for low-resource spoken language models creates a Stability-Expressivity Gap that DGSA and TDSC self-alignment close, enabling SOTA Thai TTS and first Lao zero-shot voice cloning.
-
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
A zero-shot model that generates a video of a reference face speaking user-chosen text with a reference voice timbre.
-
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.
-
DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
DiTReducio is a training-free, pattern-guided layer and branch skipping method that accelerates DiT-based TTS, reporting significant FLOP and RTF reductions with modest quality loss at tuned thresholds.
-
DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners
DeCodec learns a single neural codec that disentangles speech, background sound, semantic content, and paralinguistic style into orthogonal quantized streams, enabling reconstruction, enhancement, voice conversion, AS...
-
Entropy-based Coarse and Compressed Semantic Speech Representation Learning
Predictive entropy from a token-level speech language model finds merge boundaries, producing compressed semantic tokens that keep ASR and translation accuracy at 15 Hz while lowering latency.
-
Next Tokens Denoising for Speech Synthesis
Dragon-FM generates speech autoregressively over two-second chunks while using flow matching inside each chunk, achieving fast synthesis at 12.5 discrete audio tokens per second.
-
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.
-
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis
Reinforcement learning on duration prediction improves intelligibility and speaker similarity in a 4-step distilled text-to-speech model, and teacher-guided sampling recovers prosodic diversity.
-
Autoregressive Speech Enhancement via Acoustic Tokens
Acoustic tokens outperform semantic tokens on speaker identity in speech enhancement, an autoregressive transducer helps in some settings, but discrete representations still lag continuous ones.
-
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
SemAlignVC strips source-speaker timbre by aligning a speech semantic encoder to BERT text embeddings, then resynthesizes the content conditioned only on a target voice reference.
-
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
An iterative active learning pipeline that filters web speech data by predicted quality and redundancy produces a TTS corpus with better speaker coverage at the same size.
-
NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization
A hierarchical NVAE plug-in generates diverse pseudo-speaker embeddings that raise ASV EER above 38% on FACodec/CosyVoice2 with a controllable privacy-utility trade-off.
-
AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling
AutoSIFT disentangles text-describable style categories from residual speech styles and selectively infills only the categories the user specifies.
-
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey
The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.
-
Adaptive Duration Model for Text Speech Alignment
DurFormer, an adaptive duration prediction model with speed, scene, and semantic conditioning, improves phoneme-level duration accuracy and lowers word error rate in Mandarin text-to-speech.
-
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
IndexTTS2 achieves precise token-count-based duration control and emotion/speaker disentanglement in an autoregressive zero-shot TTS, reporting SOTA WER, speaker similarity, and emotional fidelity.
-
RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding
RT-VC converts a speaker's voice to a new target voice in real time on a CPU with 61.4ms latency, matching the quality of the current SOTA StreamVC.
-
WAKE: Watermarking Audio with Key Enrichment
WAKE embeds and decodes multiple 32-bit audio watermarks with separate 8-bit keys using an invertible neural network, avoiding the overwriting problem in existing systems.
-
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
R-VC performs zero-shot voice conversion in two sampling steps while transferring the target speaker's rhythm, matching or exceeding prior systems in naturalness and intelligibility.
-
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
A latent diffusion model with SSL-based content restoration and in-context speaker prompts improves dysarthric speech intelligibility and speaker similarity on UASpeech.
-
Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model
A cross-modal attention ensemble with balanced stacking reaches MacroF1 0.4094 on 8-class naturalistic speech emotion recognition, beating the official baseline by about 0.08.
-
DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec
DS-Codec improves low-bitrate speech codec quality by first training a mirrored codec and then switching to a non-mirrored decoder, while using product quantization to form one large codebook.
-
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
A speculative decoding variant with a tolerance-based acceptance rule speeds up CosyVoice 2 inference by 1.4x while keeping subjective quality comparable.
Discussion (0). Continue with ORCID to comment.