Pith. sign in

REVIEW 31 cited by

NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.03100 v3 pith:Y7EZTROB submitted 2024-03-05 eess.AS cs.AIcs.CLcs.LGcs.SD

classification eess.AScs.AIcs.CLcs.LGcs.SD
keywords speechfactorizednaturalspeechprosodyattributesdiffusiongeneratemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall short in speech quality, similarity, and prosody. Considering speech intricately encompasses various attributes (e.g., content, prosody, timbre, and acoustic details) that pose significant challenges for generation, a natural idea is to factorize speech into individual subspaces representing different attributes and generate them individually. Motivated by it, we propose NaturalSpeech 3, a TTS system with novel factorized diffusion models to generate natural speech in a zero-shot way. Specifically, 1) we design a neural codec with factorized vector quantization (FVQ) to disentangle speech waveform into subspaces of content, prosody, timbre, and acoustic details; 2) we propose a factorized diffusion model to generate attributes in each subspace following its corresponding prompt. With this factorization design, NaturalSpeech 3 can effectively and efficiently model intricate speech with disentangled subspaces in a divide-and-conquer way. Experiments show that NaturalSpeech 3 outperforms the state-of-the-art TTS systems on quality, similarity, prosody, and intelligibility, and achieves on-par quality with human recordings. Furthermore, we achieve better performance by scaling to 1B parameters and 200K hours of training data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    MAVIN proposes boundary-aware attention, ID-aware propagation, a multi-agent scripting pipeline, and the MAVINSet dataset as the first framework for multi-shot audio-visual generation with narrative control, claiming ...

  2. DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech

    eess.AS 2026-07 conditional novelty 6.5 of 10

    A LoRA-plus-convolution adaptation converts an AR TTS backbone into a confidence-ordered discrete diffusion model that improves WER and speed on limited data.

  3. MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation

    cs.SD 2026-08 conditional novelty 6.0 of 10

    MeloCodec quantizes chromagram-derived melody tokens and fuses them with acoustic tokens using a two-stage training scheme, improving pitch consistency and enabling controllable pitch shifting in a singing-voice codec.

  4. Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Autoregressive TTS from 8-Hz, 768-dimensional continuous tokens works when the tokenizer shapes its latent space with a low-dimensional core and an energy hierarchy, and the generator separates guidance into local, se...

  5. StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A mobile-oriented 83M-parameter masked transformer with sparse phone-anchored temporal embeddings achieves RTF 0.08 and lower WER than MaskGCT/F5-TTS on Seed-TTS test sets.

  6. AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

    cs.SD 2026-07 conditional novelty 6.0 of 10

    AuEmoChat's learned 1,000-code emotion token space, combined with emotion-guided token merging and classifier-guided flow matching, yields higher naturalness and emotion scores than four CSS baselines on NCSSD-EmCap.

  7. TokAN: Accent Normalization Using Self-Supervised Speech Tokens

    cs.SD 2026-07 accept novelty 6.0 of 10

    TokAN maps L2 speech tokens to L1-like tokens via an autoregressive converter plus GRPO rewards, cutting WER to 9.23% on seven English accents without natural parallel L1-L2 recordings.

  8. Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

    cs.CL 2026-04 conditional novelty 6.0 of 10

    Synthetic data for low-resource spoken language models creates a Stability-Expressivity Gap that DGSA and TDSC self-alignment close, enabling SOTA Thai TTS and first Lao zero-shot voice cloning.

  9. OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model

    cs.SD 2026-02 conditional novelty 6.0 of 10

    A zero-shot model that generates a video of a reference face speaking user-chosen text with a reference voice timbre.

  10. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

    eess.AS 2025-10 conditional novelty 6.0 of 10

    A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.

  11. DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration

    cs.SD 2025-09 conditional novelty 6.0 of 10

    DiTReducio is a training-free, pattern-guided layer and branch skipping method that accelerates DiT-based TTS, reporting significant FLOP and RTF reductions with modest quality loss at tuned thresholds.

  12. DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners

    cs.SD 2025-09 conditional novelty 6.0 of 10

    DeCodec learns a single neural codec that disentangles speech, background sound, semantic content, and paralinguistic style into orthogonal quantized streams, enabling reconstruction, enhancement, voice conversion, AS...

  13. Entropy-based Coarse and Compressed Semantic Speech Representation Learning

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Predictive entropy from a token-level speech language model finds merge boundaries, producing compressed semantic tokens that keep ASR and translation accuracy at 15 Hz while lowering latency.

  14. Next Tokens Denoising for Speech Synthesis

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Dragon-FM generates speech autoregressively over two-second chunks while using flow matching inside each chunk, achieving fast synthesis at 12.5 discrete audio tokens per second.

  15. SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods

    cs.SD 2025-07 conditional novelty 6.0 of 10

    SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.

  16. DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

    eess.AS 2025-07 conditional novelty 6.0 of 10

    Reinforcement learning on duration prediction improves intelligibility and speaker similarity in a 4-step distilled text-to-speech model, and teacher-guided sampling recovers prosodic diversity.

  17. Autoregressive Speech Enhancement via Acoustic Tokens

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Acoustic tokens outperform semantic tokens on speaker identity in speech enhancement, an autoregressive transducer helps in some settings, but discrete representations still lag continuous ones.

  18. SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment

    eess.AS 2025-07 conditional novelty 6.0 of 10

    SemAlignVC strips source-speaker timbre by aligning a speech semantic encoder to BERT text embeddings, then resynthesizes the content conditioned only on a target voice reference.

  19. Active Learning for Text-to-Speech Synthesis with Informative Sample Collection

    cs.SD 2025-07 conditional novelty 6.0 of 10

    An iterative active learning pipeline that filters web speech data by predicted quality and redundancy produces a TTS corpus with better speaker coverage at the same size.

  20. NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization

    eess.AS 2026-07 conditional novelty 5.5 of 10

    A hierarchical NVAE plug-in generates diverse pseudo-speaker embeddings that raise ASV EER above 38% on FACodec/CosyVoice2 with a controllable privacy-utility trade-off.

  21. AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling

    cs.SD 2026-07 unverdicted novelty 5.0 of 10

    AutoSIFT disentangles text-describable style categories from residual speech styles and selectively infills only the categories the user specifies.

  22. Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.

  23. Adaptive Duration Model for Text Speech Alignment

    cs.SD 2025-07 conditional novelty 5.0 of 10

    DurFormer, an adaptive duration prediction model with speed, scene, and semantic conditioning, improves phoneme-level duration accuracy and lowers word error rate in Mandarin text-to-speech.

  24. IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

    cs.CL 2025-06 conditional novelty 5.0 of 10

    IndexTTS2 achieves precise token-count-based duration control and emotion/speaker disentanglement in an autoregressive zero-shot TTS, reporting SOTA WER, speaker similarity, and emotional fidelity.

  25. RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding

    eess.AS 2025-06 conditional novelty 5.0 of 10

    RT-VC converts a speaker's voice to a new target voice in real time on a CPU with 61.4ms latency, matching the quality of the current SOTA StreamVC.

  26. WAKE: Watermarking Audio with Key Enrichment

    cs.SD 2025-06 conditional novelty 5.0 of 10

    WAKE embeds and decodes multiple 32-bit audio watermarks with separate 8-bit keys using an invertible neural network, avoiding the overwriting problem in existing systems.

  27. Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching

    eess.AS 2025-06 conditional novelty 5.0 of 10

    R-VC performs zero-shot voice conversion in two sampling steps while transferring the target speaker's rhythm, matching or exceeding prior systems in naturalness and intelligibility.

  28. DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model

    cs.SD 2025-05 conditional novelty 5.0 of 10

    A latent diffusion model with SSL-based content restoration and in-context speaker prompts improves dysarthric speech intelligibility and speaker similarity on UASpeech.

  29. Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A cross-modal attention ensemble with balanced stacking reaches MacroF1 0.4094 on 8-class naturalistic speech emotion recognition, beating the official baseline by about 0.08.

  30. DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec

    cs.SD 2025-05 conditional novelty 4.0 of 10

    DS-Codec improves low-bitrate speech codec quality by first training a mirrored codec and then switching to a non-mirrored decoder, while using product quantization to form one large codebook.

  31. Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding

    cs.SD 2025-05 conditional novelty 4.0 of 10

    A speculative decoding variant with a tolerance-based acceptance rule speeds up CosyVoice 2 inference by 1.4x while keeping subjective quality comparable.

Pith tools