Pith. sign in

REVIEW 14 cited by

A Survey on Neural Speech Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.15561 v3 pith:E4SNF44Q submitted 2021-06-29 eess.AS cs.CLcs.LGcs.MMcs.SD

classification eess.AScs.CLcs.LGcs.MMcs.SD
keywords speechneuralresearchsurveytextfutureincludingindustry
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text to speech (TTS), or speech synthesis, which aims to synthesize intelligible and natural speech given text, is a hot research topic in speech, language, and machine learning communities and has broad applications in the industry. As the development of deep learning and artificial intelligence, neural network-based TTS has significantly improved the quality of synthesized speech in recent years. In this paper, we conduct a comprehensive survey on neural TTS, aiming to provide a good understanding of current research and future trends. We focus on the key components in neural TTS, including text analysis, acoustic models and vocoders, and several advanced topics, including fast TTS, low-resource TTS, robust TTS, expressive TTS, and adaptive TTS, etc. We further summarize resources related to TTS (e.g., datasets, opensource implementations) and discuss future research directions. This survey can serve both academic researchers and industry practitioners working on TTS.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 185 citations worldwide. Full citation record

  1. Probing Speaker Identity Sensitivity in Audio Deepfake Detectors

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A new Identity Sensitivity Score flags misclassified audio deepfake detections with AUC up to 0.954, but its claim to isolate speaker-identity behavior from plain confidence is not yet controlled.

  2. NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

    cs.SD 2025-08 unverdicted novelty 6.0 of 10

    A Mandarin speech pipeline that tags non-word vocalizations (laughter, breath, filled pauses) for ASR and controls their generation in TTS, backed by a claimed 573-hour, 174,179-utterance word-level annotated corpus.

  3. SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods

    cs.SD 2025-07 conditional novelty 6.0 of 10

    SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.

  4. SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch

    eess.AS 2025-07 conditional novelty 6.0 of 10

    SLASH adds DSP-derived absolute pitch objectives, including direct spectrogram generation from F0, to self-supervised pitch estimation and beats DSP and SSL baselines on MIR-1K.

  5. Multi-interaction TTS toward professional recording reproduction

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Multi-turn textual directions can iteratively refine the speaking style of synthesized speech through a learned embedding refiner, with modest but measurable alignment to the directions.

  6. Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study

    cs.CL 2026-08 conditional novelty 5.0 of 10

    No single Urdu TTS system wins across all metrics, and the system listeners liked most, Google Gemini, was the furthest from reference audio on objective acoustic scores.

  7. Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model

    eess.AS 2026-07 conditional novelty 5.0 of 10

    A lightweight face adapter plus soft-tuning aligns face embeddings to a frozen StyleTTS 2 style space, yielding natural zero-shot face-to-speech and language-agnostic transfer to Spanish.

  8. Position: Towards Responsible Evaluation for Text-to-Speech

    eess.AS 2025-10 conditional novelty 5.0 of 10

    A call to reform text-to-speech evaluation around a three-level framework covering metric fidelity, comparability, and ethical oversight.

  9. Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model

    eess.AS 2025-08 conditional novelty 5.0 of 10

    Speaker-conditioned phrasing with phoneme-level PLMs (MP BERT) improves pause prediction from F0.5 0.3719 to 0.4991, and a few-shot adapter generalizes to unseen speakers.

  10. Traceable TTS: Toward Watermark-Free TTS with Strong Traceability

    eess.AS 2025-07 reject novelty 5.0 of 10

    A joint training loop makes an F5-TTS model produce audio that a paired wav2vec 2.0/LCNN discriminator can recognize, enabling watermark-free attribution; however, the reported generalization gain is not isolated from...

  11. Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A speech-deepfake dataset for ten public figures built with transcription-based segmentation reports high synthetic naturalness (NISQA 3.69) and a human misclassification rate of 61.9%.

  12. UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

    cs.SD 2025-06 conditional novelty 5.0 of 10

    UmbraTTS jointly synthesizes speech and environmental audio via conditional flow matching, conditioned on text and acoustic context, with controllable background volume.

  13. BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing

    eess.AS 2025-06 conditional novelty 5.0 of 10

    Quantization-aware training with ternary weights plus base-3 weight indexing reduces a JETS/HiFi-GAN TTS model from 25.66 MB to 4.39 MB while keeping naturalness MOS around 3.1 to 3.3.

  14. Marco-Voice Technical Report

    cs.CL 2025-08 reject novelty 4.0 of 10

    Marco-Voice is a TTS system combining voice cloning and emotional speech generation via speaker-emotion disentanglement, contrastive learning, and a new Mandarin emotional dataset, with claimed quality gains over Cosy...

Pith tools