REVIEW 14 cited by
A Survey on Neural Speech Synthesis
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text to speech (TTS), or speech synthesis, which aims to synthesize intelligible and natural speech given text, is a hot research topic in speech, language, and machine learning communities and has broad applications in the industry. As the development of deep learning and artificial intelligence, neural network-based TTS has significantly improved the quality of synthesized speech in recent years. In this paper, we conduct a comprehensive survey on neural TTS, aiming to provide a good understanding of current research and future trends. We focus on the key components in neural TTS, including text analysis, acoustic models and vocoders, and several advanced topics, including fast TTS, low-resource TTS, robust TTS, expressive TTS, and adaptive TTS, etc. We further summarize resources related to TTS (e.g., datasets, opensource implementations) and discuss future research directions. This survey can serve both academic researchers and industry practitioners working on TTS.
Forward citations
Cited by 14 Pith papers
-
Probing Speaker Identity Sensitivity in Audio Deepfake Detectors
A new Identity Sensitivity Score flags misclassified audio deepfake detections with AUC up to 0.954, but its claim to isolate speaker-identity behavior from plain confidence is not yet controlled.
-
NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations
A Mandarin speech pipeline that tags non-word vocalizations (laughter, breath, filled pauses) for ASR and controls their generation in TTS, backed by a claimed 573-hour, 174,179-utterance word-level annotated corpus.
-
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.
-
SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch
SLASH adds DSP-derived absolute pitch objectives, including direct spectrogram generation from F0, to self-supervised pitch estimation and beats DSP and SSL baselines on MIR-1K.
-
Multi-interaction TTS toward professional recording reproduction
Multi-turn textual directions can iteratively refine the speaking style of synthesized speech through a learned embedding refiner, with modest but measurable alignment to the directions.
-
Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study
No single Urdu TTS system wins across all metrics, and the system listeners liked most, Google Gemini, was the furthest from reference audio on objective acoustic scores.
-
Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model
A lightweight face adapter plus soft-tuning aligns face embeddings to a frozen StyleTTS 2 style space, yielding natural zero-shot face-to-speech and language-agnostic transfer to Spanish.
-
Position: Towards Responsible Evaluation for Text-to-Speech
A call to reform text-to-speech evaluation around a three-level framework covering metric fidelity, comparability, and ethical oversight.
-
Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
Speaker-conditioned phrasing with phoneme-level PLMs (MP BERT) improves pause prediction from F0.5 0.3719 to 0.4991, and a few-shot adapter generalizes to unseen speakers.
-
Traceable TTS: Toward Watermark-Free TTS with Strong Traceability
A joint training loop makes an F5-TTS model produce audio that a paired wav2vec 2.0/LCNN discriminator can recognize, enabling watermark-free attribution; however, the reported generalization gain is not isolated from...
-
Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges
A speech-deepfake dataset for ten public figures built with transcription-based segmentation reports high synthetic naturalness (NISQA 3.69) and a human misclassification rate of 61.9%.
-
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
UmbraTTS jointly synthesizes speech and environmental audio via conditional flow matching, conditioned on text and acoustic context, with controllable background volume.
-
BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing
Quantization-aware training with ternary weights plus base-3 weight indexing reduces a JETS/HiFi-GAN TTS model from 25.66 MB to 4.39 MB while keeping naturalness MOS around 3.1 to 3.3.
-
Marco-Voice Technical Report
Marco-Voice is a TTS system combining voice cloning and emotional speech generation via speaker-emotion disentanglement, contrastive learning, and a new Mandarin emotional dataset, with claimed quality gains over Cosy...
Discussion (0). Sign in to comment.