REVIEW 3 major objections 5 minor 4 cited by
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper claims to be the first comprehensive survey of controllable text-to-speech, organizing the field into model architectures, control strategies, and feature representations, and reporting that a Gemini-based evaluator aligns with…
desk verdict Useful, well-organized survey of controllable TTS; the 'first comprehensive' claim is plausible but not backed by a reproducible search protocol, and the Gemini evaluation is a clearly labeled pilot. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is a three-axis taxonomy: architecture (autoregressive vs non-autoregressive), control strategy (style tagging, reference speech prompt, natural language description, instruction-guided control/editing), and feature representation (continuous vs discrete tokens). The taxonomy is what turns a list of papers into a map, letting the survey claim comprehensiveness and letting readers locate any method by its position in the three dimensions. The evaluation pipeline is the second piece of machinery: a fixed prompt given to Gemini asking for 1–5 ratings on instruction following, naturalness, and expressiveness, used to rank ten systems and to compute Pearson correlations against human ratings in a 96-sample comparison.
What would settle it
Run a systematic search for controllable TTS papers published before the survey's cutoff that are absent from its taxonomy; if a coherent line of work (e.g., a pre-2024 method family) is missing, the 'first comprehensive' claim fails. Alternatively, re-run the Appendix A.5 evaluation with human raters on 100+ samples per model; if NISQA or UTMOS matches human preference as well as or better than Gemini on instruction following, the claimed advantage of the proposed pipeline collapses.
Extended reading notes
Core claim
The central claim is that controllable TTS is now a distinct, rapidly growing research area with an identifiable history and structure, and that this paper is the first survey to cover it comprehensively. On the paper's own terms, the key discovery is a three-part organization — autoregressive versus non-autoregressive architectures, four to five control strategies ranging from style tagging to instruction-guided editing, and continuous versus discrete feature representations — that places every major controllable TTS method since 2018. The secondary discovery is empirical: a Gemini-based evaluation of ten TTS systems across two tasks shows that instruction-based methods outperform zero-shot methods on all measured dimensions, and that the proposed MLLM judge correlates with human preference more strongly than NISQA or UTMOS on instruction following, naturalness, and expressiveness.
Load-bearing premise
The survey's 'first comprehensive' claim rests on the assumption that its literature search and three-axis taxonomy capture essentially all major controllable-TTS work, and the evaluation claim rests on the assumption that 20 samples per model with single-pass Gemini ratings represent human judgment.
Editorial extensions
If this is right
- A researcher can locate any controllable TTS method quickly using the three-axis taxonomy and the accompanying method table.
- The reported results imply that instruction-guided TTS models currently beat zero-shot cloning models on controllability and expressiveness, with CosyVoice and MiniMax TTS leading the tested systems.
- The evaluation results imply that a single multimodal LLM can serve as a low-cost proxy for human listeners on controllability dimensions, though the paper's own numbers show only modest absolute correlations.
- The survey identifies open problems — fine-grained attribute control, feature disentanglement, dataset scarcity, and long emotional speech — as the likely next battlegrounds.
Reading between the lines
- If the Gemini-based pipeline generalizes beyond the 20-sample-per-model setup, it could replace costly MOS and CMOS collection for controllability benchmarks; this is an extension the paper only hints at.
- The taxonomy's four control strategies could be compressed into a single spectrum from explicit (tags) to implicit (free instructions), which might make the survey's own 'instruction-guided editing' category a special case of description-based control.
- A natural stress test is to check whether the survey's coverage misses pre-2024 work: the 'first comprehensive' claim would weaken if a significant controllable-TTS line of work predating the listed entries is absent.
- Because attributes are correlated (changing pitch shifts emotion), a testable extension is an interaction-aware control benchmark that the survey explicitly leaves unexplored in its limitations.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of controllable text-to-speech (TTS), organized around three axes: model architectures (autoregressive and non-autoregressive), control strategies (style tagging, reference speech prompting, natural-language descriptions, instruction-guided synthesis, and editing), and feature representations (continuous vs. discrete). It claims to be the first comprehensive survey dedicated to controllable TTS, and it includes a bonus Gemini-based evaluation of TTS controllability along three dimensions (instruction following, naturalness, expressiveness). The survey covers a large number of recent methods, datasets, and metrics, and points to a GitHub repository for a paper list.
Significance. If the comprehensiveness claim is substantiated, this survey would provide a useful entry point to a rapidly growing field and a practical taxonomy for organizing methods. The three-axis framing (architecture, control strategy, feature representation) is a reasonable organizing principle, and the emphasis on LLM-era prompting and instruction-based control is timely. The Gemini-based evaluation is a creative idea with potential practical value for cheap, automated assessment of controllability, though its reliability is not yet established. The manuscript also ships an open GitHub resource, which aids reproducibility and community utility.
major comments (3)
- [Section 8, Figure 1, Section 6] The 'first comprehensive survey' claim (Abstract and Section 6) is not verifiable because the literature search is not described reproducibly. Section 8 lists sources (Google Scholar, arXiv, DBLP, Scopus, ChatGPT) but gives no query strings, date ranges, inclusion/exclusion criteria, or screening workflow. Figure 1's own captions state the statistics are 'incomplete' (twice), and Section 7 explicitly excludes several related areas. The authors should provide a transparent search protocol or soften the strong novelty claim.
- [Table 1 versus Figure 4] Several methods that appear in the control-strategy taxonomy are missing from the summary table of controllable neural TTS methods. For example, Parler-TTS, PromptSpeaker, InstructSpeech, AudioGPT, FunAudioLLM, and SpeechGPT are listed in Figure 4 but do not appear in Table 1. Since Table 1 is described as a summary of existing controllable neural-based methods, these omissions are concrete counterexamples to comprehensiveness; the authors should either add the missing entries or explicitly define the scope of Table 1.
- [Appendix A.5.3, Table 7] The claim that the Gemini-based evaluation 'consistently outperforms both NISQA and UTMOS across all three evaluation dimensions' is not fully supported. NISQA and UTMOS have no instruction-following score, so the comparison is incomplete for that dimension. The reported Pearson correlations are small (0.12, 0.17, 0.14), and no significance tests, confidence intervals, or details of the human-rater protocol (number of raters, rating instructions) are given. The sample size (96 samples) is small. The conclusion should be phrased as preliminary and supported with statistical details.
minor comments (5)
- [Table 6] There is a typo: 'Gemeni' should be 'Gemini'.
- [Table 1 caption] The word 'Controlability' in the table header should be spelled 'Controllability'.
- [Throughout] The model name 'VoxInstruct' is repeatedly typeset as 'V oxInstruct' with an extra space; this should be fixed.
- [Section 4.2.2 and Appendix A.5.1] The evaluation sample-size description is inconsistent: A.5.1 says 20 samples per model per task, while A.5.3 reports using 96 samples; please clarify the actual number used and why the subset was chosen.
- [References] The references list contains duplicate entries for Defossez et al. (2023a and 2023b), which refer to the same paper.
Circularity Check
No significant circularity: the survey is expository, and the Gemini evaluation is validated against external human ratings and compared with independent metrics.
full rationale
The paper is a literature survey; its central product is a taxonomy and a curated method list, not a derived quantitative claim whose inputs are its own outputs. The only quantitative contribution is the Gemini-based evaluation in Appendix A.5. That evaluation is validated by computing Pearson correlations with human ratings on 96 samples and by comparing with the external metrics NISQA and UTMOS (Table 7). Because the human ratings and the NISQA/UTMOS scores are external to the paper's own fitted values, the evaluation is not self-referential: no parameter is fitted to the human labels and then reported as a prediction. The survey's self-citations (Rong and Liu 2025 for image-conditioned TTS; Rong et al. 2025 for storytelling and GPT-based evaluation) are used as examples of existing work, not as authority for the survey's taxonomy or for its 'first comprehensive survey' claim. The admitted incompleteness of the statistics in Figure 1 and the exclusions listed in Section 7 are completeness limitations, not circular reasoning. The 'first comprehensive' novelty claim is unverified and hard to falsify as stated, but that is a substantiveness and correctness concern, not a circularity of the derivation chain.
Assumptions & free parameters
assumptions (3)
- ad hoc to paper The taxonomy of three axes (model architecture, control strategy, feature representation) is a valid and sufficient organizing framework for controllable TTS methods.
- domain assumption The literature search via Google Scholar, arXiv, DBLP, Scopus, and ChatGPT returned a representative sample of the field.
- domain assumption The Gemini ratings reflect human perception of instruction following, naturalness, and expressiveness.
Cite this review
Pith. "Pith review of Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey." pith.science (2026). https://pith.science/paper/5YBVB564
@misc{pith2026241206602,
author = {Pith},
title = {Pith review of: Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/5YBVB564}},
note = {Machine review of arXiv:2412.06602}
}
read the original abstract
Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style. Driven by rising industrial demand and breakthroughs in deep learning, e.g., diffusion and large language models (LLMs), controllable TTS has become a rapidly growing research area. This survey provides the first comprehensive review of controllable TTS methods, from traditional control techniques to emerging approaches using natural language prompts. We categorize model architectures, control strategies, and feature representations, while also summarizing challenges, datasets, and evaluations in controllable TTS. This survey aims to guide researchers and practitioners by offering a clear taxonomy and highlighting future directions in this fast-evolving field. One can visit https://github.com/imxtx/awesome-controllabe-speech-synthesis for a comprehensive paper list and updates.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 4 Pith papers
-
Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models
Speech DF Arena standardizes audio deepfake detection benchmarking across 14 datasets and 15 systems, showing that most open-source detectors have high error rates on out-of-domain attacks.
-
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation
A training-free multi-agent framework that decomposes multimodal inputs into audio events, selects specialized generators, and self-corrects outputs to produce multiple audio types.
-
Position: Towards Responsible Evaluation for Text-to-Speech
A call to reform text-to-speech evaluation around a three-level framework covering metric fidelity, comparability, and ethical oversight.
-
Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis
Rectified flows with a tunable sampling temperature offer the best naturalness-diversity trade-off among stochastic prosody predictors for text-to-speech.
Reference graph
Works this paper leans on
-
[1]
• 2 points: It loosely follows the instruc- tions but misses key elements or timing in parts
Instruction Following (1–5): • 1 point: The audio completely ignores the instruction; it does not follow the in- tended timing, emphasis, or pacing. • 2 points: It loosely follows the instruc- tions but misses key elements or timing in parts. • 3 points: Generally follows the instruc- tion with minor lapses in emphasis or pacing. • 4 points: Clearly follo...
-
[2]
• 2 points: Noticeably synthetic; some un- natural artifacts remain
Naturalness (1–5): • 1 point: The audio sounds fully synthetic or robotic; extremely unnatural. • 2 points: Noticeably synthetic; some un- natural artifacts remain. • 3 points: Moderately natural with occa- sional synthetic artifacts. • 4 points: Largely natural sounding with minor imperfections. • 5 points: Completely natural; indistin- guishable from a ...
-
[3]
Expressiveness (1–5): • 1 point: The audio is flat and monotone; no emotional variation. • 2 points: Minimal expressiveness; emo- tions are weak or inconsistent. • 3 points: Reasonably expressive with some highlights, but could be stronger. • 4 points: Clearly expressive with only slight under- or over-emphasis. • 5 points: Exceptionally expressive; full ...
-
[6]
In Pro- ceedings of the 32nd ACM International Conference on Multimedia, pages 1255–1264
SpeechCraft: A fine-grained expressive speech dataset with natural language description. In Pro- ceedings of the 32nd ACM International Conference on Multimedia, pages 1255–1264. Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan, De- tai Xin, Dongchao Yang, Eric Liu, Yichong Leng, Kaitao Song, Siliang Tang, Zhizheng Wu, Tao Qin, Xiangyang Li, Wei Ye, Shikun Zha...
work page 2024
-
[8]
Advances in Neural Information Processing Systems, 36
V oicebox: Text-guided multilingual univer- sal speech generation at scale. Advances in Neural Information Processing Systems, 36. Keon Lee, Dong Won Kim, Jaehyeon Kim, Seungjun Chung, and Jaewoong Cho. 2025. DiTTo-TTS: Dif- fusion transformers for scalable text-to-speech with- out domain-specific factors. In The Thirteenth Inter- national Conference on L...
arXiv 2025
-
[9]
Alexa vs. Siri vs. Cortana vs. Google assistant: A comparison of speech-based natural user interfaces. In Advances in Human Factors and Systems Interac- tion, pages 241–250. Hui Lu, Xixin Wu, Zhiyong Wu, and Helen Meng
-
[10]
In Proceedings of the 31st ACM In- ternational Conference on Multimedia, pages 2829– 2837
SpeechTripleNet: End-to-end disentangled speech representation learning for content, timbre and prosody. In Proceedings of the 31st ACM In- ternational Conference on Multimedia, pages 2829– 2837. Junchen Lu, Berrak Sisman, Rui Liu, Mingyang Zhang, and Haizhou Li. 2022. VisualTTS: TTS with accu- rate lip-speech synchronization for automatic voice over. In ...
arXiv 2022
-
[11]
Advances in Neural Information Processing Systems, 36:53728– 53741
Direct preference optimization: Your lan- guage model is secretly a reward model. Advances in Neural Information Processing Systems, 36:53728– 53741. Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the lim- its of transfer learning with a unified text-to-text tran...
work page 2020
Show all 24 references
-
[14]
In IEEE Spoken Lan- guage Technology Workshop, pages 595–602
Predicting expressive speaking style from text in end-to-end speech synthesis. In IEEE Spoken Lan- guage Technology Workshop, pages 595–602. IEEE. Youcef Tabet and Mohamed Boughazi. 2011. Speech synthesis techniques. a survey. In International Work- shop on Systems, Signal Pro...
2011 arXiv
-
[15]
arXiv preprint arXiv:2404.03204
RALL-E: Robust codec language modeling with chain-of-thought prompting for text-to-speech synthesis. arXiv preprint arXiv:2404.03204. 19 Junichi Yamagishi, Takashi Nose, Heiga Zen, Zhen-Hua Ling, Tomoki Toda, Keiichi Tokuda, Simon King, and Steve Renals. 2009. Robust speaker-a...
2009 arXiv
-
[17]
In IEEE International Conference on Acoustics, Speech and Signal Processing, pages 6945–6949
Learning latent representations for style con- trol and transfer in end-to-end speech synthesis. In IEEE International Conference on Acoustics, Speech and Signal Processing, pages 6945–6949. Yongmao Zhang, Guanghou Liu, Yi Lei, Yunlin Chen, Hao Yin, Lei Xie, and Zhifei Li. 202...
2023 arXiv
-
[19]
or word embeddings (Almeida and Xexéo,
-
[20]
Speech vocoder is the last com- ponent that converts the intermediate acoustic fea- tures into a waveform that can be played back
as input, which is much more efficient than previous methods. Speech vocoder is the last com- ponent that converts the intermediate acoustic fea- tures into a waveform that can be played back. This step bridges the gap between the acoustic features and the actual sounds produc...
2021
-
[2001]
In 2001 IEEE International Conference on Acoustics, Speech, and Signal Pro- cessing
Perceptual evaluation of speech quality-a new method for speech quality assessment of telephone networks and codecs. In 2001 IEEE International Conference on Acoustics, Speech, and Signal Pro- cessing. Proceedings, pages 749–752. Yan Rong and Li Liu. 2025. Seeing your speech s...
2001 arXiv
-
[2008]
A lower MCD value in- dicates a higher similarity between synthesized and reference speech, meaning better speech synthesis quality
measures the spectral distance between synthesized and reference speech, reflecting how closely the generated audio matches the target in terms of acoustic features. A lower MCD value in- dicates a higher similarity between synthesized and reference speech, meaning better spee...
2020
-
[2015]
In IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing, pages 4475–4479
Multi-speaker modeling and speaker adapta- tion for DNN-based TTS synthesis. In IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing, pages 4475–4479. Yuchen Fan, Yao Qian, Feng-Long Xie, and Frank K Soong. 2014. TTS synthesis with bidirectional LSTM base...
2014 arXiv
-
[2016]
speak calmly
and Tacotron (Wang et al., 2017) demon- strated the potential for prosody control through explicit conditioning (Shen et al., 2018; Ren et al., 21 Control Strategy Core Idea Key Features Pros & Cons Style Tagging Control specific attributes using predefined tags. Direct contro...
2017
-
[2018]
In International Conference on Learning Representations
Adversarial audio synthesis. In International Conference on Learning Representations. Zhihao Du, Qian Chen, Shiliang Zhang, Kai Hu, Heng Lu, Yexin Yang, Hangrui Hu, Siqi Zheng, Yue Gu, Ziyang Ma, and 1 others. 2024. CosyV oice: A scal- able multilingual zero-shot text-to-speec...
2024 arXiv
-
[2019]
In ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5901–5905
Disentangling correlated speaker and noise for speech synthesis via data augmentation and adver- sarial factorization. In ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5901–5905. Wei-Ning Hsu, Yu Zhang, Ron J Weiss,...
2019 arXiv
-
[2020]
In IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing, pages 6199–6203
Parallel WaveGAN: A fast waveform genera- tion model based on generative adversarial networks with multi-resolution spectrogram. In IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing, pages 6199–6203. Dongchao Yang, Rongjie Huang, Yuanyuan Wang, Hao- ha...
-
[2021]
In Conference of the International Speech Communication Association, pages 2756–2760
AISHELL-3: A multi-speaker mandarin TTS corpus. In Conference of the International Speech Communication Association, pages 2756–2760. Reo Shimizu, Ryuichi Yamamoto, Masaya Kawamura, Yuma Shirahata, Hironori Doi, Tatsuya Komatsu, and Kentaro Tachibana. 2024. PromptTTS++: Contro...
2024 arXiv
-
[2023]
In IEEE International Conference on Acoustics, Speech and Signal Processing, pages 1–5
Grad-StyleSpeech: Any-speaker adaptive text- to-speech synthesis with diffusion models. In IEEE International Conference on Acoustics, Speech and Signal Processing, pages 1–5. Hideki Kawahara, Ikuyo Masuda-Katsuse, and Alain De Cheveigne. 1999. Restructuring speech rep- resent...
1999 arXiv
-
[2024]
arXiv preprint arXiv:2412.12498
Hierarchical control of emotion rendering in speech synthesis. arXiv preprint arXiv:2412.12498. Fumitada Itakura. 1975. Line spectrum representa- tion of linear predictor coefficients of speech signals. The Journal of the Acoustical Society of America , 57(S1):S35–S35. Won Jan...
1975 arXiv
-
[2025]
In IEEE International Conference on Acoustics, Speech and Signal Processing, pages 1–5
DrawSpeech: Expressive speech synthesis using prosodic sketches as control conditions. In IEEE International Conference on Acoustics, Speech and Signal Processing, pages 1–5. Wenxi Chen, Ziyang Ma, Ruiqi Yan, Yuzhe Liang, Xi- quan Li, Ruiyang Xu, Zhikang Niu, Yanqiao Zhu, Yifa...
2015 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.