{"id":"ce1e8c4d-bb6b-4d18-bd2c-2d5b6b4dc766","arxiv_id":"2412.06602","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of controllable text-to-speech methods with a small pilot study using Gemini to automatically rate instruction following, naturalness, and expressiveness.","lead":"This paper surveys many text-to-speech systems that let users control emotion, timbre, style, and other voice attributes, and sorts them into a taxonomy. It also tests whether Google Gemini can automatically judge how well synthesized speech follows instructions.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'first comprehensive systematic survey' claim is undercut by the paper's own incomplete statistics and non-systematic search protocol; no verifiable completeness criteria are given.","rationale":"The paper's title and abstract frame the contribution as a 'systematic survey' and the first comprehensive review of controllable TTS. That is the central claim most likely to determine the paper's value to the community. The Gemini-based evaluation in Appendix A.5 is explicitly exploratory ('To some extent,' 'future work'), so its statistical weaknesses are secondary. The survey's comprehensiveness claim, in contrast, is the paper's primary raison d'être. It is also the least secure: the authors list search sources but no protocol, their own Figure 1 is labeled 'incomplete,' and Section 7 excludes entire related areas. Any of these features could be acceptable in a non-systematic overview, but they directly undermine a 'first comprehensive systematic survey' claim. The reader's weakest assumption already identified incompleteness as the key risk; I agree with that assessment. The proposed concrete test would settle it by checking for prior dedicated surveys and quantifying omitted works. Until such a check is run, conditional acceptance with a requirement to clarify completeness criteria remains the appropriate verdict, so the reader's CONDITIONAL recommendation stands unchanged.","tokens_in":38707,"tokens_out":4064,"duration_ms":41539,"concrete_test":"Perform a pre-registered systematic search in Google Scholar, arXiv, DBLP, and Scopus with query variants: 'controllable text-to-speech' OR 'controllable speech synthesis' AND (survey OR review OR overview), restricted to December 2024 or earlier. Screen titles and abstracts for any prior work dedicated to controllable TTS. Separately, reconstruct the inclusion criteria by having two independent annotators re-code all papers in Figure 1 and Table 1 against the survey's three-axis taxonomy, and record any candidate methods (e.g., leading controllable TTS systems before December 2024) that are absent from the GitHub list. If a prior dedicated survey is found, or if more than 5% of candidates are missing, the 'first comprehensive' claim should be downgraded to a useful but non-exhaustive overview.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that this is the first comprehensive survey of controllable TTS (Section 6). For this to hold, the literature selection must be both complete and reproducible. The paper does not provide a systematic search protocol: Section 8 lists sources (Google Scholar, arXiv, DBLP, Scopus, ChatGPT) but no query strings, date ranges, inclusion/exclusion criteria, or screening workflow, despite the title's word 'Systematic.' More tellingly, Figure 1's own captions state the statistics are 'incomplete' (twice), and Section 7 explicitly excludes related areas (speech enhancement, separation, pretraining, and speech-to-speech translation). The GitHub repository is described as 'comprehensive,' but the figure's incompleteness suggests the inventory is a convenience sample, not a systematic enumeration. If a prior dedicated survey exists (e.g., any 2019–2024 review with 'controllable TTS' in scope), or if major methods are missing from Table 1 and Figure 1, the novelty claim collapses. The paper itself provides no falsifiable criterion for 'comprehensive,' so the claim is currently unverifiable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of controllable text-to-speech (TTS), organized around three axes: model architectures (autoregressive and non-autoregressive), control strategies (style tagging, reference speech prompting, natural-language descriptions, instruction-guided synthesis, and editing), and feature representations (continuous vs. discrete). It claims to be the first comprehensive survey dedicated to controllable TTS, and it includes a bonus Gemini-based evaluation of TTS controllability along three dimensions (instruction following, naturalness, expressiveness). The survey covers a large number of recent methods, datasets, and metrics, and points to a GitHub repository for a paper list.","tokens_in":1164,"tokens_out":3543,"duration_ms":89224,"significance":"If the comprehensiveness claim is substantiated, this survey would provide a useful entry point to a rapidly growing field and a practical taxonomy for organizing methods. The three-axis framing (architecture, control strategy, feature representation) is a reasonable organizing principle, and the emphasis on LLM-era prompting and instruction-based control is timely. The Gemini-based evaluation is a creative idea with potential practical value for cheap, automated assessment of controllability, though its reliability is not yet established. The manuscript also ships an open GitHub resource, which aids reproducibility and community utility.","major_comments":[{"comment":"The 'first comprehensive survey' claim (Abstract and Section 6) is not verifiable because the literature search is not described reproducibly. Section 8 lists sources (Google Scholar, arXiv, DBLP, Scopus, ChatGPT) but gives no query strings, date ranges, inclusion/exclusion criteria, or screening workflow. Figure 1's own captions state the statistics are 'incomplete' (twice), and Section 7 explicitly excludes several related areas. The authors should provide a transparent search protocol or soften the strong novelty claim.","section":"Section 8, Figure 1, Section 6"},{"comment":"Several methods that appear in the control-strategy taxonomy are missing from the summary table of controllable neural TTS methods. For example, Parler-TTS, PromptSpeaker, InstructSpeech, AudioGPT, FunAudioLLM, and SpeechGPT are listed in Figure 4 but do not appear in Table 1. Since Table 1 is described as a summary of existing controllable neural-based methods, these omissions are concrete counterexamples to comprehensiveness; the authors should either add the missing entries or explicitly define the scope of Table 1.","section":"Table 1 versus Figure 4"},{"comment":"The claim that the Gemini-based evaluation 'consistently outperforms both NISQA and UTMOS across all three evaluation dimensions' is not fully supported. NISQA and UTMOS have no instruction-following score, so the comparison is incomplete for that dimension. The reported Pearson correlations are small (0.12, 0.17, 0.14), and no significance tests, confidence intervals, or details of the human-rater protocol (number of raters, rating instructions) are given. The sample size (96 samples) is small. The conclusion should be phrased as preliminary and supported with statistical details.","section":"Appendix A.5.3, Table 7"}],"minor_comments":[{"comment":"There is a typo: 'Gemeni' should be 'Gemini'.","section":"Table 6"},{"comment":"The word 'Controlability' in the table header should be spelled 'Controllability'.","section":"Table 1 caption"},{"comment":"The model name 'VoxInstruct' is repeatedly typeset as 'V oxInstruct' with an extra space; this should be fixed.","section":"Throughout"},{"comment":"The evaluation sample-size description is inconsistent: A.5.1 says 20 samples per model per task, while A.5.3 reports using 96 samples; please clarify the actual number used and why the subset was chosen.","section":"Section 4.2.2 and Appendix A.5.1"},{"comment":"The references list contains duplicate entries for Defossez et al. (2023a and 2023b), which refer to the same paper.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The survey covers a broad and timely topic and contains a lot of useful material, but the central novelty claim (first comprehensive survey) is not backed by a reproducible search protocol, and the Gemini evaluation's conclusion is stronger than the evidence supports. I would recommend major revision, asking the authors to either document the search methodology fully or rephrase the claim, and to tighten the evaluation conclusions. The Gemini evaluation, while interesting, may be better placed in a separate methods paper rather than embedded in a survey."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick take: this is a genuinely useful survey of controllable TTS, and the three-axis taxonomy (architecture, control strategy, feature representation) is well thought out. The tables are dense and current through early 2025. The \"first comprehensive survey\" claim is plausible but not fully nailed down, and the search protocol isn't systematic in the sense that word usually implies. The Gemini evaluation is a small pilot, but it's clearly labeled as such.\n\nWhat's new: the organization itself is the contribution. The discussion of control strategies—style tags, reference speech, descriptions, instructions—is the best part, and the limitations section is honest. They admit skipping enhancement, separation, pretraining, and speech-to-speech translation, and they flag that attribute interactions and societal risks are out of scope.\n\nSoft spots: the title says \"systematic survey,\" but Section 8 lists sources without query strings, inclusion/exclusion criteria, or a screening workflow. That makes \"systematic\" a stretch and \"first comprehensive\" hard to verify. Figure 1's own captions say \"incomplete\" twice, which undercuts the trend narrative. They could fix this by stating exactly what they searched, what they included, and what they knowingly left out. The Gemini evaluation in Appendix A.5 is a nice experiment—20 samples per model per task, one pass from Gemini 2.5 Flash—but it's a pilot. The correlation coefficients (0.12 to 0.17) are modest, and there's no significance testing. To their credit, they don't overclaim: it's explicitly preliminary and they promise a fuller benchmark later.\n\nBottom line: this is a reference survey, not a landmark result. It will help anyone entering the field and serves as a useful checklist for practitioners. For peer review, I'd send it out with a request to either temper the \"systematic\"/\"first comprehensive\" language or back it with a reproducible search protocol and a completeness audit. The authors seem capable of that.\n\nRecommendation: yes, engage with it.","headline":"Useful, well-organized survey of controllable TTS; the 'first comprehensive' claim is plausible but not backed by a reproducible search protocol, and the Gemini evaluation is a clearly labeled pilot.","tokens_in":39421,"tokens_out":2174,"would_cite":true,"duration_ms":23129,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims to be the first comprehensive survey of controllable text-to-speech, organizing the field into model architectures, control strategies, and feature representations, and reporting that a Gemini-based evaluator aligns with…","keywords":["controllable text-to-speech","large language models","speech synthesis survey","natural language prompting","instruction-guided synthesis","evaluation metrics","voice cloning","prosody control"],"falsifier":"Run a systematic search for controllable TTS papers published before the survey's cutoff that are absent from its taxonomy; if a coherent line of work (e.g., a pre-2024 method family) is missing, the 'first comprehensive' claim fails. Alternatively, re-run the Appendix A.5 evaluation with human raters on 100+ samples per model; if NISQA or UTMOS matches human preference as well as or better than Gemini on instruction following, the claimed advantage of the proposed pipeline collapses.","tokens_in":38552,"feed_emoji":"🗣️","tokens_out":4555,"duration_ms":42424,"temperature":0.7,"pith_summary":"This survey sets out to be the first comprehensive review of controllable text-to-speech (TTS), which lets users steer attributes such as emotion, timbre, and speaking style rather than just generating natural speech. It claims that previous TTS surveys overlooked controllability, and it organizes the field into three axes: model architectures, control strategies, and feature representations. It also proposes a Gemini-based evaluation pipeline for instruction following, naturalness, and expressiveness, reporting that it aligns with human ratings better than NISQA and UTMOS across all three dimensions. A sympathetic reader would take the survey's value as the taxonomy and the comparison table of methods, plus the demonstration that multimodal LLM judges are a promising low-cost evaluation route.","feed_headline":"First comprehensive survey maps controllable TTS methods","feed_subtitle":"An LLM-based evaluator that tracks human judgment on instruction following, naturalness, and expressiveness.","key_machinery":"The organizing device is a three-axis taxonomy: architecture (autoregressive vs non-autoregressive), control strategy (style tagging, reference speech prompt, natural language description, instruction-guided control/editing), and feature representation (continuous vs discrete tokens). The taxonomy is what turns a list of papers into a map, letting the survey claim comprehensiveness and letting readers locate any method by its position in the three dimensions. The evaluation pipeline is the second piece of machinery: a fixed prompt given to Gemini asking for 1–5 ratings on instruction following, naturalness, and expressiveness, used to rank ten systems and to compute Pearson correlations against human ratings in a 96-sample comparison.","core_discovery":"The central claim is that controllable TTS is now a distinct, rapidly growing research area with an identifiable history and structure, and that this paper is the first survey to cover it comprehensively. On the paper's own terms, the key discovery is a three-part organization — autoregressive versus non-autoregressive architectures, four to five control strategies ranging from style tagging to instruction-guided editing, and continuous versus discrete feature representations — that places every major controllable TTS method since 2018. The secondary discovery is empirical: a Gemini-based evaluation of ten TTS systems across two tasks shows that instruction-based methods outperform zero-shot methods on all measured dimensions, and that the proposed MLLM judge correlates with human preference more strongly than NISQA or UTMOS on instruction following, naturalness, and expressiveness.","pith_inferences":["If the Gemini-based pipeline generalizes beyond the 20-sample-per-model setup, it could replace costly MOS and CMOS collection for controllability benchmarks; this is an extension the paper only hints at.","The taxonomy's four control strategies could be compressed into a single spectrum from explicit (tags) to implicit (free instructions), which might make the survey's own 'instruction-guided editing' category a special case of description-based control.","A natural stress test is to check whether the survey's coverage misses pre-2024 work: the 'first comprehensive' claim would weaken if a significant controllable-TTS line of work predating the listed entries is absent.","Because attributes are correlated (changing pitch shifts emotion), a testable extension is an interaction-aware control benchmark that the survey explicitly leaves unexplored in its limitations."],"forward_implications":["A researcher can locate any controllable TTS method quickly using the three-axis taxonomy and the accompanying method table.","The reported results imply that instruction-guided TTS models currently beat zero-shot cloning models on controllability and expressiveness, with CosyVoice and MiniMax TTS leading the tested systems.","The evaluation results imply that a single multimodal LLM can serve as a low-cost proxy for human listeners on controllability dimensions, though the paper's own numbers show only modest absolute correlations.","The survey identifies open problems — fine-grained attribute control, feature disentanglement, dataset scarcity, and long emotional speech — as the likely next battlegrounds."],"supporting_citations":[{"why":"Earlier TTS review focused on text analysis, used to show prior work did not center controllability.","marker":"Klatt (1987)"},{"why":"Neural TTS survey cited as overlooking controllability.","marker":"Ning et al. (2019)"},{"why":"Neural speech synthesis survey, the baseline the paper distinguishes itself from.","marker":"Tan et al. (2021b)"},{"why":"Audio diffusion survey, likewise lacking controllability focus.","marker":"Zhang et al. (2023a)"},{"why":"PromptTTS, the key description-based method the survey positions as a new control paradigm.","marker":"Guo et al. (2023)"},{"why":"VoxInstruct, the instruction-guided paradigm that anchors the latest control strategy.","marker":"Zhou et al. (2024)"},{"why":"NISQA, the baseline model-based evaluator the Gemini pipeline is compared against.","marker":"Mittag et al. (2021)"},{"why":"UTMOS, the second baseline evaluator in the comparison.","marker":"Saeki et al. (2022)"}],"fun_headline_variants":["First comprehensive survey of controllable TTS","Survey maps controllable TTS from style to LLM prompts","LLM-based evaluator improves TTS quality assessment","Controllable TTS: first survey and LLM judge","Survey: controllable TTS methods, LLM-based evaluation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's 'first comprehensive' claim rests on the assumption that its literature search and three-axis taxonomy capture essentially all major controllable-TTS work, and the evaluation claim rests on the assumption that 20 samples per model with single-pass Gemini ratings represent human judgment.","fun_headline_variants_meta":{"raw":{"variants":["First comprehensive survey of controllable TTS","Survey maps controllable TTS from style to LLM prompts","LLM-based evaluator improves TTS quality assessment","Controllable TTS: first survey and LLM judge","Survey: controllable TTS methods, LLM-based evaluation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000548,"raw_usage":{"total_tokens":2565,"prompt_tokens":838,"completion_tokens":1727,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":454,"completion_tokens_details":{"reasoning_tokens":1650}},"tokens_in":454,"tokens_out":1727,"duration_ms":13068,"temperature":1.0,"reasoning_tokens":1650,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:28:18.402258+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a systematic search for controllable TTS papers published before the survey's cutoff that are absent from its taxonomy; if a coherent line of work (e.g., a pre-2024 method family) is missing, the 'first comprehensive' claim fails. Alternatively, re-run the Appendix A.5 evaluation with human raters on 100+ samples per model; if NISQA or UTMOS matches human preference as well as or better than Gemini on instruction following, the claimed advantage of the proposed pipeline collapses.","supporting_citations":[],"review_version":1}