{"id":"d2ccb96f-555e-4165-92c8-f97a29160f44","arxiv_id":"2501.06276","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"PROEMO extends FastSpeech 2 with HuBERT-based emotion and intensity encoders plus GPT-4 prompt scaling to generate multi-speaker emotional speech with controllable intensity.","lead":"This paper combines a FastSpeech 2 text-to-speech engine with emotion and intensity encoders and a GPT-4 prompt stage that adjusts pitch, duration and energy. The result is multi-speaker expressive speech whose emotional intensity can be varied, with evaluations on the ESD corpus.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The intensity-control claim rests on a circular PIR test: the reference labels come from the same learned rank function used to train and condition the model, so the reported 72% accuracy does not independently validate perceived emotional intensity.","rationale":"The most load-bearing weakness is in the intensity-control evaluation, not in GPT-4 behavior. The GPT-4 prompt-control component is unvalidated, but its effect is at least indirectly observable in the ECA table: local-level prompt control improves classification accuracy for both FS2w/Emo and FS2w/Emo&Int. The intensity-control claim, by contrast, has no external anchor: the learned rank function defines the training target, the sampling of Low/Medium/High test items, and the PIR reference labels. The 72% PIR agreement is therefore not an independent confirmation that listeners perceive the intended emotion-intensity levels. This self-referential evaluation should block acceptance of a central contribution unless a human-validated intensity scale is supplied. The reader's weakest_assumption concerned GPT-4 stability, but the reader's rationale did mention self-referential PIR labeling, so my emphasis is partial agreement. A CONDITIONAL verdict remains appropriate, contingent on the independent human-rating check.","tokens_in":8278,"tokens_out":6313,"duration_ms":63666,"concrete_test":"Run an independent human-rating study on ESD evaluation utterances: have at least five listeners rate each utterance's emotional intensity on a continuous scale, compute Spearman correlation between mean human ratings and the learned rank function r(xA), and then re-run the PIR test using human-derived Low/Medium/High thresholds instead of rank-function labels. If the correlation is below roughly 0.5 or the PIR accuracy drops substantially, the intensity-control claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2 states: 'Since ground truth intensity annotations are lacking, we derive them using the learned ranked function, detailed in Section 3.3, for the evaluation dataset. These annotations serve as a reference during the PIR test.' Section 5.2 then reports that participants' rankings were 'compared with intensity annotations derived from the learned rank function' and achieves approximately 72% accuracy. This is circular: the intensity encoder is trained by regression to the same rank function, the Low/Medium/High test samples are generated from those labels, and the PIR reference labels are produced by that same function. No independent human annotation validates either the rank function or the perceived intensity of the generated samples. If the rank function mostly tracks loudness or arousal rather than emotion-specific intensity, the 72% figure only shows that listeners can hear the acoustic manipulation, not that the model controls emotional intensity. The ECA improvements from prompt control are separate and may stand, but the contribution claim of 'generate multi-speaker expressive speech with varying emotional intensity' is not established by the reported evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces PROEMO, an expressive text-to-speech framework built on FastSpeech 2. It adds two HuBERT-based auxiliary encoders for emotion classification and intensity regression, and a GPT-4 prompt-control module that rescales pitch, duration, and energy at global and word levels during inference. Experiments on the ESD dataset report objective metrics (emotion classification accuracy, MCD, WER, CER) and subjective tests (MOS and perceptual intensity ranking), with the central claims that prompt control, especially local control, improves emotional expressiveness and that the framework generates multi-speaker speech with varying emotional intensity.","tokens_in":8495,"tokens_out":4747,"duration_ms":46007,"significance":"If the claims are substantiated, the paper offers a practical and modular extension of FastSpeech 2: emotion and intensity conditioning without style-prompt annotated data, using only LibriTTS for pretraining and ESD for fine-tuning. The design is clear and reproducible, and the prompt template with step-by-step reasoning is a potentially useful engineering artifact. The most credible evidence is the consistent ECA improvement with local prompt control in the FS2w/E and FS2w/E&I rows of Table 1 (e.g., 79.72% vs. 74.80% for FS2w/E&I). However, the intensity-control claim is not independently validated because the PIR reference labels and the intensity encoder's training target come from the same learned ranking function; the GPT-4 module's behavior is not quantitatively characterized; and the subjective tests lack inferential statistics. These gaps materially weaken the paper's central contribution as currently stated.","major_comments":[{"comment":"The PIR validation is circular. Section 4.2 states that intensity annotations are derived from the learned rank function r(x) described in Section 3.3, and Section 3.3 supervises the intensity encoder by regression to that same function. The Low/Medium/High test samples are generated using these labels, and the PIR reference is produced by the same function. Therefore the reported ~72% participant accuracy shows only that listeners can perceive acoustic differences that align with the model's internal ordering; it does not demonstrate control of emotional intensity on an independent scale. The authors should validate perceived intensity against human ratings of natural reference utterances, externally annotated intensity labels, or a forced-choice design built on human-ordered natural speech.","section":"Section 4.2 / Section 3.3"},{"comment":"The GPT-4 prompt-control module is the core mechanism claimed to improve expressiveness, but the paper provides no quantitative analysis of GPT-4's outputs. The authors note that the prompt from [16] produced unstable output and was redesigned, yet no data are reported on the distribution of predicted scaling factors, failure rate, stability across repeated API calls, or sensitivity to prompt wording. At minimum, an ablation with random scaling factors sampled from the same ranges (Eqs. 1-3) would show whether the ECA improvements come from the prompt semantics or merely from the perturbation of prosodic features.","section":"Section 3.4"},{"comment":"The human subjective results lack error bars and significance testing. MOS values in Table 1 (e.g., 3.408 vs. 3.728 for FS2w/E&I with L vs. G&L control) are reported for 20 participants without confidence intervals or pairwise tests, and the PIR result is reported only as 'approximately 72%' with no per-condition breakdown or chance-level comparison. The authors should report per-item standard errors, confidence intervals, and appropriate statistical tests (e.g., Wilcoxon signed-rank for MOS, binomial or permutation test for PIR), ideally with per-emotion breakdowns.","section":"Section 5.2"},{"comment":"The claim that local-level prompt control 'consistently' improves ECA over no prompt control is too broad. In Table 1, Daft-Exprt with local control has ECA 0.481 versus 0.663 without prompt control, and FS2 with local control (0.296) is barely different from FS2 without control (0.247). The consistent trend actually holds only for the FS2w/E and FS2w/E&I models. The statement should be explicitly restricted to those configurations.","section":"Section 5.1, Table 1"}],"minor_comments":[{"comment":"Typo: 'it failes to converge' should be 'it fails to converge'.","section":"Section 4.1"},{"comment":"Typo: 'classification accuracy improves to79.72%' is missing a space before the number.","section":"Section 5.1"},{"comment":"The sentence 'These participants from diverse geographical regions are expertise in speech and NLP' should read 'are experts in speech and NLP.'","section":"Section 5.2"},{"comment":"The complete prompt template is referenced in Figure 2 but not fully shown in the paper; including the full prompt in an appendix would improve reproducibility.","section":"Figure 2 / Section 3.4"},{"comment":"Pitch is scaled additively while duration and energy are scaled multiplicatively; a brief justification for this asymmetry would help readers interpret the scaling factors produced by GPT-4.","section":"Equations (1)-(3)"}],"recommendation":"major_revision","confidential_remarks":"The central architecture is plausible and the ECA evidence for prompt control in the FS2-based models is worth reporting, but the intensity-control claim cannot be assessed without breaking the circularity of the PIR evaluation. I would encourage the editor to require independent human intensity annotations or an external intensity benchmark, plus a quantitative characterization of the GPT-4 module, before considering acceptance. The manuscript is within scope for a speech/audio venue and the engineering contribution is potentially useful."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should read this as a solid, honest incremental paper, not a breakthrough. What is actually new: they extend the Sigurgeirsson and King prompt-scaling idea to a multi-speaker FS2 backbone, add a HuBERT-based emotion encoder and a per-speaker, per-emotion intensity rank function, and show that word-level LLM-prompted prosody scaling consistently nudges emotion classification accuracy up (e.g., 74.80 to 79.72 for the full model). The objective tables are reported with enough detail to reproduce the trend, and the architecture is a sensible integration of known parts. The paper does not oversell its novelty relative to [16]; it explicitly frames its contribution as the combination rather than a new paradigm.\n\nThe soft spots are real but not fatal. The biggest is the circular PIR test. The paper states in Section 4.2 that, lacking ground-truth intensity annotations, they derive reference labels from the learned rank function r(x) described in Section 3.3, and Section 5.2 shows participants agreeing with those labels at about 72%. That accuracy measures how well listeners hear the acoustic differences the model was trained to produce, not whether the intensity levels correspond to any external or human notion of emotional intensity. The authors do not hide this; they disclose it explicitly. Still, it means the contribution claim 'generate multi-speaker expressive speech with varying emotional intensity' is weaker than the abstract implies. The ECA evidence for emotion control stands independently and is the stronger part of the paper.\n\nSecond, the inference-time reliance on GPT-4 with no quantitative analysis of its outputs is a genuine gap. The authors note the original prompt was unstable and had to be redesigned, but they give no failure rate, no output distribution stats, and no sensitivity analysis to prompt wording. Since all prompt-control gains flow through GPT-4, this is a load-bearing unvalidated component. It is a minor-to-moderate concern because the objective trend is consistent across two model variants and three settings, but a reviewer should ask for at least a small stability study.\n\nMinor issues: no error bars or significance tests on the 20-participant MOS/PIR results; the Daft-Exprt comparison is weakened by its known single-speaker pretraining and the authors acknowledge convergence problems; the reader's circularity concern holds up on reading the paper, but calling it a 'fatal' flaw would be too strong.\n\nVerdict: worth refereeing. The empirical trend is coherent, the writing is clear, and the limitations are mostly disclosed in-text. A serious referee should ask for an independent human intensity annotation set, GPT-4 output statistics, and significance testing. I would not cite it this year for the intensity claim, but I would cite it as an example of the prompt-scaling-plus-intelligence approach if that becomes relevant. Bring it to reading group if you want a concrete case of circular evaluation in affective computing; otherwise it is a competent workshop-to-conference-level paper.","headline":"PROEMO is a decent incremental expressive-TTS paper with a real advance in multi-speaker emotion plus intensity control, but its headline intensity claim rests on a circular evaluation that the authors describe openly in the text.","tokens_in":9018,"tokens_out":725,"would_cite":false,"duration_ms":9237,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PROEMO claims that a GPT-4 prompt-control module on top of FastSpeech 2 with emotion and intensity encoders lets TTS generate multi-speaker expressive speech with controllable emotion and intensity, improving emotion classification…","keywords":["text-to-speech","emotional speech synthesis","emotion intensity control","prompt-based prosody control","large language models","FastSpeech 2","HuBERT","expressive speech synthesis"],"falsifier":"Rerun the evaluation with GPT-4's scaling factors replaced by random draws from the same mapped ranges (energy 0.5-2, duration 0.74-1.34, pitch scaled within the predicted range); if emotion classification accuracy does not drop below the reported 79.72 percent level, the prompt-control module is not the mechanism driving the gain.","tokens_in":8093,"feed_emoji":"🗣️","tokens_out":7770,"duration_ms":65279,"temperature":0.7,"pith_summary":"PROEMO is an attempt to give text-to-speech systems fine-grained emotional control without retraining for each emotion. The paper extends FastSpeech 2 with two HuBERT-based encoders, one that classifies emotion and one that regresses emotional intensity, and then uses GPT-4 during inference to propose global and per-word scaling factors for pitch, energy, and duration. The claimed effect is that synthesized speech becomes expressive in a controllable way: word-level prompt control raises emotion classification accuracy from 74.80 percent to 79.72 percent on the best model, and listeners sort generated samples into low, medium, and high intensity with roughly 72 percent accuracy. A sympathetic reader would take the central claim to be that prompt-driven prosody scaling, on top of learned emotion and intensity embeddings, is sufficient to steer multi-speaker expressive speech.","feed_headline":"GPT-4 prompts steer synthetic speech emotion to 79.7 percent accuracy","feed_subtitle":"Per-word edits to pitch, energy, and duration let each TTS voice express five emotions at a chosen intensity.","key_machinery":"The object that carries the argument is the modified variance adapter of FastSpeech 2, now conditioned by two HuBERT-based encoders and then rescaled at inference by LLM-chosen factors. The emotion encoder adds a classification head to HuBERT; the intensity encoder adds a regression head trained on continuous intensity scores produced by a learned per-speaker, per-emotion relative ranking function $r(x_A)=W x_A$ over openSMILE features. The prompt-control step asks GPT-4 to output scaling values for pitch, energy, and duration, maps them through a quadratic function onto preset ranges (duration 0.74-1.34, energy 0.5-2, pitch within the predicted range), and applies them via $d'_i = d_i G_d \\sigma_i$, $e'_i = e_i G_e \\epsilon_i$, $p'_i = p_i + G_p + \\pi_i$. This is what lets the same trained model shift emotional tone and intensity at inference without additional training.","core_discovery":"The paper's central claim is that a TTS pipeline can generate multi-speaker expressive speech with controllable emotion category and emotional intensity by combining learned emotion and intensity representations with inference-time prosody prompting. On top of a FastSpeech 2 backbone, an emotion encoder and an intensity encoder inject emotional content through the variance adapter, while GPT-4 proposes scaling factors that modify predicted pitch, energy, and duration at both the utterance and word level. The paper reports that local (word-level) prompt control consistently yields higher emotion classification accuracy than global-only or no prompt control, and that the full model with both encoders plus global and local control receives the highest mean opinion score while word error rate and character error rate stay close to the no-prompt baseline.","pith_inferences":["The same variance-adapter scaling trick could in principle transfer to other variance-adapter-based TTS backbones with only the allowed ranges re-tuned, since the prompt-control module never touches the learned weights.","A cheaper deployed system could distill GPT-4's scaling-factor behavior into a small prosody-prediction network; comparing distilled versus live-LLM outputs would quantitatively separate the prompt-following contribution from the learned encoders' contribution.","Because intensity rankings are learned per speaker and per emotion, the current intensity scale is relative rather than globally calibrated; testing whether listeners agree across speakers would show whether a shared intensity axis is needed for cross-speaker intensity transfer.","The 72 percent PIR accuracy could be broken down by emotion to see whether intensity is easier to perceive in some emotions (for example, anger) than others; the paper does not report such a breakdown."],"forward_implications":["Word-level prompt control, not global scaling, is the operation that most reliably improves emotion classification accuracy across FS2w/Emo and FS2w/Emo&Int.","The full FS2w/Emo&Int system with global and local prompt control yields the highest MOS, so combining learned emotion and intensity embeddings with LLM-scaled prosody improves perceived expressiveness without increasing word error rate.","A listener-based test places generated low, medium, and high intensity samples into the correct category about 72 percent of the time, which the paper takes as evidence that the intensity encoder produces perceptually meaningful intensity ordering.","Because MCD, WER, and CER remain close to baseline when prompt control is applied, the expressiveness gain is not purchased with a large loss in intelligibility or acoustic fidelity."],"supporting_citations":[{"why":"Supplies the FastSpeech 2 backbone whose variance adapter the paper modifies for prosody control.","marker":"[2]"},{"why":"Provides the original LLM prompt-control idea and the global/local scaling formulation that PROEMO extends to multi-speaker emotion and intensity.","marker":"[16]"},{"why":"Supplies the HuBERT features used in both the emotion and intensity encoders.","marker":"[22]"},{"why":"Introduces relative-attribute labeling for emotion intensity, the basis of the learned ranking function.","marker":"[23]"},{"why":"Provides the relative-attributes SVM-style computation used to train the per-speaker, per-emotion intensity ranking function.","marker":"[26]"},{"why":"Supplies the LibriTTS corpus used to pre-train the multi-speaker backbone.","marker":"[28]"},{"why":"Supplies the GE2E-based speaker embeddings used in the multi-speaker backbone.","marker":"[31]"}],"fun_headline_variants":["LLM prompts fine-tune TTS emotion and intensity per word","Word-level GPT-4 cues shape synthetic speech emotion","Prompts tune TTS pitch and energy for five emotions","Emotion intensity control via GPT-4 prosody prompts","LLM steers TTS emotion and intensity with word-level cues"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that GPT-4, called at inference time without fine-tuning and with no reported quantitative check of its outputs, will reliably produce scaling factors for pitch, energy, and duration that match the intended emotion and stay stable; the paper itself notes that the original prompt from [16] produced unstable output and had to be redesigned.","fun_headline_variants_meta":{"raw":{"variants":["LLM prompts fine-tune TTS emotion and intensity per word","Word-level GPT-4 cues shape synthetic speech emotion","Prompts tune TTS pitch and energy for five emotions","Emotion intensity control via GPT-4 prosody prompts","LLM steers TTS emotion and intensity with word-level cues"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000428,"raw_usage":{"total_tokens":2127,"prompt_tokens":818,"completion_tokens":1309,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":434,"completion_tokens_details":{"reasoning_tokens":1225}},"tokens_in":434,"tokens_out":1309,"duration_ms":8942,"temperature":1.0,"reasoning_tokens":1225,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:05:31.682198+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the evaluation with GPT-4's scaling factors replaced by random draws from the same mapped ranges (energy 0.5-2, duration 0.74-1.34, pitch scaled within the predicted range); if emotion classification accuracy does not drop below the reported 79.72 percent level, the prompt-control module is not the mechanism driving the gain.","supporting_citations":[{"cited_title":"An overview of affective speech synthesis and conversion in the deep learning era,","cited_arxiv_id":null,"evidence_quote":"Provides the original LLM prompt-control idea and the global/local scaling formulation that PROEMO extends to multi-speaker emotion and intensity."},{"cited_title":"We intend to improve upon the algorithm proposed in [23] for predicting emotion intensity levels using a multispeaker emotive speech dataset","cited_arxiv_id":null,"evidence_quote":"Introduces relative-attribute labeling for emotion intensity, the basis of the learned ranking function."},{"cited_title":"Fine-grained quantitative emotion editing for speech generation,","cited_arxiv_id":null,"evidence_quote":"Provides the relative-attributes SVM-style computation used to train the per-speaker, per-emotion intensity ranking function."},{"cited_title":"Learning visual attributes,","cited_arxiv_id":null,"evidence_quote":"Supplies the GE2E-based speaker embeddings used in the multi-speaker backbone."}],"review_version":1}