{"id":"2876f0b7-5bd4-45dc-976d-9b798c5fde17","arxiv_id":"2509.05617","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"The paper releases a MOS-annotated lyrics emotion dataset and reports that fine-tuned BERT/RoBERTa outperform zero-shot Grok, but the metrics are internally inconsistent and no test split is documented.","lead":"A new dataset of pop lyrics with six emotion-intensity scores from human raters is introduced, along with a comparison of fine-tuned BERT models and zero-shot Grok. The paper's aggregate error numbers contradict its own per-emotion tables, and the evaluation split is never described.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table II's fine-tuned MAE/RMSE values are internally inconsistent (RMSE < MAE), so the >30% improvement claim rests on invalid summary statistics; no held-out split is described either.","rationale":"The reader identified the absence of a train/test split as the weakest assumption, but an even more direct problem is that the central numeric evidence is internally contradictory. Table II reports RMSE values smaller than MAE values, which is mathematically impossible for any error distribution. This is not a matter of an unstated protocol detail; it invalidates the specific numbers used to support the >30% improvement claim. The reader's concern about overfitting is real and related, but the inconsistency in the reported aggregate metrics is the most load-bearing issue because it undermines the headline comparison even before considering generalization. I agree with the reader's overall REJECT verdict: the paper's central quantitative conclusion cannot be accepted as written. The comparison could potentially be salvaged by providing correct aggregate metrics, a held-out evaluation, and the missing multi-LLM results, so the rejection is about the current evidentiary basis, not the underlying research question. I mark agreement as 'partial' because the reader pointed to a different weakest assumption, though both concerns point to unreliable evaluation.","tokens_in":6368,"tokens_out":2942,"duration_ms":34651,"concrete_test":"Recompute the aggregate metrics from Table III. Average the six per-emotion MAE values and the six per-emotion RMSE values for fine-tuned RoBERTa, and compare with Table II's MAE=0.33 and RMSE=0.1511. If the averages do not match, or if RMSE < MAE persists, the tables are mutually inconsistent; then request the evaluation script and raw prediction files to determine which table is correct, and rerun the comparison on a held-out split (e.g., stratified 80/10/10) before any performance claim is made.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim—that fine-tuned BERT-based models reduce error by more than 30% relative to zero-shot Grok (Section IV)—is quantified by Table II. That table is internally impossible. For fine-tuned RoBERTa it reports MAE=0.33 and RMSE=0.1511. For any set of nonnegative errors, RMSE >= MAE, since squaring amplifies large deviations and the square-root of a mean square is at least the mean absolute value. Thus RMSE < MAE cannot occur. Moreover, the per-emotion values in Table III for RoBERTa average to MAE≈0.145 and RMSE≈0.33, exactly reversing the aggregate numbers: the two tables cannot both describe the same evaluation. This suggests the columns in Table II are swapped, computed on a different subset, or otherwise erroneous. Because the headline improvement and the central fine-tuned-versus-zero-shot conclusion are derived from these numbers, the quantitative comparison is unsupported as reported. Section III.B.1 also states only that the model was 'trained and validated on the annotated dataset', with no explicit train/test split or cross-validation; if the reported errors are training-set errors, even corrected values would not estimate generalization. Additionally, the conclusion mentions DeepSeek-R1's superior zero-shot performance, but no DeepSeek-R1 results appear in Tables II or III, so the claimed multi-LLM comparison is not actually shown.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a manually annotated dataset of pop song lyrics with Mean Opinion Score (MOS) labels for six Ekman emotions (joy, sadness, anger, fear, surprise, disgust) and evaluates both fine-tuned BERT/RoBERTa models and the zero-shot LLM Grok 3 on the task of predicting the six emotion intensities. The headline claim is that fine-tuning BERT-based models reduces error by more than 30% relative to zero-shot Grok, and that these fine-tuned models are preferable for lyric emotion attribution when labeled data are available. The dataset is released publicly.","tokens_in":6767,"tokens_out":2004,"duration_ms":20992,"significance":"A public MOS-labeled lyrics emotion dataset would be a useful resource for affective computing and music information retrieval, and a clean comparison of fine-tuned versus zero-shot LLMs on such a task would inform practical model selection. However, as presented, the quantitative results are internally inconsistent and the evaluation protocol is not described with sufficient detail to support the central claim. The contribution therefore cannot be accepted on the evidence currently in the manuscript.","major_comments":[{"comment":"The aggregated metrics in Table II are internally impossible and contradict Table III. For any set of nonnegative errors, RMSE >= MAE, yet Table II reports fine-tuned BERT MAE=0.33 / RMSE=0.1507 and fine-tuned RoBERTa MAE=0.33 / RMSE=0.1511. Averaging the per-emotion values in Table III for fine-tuned RoBERTa gives MAE ≈ 0.145 and RMSE ≈ 0.33, the reverse pattern. The zero-shot Grok row (MAE=0.50, RMSE=0.367) also disagrees with the per-emotion means in Table III (MAE ≈ 0.40, RMSE ≈ 0.51). Since Section IV's 'more than 30%' improvement claim is derived from Table II, the central quantitative comparison is unsupported as reported.","section":"Table II vs. Table III"},{"comment":"The fine-tuning procedure is described only as 'trained and validated on the annotated dataset'. No train/test split, cross-validation, or held-out lyrics are described. If the reported errors are computed on the same examples used for fitting, they do not estimate generalization, and the fine-tuned versus zero-shot comparison is confounded by potential train/evaluation overlap. This is a load-bearing methodological omission: the main conclusion depends on the fine-tuned errors being generalization errors.","section":"Section III.B.1"},{"comment":"The conclusion states that 'Among the zero-shot models, DeepSeek-R1 demonstrated stronger generalization capabilities', but no DeepSeek-R1 results appear in Tables II or III or anywhere in Section IV. The claimed multi-LLM comparison is thus not actually shown in the manuscript. This missing evidence directly affects the conclusion's breadth.","section":"Section V"}],"minor_comments":[{"comment":"The number of songs/lyrics in the dataset is never stated. Table I shows four annotators per item, but the text only says 'a diverse committee'. The dataset size, annotation scale, and inter-annotator agreement (e.g., variance or Krippendorff's alpha) should be reported to establish the reliability of the MOS labels.","section":"Section III.A"},{"comment":"The zero-shot section mentions 'models' in plural but only names Grok 3. The prompt template used for zero-shot evaluation is not reproduced, which hinders reproducibility, especially because prompt phrasing can strongly affect LLM outputs.","section":"Section III.B.2"},{"comment":"Fine-tuning hyperparameters (learning rate, epochs, batch size, validation split) are not reported. For a benchmark paper, these details are necessary for others to reproduce or build on the results.","section":"General"},{"comment":"The Disgust row for zero-shot Grok reads '0.45/057' — likely a typo for '0.45/0.57'. This should be corrected.","section":"Table III"},{"comment":"Some references appear incomplete or inconsistent (e.g., 'Huang et al. (2021)' is mentioned in the literature review but is not in the reference list). The reference list also mixes conference papers and arXiv preprints without full publication details.","section":"References"}],"recommendation":"reject","confidential_remarks":"The manuscript is very early-stage and the core quantitative claim is invalidated by the manuscript's own data. The Table II/Table III inconsistency is not a mere typo: it affects both fine-tuned models and the zero-shot baseline, and it directly feeds the headline 30% improvement. The absence of any train/test split description means even corrected numbers would not establish generalization. These are not presentation issues; they require re-running the experiments. I would recommend reject, though I would encourage the authors to resubmit after a thorough revision with a proper evaluation protocol, corrected tables, and full reproducibility details."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, quick take: the dataset is worth knowing about; the evaluation as reported is not. Table II lists fine-tuned MAE 0.33 and RMSE 0.15, which is mathematically impossible (RMSE can't be less than MAE), and the per-emotion mean in Table III reverses those numbers. So the '>30% improvement' claim is derived from bad summary statistics. On top of that, Section III.B.1 only says the model was 'trained and validated on the annotated dataset' with no split described, so the fine-tuned numbers may just be training fit. And the conclusion credits DeepSeek-R1 with superior zero-shot performance, but no DeepSeek-R1 results appear anywhere.\n\nWhat's genuinely new: the MOS-annotated lyric dataset itself, with six Ekman scores per song and a posted GitHub link. If the data is actually usable, that is a real resource for the emotion-in-lyrics community. The figures showing score distributions and genre correlations are informative, and using a committee of four per item is reasonable. The literature review is competent and covers the relevant work.\n\nSoft spots, in order: the internal table contradiction, missing train/test split, and the mismatch between the claimed multi-LLM evaluation and the single Grok model reported. These are load-bearing, not cosmetic. The qualitative headline (fine-tuned BERT beats zero-shot LLM) is also already established in the cited Boitel et al. work, so even a clean evaluation would be mostly confirmatory.\n\nThe citation pattern is mostly fine; no obvious invented references. I'd put this in the 'promising resource, broken evaluation' category. For peer review: I would not desk-reject it if the dataset link checks out—send it to a referee with the explicit charge to verify the split and the arithmetic. But the current version cannot support its conclusions, and the authors should be asked to fix the tables, disclose the split, and either report all zero-shot models or stop claiming to have evaluated several.","headline":"New MOS-labeled lyrics dataset is the only solid contribution; the evaluation tables are internally inconsistent and the central claim is unsupported.","tokens_in":7185,"tokens_out":2950,"would_cite":false,"duration_ms":30275,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that fine-tuning BERT-based models on a new MOS-labeled lyric corpus reduces emotion-score prediction error by more than 30% compared with zero-shot Grok 3.","keywords":["multi-label emotion classification","song lyrics","mean opinion score","BERT fine-tuning","zero-shot LLM","Ekman emotions","affective computing","music information retrieval"],"falsifier":"Re-run the same comparison on a held-out 20% of the released MOS-labeled dataset and on fresh lyrics never seen during fine-tuning, reporting per-emotion MAE and RMSE. If the fine-tuned BERT and RoBERTa models do not beat zero-shot Grok 3 by roughly 30% on unseen lyrics, the headline claim fails. A second check is to re-label a sample with a different four-rater committee; if the MOS scores change substantially, the ground truth itself is unstable.","tokens_in":6325,"feed_emoji":"🎵","tokens_out":7670,"duration_ms":74380,"temperature":0.7,"pith_summary":"The paper builds a benchmark for estimating emotional intensity in pop song lyrics. Human raters scored each lyric from 0 to 5 on six emotions, and the average of four raters became the ground-truth label. The authors then compare zero-shot prompting of a large language model with fine-tuned BERT and RoBERTa regression models. They report that both fine-tuned models cut error by more than 30% relative to the zero-shot model, with consistent gains across all six emotions. The paper's conclusion is that when labeled data is available, task-specific fine-tuning beats zero-shot prompting for continuous lyric-emotion prediction.","feed_headline":"Fine-tuned BERT cuts lyric-emotion error 30% over Grok 3","feed_subtitle":"A four-rater MOS lyric benchmark says task-specific training beats zero-shot LLMs on all six emotions.","key_machinery":"The central machinery is the MOS label: each lyric receives a continuous 0 to 5 intensity score per emotion as the arithmetic mean of four independent human ratings, converting subjective annotation into a regression target. On top of this, the paper uses a fine-tuned BERT or RoBERTa encoder with a six-output regression head trained with mean squared error, compared against a zero-shot LLM prompted to return the six numeric scores. The MOS target is what makes the fine-tuning comparison possible, and the regression setup is what lets the paper claim continuous emotion-intensity prediction rather than discrete classification.","core_discovery":"The paper constructs a manually labeled dataset of pop song lyrics annotated by four raters per lyric on six Ekman emotions—joy, sadness, anger, fear, surprise, and disgust—averaged into continuous mean-opinion scores from 0 to 5. Using that dataset, it evaluates zero-shot prompting of Grok 3 and fine-tuned regression heads on BERT and RoBERTa. The aggregated results show both fine-tuned models reaching MAE 0.33 and RMSE about 0.151, versus MAE 0.50 and RMSE 0.367 for zero-shot Grok, which the paper states is a reduction in error of more than 30%. The improvement is reported across all six emotions, with the smallest MAEs for surprise (0.10) and fear (0.13). The paper concludes that fine-tun","pith_inferences":["The paper does not report a train/test split; if the reported errors are in-sample, the central gap is likely overstated. My inference is that a held-out evaluation is the first thing to check before relying on the 30% figure.","The zero-shot comparison uses a single prompt template and a single frontier model, so it may understate what LLMs can do with structured output or chain-of-thought prompting.","With only four annotators per lyric and no reported inter-annotator agreement, MOS labels may be noisy; a useful extension would compare these labels with Best-Worst Scaling or more raters on the same lyrics.","If the gap survives held-out evaluation, the practical consequence is that lightweight BERT-size models can power emotion-aware lyric tagging cheaply, which the paper leaves implicit."],"forward_implications":["Task-specific fine-tuning on MOS-labeled lyrics yields materially better continuous emotion-intensity estimates than zero-shot prompting of a frontier LLM, with error reductions above 30% in aggregate MAE and RMSE.","The released dataset gives the field a shared benchmark with continuous ground-truth scores for six Ekman emotions across pop lyrics, enabling direct comparisons of future models.","The fine-tuned advantage is not confined to one feeling: it holds for joy, sadness, anger, fear, surprise, and disgust.","Zero-shot prompting remains a usable fallback when labeled data are unavailable, but the paper positions it as inferior for affective computing and music information retrieval applications.","The paper's result extends the fine-tuned-versus-LLM comparison to lyric emotion, aligning with prior evidence that in-domain fine-tuning beats transferring models trained on unrelated text."],"supporting_citations":[{"why":"Supplies the six basic emotions that define the label space.","marker":"(Ekman, 1992)"},{"why":"Introduced GPT-3 prompting, the paradigm used for zero-shot evaluation.","marker":"(Brown et al., 2020)"},{"why":"Created GoEmotions and fine-tuned BERT for multi-label emotion, grounding the fine-tuning approach.","marker":"(Demszky et al., 2020)"},{"why":"Showed in-domain lyric fine-tuning outperforms transfer from other text, motivating the fine-tuning comparison.","marker":"(Edmonds & Sedoc, 2021)"},{"why":"Direct comparison of GPT-3 and BERT-based models for emotion recognition, the immediate context for this paper's result.","marker":"(Boitel et al., 2024)"},{"why":"Supplies the mean opinion score averaging method used to create ground-truth labels.","marker":"(Chen et al., 2020)"},{"why":"Establishes MOS as a standard for perceived emotional strength, supporting the labeling methodology.","marker":"(Truong & van Leeuwen, 2007)"}],"fun_headline_variants":["Fine-tuned BERT beats Grok 3 by 30% on lyric emotion error","New MOS benchmark: fine-tuned models outperform zero-shot LLMs","Four-rater labels reveal fine-tuning wins for song emotion","Lyric emotion: fine-tuned BERT cuts error 30% over Grok 3","Fine-tuned models slash lyric emotion error vs zero-shot LLMs"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The fine-tuned models are described only as 'trained and validated on the annotated dataset', with no explicit train/test split, cross-validation, or held-out lyrics; if the reported errors come from the same examples used for fitting, the central performance gap is inflated by overfitting.","fun_headline_variants_meta":{"raw":{"variants":["Fine-tuned BERT beats Grok 3 by 30% on lyric emotion error","New MOS benchmark: fine-tuned models outperform zero-shot LLMs","Four-rater labels reveal fine-tuning wins for song emotion","Lyric emotion: fine-tuned BERT cuts error 30% over Grok 3","Fine-tuned models slash lyric emotion error vs zero-shot LLMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000278,"raw_usage":{"total_tokens":1489,"prompt_tokens":741,"completion_tokens":748,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":648}},"tokens_in":485,"tokens_out":748,"duration_ms":8078,"temperature":1.0,"reasoning_tokens":648,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T05:17:56.683194+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same comparison on a held-out 20% of the released MOS-labeled dataset and on fresh lyrics never seen during fine-tuning, reporting per-emotion MAE and RMSE. If the fine-tuned BERT and RoBERTa models do not beat zero-shot Grok 3 by roughly 30% on unseen lyrics, the headline claim fails. A second check is to re-label a sample with a different four-rater committee; if the MOS scores change substantially, the ground truth itself is unstable.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the six basic emotions that define the label space."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Created GoEmotions and fine-tuned BERT for multi-label emotion, grounding the fine-tuning approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Showed in-domain lyric fine-tuning outperforms transfer from other text, motivating the fine-tuning comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Direct comparison of GPT-3 and BERT-based models for emotion recognition, the immediate context for this paper's result."}],"review_version":1}