{"id":"5c5cfc6c-37c5-49e4-ae30-4ccd3f22c00b","arxiv_id":"2505.24493","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"GPT-4o text-only prompts can re-annotate a Friends-based emotion dataset, and models trained on those labels outperform models trained on the original human labels in cross-corpus speech emotion recognition.","lead":"The authors used GPT-4o to re-label 8,821 lines from the sitcom Friends using only text, producing the MELT emotion dataset with no human annotation labor. Human raters preferred the GPT-4o labels over the original human labels, and speech emotion recognition models trained on them scored higher on several test sets.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper never checks whether GPT-4o memorized MELD's public labels; the prompt's episode/line identifiers make this plausible, so the claimed annotation quality and SER gains may be partly circular.","rationale":"The paper's contribution is an annotation method that uses an LLM's embedded knowledge to label multimodal emotion data from text alone. The method's value depends on the labels being generated, not recalled. Since the prompt identifies the exact episode and line, and the model's training window includes MELD, there is a plausible direct route from the human-annotated benchmark to the model's outputs. The paper's own Section 3 quantification of label changes is not a leakage check: a model with partial memorization would be expected to produce exactly the observed mix of unchanged and changed labels, especially if memorization is stronger for famous scenes. The MOS and fine-tuning experiments do not control for this, because both use the same MELT labels whose provenance is in question. A control corpus from the same show but absent from MELD (or a matched unannotated sitcom) would separate general knowledge of Friends and emotion from recall of MELD-specific rows. If the control fails, the central claims—'fully annotated by GPT-4o', 'align more closely with human preferences', and 'improved generalization'—would need substantial qualification; at minimum, the paper must disclose the contamination risk and show the method works on data the model has not seen in labeled form. This does not change the reader's CONDITIONAL verdict; it reinforces the condition.","tokens_in":8766,"tokens_out":10918,"duration_ms":145252,"concrete_test":"Build a control set of dialogue lines from the same Friends episodes that are not included in MELD (or, if unavailable, from a matched sitcom with no published emotion labels), run the exact Sec. 2.3 prompt on them, and then (i) compare label distributions and transition patterns to the MELT results and (ii) fine-tune the same four SSL backbones on this control annotation and evaluate on IEMOCAP, TESS, RAVDESS, and CREMA-D. If the control set reproduces similar human-preference and SER advantages over MELD-trained models, the MELT results are not an artifact of MELD memorization; if the advantages disappear or reverse, the central claim is contaminated and must be revised to exclude memorized benchmarks.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on GPT-4o's MELT labels being new annotations derived from general cultural/linguistic knowledge, not retrieval of MELD's existing human labels. The prompt in Sec. 2.3 provides the exact speaker, season, episode, and utterance, enabling lookup of any Friends line; Sec. 2.2 states GPT-4o's knowledge cutoff is October 2023, while MELD was published in 2019 and is a widely used public benchmark whose label files appear in many public repositories. Section 3 reports 46.43% and 47.52% label-change rates but includes no test for whether the unchanged labels or the transition patterns in Fig. 1 are explained by memorization of MELD. If MELD labels are in GPT-4o's training data, the MOS preference results and the SER gains are partly circular: the model is not annotating from knowledge but reproducing a human-annotated dataset it has already seen, so the claimed scalability to new unannotated multimodal data is not established. This is an identifiability problem with the experimental design, not an accusation of misconduct.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MELT, a re-annotation of a filtered subset of the MELD dataset (8,821 utterances from 42 Friends characters) using GPT-4o with text-only prompts that include character, episode, and utterance context. The authors report that MELT labels differ from MELD on about half the utterances, are preferred by human raters in a MOS experiment, and improve fine-tuned SSL-based SER performance on IEMOCAP, TESS, RAVDESS, and CREMA-D relative to training on MELD. They also report that GPT-4o's perceived pitch/loudness descriptions correlate with eGeMAPS-derived categories above chance. The central claim is that this demonstrates that LLM embedded cultural knowledge can automatically annotate multimodal emotion data without human labor.","tokens_in":9051,"tokens_out":7322,"duration_ms":81901,"significance":"The work is potentially significant: if the annotations are genuinely produced from general knowledge rather than memorized benchmark labels, the approach offers a very low-cost, scalable alternative to crowdsourced emotion annotation and provides a public dataset (MELT) with audio-attribute pseudo-labels. Strengths include the public release of the dataset, a clearly described prompting framework, four SSL backbones, four external SER corpora, and a human preference study. The cross-corpus design partially grounds the evaluation in independent benchmarks, which is a positive feature. However, the validity of the central claim depends on ruling out memorization of MELD, and the empirical comparisons lack statistical support. These issues are addressable but currently open.","major_comments":[{"comment":"The paper does not address the possibility that GPT-4o has memorized MELD's published emotion labels. The prompt in Sec. 2.3 reveals the exact speaker, season, episode, and utterance text; MELD was released in 2019 and its label files are widely available, well before the model's October 2023 knowledge cutoff stated in Sec. 2.2. If GPT-4o reproduces memorized MELD labels, then the 46.43%/47.52% label-change rates in Sec. 3, the MOS preferences in Fig. 3, and the SER gains in Table 3 do not establish annotation from embedded cultural knowledge and do not support scalability to unannotated data. This is an identifiability problem in the experimental design. A concrete remedy is to run control annotations in which episode/character identifiers are removed or paraphrased, and on utterances from a different sitcom not in MELD, and to compare agreement with MELD across conditions; the authors could also directly prompt GPT-4o to reproduce MELD labels for a sample and quantify overlap. Without such a check, the headline claim is not established.","section":"Sec. 2.3 / Sec. 3"},{"comment":"The objective experiments are reported without variance or significance testing. Each cell in Table 3 appears to come from a single training run, and the claimed 'consistently improved' performance is contradicted by several cells, e.g., wav2vec 2.0 Aud on TESS has F1 0.1926 (MELT) vs 0.2063 (MELD), and WavLM Base+ on TESS has F1 0.2372 vs 0.2405. Without repeated seeds, standard deviations, or paired tests, the cross-corpus advantage of MELT cannot be distinguished from noise. Please report means and standard deviations over at least 3-5 seeds and appropriate significance tests (e.g., paired t-test or Wilcoxon test) on the per-utterance or per-run scores.","section":"Sec. 4.2 / Table 3"},{"comment":"The MOS experiment is described only as 20 participants choosing between MELT and MELD annotations on video clips. The paper reports aggregate preference percentages per emotion but omits the number of items rated, the number of raters per item, inter-rater agreement, and any statistical test. Since the subjective preference is a headline result supporting the claim that MELT 'aligns more closely with human preferences,' the experiment needs per-item details and a significance test (e.g., Wilcoxon signed-rank test over items or raters).","section":"Sec. 4.1 / Fig. 3"},{"comment":"Annotations are sampled at temperature 1.0, yet the paper claims a prompt framework ensuring 'stability and reproducibility' in Sec. 2.3. No repeated annotation runs, sampling strategy, or consistency metrics are reported. If the labels are not stable across samples, the released dataset and downstream fine-tuning results are not reproducible. Please report the number of samples per utterance, the aggregation rule (if any), and agreement across repeated API calls.","section":"Sec. 2.2"},{"comment":"The filtering rule for training data is underspecified for partially overlapping test sets. The paper says 'only emotion categories present in the test set are retained,' but it does not state how MELT/MELD's seven labels are mapped or filtered for IEMOCAP (4 classes) and CREMA-D (6 classes), nor how the classification head is set up when the training and test label sets differ. Please specify the exact label filtering/mapping per dataset, including any mapping of 'joy' to 'happy' or 'pleasant surprise' to 'surprise,' and confirm that the same procedure is applied to both MELT and MELD training conditions.","section":"Sec. 4.2.2"}],"minor_comments":[{"comment":"The total number of MELD utterances is given as 13,708 in Sec. 2.1, but Table 1 lists 9,989 + 2,610 = 12,599 for the train and test splits; please clarify that the 13,708 includes a development split not used in this work.","section":"Sec. 2.1 / Table 1"},{"comment":"The abstract contains 'consistence performance improvement,' which should be 'consistent performance improvement,' and the index terms contain 'affective compution,' which should be 'affective computing.'","section":"Abstract / Index Terms"},{"comment":"The caption refers to a 'confusion matrix and inter-label transition matrix,' but both panels appear to be transition matrices from MELD to MELT; please use consistent terminology.","section":"Fig. 1"},{"comment":"The individual prompt components (context, Chain-of-Thought, cross-validation, and prefilling) are not ablated, so the contribution of each to annotation quality is unverified; a small ablation study would strengthen the prompt-design contribution.","section":"Sec. 2.3"},{"comment":"The pitch and loudness predictions are evaluated against binned eGeMAPS attributes with bin boundaries defined by distribution percentiles, but it is not stated whether the percentiles are estimated on the training split, test split, or both; please clarify to avoid possible leakage.","section":"Sec. 5.2 / Table 4"},{"comment":"The cost claim of '≤$10' is not supported by a calculation; please provide the number of API calls, input/output token counts, and the pricing used to obtain the estimate.","section":"Sec. 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is suitable for a speech/affective computing venue, but the central claim of scalable LLM-based annotation depends on ruling out memorization of MELD. I recommend asking the authors to add control experiments for memorization and to report variance/significance for the objective and subjective comparisons before publication. The dataset release is a useful contribution, and the cross-corpus evaluation is a good starting point, but the current evidence does not yet support the strong scalability claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you work in affective computing. The paper builds a new emotion dataset (MELT) by re-annotating 8,821 utterances from the Friends-based MELD dataset using GPT-4o with text-only prompts. The contribution is concrete: they release the dataset, a prompt framework with CoT and cross-validation, and show that fine-tuning SSL backbones on MELT beats fine-tuning on the original human-labeled MELD across IEMOCAP, TESS, RAVDESS, and CREMA-D. The cross-corpus gains are the strongest part of the paper, since those test sets are independent of the training data.\n\nThe soft spot is real and the stress-test note is right: the paper's central claim is that GPT-4o produces these labels from its general cultural knowledge, not by retrieving existing labels. But the prompt gives the exact speaker, season, episode, and utterance, and MELD has been public since 2019. GPT-4o's training cutoff is October 2023. The paper never checks whether GPT-4o memorized MELD's labels. The 46-47% label change rates show substantial disagreement, which is evidence against pure memorization, but it's not a test of the independence assumption. If GPT-4o is partly reproducing MELD labels, then the MOS preference and the in-domain SER improvements are partly circular, and the scalability claim for unannotated data is not established. This is an identifiability problem with the experimental design, not misconduct.\n\nOther soft spots are more moderate. Table 3 has no error bars or significance tests; the results are consistent across most cells, but a few cells (e.g., TESS F1) go slightly negative. The dataset is filtered to 42 speakers and about 70% of MELD, so the comparison to MELD is only valid for that subset. There are also no ablations of the prompt components, so we don't know which parts matter. The cost claim (<$10) is nice but not central.\n\nWho is this for? People building emotion datasets or using LLMs as annotators. It's a legitimate empirical paper with a useful new resource. It deserves peer review. The reviewer should push for a memorization check (e.g., hold out some MELD utterances, compare GPT-4o's labels to the original and to a new human annotation), plus error bars on Table 3. I'd take it under conditional acceptance, not desk reject.\n\nRecommendation: send it to review, but be clear the leakage question is the load-bearing issue.","headline":"Solid empirical contribution with a real caveat: the paper never checks whether GPT-4o memorized MELD's public labels, which weakens its central claim of annotation from embedded knowledge.","tokens_in":9547,"tokens_out":3281,"would_cite":true,"duration_ms":31859,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GPT-4o can annotate emotional speech from transcripts alone, and its labels beat human labels in preference tests and improve downstream speech emotion recognition.","keywords":["speech emotion recognition","LLM-based annotation","GPT-4o","multimodal emotion dataset","MELT","prompt engineering","self-supervised learning","affective computing"],"falsifier":"Take a freshly recorded emotional speech dataset that has never been published or discussed online, run the paper's exact prompt on its transcripts, and have human raters compare the GPT-4o labels against new human labels. If the model's agreement, preference scores, and downstream SER gains fall to near chance, the reported advantage comes from memorized knowledge of MELD rather than a general ability to annotate emotional speech from text.","tokens_in":8609,"feed_emoji":"🎭","tokens_out":10334,"duration_ms":115612,"temperature":0.7,"pith_summary":"The paper sets out to show that a large language model can annotate emotional speech data without hearing a single audio file, by drawing on the cultural knowledge it absorbed during pretraining. The authors take the MELD corpus, 13,708 utterances from the TV show Friends, filter it to 8,821 utterances, and have GPT-4o re-annotate each line from a text prompt that names the speaker, episode, and dialogue and asks for the emotion and voice qualities in a fixed JSON format. In a blind comparison, human raters preferred GPT-4o's labels to MELD's original human labels, and speech emotion recognition models fine-tuned on the new labels generally transferred better to four held-out emotion datasets. If these results hold, LLM embedded knowledge becomes a viable substitute for costly, inconsistent human annotation in affective computing.","feed_headline":"Text-only GPT-4o labels beat human emotion labels in test","feed_subtitle":"GPT-4o remakes the Friends emotion dataset from text alone; speech models trained on it generalize better.","key_machinery":"The load-bearing mechanism is GPT-4o's embedded knowledge of popular culture, activated by a purpose-built prompting framework. Each prompt supplies the speaker, season/episode, utterance text, and a seven-emotion taxonomy, then asks the model to reason through how the voice would sound — emotion, loudness, pitch, rhythm speed, emotional impact — before emitting a structured JSON answer; Chain-of-Thought prompting, cross-validation questions, and pre-filled output structure keep the annotations grounded and consistent. The textual context is the only multimodal input the model receives: the audio and video of Friends are never shown, so the acoustic and situational commonsense must come from what GPT-4o learned at scale.","core_discovery":"The central claim is that MELT is a multimodal emotion dataset fully annotated by GPT-4o, and that its labels are both closer to human preference and more useful for training speech emotion recognition systems than the human majority-vote labels of MELD. The paper reports that 46.43% of training labels and 47.52% of test labels changed relative to MELD, that aggregate MOS ratings from 20 blind raters favored MELT (with the largest gaps for anger and surprise), and that fine-tuning four self-supervised backbones on MELT produced higher UAR, accuracy, and F1 than training on MELD in most configurations across IEMOCAP, TESS, RAVDESS, and CREMA-D. The authors frame this as the first evidence that GPT-class models can act as annotators for multimodal emotion data, using only text plus the knowledge the model has internalized about a well-known television series.","pith_inferences":["Editorial extension: Because the prompt reveals the exact show, season, episode, and line, and MELD is a public dataset, some of the agreement and preference scores could be inflated by GPT-4o having memorized MELD's published labels; the paper does not test this.","Editorial extension: The method is likely to transfer only to content that is densely present in web-scale training data; for private conversations, low-resource languages, or novel recordings the embedded-knowledge advantage may shrink or disappear.","Editorial extension: Adding a brief audio-derived description to the prompt, rather than only transcripts, could let the pipeline annotate never-seen content while keeping LLM reasoning, a testable variant the paper does not run."],"forward_implications":["Emotion annotation for widely known media can be produced from transcripts alone for roughly $10, replacing multi-annotator campaigns.","SER models fine-tuned on LLM labels should generalize to unseen emotion corpora at least as well as, and often better than, models trained on human majority-vote labels.","The same structured prompting template can be applied to other scripted dialogue datasets to obtain consistent, context-aware emotion labels without additional human labor.","LLM re-annotation yields a more balanced emotion distribution than the original human labels, which can reduce majority-class bias in downstream training."],"supporting_citations":[{"why":"Documents GPT-4o's web-scale training and knowledge cutoff, the embedded-knowledge premise that the annotation method relies on.","marker":"[8]"},{"why":"Supplies the Friends-based MELD dialogues and their human majority-vote labels, from which MELT is filtered and re-annotated.","marker":"[16]"},{"why":"Provides the wav2vec 2.0 base self-supervised backbone fine-tuned on MELT versus MELD in the SER evaluation.","marker":"[20]"},{"why":"Provides the emotion-pretrained wav2vec 2.0 backbone used as another SER evaluation model.","marker":"[21]"},{"why":"Provides the HuBERT base backbone used as another SER evaluation model.","marker":"[22]"},{"why":"Provides the WavLM base+ backbone used as another SER evaluation model.","marker":"[23]"},{"why":"Acts as the IEMOCAP out-of-domain test set for measuring cross-corpus generalization.","marker":"[24]"},{"why":"Acts as the TESS out-of-domain test set sharing the same seven emotion classes.","marker":"[25]"},{"why":"Acts as the RAVDESS out-of-domain test set for measuring generalization.","marker":"[26]"},{"why":"Acts as the CREMA-D out-of-domain test set for measuring generalization.","marker":"[27]"}],"fun_headline_variants":["GPT-4o text-only labels beat human emotion annotations","LLM labels from text alone boost speech emotion recognition","MELT: GPT-4o annotates emotions better than humans from text","Text-based GPT-4o annotation improves speech emotion models","AI annotator beats humans on emotion data without audio"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes GPT-4o's emotion judgments come from its general cultural knowledge and are not answers memorized from MELD, the public dataset whose episode-and-utterance prompts likely appear in the model's training data.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4o text-only labels beat human emotion annotations","LLM labels from text alone boost speech emotion recognition","MELT: GPT-4o annotates emotions better than humans from text","Text-based GPT-4o annotation improves speech emotion models","AI annotator beats humans on emotion data without audio"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000165,"raw_usage":{"total_tokens":1249,"prompt_tokens":944,"completion_tokens":305,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":560,"completion_tokens_details":{"reasoning_tokens":221}},"tokens_in":560,"tokens_out":305,"duration_ms":3825,"temperature":1.0,"reasoning_tokens":221,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:20:32.464606+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a freshly recorded emotional speech dataset that has never been published or discussed online, run the paper's exact prompt on its transcripts, and have human raters compare the GPT-4o labels against new human labels. If the model's agreement, preference scores, and downstream SER gains fall to near chance, the reported advantage comes from memorized knowledge of MELD rather than a general ability to annotate emotional speech from text.","supporting_citations":[{"cited_title":"Be- yond deep learning: Charting the next frontiers of affective com- puting,","cited_arxiv_id":null,"evidence_quote":"Documents GPT-4o's web-scale training and knowledge cutoff, the embedded-knowledge premise that the annotation method relies on."},{"cited_title":"Secap: Speech emotion captioning with large language model,","cited_arxiv_id":null,"evidence_quote":"Provides the emotion-pretrained wav2vec 2.0 backbone used as another SER evaluation model."},{"cited_title":"Meld: A multimodal multi-party dataset for emo- tion recognition in conversations,","cited_arxiv_id":null,"evidence_quote":"Provides the WavLM base+ backbone used as another SER evaluation model."},{"cited_title":"On the time course of vocal emotion recognition,","cited_arxiv_id":null,"evidence_quote":"Acts as the IEMOCAP out-of-domain test set for measuring cross-corpus generalization."},{"cited_title":"Applying tdnn architectures for an- alyzing duration dependencies on speech emotion recognition","cited_arxiv_id":null,"evidence_quote":"Acts as the TESS out-of-domain test set sharing the same seven emotion classes."},{"cited_title":"A wide evaluation of chatgpt on affective computing tasks,","cited_arxiv_id":null,"evidence_quote":"Acts as the RAVDESS out-of-domain test set for measuring generalization."}],"review_version":1}