{"id":"935ceb5e-87da-4e16-9215-dcc083b6ad76","arxiv_id":"2504.18673","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Third-party emotion annotations, both human and LLM, show low to fair agreement with authors' self-reported emotions, with LLMs outperforming humans but still misaligning substantially.","lead":"Researchers asked social media users to label their own posts with emotions, then had crowd workers and five LLMs label the same posts. The labels rarely agreed, and while LLMs were closer than humans, both missed the authors' reported emotions, suggesting emotion datasets built on third-party labels may not capture what authors actually feel.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The gold-standard assumption is the weakest link: 9.2% of first-party labels were flagged spurious yet retained, and no sensitivity analysis shows that criterion noise does not drive the low third-party alignment.","rationale":"The reader's weakest assumption and my concern coincide: the evaluation treats first-party self-reported labels as the gold standard for private states, and the paper itself disclaims external verifiability. I sharpen this into a concrete missing analysis: Appendix D documents that 9.2% of first-party labels were flagged as spurious yet retained, and Section 7 asserts minimal impact without demonstrating it. This is load-bearing because every quantitative result in the paper, including the headline kappa range of 0 to 0.45 and macro F1 of 0.28 to 0.40, is computed against this criterion. If those labels are noisy, the measured third-party misalignment is inflated. The proposed re-analysis is feasible with data already in hand, because the spuriousness flags exist; it would settle whether this admitted noise changes the conclusion. The paper has real strengths: an IRB-approved human-subjects design, multiple quality checks, five LLMs, direct comparison of human and machine annotators, and appropriate mixed-effects models for the annotator-level comparison. The narrower empirical claim, that third-party labels diverge from first-party self-reports on this task and population, is well supported. What is not supported without the sensitivity analysis is the broader inference about private states. This does not require rejection; it requires making the sensitivity analysis and a softened private-state claim explicit conditions of acceptance, which matches the reader's CONDITIONAL verdict.","tokens_in":17932,"tokens_out":6187,"duration_ms":68924,"concrete_test":"Re-run the full RQ1/RQ2/RQ3 analysis on the subset of posts whose first-party labels were not flagged as spurious by either quality checker or the senior adjudicator, and compare macro-average Cohen's kappa and F1 with the full-sample results. Also compute a worst-case bound by treating the 9.2% flagged labels as random draws from the observed emotion distribution. If excluding flagged posts raises human or LLM macro-average kappa above roughly 0.6, or changes the LLM-vs-human ordering, the headline claim is partly an artifact of criterion noise; if the metrics remain in the low-to-fair range, the measured misalignment is robust to this admitted label noise.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is not merely that third-party labels disagree with first-party self-reports; it is that third-party annotations 'fail to faithfully represent authors' private states' (Abstract; Section 6). That inference requires first-party self-reports to be a valid operationalization of the private state. The paper's own Section 7 concedes that 'whether first-party labels faithfully represent first-party internal emotion can not be externally verifiable,' and Appendix D reports that 9.2% of first-party labels were flagged as spurious by two independent reviewers yet retained. If those labels reflect careless or strategic responding rather than felt emotion, they add noise to the criterion, mechanically suppressing Cohen's kappa and F1 for all third-party groups. The paper asserts that retaining them 'is expected to have minimal impact on the overall findings' but provides no re-analysis. A further ambiguity is that authors labeled posts retrospectively ('emotions they believed were expressed in the post'), which is itself an interpretive act subject to self-presentation, social desirability, and memory distortion. Low first-vs-third-party agreement may therefore reflect unreliability of the criterion as much as third-party limitations, leaving the private-state conclusion underdetermined by the measured gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a set of human-subject experiments comparing first-party (author-provided) emotion labels on original social media posts with third-party annotations from crowd workers and five large language models. The authors measure agreement using Cohen's kappa, F1, recall, and precision, and report that all third-party annotators achieve only low-to-fair alignment with first-party labels (kappa 0 to 0.45), with LLMs outperforming human annotators on most emotions. They further test whether demographic similarity between human annotators and authors improves alignment and whether adding demographic information to LLM prompts helps. The central claim is that third-party annotations—human or machine—fail to faithfully represent authors' private states, and the paper proposes a framework for evaluating such limitations.","tokens_in":18133,"tokens_out":3474,"duration_ms":37149,"significance":"If valid, the result challenges a widespread assumption in subjective NLP that third-party labels are an acceptable gold standard for private states, with direct implications for emotion recognition, sentiment analysis, and related high-stakes applications. The study's strengths include a direct first-party-versus-third-party design, a reasonably large author-recruited corpus (729 posts from 123 participants), multiple LLMs and human annotators, a detailed emotion taxonomy, and both quantitative and qualitative analyses. The paper is honest about many limitations, including the unverifiability of first-party reports. However, the interpretive claim that low agreement demonstrates failure to model 'authors' private states' depends crucially on treating first-party self-reports as a valid ground truth, and the paper does not provide a sensitivity analysis to show that noise in that criterion does not drive the reported low agreement.","major_comments":[{"comment":"The gold-standard assumption is load-bearing for the paper's central claim, and the manuscript provides no sensitivity analysis to support it. Appendix D states that 9.2% of first-party labels were flagged as spurious by two independent reviewers yet retained, and Section 5 reports that most posts with spurious first-party labels appear in the low-alignment subset. This pattern is exactly what would be observed if criterion noise, rather than third-party limitation alone, were suppressing the measured kappa and F1. The claim in Section 7 that retention 'is expected to have minimal impact on the overall findings' needs quantitative support: the analysis should be rerun after excluding flagged posts, or with a label-noise model, to show that the main conclusions survive.","section":"Section 7 and Appendix D"},{"comment":"The text overstates the significance of the in-group advantage. It says the Wilcoxon tests show that differences 'measured based on F1, recall, precision, and Cohen's kappa, are statistically significant,' but Table 1 reports precision p = 0.050 without an asterisk and Table 2 reports annotator-level precision p = 0.052, which is not significant at the 0.05 level. Since RQ2 is central to the demographic-similarity claim, the reporting should be corrected to distinguish recall/F1 benefits from the nonsignificant precision difference.","section":"Section 4.3 and Table 1"},{"comment":"Multiple Wilcoxon signed-rank tests are reported without any multiple-comparison correction. This matters most for the marginal results: RQ2 precision (p = 0.050), annotator-level precision (p = 0.052), and RQ3 F1 (p = 0.0095) and precision (p = 0.0004) could be affected by correction, even though the very small p-values in Table 3 are robust. The authors should either apply a correction such as Benjamini–Hochberg and report which conclusions survive, or explicitly justify the uncorrected interpretation.","section":"Sections 4.3, 4.4, and 4.5"}],"minor_comments":[{"comment":"The first-party instruction asked participants to select emotions 'they believed were expressed in the post,' which is itself an interpretive judgment rather than direct access to felt emotion. This wording should be acknowledged more prominently in the limitations, since it narrows the gap between first- and third-party interpretations.","section":"Section 3.2"},{"comment":"The sentence listing metrics says 'measured by F1, recall, precision, and recall,' duplicating recall; one of these should be Cohen's kappa.","section":"Section 4.3"},{"comment":"The word 'persevered' in 'we persevered individual third-party annotations' should be 'preserved.'","section":"Section 4.3"},{"comment":"The table header mixes median and mean labels: the left panel is labeled 'In-group Median' but the right panel says 'Out-group Mean,' and both columns contain a statistic and a p-value. The presentation should be made consistent and clearly defined.","section":"Table 3"},{"comment":"The F1 thresholds for high- and low-alignment cases (0.6 and 0.2) are described as 'determined based on the distribution of F1-scores,' which is post hoc. A brief justification or a statement that the qualitative patterns are robust to threshold choice would strengthen the analysis.","section":"Section 5"},{"comment":"The main F1 results are presented only in a figure at the first occurrence, with exact numeric macro-averages given in the text. Adding a table with per-emotion F1, precision, and recall for in-group, out-group, and LLMs would improve reproducibility and readability.","section":"Section 4.2 and Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The core empirical contribution—a direct comparison of first-party and third-party emotion labels—is valuable and within the scope of the journal. The main risk is interpretive: the paper's strongest conclusion ('third-party annotations fail to faithfully represent authors' private states') goes beyond what the current analysis can establish without a sensitivity analysis addressing the 9.2% flagged first-party labels and the retrospective, belief-based wording of the first-party labeling task. I would encourage the editor to request a specific rerun excluding or reclassifying the flagged posts, along with corrected multiple-testing procedures, before considering acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know one thing up front: this is the most direct comparison I’ve seen of third-party emotion annotations against authors’ own labels on real social media posts. The gap is large and consistent — Cohen’s kappa between 0 and 0.45, and LLMs beat humans on most emotions. That finding will survive scrutiny. The paper’s broader claim that third-party annotations “fail to faithfully represent authors’ private states” is a step beyond what the data can prove, and the paper half-admits this itself.\n\nWhat’s genuinely new: the design. First-party authors contribute their own posts and label them; then in-group (age/gender/race matched) and out-group human annotators plus five LLMs label the same posts. That lets them separate demographic similarity effects from model effects. The demographic prompting result is marginal as they say, but the in-group advantage in recall is a real contribution. The qualitative misalignment analysis — spotting posts where the author says “neutral” despite emotional language, or where the emotion is carried by context a third party can’t see — is useful for anyone building annotation guidelines.\n\nSoft spots, in order of importance. First, the gold standard. First-party labels are retrospective self-reports (“emotions they believed were expressed in the post”), not necessarily felt emotion, and 9.2% were flagged as spurious by the authors’ own reviewers yet retained. The paper says this “is expected to have minimal impact” but gives no sensitivity analysis. That doesn’t sink the narrow agreement result — even 9% noise can’t turn a kappa of 0.2 into 0.6 — but it does mean the “private state” language is stronger than the evidence supports. Second, Buechel and Hahn (2017) is cited but never discussed, even though its title is literally “Readers vs. writers vs. texts” on emotion annotation. They need to engage with it. Third, the multiple Wilcoxon tests are uncorrected; minor, since the effects are consistent. Fourth, exact LLM prompts appear only as screenshots, which is avoidable and matters for replication.\n\nWho benefits: anyone training or evaluating emotion models, anyone using LLMs as annotators for subjective constructs, and anyone writing guidelines for private-state annotation. It deserves a serious referee. I’d recommend acceptance with revisions: soften the private-state inference, add the sensitivity analysis dropping flagged posts, put prompts in the appendix as text, and fix the Buechel and Hahn omission. I’d bring it to our reading group.","headline":"Solid empirical study showing third-party emotion labels diverge sharply from author self-reports; the 'private state' interpretation outruns the gold-standard validity, but the core finding and LLM-vs-human comparison deserve peer review.","tokens_in":18680,"tokens_out":3466,"would_cite":true,"duration_ms":32328,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Third-party emotion labels—human or LLM—reach only low to fair agreement with authors' self-reported emotions.","keywords":["emotion recognition","first-party labels","third-party annotation","large language models","demographic similarity","private states","annotation reliability","social media posts"],"falsifier":"One decisive check would be to recompute the alignment after removing the 9.2% of first-party labels that the authors' own quality review flagged as spurious; if Cohen's $\\kappa$ and $F_1$ then rose from low/fair to substantial (say $\\kappa > 0.6$), the central claim would weaken into a self-report noise artifact, but if the scores remained in the same low range, the claim would stand.","tokens_in":17741,"feed_emoji":"💭","tokens_out":7730,"duration_ms":72344,"temperature":0.7,"pith_summary":"The paper asks whether an outsider—a human crowd worker or a large language model—can correctly identify the emotions an author says they felt in a social media post. It reports that they largely cannot: on 729 posts labeled by the authors themselves and independently annotated by human annotators and five LLMs, third-party agreement with first-party labels reaches only low to fair levels (Cohen's $\\kappa$ from 0 to 0.45). LLMs outperform humans on most emotions, and human annotators who share the author's age, gender, and race outperform those who do not. Adding author demographics to LLM prompts improves agreement only marginally. The authors' conclusion is that third-party-annotated datasets are not trustworthy ground truth for private states, and systems built on them risk misreading the people they are meant to model.","feed_headline":"Third-party emotion labels only weakly match what authors feel","feed_subtitle":"Across 729 self-labeled posts, human and LLM annotators reach at most fair agreement with authors' own emotions.","key_machinery":"The central object is a paired first-party/third-party annotation study: social media users supplied their own posts and labeled them with the 27-emotion-plus-neutral taxonomy, then three demographically matched humans, three unmatched humans, and five LLMs labeled the same screenshots. Alignment is measured by Cohen's $\\kappa$ and by $F_1$, recall, and precision, with majority voting to aggregate each group's labels. The design holds the text constant and varies only who labels it, so any measured gap is attributed to the annotator's position as an outsider rather than to the stimulus.","core_discovery":"On its own terms, the paper's central discovery is that self-reported emotions and third-party judgments are systematically misaligned, and the gap is not limited to fine-grained label confusion. When the 28 emotion categories are collapsed into seven broad groups, human and LLM annotators still reach only low to moderate agreement (macro $F_1$ around 0.45–0.54), and at the fine-grained level LLM $F_1$ scores range from 0.2 to 0.6 while human scores range from 0.1 to 0.5. Treating the author's self-report as ground truth, the best third-party annotators—the LLMs—are better at saying what a text expresses than at recovering what the author says they felt. The claim that follows is that emotion-recognition models trained or evaluated on third-party labels are modeling a reader's perception of emotion, not the author's private state.","pith_inferences":["An extension the authors do not draw: if the same first-party/third-party gap holds for other private states such as stance or opinion, many existing NLP benchmarks for those tasks may be measuring reader perception rather than author belief; testing that transfer is a natural next study.","A testable extension would compare first-party labels collected immediately after posting with labels collected retrospectively; if the delayed labels agree more closely with third parties, part of the measured gap is memory or self-report timing, not inability to infer emotion.","Because the retained spurious labels suggest a conservative estimate, a follow-up that reports alignment both with and without those labels would clarify how much of the gap is annotation error versus an unobservable private state.","In practice, the results imply that high-stakes systems should treat emotion predictions as uncertain inferences and should surface that uncertainty or ask the user directly; this is a design implication, not a result the paper proves."],"forward_implications":["Emotion recognition systems trained on third-party labels will tend to encode a reader's guess rather than the author's state, and are likely to misclassify the same people in high-stakes applications.","LLM annotators' higher agreement with first-party labels does not make them a faithful replacement for first-party data, since their agreement is still only fair.","Demographic matching of human annotators improves performance but leaves alignment far from reliable, so matching alone will not solve the private-state annotation problem.","Because misalignment persists on coarse emotion groups, the failure is not merely confusion between near-synonymous labels; third parties pick substantially different emotion categories.","Datasets and benchmarks that use third-party labels for emotion should be treated as measuring perceived emotion rather than felt emotion."],"supporting_citations":[{"why":"Supplies the 27-emotion plus neutral taxonomy and the emotion recognition task the study adapts.","marker":"Demszky et al., 2020"},{"why":"Earlier evidence that author-intended sarcasm and third-party-perceived sarcasm diverge, motivating the first-party comparison.","marker":"Oprea and Magdy, 2019"},{"why":"Shows third-party inferences of stance misalign with self-reported survey responses, suggesting the limitation extends across private states.","marker":"Joseph et al., 2021"},{"why":"Provides the in-group advantage result in emotion recognition that the demographic similarity analysis builds on.","marker":"Elfenbein and Ambady, 2002b"},{"why":"Establishes crowdsourced non-expert annotation as a standard data-collection method, the practice the paper interrogates.","marker":"Snow et al., 2008"},{"why":"Evidence that LLMs can match or outperform crowd workers, framing the expectation that LLM annotations would be strong.","marker":"Gilardi et al., 2023"}],"fun_headline_variants":["Third-party emotion reads miss what authors feel","LLMs beat humans at reading emotions, but both miss authors","Emotion labels from outsiders fail to match authors' own","Why third-party emotion labels diverge from authors' private states","Outsiders can't reliably read your emotions from text"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that an author's self-reported emotion label faithfully captures their internal emotional state; the paper itself concedes that this cannot be externally verified and that 9.2% of first-party labels were flagged as spurious yet kept in the analysis.","fun_headline_variants_meta":{"raw":{"variants":["Third-party emotion reads miss what authors feel","LLMs beat humans at reading emotions, but both miss authors","Emotion labels from outsiders fail to match authors' own","Why third-party emotion labels diverge from authors' private states","Outsiders can't reliably read your emotions from text"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000172,"raw_usage":{"total_tokens":1260,"prompt_tokens":912,"completion_tokens":348,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":268}},"tokens_in":528,"tokens_out":348,"duration_ms":3306,"temperature":1.0,"reasoning_tokens":268,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:12:48.083478+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One decisive check would be to recompute the alignment after removing the 9.2% of first-party labels that the authors' own quality review flagged as spurious; if Cohen's $\\kappa$ and $F_1$ then rose from low/fair to substantial (say $\\kappa > 0.6$), the central claim would weaken into a self-report noise artifact, but if the scores remained in the same low range, the claim would stand.","supporting_citations":[{"cited_title":"iSarcasm: A Dataset of Intended Sarcasm","cited_arxiv_id":"1911.03123","evidence_quote":"Earlier evidence that author-intended sarcasm and third-party-perceived sarcasm diverge, motivating the first-party comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows third-party inferences of stance misalign with self-reported survey responses, suggesting the limitation extends across private states."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes crowdsourced non-expert annotation as a standard data-collection method, the practice the paper interrogates."}],"review_version":1}