{"id":"97527b6b-ade3-49ab-91d7-4d7b4b8c8a05","arxiv_id":"2502.01991","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"GPT-4o few-shot prompting with explanations identified moral frames in vaccine tweets with 90.79 percent agreement from nine annotators, who also reported lower cognitive load and difficulty.","lead":"This paper tests whether large language models, prompted with a few examples and written explanations, can help people label moral messages in COVID-19 vaccine tweets on social media. Nine annotators agreed with the AI labels about 91 percent of the time and said the AI explanations made the task easier, faster, and less mentally draining.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 90.79% accuracy is measured against a majority vote taken after annotators viewed the model's labels and explanations, so it may measure endorsement of the AI suggestion rather than correctness; no independent gold standard is established.","rationale":"The reader's weakest assumption is exactly the load-bearing concern: the evaluation uses a majority vote taken after annotators have seen the model's output as ground truth, making the central accuracy claim self-confirming. The paper's own examples of three-way disagreement in Sec 5.1 show that many items do not have an obvious label, so a majority of nine annotators who just read the model's explanation is not a trustworthy oracle. The high Krippendorff alpha of 0.979 is a red flag rather than evidence of reliability, because it is consistent with most annotators simply accepting the displayed AI label. The without-explanation ablation in Sec 4.5 does not rescue the result because the two conditions are not controlled for order or participant identity. The survey-based claims about cognitive load and time reduction are secondary; even if those self-reports are sincere, they do not support the abstract's claim that the LLM 'enhances accuracy.' The paper does contain useful qualitative observations and a functional annotation interface, but the headline quantitative claim should not be accepted without an independent, blind gold standard. A blind re-annotation study on the same 150 tweets is the direct test that would settle whether the 90.79% figure reflects genuine LLM accuracy or anchoring by human annotators.","tokens_in":16401,"tokens_out":4924,"duration_ms":49972,"concrete_test":"Recruit a new set of at least nine annotators with the same training but have them label the same 150 tweets without ever seeing GPT-4o's labels or explanations (or use a counterbalanced design in which they record their own judgment before revealing the AI suggestion). Treat the blind majority vote as gold and recompute GPT-4o's accuracy and macro-F1. If the resulting accuracy is materially below 90.79% (for example, near the 64.81% figure), the reported accuracy is an anchoring artifact rather than evidence of LLM competence; also report blind majority reliability (Krippendorff's alpha) on the same items.","verdict_should_be":"REJECT","load_bearing_attack":"The load-bearing step is the definition of ground truth in Secs. 4.2 and 4.3. Annotators are shown GPT-4o's moral-foundation label, actor/target roles, and explanations, and are asked to click 'yes' if they agree; the LLM prediction is counted as correct when the majority of the nine annotators click 'yes'. Because the judgment is elicited after the answer is displayed, the resulting 90.79% accuracy (Table 2) conflates correctness with acquiescence to a plausible AI suggestion. This is not a minor measurement detail: in a subjective psycholinguistic task with known three-way disagreements (Sec 5.1), showing the model's answer before collecting the human judgment can anchor annotators, and the reported inter-annotator agreement of alpha=0.979 is exactly what one would expect when most annotators are agreeing with a shared displayed label. The without-explanation comparison in Sec 4.5 is also confounded: annotators cannot be blind to the presence or absence of the AI label, and the paper does not state whether the same annotators scored both conditions or in what order, so the 64.81% figure may reflect task order, fatigue, or different participants. The survey claims in Sec 5.2 (time reduction, exhaustion decrease) are self-reports with no statistical test and no control condition. Consequently, the central claim that LLMs identify morality frames at 90.79% accuracy is not established; the experiment establishes, at most, that annotators usually accept GPT-4o's labels in this interface.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using GPT-4o with few-shot prompting and in-context explanations to generate morality frame annotations (moral foundation, actor/target roles with polarity, and explanations) for COVID-19 vaccine tweets, and it evaluates these outputs with nine human annotators through a purpose-built web interface. The reported results include an overall morality-frame accuracy of 90.79% (Table 2), an inter-annotator agreement of Krippendorff's alpha = 0.979, and survey responses suggesting that LLM-generated labels and explanations reduce task difficulty, cognitive load, and annotation time. The authors argue that this demonstrates a promising human-AI collaborative annotation framework for complex psycholinguistic tasks.","tokens_in":16577,"tokens_out":4952,"duration_ms":49539,"significance":"If the central claims were valid, the paper would offer a practical contribution to human-AI collaboration for subjective annotation tasks, with a concrete tool and an evaluation on a socially relevant corpus. Strengths include the use of an existing COVID-19 vaccine tweet dataset, a relatively detailed description of the annotation interface and participant recruitment, an ablation comparison, and a candid Limitations section. However, the main empirical claim is not established by the current experimental design: the accuracy metric is computed against the majority vote of annotators who were shown the LLM's labels and explanations before judging them, so the reported accuracy conflates correctness with annotator endorsement. This is a load-bearing issue for the paper's central conclusion and cannot be resolved by a purely local revision.","major_comments":[{"comment":"The evaluation protocol makes the headline accuracy measure endogenous. Annotators are shown the LLM's predicted moral foundation, actor/target roles, and explanation, and are asked to click 'yes' if they agree; the LLM prediction is counted as a 'win' when a majority of the nine annotators click 'yes.' This measures acceptance of AI-generated suggestions, not correctness against an independent gold standard. Because moral frame identification is a subjective psycholinguistic task, and the paper itself reports a three-way disagreement in Sec. 5.1, displaying the model's answer before eliciting the human judgment can anchor annotators. The claim in Sec. 5.1 that 'LLMs can identify morality frames ... with an overall accuracy of 90.79%' is therefore not supported by Table 2. The paper should either use an independent blind annotation as ground truth or compare against a pre-existing gold-standard set, such as the labels in the source dataset of Pacheco et al. [57].","section":"Sec. 4.2 and Sec. 4.3"},{"comment":"The reported Krippendorff's alpha of 0.979 requires clarification and is difficult to reconcile with the described annotation task. If alpha is computed over annotators' yes/no responses to the displayed LLM label, it measures agreement with a shared AI suggestion rather than independent coding reliability. If alpha is computed over moral foundation labels, a value of 0.979 is implausibly high given that Sec. 5.1 describes annotators disagreeing among 'sanctity/degradation,' 'none,' and 'care/harm' for at least one tweet. The manuscript should specify the unit of coding, the number of coders, and the disagreement metric used; as reported, the number cannot serve as evidence of reliable annotation.","section":"Sec. 4.2"},{"comment":"The ablation comparison confounds the prompting condition with what the annotators could see. The 64.81% figure in the 'few-shot w/o expl' row is described as resulting from annotators providing judgments 'without looking at explanations,' but it is not stated whether these annotators also saw the LLM labels, whether the same annotators scored the same texts in both conditions, or in what order the conditions were administered. Without a counterbalanced or within-subject design, the 26-percentage-point gap between rows in Table 2 cannot be attributed to the presence of explanations rather than to task order, fatigue, learning, or differences between participant groups. The ablation also does not report inter-annotator agreement for the no-explanation condition.","section":"Sec. 4.5"},{"comment":"The survey-based claims about time reduction and exhaustion are not supported by the data presented. Table 3 reports only a single 'Avg. Time/Batch (min)' column, with no paired measurement for the condition without LLM labels and explanations, no statistical test, and no definition of how the claimed 'at least 50%' time reduction or 'over 60%' decrease in exhaustion were computed. These statements should be presented either as qualitative participant opinions clearly separated from quantitative findings, or substantiated with paired measurements and appropriate tests.","section":"Sec. 5.2 and Table 3"},{"comment":"The accuracy and F1 definitions are underspecified. It is unclear how 'overall accuracy' combines moral foundation correctness with actor-target polarity correctness, whether partial credit is allowed, how the majority vote over corrected annotator labels is aggregated when annotators disagree, and over which classes the macro F1 is computed. The paper should provide a precise scoring rule and a confusion matrix or per-component breakdown, especially because Sec. 5.1 states that some annotators rejected LLM predictions due to role/polarity errors despite correct moral foundations.","section":"Sec. 4.3 and Table 2"}],"minor_comments":[{"comment":"There is a typo in the Loyalty/Betrayal row: 'for the froup' should read 'for the group.'","section":"Table 1"},{"comment":"It is unclear whether the pilot test with a separate batch of 10 tweets was fully excluded from the analysis and whether any pilot feedback changed the annotation interface or instructions before the main 150-tweet evaluation.","section":"Sec. 4.2.1"},{"comment":"The correlation heatmaps in Fig. 6 are reported without significance levels or multiple-comparison corrections; with 150 tweets, many of the displayed correlations may be unstable, and the text should acknowledge this.","section":"Sec. 4.6"},{"comment":"The abstract and introduction describe a 'think-aloud' tool, but no think-aloud verbal protocol data are reported in the results; the paper should either describe how think-aloud responses were collected and analyzed or remove the term.","section":"Sec. 3.2 and Sec. 5"},{"comment":"The claim that 'LLMs can identify moral cases better than non-moral ('none') cases' is not quantified; per-class accuracy or a confusion matrix for the 'none' class should be reported to support this statement.","section":"Sec. 5.1"},{"comment":"The Limitations section candidly notes novelty bias and voluntary participation, but the manuscript does not describe any procedural controls (e.g., debriefing questions, attention checks, or a control condition) that would mitigate these threats.","section":"Sec. 7"}],"recommendation":"reject","confidential_remarks":"The core problem is that the paper's central accuracy claim measures annotators' acceptance of AI-generated labels rather than correctness against an independent gold standard. The authors' own account of three-way disagreement in Sec. 5.1 confirms the subjectivity of the task, which makes the displayed-label protocol especially problematic. I do not see a modest revision that would fix this within the current manuscript; the evaluation would need to be redesigned with blind annotation or an external gold standard. The paper might be suitable for a venue that explicitly targets studies of annotator behavior and human-AI endorsement, but as a claim about LLM accuracy in morality frame identification it is not supportable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the 90.79% headline is not a model accuracy; it is an endorsement rate. Annotators were shown GPT-4o's label and explanation and then asked whether they agreed, and a majority \"yes\" was scored as a model win. That design makes the ground truth partly caused by the prediction, so the central claim is not supported.\n\nWhat is genuinely useful: the paper identifies a real bottleneck—morality-frame annotation is expensive and cognitively heavy—and it builds a sensible web tool that gives annotators a structured way to react to LLM outputs. The few-shot-with-explanations setup is borrowed from Lampinen et al., and morality-frame identification itself is a continuation of the authors' own earlier work, so the novelty is mainly the case study plus the human-evaluation angle. The writeup is clear and the qualitative feedback from nine annotators is plausible and worth reading. The citation pattern is heavy on the authors' prior work, but that is because the formalism and the tweet corpus come from that line; by itself that is not a flaw.\n\nThe soft spots are not minor. First, there is no independent gold standard. The Pentagon example in Sec 5.1 shows the problem: the model calls a neutral factual tweet authority/subversion, and an annotator looking at that explanation is being asked to judge the model's own framing. The reported Krippendorff's alpha of 0.979 is inconsistent with the documented three-way disagreements in the same section; high agreement with a shared displayed label is exactly what anchoring would produce. Second, the ablation in Sec 4.5 confounds the LLM condition with what annotators could see: they cannot be blind to whether an explanation is present, and the paper does not say whether the same nine people did both conditions or in what order, so the 64.81% figure could reflect order or fatigue. Third, the survey claims about 50% time reduction and 60% less exhaustion are self-reports with no statistical test and no control condition. The correlation analysis in Sec 4.6 leans on previously annotated stances and reasons, so it does not shore up the accuracy claim. The Limitations section acknowledges bias, multi-label omission, and participant selection, but it never mentions that the evaluation design lets the model's own output influence the gold standard.\n\nWho this is for: people working on human-AI annotation interfaces, but as a demonstration of a possible tool, not as evidence for the accuracy claim. It deserves a serious referee, because the design flaw is fixable and the research question is real; but in current form the headline result should not be taken at face value. A revised version with an independent gold standard, blinded or counterbalanced annotation, pre-registered analysis, and released artifacts would be worth engaging with.","headline":"The 90.79% headline measures annotator endorsement of GPT-4o's suggestions, not objective accuracy, because the human judges saw the model's label and explanation before saying yes.","tokens_in":17262,"tokens_out":2894,"would_cite":false,"duration_ms":28228,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Few-shot LLM prompts with explanations identify morality frames in vaccine tweets at 90.79% accuracy.","keywords":["morality frames","moral foundations","few-shot prompting","in-context learning","LLM-assisted annotation","vaccine debate","social media","cognitive load"],"falsifier":"Have a fresh set of annotators label the same 150 tweets without ever seeing the LLM outputs; if their independent labels agree with GPT-4o far below 90.79%, or if a control condition in which the suggested labels are deliberately wrong still draws majority 'yes' responses, then the reported accuracy is endorsement rather than correctness.","tokens_in":16025,"feed_emoji":"💉","tokens_out":7452,"duration_ms":68597,"temperature":0.7,"pith_summary":"The paper aims to show that large language models can assist human annotators in identifying morality frames—the moral foundation expressed in a text plus the actor and target roles with positive or negative polarity—in COVID-19 vaccine debates on social media. Using GPT-4o with seven in-context examples and explanations, the authors report an overall accuracy of 90.79% against the majority vote of nine annotators, with 92.67% accuracy and 93.51% F1 on moral foundations alone. When annotators did not see the LLM's explanations, endorsement dropped to 64.81%. Survey responses indicate the AI labels and explanations made the task easier, reduced cognitive load, and roughly halved annotation time. The paper concludes that integrating LLM-generated labels and explanations into annotation improves accuracy and reduces burden.","feed_headline":"Study: AI labels morality frames in vaccine tweets at 90.79% accuracy","feed_subtitle":"Nine annotators endorsed the labels nine times out of ten and said explanations cut difficulty and cognitive load.","key_machinery":"The load-bearing mechanism is few-shot prompting with explanations: the prompt begins with task instructions and definitions of the six moral foundations and actor-target roles, then provides seven examples covering all six foundations plus a non-moral case, each with its label followed by an 'Explanation:' line. The model is forced to choose from the given moral foundation categories, and GPT-4o generates both the moral foundation label and the actor-target polarity roles with explanations. A web-based think-aloud annotation tool presents the LLM output, records yes or no judgments and corrections, and the study aggregates results by majority vote, reporting an inter-annotator agreement coefficient of 0.979.","core_discovery":"The central claim is that few-shot prompting with explanations enables an LLM to produce morality frame annotations that human annotators endorse. On 150 randomly selected tweets from the COVID-19 vaccine debate dataset, GPT-4o prompted with seven examples and explanations achieved 90.79% overall accuracy by the annotators' majority vote, with 92.67% accuracy and 93.51% macro F1 for moral foundation prediction alone. When annotators judged the LLM labels without seeing the explanations, endorsement fell to 64.81%, which the paper interprets as evidence that explanations are what make the collaboration work. All nine annotators reported that the explanations were helpful and reduced cognitive load, and the paper argues this supports a human-AI collaborative model for psycholinguistic annotation.","pith_inferences":["If the anchoring concern is real, the with-versus-without-explanation gap may partly reflect trust in the AI rather than improved understanding; a control that deliberately injects wrong labels would separate deference from learning.","The 90.79% figure is an endorsement rate from nine volunteer annotators with academic degrees who reside in the United States, so generalization to other annotator populations or cultural contexts is untested.","The tool's correction records could be mined as a human-only benchmark for morality frames on the subset of tweets where a majority rejected the AI label, providing a check on LLM bias that the paper does not perform."],"forward_implications":["LLM-generated explanations, not just labels, are what push endorsement from 64.81% to 90.79%, so explanation quality is central to the assistive benefit.","Integrating LLM outputs into annotation can cut per-batch time by at least 50% and reduce self-reported exhaustion by over 60%, making larger psycholinguistic datasets more feasible to build.","Even when the LLM's label was wrong, annotators reported that the explanation helped them arrive at the correct annotation, supporting a collaborative rather than fully automated workflow.","Human oversight remains necessary because LLMs sometimes classify factual statements as moral and can exhibit bias in explanations, as the Pentagon vaccine mandate example shows.","The framework is presented as domain-agnostic and could be applied to other polarized debates, such as political discourse or climate discussion."],"supporting_citations":[{"why":"Supplies the 750-tweet COVID-19 vaccine dataset and the morality frame formalism with actor-target roles and polarity that the paper annotates.","marker":"[57]"},{"why":"Introduces the morality frames task and relational-learning baseline that the paper's prompting approach extends.","marker":"[66]"},{"why":"Shows that placing explanations after answers in context helps LLMs learn; the paper adopts this explanation-placement design.","marker":"[43]"},{"why":"Establishes few-shot in-context learning as the capability the paper exploits for prompting GPT-4o.","marker":"[11]"},{"why":"Identifies GPT-4o as the LLM used in all prompting and explanation generation.","marker":"[54]"},{"why":"Provide Moral Foundation Theory and the six moral foundations that define the label set.","marker":"[25, 26]"},{"why":"Supplies the inter-annotator agreement measure reported as 0.979.","marker":"[42]"}],"fun_headline_variants":["LLMs help humans tag vaccine morality with 90.8% accuracy","AI-assisted annotation hits 90.8% on vaccine morality frames","Explanations key: LLM labels vaccine morality frames well","Human-LLM collab labels vaccine tweets: 90.8% accurate","AI cuts cognitive load in morality labeling of vaccine debates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole result rests on the premise that annotators who have just seen the AI's suggested label and explanation are still able to judge that label independently; if they simply defer to the AI suggestion, the 90.79% figure measures agreement rather than correctness.","fun_headline_variants_meta":{"raw":{"variants":["LLMs help humans tag vaccine morality with 90.8% accuracy","AI-assisted annotation hits 90.8% on vaccine morality frames","Explanations key: LLM labels vaccine morality frames well","Human-LLM collab labels vaccine tweets: 90.8% accurate","AI cuts cognitive load in morality labeling of vaccine debates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000844,"raw_usage":{"total_tokens":3642,"prompt_tokens":882,"completion_tokens":2760,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":2669}},"tokens_in":498,"tokens_out":2760,"duration_ms":20111,"temperature":1.0,"reasoning_tokens":2669,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T13:46:00.216686+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a fresh set of annotators label the same 150 tweets without ever seeing the LLM outputs; if their independent labels agree with GPT-4o far below 90.79%, or if a control condition in which the suggested labels are deliberately wrong still draws majority 'yes' responses, then the reported accuracy is endorsement rather than correctness.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 750-tweet COVID-19 vaccine dataset and the morality frame formalism with actor-target roles and polarity that the paper annotates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the morality frames task and relational-learning baseline that the paper's prompting approach extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that placing explanations after answers in context helps LLMs learn; the paper adopts this explanation-placement design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Identifies GPT-4o as the LLM used in all prompting and explanation generation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the inter-annotator agreement measure reported as 0.979."}],"review_version":1}