{"id":"8459a28a-c843-475f-a708-63a08179d1a4","arxiv_id":"2601.16800","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"LLMs can serve as moderate-fidelity annotators and adjudicators for ASTE/ACOS opinion labels, but exact-match triplet/quadruple accuracy remains far below human relational precision.","lead":"The paper tests whether LLMs can label fine-grained opinions—aspect, sentiment, and opinion span—in reviews, and whether an LLM can merge multiple annotation attempts into a final label set. It finds LLMs are fairly good at locating opinion spans but often fail to link them into the correct triplets or quadruples, making them useful assistants rather than full replacements for human annotators.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Public test-set contamination may inflate span-level F1: ASTE/ACOS evaluation instances predate model pretraining, so the central annotation-generality claim is not yet measured.","rationale":"The reader's weakest assumption (LLM-LLM IAA lacking a human anchor) is real but secondary: the 'span-level reliable' half of the central claim is already anchored by element-wise F1 against human gold, so systematic shared LLM bias would show up as uniformly poor F1, which is not what Tables 4-5 show. The unaddressed threat is benchmark contamination. The exact test instances are decades-old public data, and the models are contemporary; this is a concrete external-validity risk. The paper's own evidence—no contamination analysis, no fresh-data evaluation, no reporting of knowledge cutoffs—leaves the central claim unvalidated for novel text. This is not an accusation of misconduct; it identifies a missing control. The rule-based adjudication baseline is also missing, but that affects only the 'adjudicator' novelty; contamination affects the entire feasibility result. The verdict remains conditional: the pipeline is a reasonable proposal if its benchmark numbers survive a contamination check.","tokens_in":11697,"tokens_out":6029,"duration_ms":73473,"concrete_test":"Build a fresh evaluation set of opinion-bearing sentences from post-2025 reviews (e.g., recent Amazon/Yelp data), double-annotated by trained human annotators per ASTE/ACOS guidelines, and run the identical DSPy pipeline (§4.2, §5) with the same models and prompt settings. Compare span-level element-wise F1 and exact-match triplet/quadruple F1 against Tables 4-5. If the new-set span F1 is materially lower (say >10 points) than the reported benchmark numbers, contamination is the likely cause and the central generalizability claim fails; if it matches, the span-level reliability claim is supported for novel text.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that LLMs are 'reliable at the span level' rests on element-wise F1 against human gold labels (Tables 4-5) computed on the public ASTE (lap14/res14/res15/res16) and ACOS (laptop/restaurant) test splits described in §3.2. These SemEval-derived instances (2014-2016) predate the training data of every model in Table 1 (Qwen3, MiniCPM3, Phi-4, DeepSeek-R1, gpt-oss; 2024-2025 releases). The paper neither reports model knowledge cutoffs nor tests for contamination. With temperature 0 and ICL examples drawn from the same public dev splits (§4.2), high span-level F1 can reflect memorized benchmark instances rather than an ability to annotate unseen opinion text. This is the load-bearing condition for 'high-fidelity annotation assistants' and 'data augmentation tools' in the abstract: those uses are for new, domain-specific text, not for recalling old benchmarks.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates the use of LLMs as automatic annotators and adjudicators for fine-grained opinion analysis, specifically for Aspect Sentiment Triplet Extraction (ASTE) and Aspect-Category-Opinion-Sentiment (ACOS) quadruple extraction. It proposes a declarative DSPy-based annotation pipeline with multiple LLMs as annotators and a separate LLM as an adjudicator that merges redundant annotations. Experiments on six public benchmark datasets across three model-size classes report exact-match and element-wise F1 scores against human gold labels, plus Krippendorff α among the LLM annotators. The central claim is that LLMs achieve high inter-annotator agreement and are reliable at identifying opinion spans, but struggle to reproduce the relational structures linking spans, positioning LLMs as annotation assistants rather than full replacements for human annotators.","tokens_in":11995,"tokens_out":3245,"duration_ms":37410,"significance":"If valid, the result would support a practical pipeline for lowering the cost of creating fine-grained opinion datasets, with humans verifying relational links rather than annotating from scratch. The paper has notable strengths: it evaluates across multiple model families and sizes, provides element-wise F1 decompositions that localize errors to specific span relations, reports IAA for LLM annotators, and includes a qualitative error analysis. However, the central generalization claim is currently under-supported because the evaluation uses public benchmark test splits that predate model training, the IAA result lacks any human-human baseline, and the abstract promises a rule-based voting comparison that never appears in the body. These issues are fixable within the scope of the paper, but they are load-bearing for the stated conclusions.","major_comments":[{"comment":"The abstract states that the adjudication methodology is 'benchmarked against exact, flexible, and element-wise variants of a rule-based voting aggregator,' but no such rule-based voting baseline appears anywhere in the body. Tables 2 and 3 compare only individual annotators and the LLM adjudicator. This is a substantive missing comparison for the adjudication contribution; either add the baseline or revise the abstract.","section":"Abstract; Sections 2–7"},{"comment":"The test splits come from SemEval 2014–2016 (lap14, res14, res15, res16) and the ACOS splits from the same period, while all evaluated models (Qwen3, MiniCPM3, Phi-4, DeepSeek-R1, gpt-oss) were released in 2024–2025. No model knowledge cutoffs are reported and no contamination check is performed. Because temperature is 0 and in-context examples are drawn from the same public dev splits, high span-level F1 could reflect memorization of public benchmark instances rather than an ability to annotate unseen opinion text. Since the abstract and conclusion frame the contribution as reducing annotation cost for new, domain-specific data, this is load-bearing. Please report knowledge cutoffs, test for overlap, or evaluate on text released after model training.","section":"§3.2, §6, Tables 2–5"},{"comment":"The reliability claim is based on Krippendorff α among the three LLM annotators. If the LLMs share systematic biases—likely given overlapping training data—high LLM-LLM agreement can occur while all three disagree with humans. No human-human α on the same datasets is reported, so 'highly reliable' has no external anchor. Additionally, the computation of α for overlapping span annotations is not described (unitizing, annotation units, software used). Please report human-human agreement or validate LLM agreement against human labels.","section":"§7.1, Table 6"},{"comment":"The adjudicator is the best-performing annotator, A1, and the input to adjudication includes A1's own annotations. This self-referential design can inflate adjudicated agreement relative to a scenario where the adjudicator is independent of the annotators. The effect is not quantified. Please either use an independent adjudicator, remove A1's output from the adjudicator input, or analyze the extent of this inflation.","section":"§5, §6"}],"minor_comments":[{"comment":"The bottom half of Table 6 is mislabeled as 'Performance scores for ACOS tasks' but the values appear to be Krippendorff α values. In addition, the Mini laptop row contains '36.75' where an α value between 0 and 1 is expected; this appears to be a typographical error.","section":"Table 6"},{"comment":"The paper spells 'Krippendorff' as 'Kirppendorff' in the conclusion. Please correct.","section":"§8"},{"comment":"The annotator assignment line reads 'A1, A1, A3' instead of 'A1, A2, A3'.","section":"§5"},{"comment":"In the error analysis discussion, the text says 'a minor mistake in identifying the aspect term (ac)' — in ASTE the aspect term is abbreviated 'at', not 'ac', which is the ACOS category abbreviation.","section":"§7.2"},{"comment":"No code, prompts, or data are released. Given the claim of a declarative DSPy pipeline, releasing the optimized prompts and scripts would materially improve reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper's core idea is timely and the experimental design is a reasonable first step, but the current manuscript overclaims relative to the evidence. The missing rule-based voting baseline promised in the abstract is a serious omission, and the contamination and human-anchor issues undermine the central 'reliable annotator' claim. These are addressable with additional experiments or careful rewording, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a useful empirical paper, not a breakthrough, and it has reporting problems that need fixing before I'd trust the conclusions. The core contribution is the element-wise breakdown: for ASTE/ACOS, LLMs get sentiment and individual spans mostly right, but the joint relations—aspect-opinion pairs, and especially full quadruples—fall apart. That span-versus-relational bifurcation is consistent across models and datasets, and it is the kind of concrete diagnostic that makes the paper worth reading. The DSPy pipeline is competently executed, and the adjudicator, while simple, does improve results most of the time.\n\nWhere I part ways with the abstract: \"high Inter-Annotator Agreement across individual LLM-based annotators\" is measured between LLMs, not against humans. If three models share the same biases—likely, given overlapping training data—high alpha can coexist with all three being wrong. The F1 against human gold anchors quality, so it's not fatal, but without human-human alpha on the same data, \"reliable\" is unanchored.\n\nMore concerning is that all evaluation uses the public ASTE/ACOS test splits from 2014–2016, while the models are from 2024–2025. The paper never checks for contamination or reports knowledge cutoffs. At temperature zero, with ICL examples drawn from the same public dev splits, the span-level F1 could partly be memory of benchmark instances. That doesn't sink the paper—the span-versus-relation pattern would still need explaining—but it does mean the \"high-fidelity annotation assistant for new text\" claim is not actually measured.\n\nThe abstract also promises a benchmark against exact, flexible, and element-wise rule-based voting aggregators, and I could not find that comparison anywhere in the body. That is a missing promised result. There are also small but real errors: Table 6 has an ACOS value of 36.75 that should be 0.3675, and one F1 in Table 2 doesn't match its P and R. No error bars or significance tests either, though with temperature zero that is less critical.\n\nNet: the main diagnostic is solid enough to build on, but the paper overstates reliability and under-delivers on its own promised comparison. I would send it to peer review, but only with a request that the authors run a contamination check, add a human-human IAA baseline, fix the missing rule-based table, and tone down the reliability language.\n\nBest,\n[You]","headline":"Useful element-wise diagnostic that LLMs get opinion spans but not relations; needs a contamination check and a cleaner IAA story before the 'high reliability' claim can stand.","tokens_in":12444,"tokens_out":2044,"would_cite":true,"duration_ms":26055,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large language models can annotate opinion spans reliably, but the relational links between spans are where they lose fidelity.","keywords":["LLM annotation","fine-grained opinion analysis","ASTE","ACOS","annotation adjudication","inter-annotator agreement","sentiment analysis","data augmentation"],"falsifier":"Compare the inter-annotator agreement of the three LLM annotators to human-human agreement on a sample of the same ASTE and ACOS test texts. If human-human agreement is substantially higher than LLM-LLM agreement on the relational components (aspect–opinion link and aspect category), the claim that LLMs are reliable annotators at the relational level is falsified. A simpler check: list cases where all three LLMs agree but all disagree with the human annotation and count their frequency.","tokens_in":11641,"feed_emoji":"🤖","tokens_out":3448,"duration_ms":37065,"temperature":0.7,"pith_summary":"This paper argues that LLMs can serve as automatic annotators and adjudicators for fine-grained opinion analysis, cutting the cost of creating labeled datasets. Using a declarative annotation pipeline and an LLM-based adjudicator that merges multiple model annotations, the authors show the approach works well for identifying individual opinion spans—aspect terms, opinion phrases, sentiment polarity—but fails to faithfully reconstruct the relational structures, such as which opinion phrase modifies which target, especially in the more complex ACOS quadruple task. The upshot the authors draw is that LLMs are better positioned as high-fidelity annotation assistants and data augmentation tools, not replacements for human annotators.","feed_headline":"LLMs reliably spot opinion spans but miss the relations","feed_subtitle":"A multi-model adjudication pipeline cuts annotation cost, yet quadruple labels still need human fixing.","key_machinery":"The mechanism is a declarative annotation pipeline that programs LLMs rather than hand-crafting prompts: a data model specifies input/output structure, a small set of annotated examples is used to optimise the prompt, and three LLMs of different sizes each produce a redundant annotation set. A fourth step uses one LLM as an adjudicator, taking the redundant annotations plus the text and producing final labels—an ensemble in the spirit of stacked generalisation. The pipeline is evaluated with exact-match F1 against human annotations for ASTE triplets and ACOS quadruples.","core_discovery":"The central discovery is a performance bifurcation: at the span level, LLM annotators align closely with human annotations (sentiment polarity especially, with aspect and opinion spans not far behind), while at the relation level—pairing an aspect term with the opinion span that expresses it, and assigning the correct aspect category—their agreement with human labels drops sharply. Larger models (32B) align better than smaller ones, and an LLM adjudicator that combines redundant annotations from several models improves exact-match F1 in most settings, sometimes making a 4B ensemble outperform individual 14B models. On the ACOS quadruple task, the aspect category component is the main bottlen","pith_inferences":["The paper's reliability claim rests on inter-annotator agreement among LLMs, but if the three models share training-data biases, high agreement can coexist with systematic error; a human-human agreement baseline on the same datasets would anchor the claim.","A natural extension is to test the adjudication method on other structured annotation tasks, such as event or relation extraction, where span extraction is usually easier than link prediction.","Because the category prediction in ACOS is the main failure mode, augmenting training data with rare categories or using ontology-aware prompting could yield outsized gains.","The temperature-0 setting may understate the variability that matters in practice; a test of adjudication under sampling diversity could reveal whether the ensemble is robust to more varied annotator outputs."],"forward_implications":["LLM-generated span labels can be used to bootstrap or expand opinion datasets cheaply, with human effort redirected to correcting relational links.","Adjudication of multiple LLM annotations can yield final labels that outperform the best individual annotator, especially for smaller models.","The ACOS bottleneck on aspect categories suggests that category coverage and implicit aspect handling are the priority for improving automatic annotation.","For ASTE, the pipeline is closer to production-ready, while ACOS needs more work before LLM annotations can be trusted."],"fun_headline_variants":["LLMs spot opinion spans but fail at relations","LLM annotation: high on spans, low on relations","LLM adjudicator narrows the relation gap","Bigger LLMs and adjudication improve opinion annotation","LLMs: reliable span annotators, shaky linkers"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that high agreement among the LLM annotators signals annotation quality; without a human-human agreement baseline on the same datasets, the three models could be consistently wrong together and still show high inter-annotator agreement.","fun_headline_variants_meta":{"raw":{"variants":["LLMs spot opinion spans but fail at relations","LLM annotation: high on spans, low on relations","LLM adjudicator narrows the relation gap","Bigger LLMs and adjudication improve opinion annotation","LLMs: reliable span annotators, shaky linkers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000211,"raw_usage":{"total_tokens":1254,"prompt_tokens":747,"completion_tokens":507,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":431}},"tokens_in":491,"tokens_out":507,"duration_ms":5937,"temperature":1.0,"reasoning_tokens":431,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T06:15:32.906980+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the inter-annotator agreement of the three LLM annotators to human-human agreement on a sample of the same ASTE and ACOS test texts. If human-human agreement is substantially higher than LLM-LLM agreement on the relational components (aspect–opinion link and aspect category), the claim that LLMs are reliable annotators at the relational level is falsified. A simpler check: list cases where all three LLMs agree but all disagree with the human annotation and count their frequency.","supporting_citations":[],"review_version":2}