{"id":"9badb4ca-38d4-4062-848d-560b786114be","arxiv_id":"2412.00281","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AnnotateGPT uses GPT-4 to annotate manuscripts by review criteria, and a nine-person TAM survey reports perceived usefulness and ease of use, though annotation accuracy is never measured.","lead":"This paper presents AnnotateGPT, a browser extension that uses GPT-4 to highlight important passages of a manuscript and organize them by review criteria before a human reviewer reads it. A nine-person usability survey suggests reviewers find the tool helpful for maintaining focus, but the study does not test whether the AI's highlights are actually correct or whether they improve review quality.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LLM annotation accuracy is assumed, not measured, so the TAM results cannot establish that AnnotateGPT is a viable reviewing aid.","rationale":"The reader's weakest-assumption analysis identified LLM accuracy at identifying relevant excerpts as load-bearing, and the paper itself confirms this is an untested premise (Section 5) with acknowledged consequences (Section 6.1). My independent reading reaches the same conclusion. The central claim is not disproven, but it is not established: the TAM survey measures perception, not annotation correctness or downstream review quality, and the small self-selected sample reviewing familiar papers cannot compensate for the missing accuracy measurement. Other potential concerns, such as the lack of a control group or the overgeneralization from nine participants, are real but secondary; they all become less damaging if annotation accuracy is shown to be high. The concrete test I propose directly targets the load-bearing premise by measuring precision and recall against expert human annotations, and by comparing against human inter-annotator agreement. This check would either validate the premise or force a substantive weakening of the 'viable middle ground' claim. Since the reader already rendered a conditional verdict, my analysis supports that verdict rather than shifting it.","tokens_in":8961,"tokens_out":2202,"duration_ms":23605,"concrete_test":"Build a gold-standard evaluation set: select 10 manuscripts outside the authors' immediate research field and have at least three expert reviewers independently mark excerpts relevant to four standard criteria (originality, relevance, rigor, contribution). Run AnnotateGPT's default prompts on the same manuscripts and compare its highlights with the human majority annotations at paragraph or sentence level, computing per-criterion precision, recall, and F1. Then compare these scores to the human inter-annotator agreement on the same task. If AnnotateGPT's F1 is substantially below human inter-annotator agreement (e.g., F1 < 0.7 or below the human baseline by more than 0.15), the premise stated in Section 5 is not met and the viability claim should be revised; if F1 matches or exceeds human agreement, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that annotation is a viable middle ground for AI-human collaboration in peer review. For that claim to hold, GPT-4's excerpt highlights must be sufficiently accurate and relevant. Section 5 explicitly refuses to test this: \"This evaluation does not test the accuracy of GPT-4 in highlighting the right paragraphs... this work takes as a premise the increasing accuracy that LLMs exhibit in this task.\" Section 6.1 then concedes that false positives cause unnecessary fatigue and, more importantly, that false negatives \"could lead to overlooking relevant paragraphs.\" The only empirical evidence is a nine-participant TAM study measuring perceived usefulness and ease of use, not whether the highlights are correct or whether they improve review quality. Participants reviewed their own papers, which is likely to mask low annotation accuracy because they already know which excerpts matter. The authors themselves note in the threats to validity that perceived usefulness may be influenced by the perceived accuracy of GPT. Because the viability of the interaction model is inseparable from the reliability of the annotations it presents, the load-bearing premise is unverified. If annotation precision and recall are poor, the TAM positivity would not support the paper's conclusion that annotation is a viable middle ground; it would only show that reviewers like a tool whose core output may misdirect their attention.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes annotation-based human-AI collaboration for academic peer review: rather than generating whole reviews, an LLM (GPT-4) highlights manuscript excerpts relevant to explicit review criteria. The authors design AnnotateGPT, a Chrome extension implementing this interaction model, with color-coded criteria highlights, annotation-centric prompts, and compilation of highlights into criterion-based review reports. They motivate the design through three quality criteria for feedback (contextualized, specific, timely) and present a TAM usability survey with nine participants who reviewed their own papers. The paper concludes that annotation is a viable middle ground for AI-human collaboration and generalizes the findings to conference organizers and authors.","tokens_in":9271,"tokens_out":2249,"duration_ms":21821,"significance":"If the interaction-model claim is validated, the work offers a useful design pattern: LLM assistance that keeps the human reviewer in control while reducing the cost of locating relevant excerpts. The paper's strengths include a clear conceptual model (UML in Fig. 1), a working open-source artifact with public code and video, explicit discussion of limitations, and a framing that distinguishes augmentation from full automation. These are meaningful contributions to the design-oriented literature on AI-assisted peer review. The main significance is, however, conditional on an accuracy premise that the paper explicitly declines to test, and the empirical evidence is a nine-participant TAM survey without a control group.","major_comments":[{"comment":"The paper explicitly states: \"This evaluation does not test the accuracy of GPT-4 in highlighting the right paragraphs... this work takes as a premise the increasing accuracy that LLMs exhibit in this task.\" This premise is load-bearing for the central claim in the abstract and Section 6 that annotation is a \"viable middle ground\" for AI-human collaboration. If the highlights have low precision or recall, perceived usefulness in the TAM survey does not translate into actual review improvement; Section 6.1 itself concedes that false negatives \"could lead to overlooking relevant paragraphs.\" Since the viability of the interaction model is inseparable from the reliability of the annotations it presents, the central claim is currently supported only by assumption, not by evidence. A precision/recall check on a small sample of manuscripts, or a restriction of the claim to a design-study scope, is needed.","section":"Section 5, first paragraph"},{"comment":"The evaluation protocol asked nine participants to use AnnotateGPT to review one of their own papers, with the rationale that content familiarity enhances their ability to understand the output. This confounds the assessment: a reviewer who already knows the paper cannot experience the tool's value in focusing attention on potentially overlooked excerpts, and cannot distinguish whether highlights are useful because they are correct or merely because they are plausible. The threats-to-validity paragraph mentions that perceived usefulness may be affected by perceived GPT accuracy, but it does not address the self-paper confound. As a result, the TAM results measure an interaction model under near-ideal conditions where the reviewer does not need the assistance, which weakens the support for the \"viable middle ground\" claim.","section":"Section 5, Execution"},{"comment":"The title \"Formalization of Learning\" and the generalizations in Section 6.2 (conference organizers, authors) go beyond what a nine-participant, no-control TAM study can support. The authors themselves acknowledge in Section 7 that \"we need larger quantitative evaluations (involving more subjects) and qualitative evaluations (involving conference endowment).\" Given that admission, the current paper should not present these generalizations as findings; it should frame them as hypotheses or speculative implications. This is a scope/claim mismatch that a revision should address, either by weakening the conclusions or by adding the missing evaluation.","section":"Section 6 and Section 7"}],"minor_comments":[{"comment":"The text states that \"the impact of false negatives is limited to cause some unnecessary fatigue... by highlighting unnecessary paragraphs\" and then immediately says \"false negatives have a more significant impact as they could lead to overlooking relevant paragraphs.\" The first clause appears to describe false positives, not false negatives. This typo should be corrected.","section":"Section 6.1"},{"comment":"The term \"performant feedback\" is introduced without definition or citation, and its relation to Nicol's \"timely feedback\" is only loosely explained. Please provide a definition or a clearer contrast.","section":"Section 3"},{"comment":"The citation [9] is referred to as \"Ghosal et al.\" in the bibliography but \"Goshal et al.\" in the body text. The spelling should be made consistent.","section":"Section 2"},{"comment":"The \"Reverse Prompt Engineering\" technique is cited only to a blog post [1]. If this is a recognized method, a peer-reviewed or more stable reference would be preferable; otherwise the name should be introduced as the authors' own term.","section":"Section 4.1"},{"comment":"The figure shows TAM results but the axes or legend are not described in the caption. Please state what the numbers and colors represent, and ideally provide the exact questionnaire items used for Perceived Usefulness and Perceived Ease of Use.","section":"Figure 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is an honest design study with a working artifact, but its central claim overreaches the evidence. The gap between the unverified accuracy premise and the \"viable middle ground\" conclusion, combined with the self-paper TAM protocol, is substantial but fixable: the authors can either add a modest annotation-accuracy evaluation and compare against unannotated reading, or reframe the paper as a design exploration with a preliminary usability probe. I recommend major revision rather than rejection, since no load-bearing error appears beyond the missing evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: AnnotateGPT is a real, usable browser extension, and the interaction model — criterion-driven highlighting rather than full review generation — is a sensible middle ground. The paper is honest about what it does not test, but the evidence is too thin to carry the “viable” conclusion in the abstract.\n\nWhat’s new: the annotation-centric interaction model. Prior LLM review tools generate reports or scores; this one lets reviewers prompt for excerpts per criterion, color-code them, add manual highlights, and compile them into a review. The mapping to feedback theory (Nicol) is used thoughtfully. The artifact is public — code on GitHub, demo video, example report — which is real credit.\n\nSoft spots: the load-bearing assumption is that GPT-4 highlights are accurate enough to guide attention. Section 5 explicitly says accuracy is not tested and takes LLM accuracy as a premise. That is acceptable for a design study, but Section 6 and the abstract still claim annotation is a “viable middle ground.” The TAM results (n=9) measure perceived usefulness and ease of use, not whether reviews improve or highlights are correct. Participants reviewed their own papers, which masks accuracy problems because they already know what matters. The threats-to-validity section acknowledges the perceived-usefulness confound, but it doesn’t fix the gap.\n\nThe stress-test note is on target. The viability claim is inseparable from annotation reliability, and the paper gives no evidence for it. That doesn’t sink the paper as a proof-of-concept, but it means the conclusion should be “promising interaction model, pending accuracy evaluation” rather than “viable.”\n\nWho gets value: researchers working on LLM-assisted review or human-AI collaboration, and tool builders. A reviewer with a soundness bar will want a follow-up with external manuscripts, a control group, and at least a basic precision/recall check on highlights.\n\nRecommendation: send it to review rather than desk-reject. The artifact and design thinking deserve referee time, and the open code means the missing evaluation can be done by others. The paper needs major revision on the evidence-versus-claims mismatch, but the core idea is legitimate.","headline":"An open, honest design study whose central viability claim outruns its thin evidence; worth refereeing, but the conclusion needs to be scaled back to 'promising interaction model pending accuracy evaluation.'","tokens_in":9613,"tokens_out":1847,"would_cite":false,"duration_ms":17920,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI-generated annotations offer a middle path for peer review","keywords":["peer review","annotations","large language models","GPT-4","AI augmentation","Technology Acceptance Model","review criteria","AnnotateGPT"],"falsifier":"A controlled study in which reviewers read the same manuscript either with GPT-4's highlights or without, and then all reviews are scored for missed relevant passages and false claims, would settle the benefit claim; if annotated readers systematically overlook relevant content that unannotated readers catch, the viability claim fails. More directly, computing precision and recall of GPT-4 highlights against a human-annotated gold standard across a sample of manuscripts would test the accuracy premise.","tokens_in":8769,"feed_emoji":"🖍️","tokens_out":5074,"duration_ms":40246,"temperature":0.7,"pith_summary":"This paper argues that the right role for large language models in academic peer review is not to write entire reviews but to highlight the manuscript excerpts that bear on each review criterion. To test that idea, the authors built AnnotateGPT, a browser extension that overlays GPT-4's criterion-based highlights and comments onto a PDF manuscript, and asked nine experienced reviewers to try it. Participants saw the tool as useful for focus and criterion consistency and easy to use, leading the authors to propose annotation as a 'viable middle ground' for AI-human collaboration. The significance lies in a concrete division of labour: the LLM finds and marks evidence, while the human retains judgment about its meaning and value.","feed_headline":"GPT-4 highlights manuscripts to help reviewers focus","feed_subtitle":"Nine experienced reviewers found AnnotateGPT useful and easy to use, indicating a low-risk role for AI in peer review.","key_machinery":"The load-bearing machinery is AnnotateGPT's annotation-centric interaction model, implemented as a Chrome extension over a PDF viewer. Its conceptual core is a UML model where a Review is composed of CriterionReviews, each built from Annotations; an Annotation is an excerpt with optional comments and a sentiment (strength or weakness), and prompts ('annotate', 'compile', 'viewpoints') map onto those entities. The prompting strategy uses 'reverse prompt engineering'—showing GPT-4 desired JSON outputs with examples and letting it iteratively refine the prompt—and each criterion carries a description and actionable recommendations that are injected into the prompt. Color-coded highlighting leverages the familiar semantics of tools like NVivo or PDF Annotator, where colors now stand for review criteria.","core_discovery":"On the paper's own terms, the central claim is that annotation—specifically excerpt highlighting driven by review criteria—is a workable middle ground between fully automated review and unaided human review. AnnotateGPT embodies this by letting a reviewer pick a criterion (originality, rigor, relevance, contribution), prompting GPT-4 to return JSON with three supporting excerpts and a sentiment judgment for each, and rendering those as color-coded highlights in the PDF. The reviewer can fact-check, add comments, ask the model for alternative viewpoints, and finally compile a structured report from the annotations. A TAM questionnaire with nine participants found agreement on usefulness, especially as a focus and consistency enabler, and on ease of use, with stronger scores for seamless embedding in the PDF viewer. The authors explicitly set aside the question of GPT-4's accuracy as a premise rather than a result.","pith_inferences":["The paper's claim would be strengthened by measuring whether annotated reading actually changes review outcomes, not just self-reported acceptance; a controlled study comparing reviews of annotated vs unannotated manuscripts could test this.","The annotation-as-middle-ground idea generalizes beyond peer review to any expert reading task where criteria-based attention guidance is needed, such as grant review or legal document analysis.","The paper's own 'false negative' concern suggests a natural extension: instead of only highlighting what the LLM finds, the system could surface the parts of the manuscript it did not highlight, making the absence of evidence visible.","If open-source LLMs reach sufficient annotation accuracy, the cost barrier of GPT-4 disappears and the tool becomes commodity, which is exactly the direction the paper names as future work."],"forward_implications":["If annotation proves viable, LLMs can be embedded in review without delegating judgment, addressing ethical concerns about AI replacing human reviewers.","Reviewers can start from an annotated manuscript, potentially reducing the time spent locating relevant passages and improving criterion consistency.","The same annotation model can be tuned to conference-specific criteria, allowing organizers to distribute customized browser-based review platforms.","Authors could use the same tool to self-assess their drafts against a venue's criteria before submission.","False negatives in GPT-4's highlights—missed relevant excerpts—remain a limitation that future versions would need to mitigate."],"supporting_citations":[{"why":"Establishes that LLMs outperform humans in aspect coverage and informativeness while lacking high-level analysis, motivating the complement-not-replace stance.","marker":"[23]"},{"why":"Supplies the premise that highlighting improves readers' comprehension and attention, which the paper's benefit argument rests on.","marker":"[15]"},{"why":"Provides the Technology Acceptance Model used to operationalize perceived usefulness and ease of use in the nine-participant evaluation.","marker":"[7]"},{"why":"Contributes the three feedback criteria (contextualized, specific, timely) that structure the annotation model and its prompts.","marker":"[12]"},{"why":"Supplies the 'reverse prompt engineering' technique used to iteratively design the GPT-4 annotation prompts.","marker":"[1]"},{"why":"Supports the balanced AI-human approach with an empirical feasibility study of ChatGPT as a standalone reviewer.","marker":"[3]"}],"fun_headline_variants":["AI highlights: a low-risk middle ground for peer review","AnnotateGPT lets reviewers focus with GPT-4 highlights","Study: AI annotations aid reviewers without replacing them","Reviewers find AI annotation useful for focus and consistency"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that GPT-4's criterion-based excerpt highlights are accurate enough to guide reviewers' attention in the right direction; if they frequently miss or mis-tag evidence, the perceived usefulness found in the questionnaire would not translate into better reviews.","fun_headline_variants_meta":{"raw":{"variants":["AI highlights: a low-risk middle ground for peer review","AnnotateGPT lets reviewers focus with GPT-4 highlights","Study: AI annotations aid reviewers without replacing them","Reviewers find AI annotation useful for focus and consistency"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0003,"raw_usage":{"total_tokens":1707,"prompt_tokens":897,"completion_tokens":810,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":745}},"tokens_in":513,"tokens_out":810,"duration_ms":7840,"temperature":1.0,"reasoning_tokens":745,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:31:32.922526+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled study in which reviewers read the same manuscript either with GPT-4's highlights or without, and then all reviews are scored for missed relevant passages and false claims, would settle the benefit claim; if annotated readers systematically overlook relevant content that unannotated readers catch, the viability claim fails. More directly, computing precision and recall of GPT-4 highlights against a human-annotated gold standard across a sample of manuscripts would test the accuracy premise.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes that LLMs outperform humans in aspect coverage and informativeness while lacking high-level analysis, motivating the complement-not-replace stance."},{"cited_title":"The English Journal 93(5), 82–89 (2004)","cited_arxiv_id":null,"evidence_quote":"Supplies the premise that highlighting improves readers' comprehension and attention, which the paper's benefit argument rests on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Technology Acceptance Model used to operationalize perceived usefulness and ease of use in the nine-participant evaluation."},{"cited_title":"Assessment & Evaluation in Higher Education35(5), 501–517 (aug 2010)","cited_arxiv_id":null,"evidence_quote":"Contributes the three feedback criteria (contextualized, specific, timely) that structure the annotation model and its prompts."},{"cited_title":"https://www.allabtai","cited_arxiv_id":null,"evidence_quote":"Supplies the 'reverse prompt engineering' technique used to iteratively design the GPT-4 annotation prompts."},{"cited_title":"The Yale Journal of Biology and Medicine96(3), 415 (2023)","cited_arxiv_id":null,"evidence_quote":"Supports the balanced AI-human approach with an empirical feasibility study of ChatGPT as a standalone reviewer."}],"review_version":1}