{"id":"2fbc7516-f90f-4b72-9883-152789c5cdd7","arxiv_id":"2501.10326","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of LLM-based automated scholarly paper review, cataloging models, datasets, methods, and publisher policies as of 2023-2024.","lead":"This survey maps how large language models are being used to automate scholarly paper review, covering models, datasets, methods, source code, and publisher policies. It is a reference for researchers and publishers deciding whether AI can and should take over peer review.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The snowballing protocol is under-specified and unlogged; if recall against an independent search is low, the 'holistic view' claim is unsupported.","rationale":"The reader's weakest assumption is the representativeness of the snowballing corpus. I agree. My stress-test finds the assumption is not merely unverified; the reported protocol is too under-specified to be checked at all, since Wohlin's procedure requires iteration logging and saturation evidence. Without such a log, an independent reviewer cannot tell whether the corpus is comprehensive or a convenience sample around five seeds. This is the load-bearing condition for the central claim. The ARIES misattribution and unverified statistics are real but secondary; they affect individual entries, not the entire map. The proposed recall test would settle representativeness directly. Since the reader already assigned CONDITIONAL based on similar concerns, my read does not change the verdict.","tokens_in":30593,"tokens_out":5524,"duration_ms":54052,"concrete_test":"Use OpenAlex or Semantic Scholar to (a) enumerate all 2023-2024 works matching queries such as 'LLM AND peer review', 'automated paper review', 'GPT-4 reviewer', and 'review report generation'; (b) compute the union of this set with the forward and backward citation closure of the five seeds; (c) filter by the survey's own inclusion criteria; and (d) measure how many eligible papers are missing from the survey's reference list. If the missing count is material (for example, greater than 10% of the eligible set), the holistic claim is not supported. A second required step is to ask the authors to release the full snowballing log, including per-iteration included and excluded papers and the saturation check, then re-run the same recall computation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 1 states the literature was collected by snowballing from five seeds (Hosseini & Horbach 2023; Robertson 2023; Biswas et al. 2023; Liu & Shah 2024; Liang et al. 2024b) with forward and backward citation search following Wohlin (2014), restricted to papers from 2023-2024 that use LLMs for ASPR. The central claim—'comprehensive review'/'holistic view'—stands or falls on whether this protocol yields a representative corpus. As reported, the protocol is not auditable: no start-set inclusion rationale, no number of iterations, no list of papers screened or excluded, and no saturation criterion. Wohlin's method requires a documented, iterative process; the paper gives only a one-sentence description. The consequence is that any LLM-based ASPR work that neither cites nor is cited by the five seeds will be invisible to the search. Such disconnected works likely exist, including arXiv-only preprint review systems. If they form a material fraction of the 2023-2024 field, the survey's map of models, methods, and datasets is incomplete and the central claim fails. This is not a cosmetic issue; it is the evidential basis for the survey's usefulness as a reference.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews the state of automated scholarly paper review (ASPR) in the era of large language models, organizing the literature into the LLMs used, the technological capabilities they bring (long-text modeling, multimodality, multi-turn dialogue, knowledge acquisition), methods for generating reviews (prompting, fine-tuning, multi-agent frameworks), new datasets, source code and online systems, observed performance and limitations, publisher policies, academic suggestions, and future directions. The literature is collected by snowballing from five seed papers, restricted to works published between 2023 and 2024 that use LLMs for ASPR.","tokens_in":30772,"tokens_out":6009,"duration_ms":55278,"significance":"If the collected corpus is accepted as representative, the survey provides a useful and well-structured map of a fast-moving field, and its consolidated tables of models, datasets, source code, and publisher policies are practical resources for researchers. The authors also make a commendable effort to cover the full pipeline from screening to author response, and they identify several concrete open problems such as multimodal review input, reasoning-model integration, and privately hosted deployment. However, the central claim of a 'comprehensive' or 'holistic' view depends on the completeness and replicability of the snowballing procedure, and that procedure is currently not auditable; several specific factual errors in the survey tables and text further reduce confidence. The contribution is therefore valuable but not yet fully reliable.","major_comments":[{"comment":"The snowballing protocol is under-specified and cannot support the 'comprehensive review'/'holistic view' claim. The text names five seed papers and says the collection was expanded by forward and backward citation searches following Wohlin (2014), but it gives no inclusion/exclusion rationale for the seeds, no number of iterations, no list of screened and excluded papers, and no saturation criterion. Without such documentation, the corpus is not auditable or replicable, and any LLM-based ASPR work that neither cites nor is cited by the five seeds is systematically invisible. Please provide a protocol document (e.g., a PRISMA-style flow diagram), report the iterations and decisions, and ideally complement snowballing with an independent database search to verify recall.","section":"Section 1 (research method)"},{"comment":"Table 3 misattributes ARIES to Couto et al. (2024) in both the 'Comment generation' and 'Author response' rows. The linked repository (github.com/allenai/aries) and the reference list identify ARIES as D'Arcy et al. (2024), 'ARIES: A corpus of scientific paper edits made in response to peer reviews.' This is a factual error in a table whose purpose is to provide reproducible resources, and it is inconsistent with the same work's correct listing under author response in Table 2.","section":"Table 3 and references"},{"comment":"The statement that 'Llama 3, by incorporating 17% structured code data into its pre-training corpus, has improved its zero-shot logical reasoning capability by 23.6% compared to its predecessor' is not supported by the cited references, Touvron et al. (2023a) and Liang et al. (2023), both of which predate Llama 3 and contain no such statistics. Provide a verifiable primary source (e.g., the Llama 3 model card) or remove the quantitative claim.","section":"Section 2.1, page 3"},{"comment":"The architecture classification of Claude 3 as 'MoE-Dec' in Table 1 is not justified by the cited Anthropic source, which does not describe Claude 3 as a mixture-of-experts model. Additionally, the text in Section 2.2 attributes Gemini 1.5's MoE architecture to 'Xue et al. (2024)' but that reference is a paper on wireless distributed MoE, not the Gemini 1.5 technical report. Please correct the citation to Gemini Team (2024) and either provide a source for Claude 3's architecture or mark it as undisclosed.","section":"Section 2.2 and Table 1"},{"comment":"The survey explicitly defers all discussion of biases, fairness, ethics, accountability, and misuse to the authors' earlier work (Lin et al., 2023a). While the earlier work is relevant, this deferral is a substantive gap for a survey claiming a holistic view of LLM-driven ASPR: the new context of LLM-generated reviews introduces or sharpens several of these issues (e.g., training-data contamination, reviewer-accountability, disclosure policies), and simply referring readers elsewhere is not a substitute for an integrated treatment. Add at least a summary subsection on these issues and explain how the 2023-2024 literature addresses them.","section":"Section 1, last paragraph"}],"minor_comments":[{"comment":"The claim that Robertson (2023) 'appears to be the first research paper to employ LLMs for large-scale generation of review comments' is a strong attribution that the survey itself hedges with 'appears to be'; please either substantiate the priority claim with a systematic check or rephrase it as an early example.","section":"Section 4.1, page 11"},{"comment":"The statement that 'a growing phenomenon has emerged where different reviewers utilize LLMs to generate review comments that are strikingly homogeneous' lacks a supporting citation; please cite evidence or qualify it as an observation from the reviewed literature.","section":"Section 8.5, page 21"},{"comment":"There are two separate entries for 'D'Arcy et al. (2024)' (MARG and ARIES), which creates ambiguity in in-text citations; please disambiguate with suffixes (e.g., D'Arcy et al., 2024a, 2024b).","section":"References"},{"comment":"Please use consistent capitalization for model names, e.g., 'gpt-3.5-turbo-16k' should be 'GPT-3.5-turbo-16k' or 'gpt-3.5-turbo-16k' as an API identifier, and clarify which is intended.","section":"Section 3.1, page 9"},{"comment":"The 'ChatReviewer 2023' row in Table 3 does not list an author in the approach column; for consistency, include the citation (Ni, 2023) or a footnote.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The survey fills a clear need for a systematic overview of LLM-based automated reviewing and has useful reference tables. The main concern is methodological: the snowballing protocol needs to be documented and ideally supplemented with an independent search before the 'comprehensive review' claim can be supported. The factual errors in Tables 1 and 3, the unsourced Llama 3 statistics, and the missing ethics discussion should also be addressed. I see no grounds for rejection if these issues are fixed, but they are substantial enough to require a round of major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely useful survey of LLMs for automated scholarly paper review (ASPR) in 2023-2024, and the summary tables—models, datasets, code, publisher policies—are the real contribution. It is curation, not new research, but the curation is mostly careful and the organization is clear. I disagree with any take that calls this a weak paper; as a reference map it earns its keep. The central “holisic view” claim is broadly supported by the cited literature.\n\nThe soft spots are real but individually minor. Table 3 misattributes ARIES to Couto et al. (2024) when the ACL paper is D’Arcy et al. (2024); that’s the kind of error a careful reader will catch and it undermines trust in the table. The Llama 3 “23.6% logical reasoning gain” figure is cited to Touvron et al. (2023a) and Liang et al. (2023), but neither source actually says that, so it’s unsourced. Claude 3 is listed as MoE-Decoder without official confirmation—speculative. And the snowballing protocol is described in one sentence: five seeds, forward/backward search, 2023-2024 window, but no iteration count, screening log, or saturation criterion. The stress-test note makes a fair point: as reported, the protocol isn’t auditable, so the “holisic view” claim is weaker than it could be. But I don’t think that’s lethal—the paper covers a wide range of work beyond the seeds, including independent datasets and systems, so the map is likely reasonably complete in spirit. It’s a transparency problem, not a fatal sampling flaw.\n\nThe paper also leans on the authors’ earlier ASPR work (Lin et al. 2023a) for definitions and ethics, which is acceptable but creates a slight self-citation skew. The publisher policy table is useful and appropriately labeled where unknown.\n\nBottom line: the paper is for researchers entering this area, or for anyone wanting a quick overview of models, datasets, and publisher stances. I’d cite it, and I’d send it to peer review—it deserves referee time, but only with the expectation that the authors fix the citation errors, source or remove the Llama 3 claim, and document the search protocol more fully.","headline":"A genuinely useful survey of LLM-based automated paper review, with a few fixable citation and sourcing errors; worth peer review.","tokens_in":31318,"tokens_out":2291,"would_cite":true,"duration_ms":22317,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey tries to establish a comprehensive map of automated scholarly paper review as it has developed in the era of large language models, covering models, methods, datasets, source code, publisher policies, and open problems.","keywords":["automated scholarly paper review","large language models","peer review","LLM-based review generation","multi-agent review systems","review datasets","publisher policies","academic publishing"],"falsifier":"Search a bibliographic database for papers combining 'peer review' and 'large language model' published before January 2023, or in journals and conferences not connected to the five seed papers; finding a substantial number of such works would show the survey's corpus is incomplete and falsify the comprehensiveness claim.","tokens_in":30346,"feed_emoji":"🤖","tokens_out":12601,"duration_ms":98978,"temperature":0.7,"pith_summary":"This survey sets out to establish a comprehensive map of automated scholarly paper review (ASPR) as it exists in the era of large language models. It argues that LLMs have turned ASPR from a theoretical idea into a rapidly growing research area, one in which automated systems now coexist with and assist traditional peer review. The authors collect work from 2023 to 2024 using snowballing from five seed papers, then organize it by which models are used, which technical bottlenecks have been solved, which methods generate review reports, and which datasets and source code are available. They also catalog performance problems and publisher policies. If the map is accurate, it gives researchers a reliable reference for the tools, gaps, and open challenges in building automated paper review.","feed_headline":"Survey maps how large language models assist peer review","feed_subtitle":"It organizes models, methods, datasets, code, and publisher policies for automated scholarly paper review.","key_machinery":"The survey's organizing machinery is the ASPR pipeline—screening, main review, comment generation, review-quality assessment, and author response—used as a grid to classify papers, datasets, and code. Onto this grid the authors map a set of LLM capabilities that resolve earlier bottlenecks (long-text modeling, multimodal input, multi-turn conversation, instant knowledge retrieval) and a set of generation methods (prompting, fine-tuning, multi-agent frameworks) that produce review reports. This structure is what lets the survey turn a collection of papers into a holistic picture of the field.","core_discovery":"The central claim is that LLM-driven ASPR has reached a stage the authors call the coexistence phase, in which automated review serves as a human assistant rather than a full replacement, and that this phase is now substantial enough to survey. The paper catalogs the field across five axes: the LLMs used, with closed-source models like GPT-4 dominating usage and open-source models catching up; the technical bottlenecks solved, including long-text processing, multimodal input, multi-turn conversation, and real-time knowledge acquisition; the methods for generating review reports, namely prompt engineering, supervised fine-tuning, and multi-agent frameworks; the available datasets and source code; and the current performance issues, including insufficient comprehension, bias, hallucination, confidentiality risks, and limited customization. The paper also reports that most major publishers prohibit reviewers from using AI-generated content tools, while a few are building in-house AI review assistants, and it extracts from academia a set of recommendations including private deployment, reviewer training, transparency, and alignment with scholarly values.","pith_inferences":["A likely next step, if the documented trend continues, is that full-scale ASPR will first deploy in low-stakes tasks such as desk screening and meta-review, where a wrong decision is less costly than in final accept/reject judgments.","The survey's taxonomy implies a testable hypothesis: fine-tuned open-source reviewers will close much of the quality gap with closed-source general-purpose models once large, high-quality review datasets become available.","The survey reports author helpfulness ratings but not whether authors would consent to fully automated review; that consent question may be the true adoption bottleneck.","One missing element in the surveyed corpus is longitudinal evidence on whether LLM-assisted review actually changes editorial outcomes; a causal study comparing acceptance rates with and without LLM assistance would test the coexistence-phase premise."],"forward_implications":["A researcher new to the field can use the survey's tables to select an LLM, a review-generation method, and a dataset for a specific subtask without redoing the literature search.","The performance findings imply that current LLM-generated reviews are not yet dependable enough to replace human reviewers: they are prone to false positives in error detection, inflated scores, and hallucinated references.","Because most publishers currently prohibit AI-generated review content, scaling ASPR depends on policy change or on the in-house AI tools a few publishers are already building.","The scarcity of multimodal review datasets is the main constraint on multimodal LLM-based review, so dataset construction may be a higher-leverage activity than new model development.","The survey's catalog of open challenges—hallucination correction, multimodal generation, reasoning models, generative attacks, low-resource private deployment, and personalized review—defines a concrete research agenda."],"supporting_citations":[{"why":"Defines the ASPR concept and pipeline that this survey extends, providing the coexistence-phase framing and the list of technological bottlenecks.","marker":"Lin et al., 2023a"},{"why":"Supplies the snowballing method used to select the surveyed literature, making the corpus definition depend on this procedure.","marker":"Wohlin (2014)"},{"why":"Seed paper on LLM use in peer review; contributes the bias, confidentiality, and reviewer-accountability concerns that shape Sections 8 and 10.","marker":"Hosseini & Horbach (2023)"},{"why":"Seed paper showing prompt-based LLM review comment generation at scale and providing the author helpfulness ratings used in Section 7.1.","marker":"Robertson (2023)"},{"why":"Seed feasibility study of ChatGPT for journal reviews, cited for streamlining the publication pipeline.","marker":"Biswas et al. (2023)"},{"why":"Seed paper providing the checklist verification and error identification experiments that ground Section 7's performance claims.","marker":"Liu & Shah (2024)"},{"why":"Seed paper with an automated pipeline for review report generation and a large-scale empirical analysis of LLM feedback usefulness.","marker":"Liang et al., 2024b"},{"why":"Provides the meta-reviewer comparison of closed-source versus open-source LLMs and the ReviewCritique dataset, underpinning Section 8.1's discussion of false positives.","marker":"Du et al., 2024"}],"fun_headline_variants":["LLM peer review: a survey of tools, methods, and policies","Survey: how LLMs are changing automated scholarly review","Automated paper review enters the LLM era: a survey","LLMs for peer review: what works, what doesn't, what's next","The state of LLM-based automated scholarly paper review"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's claim to provide a holistic view depends on the assumption, stated in Section 1, that snowballing from five seed papers restricted to 2023–2024 captures all the field's significant work; if important LLM-based review systems exist outside that window or citation chain, the picture is incomplete.","fun_headline_variants_meta":{"raw":{"variants":["LLM peer review: a survey of tools, methods, and policies","Survey: how LLMs are changing automated scholarly review","Automated paper review enters the LLM era: a survey","LLMs for peer review: what works, what doesn't, what's next","The state of LLM-based automated scholarly paper review"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000856,"raw_usage":{"total_tokens":3732,"prompt_tokens":971,"completion_tokens":2761,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":2655}},"tokens_in":587,"tokens_out":2761,"duration_ms":18029,"temperature":1.0,"reasoning_tokens":2655,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:11:26.528504+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Search a bibliographic database for papers combining 'peer review' and 'large language model' published before January 2023, or in journals and conferences not connected to the five seed papers; finding a substantial number of such works would show the survey's corpus is incomplete and falsify the comprehensiveness claim.","supporting_citations":[{"cited_title":", author Dobaria, D","cited_arxiv_id":null,"evidence_quote":"Seed feasibility study of ChatGPT for journal reviews, cited for streamlining the publication pipeline."},{"cited_title":", & author Shah, N","cited_arxiv_id":null,"evidence_quote":"Seed paper providing the checklist verification and error identification experiments that ground Section 7's performance claims."}],"review_version":1}