REVIEW 6 cited by
The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance Rates
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Journals and conferences worry that peer reviews assisted by artificial intelligence (AI), in particular, large language models (LLMs), may negatively influence the validity and fairness of the peer-review system, a cornerstone of modern science. In this work, we address this concern with a quasi-experimental study of the prevalence and impact of AI-assisted peer reviews in the context of the 2024 International Conference on Learning Representations (ICLR), a large and prestigious machine-learning conference. Our contributions are threefold. Firstly, we obtain a lower bound for the prevalence of AI-assisted reviews at ICLR 2024 using the GPTZero LLM detector, estimating that at least $15.8\%$ of reviews were written with AI assistance. Secondly, we estimate the impact of AI-assisted reviews on submission scores. Considering pairs of reviews with different scores assigned to the same paper, we find that in $53.4\%$ of pairs the AI-assisted review scores higher than the human review ($p = 0.002$; relative difference in probability of scoring higher: $+14.4\%$ in favor of AI-assisted reviews). Thirdly, we assess the impact of receiving an AI-assisted peer review on submission acceptance. In a matched study, submissions near the acceptance threshold that received an AI-assisted peer review were $4.9$ percentage points ($p = 0.024$) more likely to be accepted than submissions that did not. Overall, we show that AI-assisted reviews are consequential to the peer-review process and offer a discussion on future implications of current trends
Forward citations
Cited by 6 Pith papers
-
AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification
On a new 8,790-instance benchmark built from Nature Communications review records, LLMs characterize reviewer concerns well (GPT-5.5: 0.754) but verify evidence-backed revision resolution poorly (best 0.501).
-
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
LLMs align with human moral judgments only under high consensus, concentrate on a narrow set of moral values, and the profile-based prompting method's reported improvement is evaluated in-sample.
-
ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
ResearchCodeBench, a 212-task benchmark built from 20 recent ML papers, finds that even the best LLM (Gemini-2.5-Pro-Preview) implements only 37.3% of the weighted code correctly.
-
GPT Editors, Not Authors: The Stylistic Footprint of LLMs in Academic Preprints
Across 2,408 arXiv preprints, LLM-typical word usage does not cluster in any section, indicating that AI assistance, when used, is uniform rather than limited to specific parts of a paper.
-
AI for Auto-Research: Roadmap & User Guide
The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.
-
When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
GPT-5-mini gives weaker papers systematically higher scores than human reviewers, and hidden field-specific prompts in PDFs can force it to assign perfect scores or suppress weaknesses.
Discussion (0). Continue with ORCID to comment.