Pith. sign in

REVIEW 13 cited by

ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.00622 v1 pith:BAC4OOTT submitted 2023-06-01 cs.CL cs.AIcs.DL

classification cs.CLcs.AIcs.DL
keywords errorsllmsreviewinglanguagelargemodelspairsquestion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Given the rapid ascent of large language models (LLMs), we study the question: (How) can large language models help in reviewing of scientific papers or proposals? We first conduct some pilot studies where we find that (i) GPT-4 outperforms other LLMs (Bard, Vicuna, Koala, Alpaca, LLaMa, Dolly, OpenAssistant, StableLM), and (ii) prompting with a specific question (e.g., to identify errors) outperforms prompting to simply write a review. With these insights, we study the use of LLMs (specifically, GPT-4) for three tasks: 1. Identifying errors: We construct 13 short computer science papers each with a deliberately inserted error, and ask the LLM to check for the correctness of these papers. We observe that the LLM finds errors in 7 of them, spanning both mathematical and conceptual errors. 2. Verifying checklists: We task the LLM to verify 16 closed-ended checklist questions in the respective sections of 15 NeurIPS 2022 papers. We find that across 119 {checklist question, paper} pairs, the LLM had an 86.6% accuracy. 3. Choosing the "better" paper: We generate 10 pairs of abstracts, deliberately designing each pair in such a way that one abstract was clearly superior than the other. The LLM, however, struggled to discern these relatively straightforward distinctions accurately, committing errors in its evaluations for 6 out of the 10 pairs. Based on these experiments, we think that LLMs have a promising use as reviewing assistants for specific reviewing tasks, but not (yet) for complete evaluations of papers or proposals.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 27 citations worldwide. Full citation record

  1. NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    NatureBench evaluates ten frontier AI coding agents on 90 tasks from Nature papers under web-search-disabled conditions and finds the strongest agent surpasses published SOTA on only 17.8% of tasks, succeeding mainly ...

  2. Stop Automating Peer Review Without Rigorous Evaluation

    cs.AI 2026-05 conditional novelty 7.0 of 10

    LLM paper reviewers show excessive agreement and are trivially gamed by zero-shot "paper laundering" rewrites, so they should not automate acceptance-relevant judgment without a science of evaluation.

  3. AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification

    cs.CL 2026-07 conditional novelty 6.0 of 10

    On a new 8,790-instance benchmark built from Nature Communications review records, LLMs characterize reviewer concerns well (GPT-5.5: 0.754) but verify evidence-backed revision resolution poorly (best 0.501).

  4. Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

    cs.AI 2026-04 reject novelty 6.0 of 10

    In a benchmark of four AI Scientist systems on 15 FARS proposals, LLM reviewers rated FARS's own papers about twice as high as Sakana v1/v2, CycleResearcher, and Data-to-Paper outputs, but the evaluation lacks human v...

  5. BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?

    cs.CR 2025-10 conditional novelty 6.0 of 10

    An LLM agent generating fabricated papers without experiments gets acceptance-level scores from LLM reviewers up to 82% of the time, and simple integrity-checking mitigations barely beat random.

  6. Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers

    cs.CL 2025-07 conditional novelty 6.0 of 10

    LIMITGEN evaluates LLMs on identifying paper limitations and shows strong models still miss roughly half of obvious flaws, with retrieval augmentation giving modest but consistent gains.

  7. Automatic Evaluation Metrics for Artificially Generated Scientific Research

    cs.CY 2025-02 conditional novelty 6.0 of 10

    A simple title-and-abstract model predicts citation counts better than review scores and outperforms LLM reviewers in matching human review scores, but remains below human consistency.

  8. Are Today's LLMs Ready to Explain Well-Being Concepts?

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    AI judges can score explanations of well-being concepts, and small models fine-tuned with preference data score better than larger models, although judges and explainers are all AIs.

  9. Automated Novelty Evaluation of Academic Paper: A Collaborative Approach Integrating Human and Large Language Model Knowledge

    cs.CL 2025-07 reject novelty 5.0 of 10

    Method novelty prediction from peer-review novelty sentences and ChatGPT method summaries improves accuracy on ICLR 2022 data, but the benchmark leaks reviewer opinions into the input.

  10. SC4ANM: Identifying Optimal Section Combinations for Automated Novelty Prediction in Academic Papers

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Using ICLR 2022/2023 reviewer novelty scores as labels, the study reports that the introduction, results, and discussion sections form the most effective input combination for automated novelty prediction.

  11. How Far Are AI Scientists from Changing the World?

    cs.AI 2025-07 conditional novelty 4.0 of 10

    This survey proposes a four-level capability framework for AI Scientist systems and, using an AI reviewer, finds that current systems produce papers rated well below normal scientific standards.

  12. Position: The ML Community Must Build an AI-Augmented Peer-Review Ecosystem

    cs.AI 2025-06 conditional novelty 4.0 of 10

    The paper argues that AI-assisted peer review is an urgent priority and that its success depends on collecting richer, structured peer review process data.

  13. OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models

    cs.CY 2025-05 conditional novelty 4.0 of 10

    The paper advocates protecting and leveraging OpenReview's peer review corpus as a community asset for LLM-based review assistance, benchmarks, and alignment.

Pith tools