Pith. sign in

REVIEW 6 cited by

AI-Driven Review Systems: Evaluating LLMs in Scalable and Bias-Aware Academic Reviews

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.10365 v1 pith:GEHQ4RLQ submitted 2024-08-19 cs.AI

classification cs.AI
keywords reviewshumanreviewpreferencesqualityreviewingevaluateevaluating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automatic reviewing helps handle a large volume of papers, provides early feedback and quality control, reduces bias, and allows the analysis of trends. We evaluate the alignment of automatic paper reviews with human reviews using an arena of human preferences by pairwise comparisons. Gathering human preference may be time-consuming; therefore, we also use an LLM to automatically evaluate reviews to increase sample efficiency while reducing bias. In addition to evaluating human and LLM preferences among LLM reviews, we fine-tune an LLM to predict human preferences, predicting which reviews humans will prefer in a head-to-head battle between LLMs. We artificially introduce errors into papers and analyze the LLM's responses to identify limitations, use adaptive review questions, meta prompting, role-playing, integrate visual and textual analysis, use venue-specific reviewing materials, and predict human preferences, improving upon the limitations of the traditional review processes. We make the reviews of publicly available arXiv and open-access Nature journal papers available online, along with a free service which helps authors review and revise their research papers and improve their quality. This work develops proof-of-concept LLM reviewing systems that quickly deliver consistent, high-quality reviews and evaluate their quality. We mitigate the risks of misuse, inflated review scores, overconfident ratings, and skewed score distributions by augmenting the LLM with multiple documents, including the review form, reviewer guide, code of ethics and conduct, area chair guidelines, and previous year statistics, by finding which errors and shortcomings of the paper may be detected by automated reviews, and evaluating pairwise reviewer preferences. This work identifies and addresses the limitations of using LLMs as reviewers and evaluators and enhances the quality of the reviewing process.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 8 citations worldwide. Full citation record

  1. AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

    cs.CY 2026-08 conditional novelty 6.0 of 10

    LLM-generated reviews on real pre-revision submissions are longer, more positive, and less score-calibrated than human reviews, and aggregate quality scores alone overestimate their quality.

  2. Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review

    cs.CL 2026-05 conditional novelty 6.0 of 10

    Official conference reviewer guidelines improve LLM review scores' agreement with human ratings, while LLM-distilled 'reviewer-imitating' guidelines and rubric-style scoring do worse.

  3. BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?

    cs.CR 2025-10 conditional novelty 6.0 of 10

    An LLM agent generating fabricated papers without experiments gets acceptance-level scores from LLM reviewers up to 82% of the time, and simple integrity-checking mitigations barely beat random.

  4. Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers

    cs.CL 2025-07 conditional novelty 6.0 of 10

    LIMITGEN evaluates LLMs on identifying paper limitations and shows strong models still miss roughly half of obvious flaws, with retrieval augmentation giving modest but consistent gains.

  5. CIKT: A Collaborative and Iterative Knowledge Tracing Framework with Large Language Models

    cs.AI 2025-05 conditional novelty 5.0 of 10

    CIKT, a two-part LLM framework where an Analyst writes student profiles and a Predictor uses them, reports consistent accuracy gains on three knowledge tracing datasets.

  6. When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review

    cs.CY 2025-09 conditional novelty 4.0 of 10

    GPT-5-mini gives weaker papers systematically higher scores than human reviewers, and hidden field-specific prompts in PDFs can force it to assign perfect scores or suppress weaknesses.

Pith tools