ReviewEval: An evaluation framework for AI-generated reviews

Madhav Krishan Garg, Tejash Prasad, Tanmay Singhal, Chhavi Kirtani, Murari Mandal, Dhruv Kumar · 2025 · DOI 10.18653/v1/2025.findings-emnlp.1120

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

open at publisher browse 3 citing papers

representative citing papers

What Makes a Good AI Review? Concern-Level Diagnostics for AI Peer Review

cs.AI · 2026-04-21 · unverdicted · novelty 7.0

Concern alignment is a new diagnostic framework that audits AI peer reviews at the level of individual concerns via match graphs to reveal gaps in detection, calibration, and prioritization beyond verdict agreement.

PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing

cs.AI · 2026-05-28 · unverdicted · novelty 6.0

PRAIB reveals LLM reviews are less variable, positively biased, overconfident, longer, and overlook atomic weaknesses noted by humans compared to real reviewer feedback.

PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers

cs.CL · 2026-05-26 · unverdicted · novelty 6.0

PRISM benchmark finds LLMs match or exceed humans on isolated review dimensions like novelty verification but none achieve the balanced performance of human reviewers across depth, flaw prioritization, and constructiveness.

citing papers explorer

Showing 1 of 1 citing paper after filters.

PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers cs.CL · 2026-05-26 · unverdicted · none · ref 26
PRISM benchmark finds LLMs match or exceed humans on isolated review dimensions like novelty verification but none achieve the balanced performance of human reviewers across depth, flaw prioritization, and constructiveness.

ReviewEval: An evaluation framework for AI-generated reviews

fields

years

verdicts

representative citing papers

citing papers explorer