Benchmarking and studying the llm-based code review

· 2025 · cs.SE · arXiv 2509.01494

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

open full Pith review browse 3 citing papers arXiv PDF

abstract

Automated Code Review (ACR) is crucial for software quality, yet existing benchmarks often fail to reflect real-world complexities, hindering the evaluation of modern Large Language Models (LLMs). Current benchmarks frequently focus on fine-grained code units, lack complete project context, and use inadequate evaluation metrics. To address these limitations, we introduce SWRBench , a new benchmark comprising 1000 manually verified Pull Requests (PRs) from GitHub, offering PR-centric review with full project context. SWRBench employs an objective LLM-based evaluation method that aligns strongly with human judgment (~90 agreement) by verifying if issues from a structured ground truth are covered in generated reviews. Our systematic evaluation of mainstream ACR tools and LLMs on SWRBench reveals that current systems underperform, and ACR tools are more adept at detecting functional errors. Subsequently, we propose and validate a simple multi-review aggregation strategy that significantly boosts ACR performance, increasing F1 scores by up to 43.67%. Our contributions include the SWRBench benchmark, its objective evaluation method, a comprehensive study of current ACR capabilities, and an effective enhancement approach, offering valuable insights for advancing ACR research.

citation-role summary

background 1

citation-polarity summary

background 1

representative citing papers

Code Review Agent Benchmark

cs.SE · 2026-03-24 · unverdicted · novelty 7.0

c-CRAB benchmark shows state-of-the-art code review agents solve only around 40% of tasks derived from human reviews, suggesting potential for human-AI collaboration.

On the Footprints of Reviewer Bots Feedback on Agentic Pull Requests in OSS GitHub Repositories

cs.SE · 2026-04-27 · unverdicted · novelty 6.0

Reviewer bots' higher comment volume on AI agent PRs is associated with slower resolutions and poorer average feedback quality, while feedback quality itself has no association with PR outcomes.

Rethinking Code Review in the Age of AI: A Vision for Agentic Code Review

cs.SE · 2026-05-17 · unverdicted · novelty 4.0 · 2 refs

Proposes a five-stage agentic AI framework for code review with human quality gates to maintain context, accountability, and team understanding.

citing papers explorer

Showing 3 of 3 citing papers.

Code Review Agent Benchmark cs.SE · 2026-03-24 · unverdicted · none · ref 28 · internal anchor
c-CRAB benchmark shows state-of-the-art code review agents solve only around 40% of tasks derived from human reviews, suggesting potential for human-AI collaboration.
On the Footprints of Reviewer Bots Feedback on Agentic Pull Requests in OSS GitHub Repositories cs.SE · 2026-04-27 · unverdicted · none · ref 27 · internal anchor
Reviewer bots' higher comment volume on AI agent PRs is associated with slower resolutions and poorer average feedback quality, while feedback quality itself has no association with PR outcomes.
Rethinking Code Review in the Age of AI: A Vision for Agentic Code Review cs.SE · 2026-05-17 · unverdicted · none · ref 136 · 2 links · internal anchor
Proposes a five-stage agentic AI framework for code review with human quality gates to maintain context, accountability, and team understanding.

Benchmarking and studying the llm-based code review

citation-role summary

citation-polarity summary

fields

years

verdicts

roles

polarities

representative citing papers

citing papers explorer