Pith. sign in

REVIEW 4 cited by

Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.13696 v2 pith:YO4I724J submitted 2024-07-18 cs.CL

classification cs.CL
keywords benchmarksbenchbenchbenchmarkvalidityagreementconclusionscrucialdemonstrate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in Language Models (LMs) have catalyzed the creation of multiple benchmarks, designed to assess these models' general capabilities. A crucial task, however, is assessing the validity of the benchmarks themselves. This is most commonly done via Benchmark Agreement Testing (BAT), where new benchmarks are validated against established ones using some agreement metric (e.g., rank correlation). Despite the crucial role of BAT for benchmark builders and consumers, there are no standardized procedures for such agreement testing. This deficiency can lead to invalid conclusions, fostering mistrust in benchmarks and upending the ability to properly choose the appropriate benchmark to use. By analyzing over 40 prominent benchmarks, we demonstrate how some overlooked methodological choices can significantly influence BAT results, potentially undermining the validity of conclusions. To address these inconsistencies, we propose a set of best practices for BAT and demonstrate how utilizing these methodologies greatly improves BAT robustness and validity. To foster adoption and facilitate future research,, we introduce BenchBench, a python package for BAT, and release the BenchBench-leaderboard, a meta-benchmark designed to evaluate benchmarks using their peers. Our findings underscore the necessity for standardized BAT, ensuring the robustness and validity of benchmark evaluations in the evolving landscape of language model research. BenchBench Package: github.com/IBM/BenchBench Leaderboard: hf.co/spaces/IBM/BenchBench

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fluid Language Model Benchmarking

    cs.CL 2025-09 conditional novelty 8.0 of 10

    Fluid Benchmarking, combining IRT-based ability estimation with Fisher-information-based adaptive item selection, improves LM evaluation across efficiency, validity, variance, and saturation in pretraining settings.

  2. Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Agent-safety benchmark scores are not interchangeable: R-Judge's F1 is matched by an always-unsafe baseline, rankings differ across benchmarks, and which held-out outcome you choose flips the capability–safety correlation.

  3. JuStRank: Benchmarking LLM Judges for System Ranking

    cs.CL 2024-12 conditional novelty 7.0 of 10

    JuStRank ranks AI judges by how well their aggregated scores reproduce the Chatbot Arena human system ranking, revealing that judge realization and bias, not just model size, determine ranking quality.

  4. Correlated Errors in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Large language models from different providers and architectures often make the same errors, and more accurate models are especially likely to share mistakes.

Pith tools