Pith. sign in

REVIEW 13 cited by

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.06559 v2 pith:CLRQDII6 submitted 2025-02-10 cs.AI

classification cs.AI
keywords benchmarksissuesbenchmarkingmodelspracticesquantitativebenchmarkbroader
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Quantitative Artificial Intelligence (AI) Benchmarks have emerged as fundamental tools for evaluating the performance, capability, and safety of AI models and systems. Currently, they shape the direction of AI development and are playing an increasingly prominent role in regulatory frameworks. As their influence grows, however, so too does concerns about how and with what effects they evaluate highly sensitive topics such as capabilities, including high-impact capabilities, safety and systemic risks. This paper presents an interdisciplinary meta-review of about 100 studies that discuss shortcomings in quantitative benchmarking practices, published in the last 10 years. It brings together many fine-grained issues in the design and application of benchmarks (such as biases in dataset creation, inadequate documentation, data contamination, and failures to distinguish signal from noise) with broader sociotechnical issues (such as an over-focus on evaluating text-based AI models according to one-time testing logic that fails to account for how AI models are increasingly multimodal and interact with humans and other technical systems). Our review also highlights a series of systemic flaws in current benchmarking practices, such as misaligned incentives, construct validity issues, unknown unknowns, and problems with the gaming of benchmark results. Furthermore, it underscores how benchmark practices are fundamentally shaped by cultural, commercial and competitive dynamics that often prioritise state-of-the-art performance at the expense of broader societal concerns. By providing an overview of risks associated with existing benchmarking procedures, we problematise disproportionate trust placed in benchmarks and contribute to ongoing efforts to improve the accountability and relevance of quantitative AI benchmarks within the complexities of real-world scenarios.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems

    cs.AI 2026-08 unverdicted novelty 6.0 of 10

    MAAC defines a nine-dimension framework for process-oriented cognitive evaluation of text-based AI, shifting assessment from outputs to underlying reasoning.

  2. The Foreign Policy AI Evaluation Gap

    cs.CY 2026-07 conditional novelty 6.0 of 10

    Public technical AI governance almost never evaluates real foreign-policy AI workflows; the paper maps that gap and proposes task-scoped, human-recombined evaluation instead of model leaderboards.

  3. What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles

    cs.AI 2025-08 unverdicted novelty 6.0 of 10

    TurtleSoup-Bench is a new interactive benchmark showing that LLMs struggle with imaginative reasoning compared to humans.

  4. Deprecating Benchmarks: Criteria and Framework

    cs.CY 2025-07 conditional novelty 6.0 of 10

    A framework for deprecating outdated or flawed AI benchmarks, with seven criteria and a three-phase process of assessment, reporting, and notification.

  5. Lilith: Developmental Modular LLMs with Chemical Signaling

    q-bio.NC 2025-07 reject novelty 6.0 of 10

    A conceptual framework in which untrained modular LLMs are developed through simulated life and token-based chemical signaling, with the goal of enabling empirical study of consciousness emergence via Integrated Infor...

  6. Establishing Best Practices for Building Rigorous Agentic Benchmarks

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Agentic benchmarks frequently mis-grade agents, and the new ABC checklist helps identify and correct such errors in ten popular benchmarks.

  7. Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A TEE-based protocol for cryptographically verifiable AI safety benchmark results, demonstrated on Llama-3.1 with AWS Nitro Enclaves.

  8. VLM@school -- Evaluation of AI image understanding on German middle school knowledge

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new German middle school visual question-answering benchmark shows open-weight VLMs score below 45% overall, with especially weak results in music, math, and adversarial questions.

  9. Position: Explanation Stability Is a Property of the Model Method Pair, Not the Model

    cs.CV 2026-07 conditional novelty 5.0 of 10

    On chest X-rays, LayerCAM ranks InceptionV3 as the most stable model but Grad-CAM++ ranks DenseNet201 first, showing explanation stability is a property of the model-method pair.

  10. AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

    cs.AI 2026-02 conditional novelty 5.0 of 10

    VERA-MH, an automated safety benchmark for suicide-risk chatbot conversations, agreed with clinician ratings (IRR 0.81), though the clinical reference was not fully independent.

  11. A Conceptual Framework for AI Capability Evaluations

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A descriptive conceptual framework with seven elements (target, task, subject, inputs, instance, measurement, result analysis) for systematizing analysis of AI capability evaluations.

  12. From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation

    cs.CL 2025-06 conditional novelty 4.0 of 10

    On a new 490-question Arabic depth dataset, Claude 3.5 Sonnet answered about 30 percent correctly, while GPT-4 answered about 9 percent, showing current models are weak on culturally specialized Arabic knowledge.

  13. Policy-Driven AI in Dataspaces: Taxonomy, Explainability, and Pathways for Compliant Innovation

    cs.CR 2025-07 reject novelty 2.0 of 10

    The paper is a literature review that classifies privacy-preserving AI techniques in dataspaces using a qualitative taxonomy of privacy, performance, and compliance ratings.

Pith tools