Pith. sign in

REVIEW 5 cited by

When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.01781 v2 pith:DDECFACT submitted 2024-02-01 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords benchmarkbenchmarksleaderboardsmodelrankingsselectionanswerexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Model (LLM) leaderboards based on benchmark rankings are regularly used to guide practitioners in model selection. Often, the published leaderboard rankings are taken at face value - we show this is a (potentially costly) mistake. Under existing leaderboards, the relative performance of LLMs is highly sensitive to (often minute) details. We show that for popular multiple-choice question benchmarks (e.g., MMLU), minor perturbations to the benchmark, such as changing the order of choices or the method of answer selection, result in changes in rankings up to 8 positions. We explain this phenomenon by conducting systematic experiments over three broad categories of benchmark perturbations and identifying the sources of this behavior. Our analysis results in several best-practice recommendations, including the advantage of a hybrid scoring method for answer selection. Our study highlights the dangers of relying on simple benchmark evaluations and charts the path for more robust evaluation schemes on the existing benchmarks. The code for this paper is available at https://github.com/National-Center-for-AI-Saudi-Arabia/lm-evaluation-harness.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Holding one player model fixed, changing the verdict grammar, criterion disclosure, and budget rendering moved measured honesty verdicts dramatically, so eval findings can reflect the instrument.

  2. Automatic Evaluation of Healthcare LLMs Beyond Question-Answering

    cs.CL 2025-02 reject novelty 6.0 of 10

    In healthcare LLM evaluation, multiple-choice accuracy and open-ended task scores correlate only weakly, and the paper's proposed Relaxed Perplexity metric aims to improve open-ended factuality scoring but rests on an...

  3. RoToR: Towards More Reliable Responses for Order-Invariant Inputs

    cs.CL 2025-02 conditional novelty 6.0 of 10

    RoToR makes a frozen LLM order-invariant by circularly rotating a single global sort of segment position IDs, and Selective Routing combines it with the original model for mixed lists.

  4. Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

    cs.AI 2025-02 conditional novelty 4.0 of 10

    A meta-review of about 110 critical studies finds nine systemic weaknesses in AI benchmarking and concludes that benchmarks are receiving disproportionate trust in AI governance.

  5. A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks

    cs.CL 2025-02 conditional novelty 3.0 of 10

    A narrative review of behavioral and representational Theory of Mind in LLMs, with a taxonomy of safety risks and mitigation directions.

Pith tools