Pith. sign in

REVIEW 13 cited by

Do Large Language Model Benchmarks Test Reliability?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.03461 v1 pith:RNLXS2AB submitted 2025-02-05 cs.LG cs.CL

classification cs.LGcs.CL
keywords benchmarksmodelsreliabilityfailuresllmsmodelbeenerrors
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

When deploying large language models (LLMs), it is important to ensure that these models are not only capable, but also reliable. Many benchmarks have been created to track LLMs' growing capabilities, however there has been no similar focus on measuring their reliability. To understand the potential ramifications of this gap, we investigate how well current benchmarks quantify model reliability. We find that pervasive label errors can compromise these evaluations, obscuring lingering model failures and hiding unreliable behavior. Motivated by this gap in the evaluation of reliability, we then propose the concept of so-called platinum benchmarks, i.e., benchmarks carefully curated to minimize label errors and ambiguity. As a first attempt at constructing such benchmarks, we revise examples from fifteen existing popular benchmarks. We evaluate a wide range of models on these platinum benchmarks and find that, indeed, frontier LLMs still exhibit failures on simple tasks such as elementary-level math word problems. Analyzing these failures further reveals previously unidentified patterns of problems on which frontier models consistently struggle. We provide code at https://github.com/MadryLab/platinum-benchmarks

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fluid Language Model Benchmarking

    cs.CL 2025-09 conditional novelty 8.0 of 10

    Fluid Benchmarking, combining IRT-based ability estimation with Fisher-information-based adaptive item selection, improves LM evaluation across efficiency, validity, variance, and saturation in pretraining settings.

  2. Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility

    cs.CL 2026-05 conditional novelty 7.0 of 10

    SGRE extracts a reasoning skeleton from a teacher trace, coarsens its graph, and verbalizes it densely; the final answer is preserved verbatim, and students distilled on the edited traces show large accuracy drops.

  3. Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems

    cs.CL 2025-08 conditional novelty 7.0 of 10

    A fully automated pipeline produces proof-centric math benchmarks, demonstrated on algebraic geometry with 456 items, where leading LLMs score near 60 percent.

  4. CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting

    cs.LG 2025-05 accept novelty 7.0 of 10

    Capping achievable accuracy with randomized correct answers turns any model that exceeds the cap into a detectable contamination alarm.

  5. Position: Evaluation of ECG Representations Must Be Fixed

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Current ECG representation benchmarks overstate the benefits of pretraining and produce unstable method rankings; a random encoder with linear probing is competitive on many tasks.

  6. Weight Decay Improves Language Model Plasticity

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Pretrained models trained with larger weight decay fine-tune better on downstream tasks, so the best pretraining checkpoint by loss is not always the best starting point for later training.

  7. Model soups need only one ingredient

    cs.LG 2026-02 conditional novelty 6.0 of 10

    A single checkpoint, edited by splitting each layer's update with SVD and reweighting the high- and low-energy parts, reaches soup-level OOD robustness without multi-model training.

  8. Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

    cs.CL 2025-10 conditional novelty 6.0 of 10

    An IRT-based adaptive testing framework, ATLAS, estimates LLM ability with 30-89 items per benchmark, matching whole-bank ability estimates and re-ranking 23-31% of models relative to accuracy.

  9. From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    KMMLU-Redux and KMMLU-Pro are new Korean benchmark datasets from national technical and professional licensure exams, with LLM evaluations reported against official pass thresholds.

  10. Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A comprehensive reanalysis finds that min-p sampling does not outperform top-p, top-k, or basic sampling once the original data are re-tested and hyperparameter budgets are equalized.

  11. A Sovereign, Open-Source Foundation Model for German and English

    cs.CL 2026-07 conditional novelty 5.5 of 10

    Soofi S 30B-A3B, a hybrid Mamba-MoE model pretrained on ~27T tokens with deliberately up-weighted German, reports the highest English and German aggregate scores among fully open base models in its comparison while ma...

  12. Simple Policy Gradients for Reasoning with Diffusion Language Models

    cs.LG 2025-10 reject novelty 5.0 of 10

    AGRPO makes GRPO-style policy gradients tractable for diffusion LLMs by Monte-Carlo sampling denoising timesteps, but the unbiasedness claim only holds for a step-level objective, not the token-level GRPO objective.

  13. Domain Specific Benchmarks for Evaluating Multimodal Large Language Models

    cs.LG 2025-06 conditional novelty 3.0 of 10

    A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.

Pith tools