REVIEW 13 cited by
Do Large Language Model Benchmarks Test Reliability?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
When deploying large language models (LLMs), it is important to ensure that these models are not only capable, but also reliable. Many benchmarks have been created to track LLMs' growing capabilities, however there has been no similar focus on measuring their reliability. To understand the potential ramifications of this gap, we investigate how well current benchmarks quantify model reliability. We find that pervasive label errors can compromise these evaluations, obscuring lingering model failures and hiding unreliable behavior. Motivated by this gap in the evaluation of reliability, we then propose the concept of so-called platinum benchmarks, i.e., benchmarks carefully curated to minimize label errors and ambiguity. As a first attempt at constructing such benchmarks, we revise examples from fifteen existing popular benchmarks. We evaluate a wide range of models on these platinum benchmarks and find that, indeed, frontier LLMs still exhibit failures on simple tasks such as elementary-level math word problems. Analyzing these failures further reveals previously unidentified patterns of problems on which frontier models consistently struggle. We provide code at https://github.com/MadryLab/platinum-benchmarks
Forward citations
Cited by 13 Pith papers
-
Fluid Language Model Benchmarking
Fluid Benchmarking, combining IRT-based ability estimation with Fisher-information-based adaptive item selection, improves LM evaluation across efficiency, validity, variance, and saturation in pretraining settings.
-
Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility
SGRE extracts a reasoning skeleton from a teacher trace, coarsens its graph, and verbalizes it densely; the final answer is preserved verbatim, and students distilled on the edited traces show large accuracy drops.
-
Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems
A fully automated pipeline produces proof-centric math benchmarks, demonstrated on algebraic geometry with 456 items, where leading LLMs score near 60 percent.
-
CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting
Capping achievable accuracy with randomized correct answers turns any model that exceeds the cap into a detectable contamination alarm.
-
Position: Evaluation of ECG Representations Must Be Fixed
Current ECG representation benchmarks overstate the benefits of pretraining and produce unstable method rankings; a random encoder with linear probing is competitive on many tasks.
-
Weight Decay Improves Language Model Plasticity
Pretrained models trained with larger weight decay fine-tune better on downstream tasks, so the best pretraining checkpoint by loss is not always the best starting point for later training.
-
Model soups need only one ingredient
A single checkpoint, edited by splitting each layer's update with SVD and reweighting the high- and low-energy parts, reaches soup-level OOD robustness without multi-model training.
-
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
An IRT-based adaptive testing framework, ATLAS, estimates LLM ability with 30-89 items per benchmark, matching whole-bank ability estimates and re-ranking 23-31% of models relative to accuracy.
-
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
KMMLU-Redux and KMMLU-Pro are new Korean benchmark datasets from national technical and professional licensure exams, with LLM evaluations reported against official pass thresholds.
-
Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models
A comprehensive reanalysis finds that min-p sampling does not outperform top-p, top-k, or basic sampling once the original data are re-tested and hyperparameter budgets are equalized.
-
A Sovereign, Open-Source Foundation Model for German and English
Soofi S 30B-A3B, a hybrid Mamba-MoE model pretrained on ~27T tokens with deliberately up-weighted German, reports the highest English and German aggregate scores among fully open base models in its comparison while ma...
-
Simple Policy Gradients for Reasoning with Diffusion Language Models
AGRPO makes GRPO-style policy gradients tractable for diffusion LLMs by Monte-Carlo sampling denoising timesteps, but the unbiasedness claim only holds for a step-level objective, not the token-level GRPO objective.
-
Domain Specific Benchmarks for Evaluating Multimodal Large Language Models
A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.
Discussion (0). Continue with ORCID to comment.