REVIEW 2 major objections 5 minor 14 references
Quantifying Ranking Uncertainty in LLM Benchmarks
T0 review · 2 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read LLM leaderboard rankings on MMLU are substantially less certain once between-subject variability is accounted for, and rank confidence intervals should accompany any top-model claim.
desk verdict A useful application of existing rank-CI machinery to MMLU; the qualitative message about subject-level variation is solid, but the quantitative widths rest on an untested exchangeability assumption that needs a prominent caveat. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Rank confidence intervals (CIs) constructed from directional pairwise hypothesis tests. For each pair of models, two one-sided tests are run; after FWER control (Holm's method), the counts of models significantly better and significantly worse give upper and lower bounds on a model's rank. The paper's key move is to vary the definition of the unit — subject, prompt, or observation — which changes the paired differences, the standard deviation, and the degrees of freedom, and thereby changes the interval widths. This turns the abstract statistical choice of 'population' into an observed difference in conclusions.
What would settle it
Take a different set of knowledge domains (e.g., 57 subjects randomly drawn from a broad ontology of disciplines), rerun the paired t-test rank-CI analysis, and check whether the subject-level rank intervals remain as wide as those reported on MMLU. If the intervals narrow dramatically, the reported subject variability is an artifact of MMLU's particular subject list rather than a general feature of subject populations.
Extended reading notes
Core claim
The paper's central claim is that ranking variability across MMLU subjects is substantial and must be considered when comparing LLMs or identifying top-performing models. It demonstrates this by constructing rank confidence intervals from directional pairwise hypothesis tests under three different definitions of the inferential unit: individual question-prompt observations, subjects, and prompt variants. When subjects are treated as a random sample from a population of subjects, the rank intervals become much wider than when observations or prompts are the units, with several models overlapping in the middle ranks. The paper also shows that within a single subject a model can be clearly best
Load-bearing premise
The entire subject-level analysis depends on treating MMLU's 57 subjects as a random sample from a population of subjects; in reality they are a hand-picked set of knowledge domains, so the interval widths and the conclusion that subject variability is substantial may not generalize beyond this specific benchmark.
Editorial extensions
If this is right
- Leaderboard point ranks should be accompanied by rank intervals; without them, middle-ranked models on MMLU are not reliably ordered.
- A model should not be declared 'top' unless its rank interval excludes all competitors; the subject-level analysis shows several models can claim the same rank.
- Subject-level variability exceeds prompt-level variability, so stabilizing rankings should focus on how subjects are selected rather than adding more prompt variants.
- Filtering out non-informative subjects (where models are nearly tied) concentrates comparisons on subjects that actually separate models and changes benchmark-level conclusions.
- An indifference margin of 2% accuracy makes rank intervals overlap even for units that previously separated clearly, showing many statistically significant differences are not practically meaningful.
Reading between the lines
- The subject-as-random-sample assumption is doing heavy lifting: MMLU's 57 subjects are hand-picked, so the reported interval widths are only valid for a hypothetical population of similar subjects; a genuinely random sample of knowledge domains could produce different widths.
- A natural extension is to treat whole benchmarks (rather than subjects) as the random sample when combining multiple leaderboards; the authors flag this as future work, and the same pairwise-test machinery would apply.
- A concrete testable prediction: splitting MMLU's subjects into two random halves and re-running the subject-level rank-CI analysis on each half should reproduce the wide intervals, confirming that subject heterogeneity is a real feature and not a quirk of the full subject list.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper applies rank confidence intervals (CIs), imported from Holm (2013) and Neuhof & Benjamini (2024), to the MMLU/PromptEval benchmark. It constructs pairwise hypothesis tests under three different unit definitions — individual question–prompt observations, subjects, and prompt variants — and compares the resulting rank CIs. The authors report that subject-level rank CIs are wide, prompt-level rank CIs are narrow, and per-subject analyses reveal strong domain-dependent model differences. They also demonstrate two modifications: an indifference margin and filtering of non-informative subjects. The paper concludes that ranking variability across MMLU subjects is substantial and should be reported alongside leaderboard ranks.
Significance. The paper addresses a timely problem: leaderboard ranks are typically reported as point estimates without uncertainty. The existing rank-CI machinery is sound and the paper correctly credits prior work; the main contribution is an empirical decomposition of uncertainty sources on a widely used benchmark. The public code and the use of the PromptEval correctness matrix are concrete strengths. If the main claim holds, the paper would justify a simple, useful practice: attach subject-level rank intervals to top-model claims. However, the quantitative conclusion that subject variability dominates prompt variability is not rigorously established, and the subject-level intervals rely on an exchangeability assumption that is stated but not tested. These issues currently limit the strength of the empirical claim rather than invalidating the methodology.
major comments (2)
- [§3.1, Sensitivity to Prompt Variants] The paper claims that subject-level variability dominates prompt-level variability, but this rests on comparing Figure 2 (n=57 subjects) with Figure 3 (n=100 prompt variants). Different sample sizes change both the degrees of freedom and the power of the pairwise tests, so the wider intervals in Figure 2 may reflect sample size rather than subject heterogeneity. The sentence that random subsampling of prompt variants gave 'consistent results' is not accompanied by any reported figure, table, or quantitative summary. Please report the subsampled prompt-level rank CIs at n=57, or another matched-sample comparison; without this, the dominance claim is not supported by the evidence shown.
- [§3.1, Eq. (7)–(8)] The subject-level rank CIs in Figure 2 are interpreted as inference to a population of subjects, and Eq. (7)–(8) define the estimand as an expectation over a subject population. This requires the 57 MMLU subjects to be exchangeable draws from that population. MMLU subjects are hand-picked knowledge domains, and Figure 4 itself shows that model differences are domain-dependent (clinical knowledge vs. virology). If subject-level differences are clustered by field (STEM, humanities, social sciences), the i.i.d.-subject model is misspecified, the effective sample size is smaller than 57, and the reported interval widths are miscalibrated for any real 'new subject' population. The qualitative finding that rankings vary across the observed subjects is robust from the raw per-subject accuracies, but the quantitative widths and the prescription that subject-level intervals should accompany top-m
minor comments (5)
- [§3.1] Typos and unclear wording: 'prompt variants are considered fixed and averaged are averaged within each subject' and 'the variability between questions exceed that between prompts' should be corrected. Also, the sentence beginning 'To determine whether the number of units...' is vague about what 'consistent results' means; a reference to an appendix figure would help.
- [Appendix A.2, Algorithms 1 and 2] The comments in the algorithms refer to 'c j' (e.g., 'Count models significantly worse than c j'), but c j is not defined. Presumably this is model m_j. Please fix the notation.
- [§4 and Appendix B.4] The non-informative-subject filtering rule ('14 or 15 out of 15 models rank in the top 3') is ad hoc and the alpha-split (0.3α / 0.7α) is presented without justification. Since the filtering changes the benchmark-level rank CIs, a sensitivity analysis with respect to the threshold and split proportion would strengthen the claim.
- [Appendix B.3] The indifference margin δ=0.02 is arbitrary. The text says the margin can be chosen empirically with sample splitting, but no demonstration of that procedure is given. Consider reporting results for a range of δ values.
- [Various] The Wilcoxon robustness check in Appendix B.1 is reported only for the benchmark-level analyses, not for the subject-level analysis in Figure 2, which is the main source of the paper's headline claim. A Wilcoxon version of Figure 2 would make the normality assumption concern less salient.
Circularity Check
No significant circularity: rank CI construction is explicit and data-driven; the empirical variability conclusions are observations, not artifacts.
full rationale
The paper's central derivation constructs rank CIs by counting rejected one-sided pairwise tests (Eqs. 3-4), with p-values from paired t-statistics on observed MMLU differences at three unit levels (Eqs. 5-10). Every step is explicit and data-driven: no fitted parameter is renamed as a prediction, and no estimand is defined in terms of the rank intervals it produces. The claim that ranking variability across MMLU subjects is substantial is read directly from the data (Figures 2 and 4), not manufactured by the interval-construction procedure. The self-citations to Neuhof & Benjamini (2024) for simultaneous rank CIs (Theorem A.4, Algorithm 2) cite a published, parameter-free theorem with stated assumptions and are not load-bearing for the main empirical claim, which uses Holm (2013) marginal rank CIs; thus they do not constitute circularity under the stated rules. The exchangeability assumption that MMLU subjects are a sample from a subject population (Section 3.1, Eq. 8) is a substantive modeling assumption and a generalizability limitation, but it is not circular: the reported intervals are not forced to equal the inputs by construction, and the observed subject-level variability is an empirical finding independent of the exchangeability assumption needed for extrapolation.
Assumptions & free parameters
free parameters (4)
- Indifference margin δ =
0.02
- Non-informative subject filtering threshold =
14 or 15 out of 15 models ranking in top 3
- Alpha split proportion for filtering =
0.3α / 0.7α
- Number of sampled questions per subject =
100
assumptions (4)
- domain assumption MMLU subjects are a random sample from a population of subjects.
- domain assumption Paired differences are approximately normally distributed (via CLT).
- standard math The rank CI construction theorems of Holm (2013) and Neuhof & Benjamini (2024) are valid.
- domain assumption Independence across subjects and across questions/prompts within subjects.
Cite this review
Pith. "Pith review of Quantifying Ranking Uncertainty in LLM Benchmarks." pith.science (2026). https://pith.science/paper/ZICVXITX
@misc{pith2026260716259,
author = {Pith},
title = {Pith review of: Quantifying Ranking Uncertainty in LLM Benchmarks},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZICVXITX}},
note = {Machine review of arXiv:2607.16259}
}
read the original abstract
Pretrained models are typically ranked on multi-task leaderboards to assess their effectiveness across diverse tasks. Rank confidence intervals were recently introduced as a method to quantify the uncertainty in these rankings by aggregating pairwise hypothesis tests. In this work, we analyze the sources of uncertainty in the knowledge evaluation benchmark MMLU and show how hypothesis tests can be modified to account for their effects. We demonstrate that ranking variability across MMLU subjects is substantial and should be considered when comparing LLMs or identifying the top-performing models.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Ackerman, S., Farchi, E., Raz, O., and Toledo, A. Sta- tistical multi-metric evaluation and visualization of llm system predictive performance.arXiv preprint arXiv:2501.18243,
-
[4]
The simultaneous rank CIs in Figure 6 are wider than the marginal rank CIs, as they account for multiple comparisons by considering all NM(NM −1) hypotheses as a single family (see (Al Mohamad et al., 2021; Neuhof & Benjamini, 2024)). This adjustment ensures a coverage guarantee: with probability at least 1−α , the true ranks of all selected models are co...
2021
-
[8]
Miller, E. Adding error bars to evals: A statistical ap- proach to language model evaluations.arXiv preprint arXiv:2411.00640,
-
[10]
Uncertainty in ranking.arXiv preprint arXiv:2107.03459,
Rising, J. Uncertainty in ranking.arXiv preprint arXiv:2107.03459,
-
[11]
and Soares, C
6 Quantifying Ranking Uncertainty in LLM Benchmarks Valdeira, F. and Soares, C. Ranking with confidence for large scale comparison data. InProceedings of the 2025 SIAM International Conference on Data Mining (SDM), pp. 223–232. SIAM,
2025
-
[13]
However, the distribution may not be approximately normal, or even symmetric
By the Central Limit Theorem, this average is approximately normally distributed. However, the distribution may not be approximately normal, or even symmetric. In these cases, use a more robust test or a nonparametric test, such as the Wilcoxon signed-rank test (Wilcoxon, 1945), instead of the t-test. In Figure 5, we show benchmark-level rank CIs using th...
1945
-
[1945]
Rank CIs definitions and methods A.1
7 Quantifying Ranking Uncertainty in LLM Benchmarks A. Rank CIs definitions and methods A.1. Marginal and Simultaneous Coverage Our goal is to estimate the true ranks for allNM models. Let X be an n×NM observed aggregation of scores so thatX∼P µ. Here Pµ is a distribution family indexed by the vector of meansµ∈R N M. Let ([L1(X), U1(X)], . . . ,[LNM (X), ...
2005
-
[2005]
Chandra, O. and Xie, M.-g. Finite-sample valid rank con- fidence sets for a broad class of statistical and machine learning models.arXiv preprint arXiv:2512.00316,
Show all 14 references
-
[2006]
Ranking the scores of algorithms with confidence
Foucart, A., Elskens, A., and Decaestecker, C. Ranking the scores of algorithms with confidence. InESANN 2025 proceedings, pp. 431–436,
2025
-
[2015]
v067.i01
doi: 10.18637/jss. v067.i01. URL https://www.jstatsoft.org/ index.php/jss/article/view/v067i01. Benjamini, Y . and Yekutieli, D. False discovery rate– adjusted multiple confidence intervals for selected pa- rameters.Journal of the American Statistical Association, 100(469):71–81,
-
[2021]
Al Mohamad, D., Goeman, J
doi: 10.1214/21-EJS1847. Al Mohamad, D., Goeman, J. J., and van Zwet, E. W. Simul- taneous confidence intervals for ranks with application to ranking institutions.Biometrics, 78(1):238–247,
-
[2022]
Statistical uncertainty quantification for aggregate performance met- rics in machine learning benchmarks.arXiv preprint arXiv:2501.04234,
Longjohn, R., Gopalan, G., and Casleton, E. Statistical uncertainty quantification for aggregate performance met- rics in machine learning benchmarks.arXiv preprint arXiv:2501.04234,
-
[2024]
doi: 10.1093/restud/rdad006
ISSN 0034-6527. doi: 10.1093/restud/rdad006. Neuhof, B. and Benjamini, Y . Confident feature ranking. In International Conference on Artificial Intelligence and Statistics, pp. 1468–1476. PMLR,
-
[2025]
csranks: an r package for estimation and inference involving ranks.arXiv preprint arXiv:2401.15205,
Chetverikov, D., Mogstad, M., Morgen, P., Romano, J., Shaikh, A., and Wilhelm, D. csranks: an r package for estimation and inference involving ranks.arXiv preprint arXiv:2401.15205,
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.