Pith. sign in

REVIEW 2 major objections 5 minor 14 references

Quantifying Ranking Uncertainty in LLM Benchmarks

T0 review · 2 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read LLM leaderboard rankings on MMLU are substantially less certain once between-subject variability is accounted for, and rank confidence intervals should accompany any top-model claim.

desk verdict A useful application of existing rank-CI machinery to MMLU; the qualitative message about subject-level variation is solid, but the quantitative widths rest on an untested exchangeability assumption that needs a prominent caveat. read the letter →

arxiv 2607.16259 v1 pith:ZICVXITX submitted 2026-06-28 cs.LG cs.CLstat.ML

classification cs.LGcs.CLstat.ML MSC 62F0362F0762F25
keywords rankconfidenceintervalsMMLUbenchmarkLLMevaluationhypothesistestingsubjectheterogeneityleaderboarduncertaintypairedt-testmodelranking
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Leaderboards rank large language models by a single average accuracy, but that point rank hides a substantial source of uncertainty: the choice of subjects that make up the benchmark. This paper builds rank confidence intervals from pairwise hypothesis tests and shows that when MMLU's subjects — rather than individual question-prompt observations — are treated as the unit of analysis, the intervals widen dramatically. Several middle-ranked models become statistically indistinguishable, and a model that looks clearly best on one subject can be indistinguishable from others on another. The paper concludes that subject-level variability dominates and should be a standard part of any top-model claim.

What carries the argument

Rank confidence intervals (CIs) constructed from directional pairwise hypothesis tests. For each pair of models, two one-sided tests are run; after FWER control (Holm's method), the counts of models significantly better and significantly worse give upper and lower bounds on a model's rank. The paper's key move is to vary the definition of the unit — subject, prompt, or observation — which changes the paired differences, the standard deviation, and the degrees of freedom, and thereby changes the interval widths. This turns the abstract statistical choice of 'population' into an observed difference in conclusions.

What would settle it

Take a different set of knowledge domains (e.g., 57 subjects randomly drawn from a broad ontology of disciplines), rerun the paired t-test rank-CI analysis, and check whether the subject-level rank intervals remain as wide as those reported on MMLU. If the intervals narrow dramatically, the reported subject variability is an artifact of MMLU's particular subject list rather than a general feature of subject populations.

Watch

Extended reading notes

Core claim

The paper's central claim is that ranking variability across MMLU subjects is substantial and must be considered when comparing LLMs or identifying top-performing models. It demonstrates this by constructing rank confidence intervals from directional pairwise hypothesis tests under three different definitions of the inferential unit: individual question-prompt observations, subjects, and prompt variants. When subjects are treated as a random sample from a population of subjects, the rank intervals become much wider than when observations or prompts are the units, with several models overlapping in the middle ranks. The paper also shows that within a single subject a model can be clearly best

Load-bearing premise

The entire subject-level analysis depends on treating MMLU's 57 subjects as a random sample from a population of subjects; in reality they are a hand-picked set of knowledge domains, so the interval widths and the conclusion that subject variability is substantial may not generalize beyond this specific benchmark.

Editorial extensions

If this is right

  • Leaderboard point ranks should be accompanied by rank intervals; without them, middle-ranked models on MMLU are not reliably ordered.
  • A model should not be declared 'top' unless its rank interval excludes all competitors; the subject-level analysis shows several models can claim the same rank.
  • Subject-level variability exceeds prompt-level variability, so stabilizing rankings should focus on how subjects are selected rather than adding more prompt variants.
  • Filtering out non-informative subjects (where models are nearly tied) concentrates comparisons on subjects that actually separate models and changes benchmark-level conclusions.
  • An indifference margin of 2% accuracy makes rank intervals overlap even for units that previously separated clearly, showing many statistically significant differences are not practically meaningful.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The subject-as-random-sample assumption is doing heavy lifting: MMLU's 57 subjects are hand-picked, so the reported interval widths are only valid for a hypothetical population of similar subjects; a genuinely random sample of knowledge domains could produce different widths.
  • A natural extension is to treat whole benchmarks (rather than subjects) as the random sample when combining multiple leaderboards; the authors flag this as future work, and the same pairwise-test machinery would apply.
  • A concrete testable prediction: splitting MMLU's subjects into two random halves and re-running the subject-level rank-CI analysis on each half should reproduce the wide intervals, confirming that subject heterogeneity is a real feature and not a quirk of the full subject list.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper applies rank confidence intervals (CIs), imported from Holm (2013) and Neuhof & Benjamini (2024), to the MMLU/PromptEval benchmark. It constructs pairwise hypothesis tests under three different unit definitions — individual question–prompt observations, subjects, and prompt variants — and compares the resulting rank CIs. The authors report that subject-level rank CIs are wide, prompt-level rank CIs are narrow, and per-subject analyses reveal strong domain-dependent model differences. They also demonstrate two modifications: an indifference margin and filtering of non-informative subjects. The paper concludes that ranking variability across MMLU subjects is substantial and should be reported alongside leaderboard ranks.

Significance. The paper addresses a timely problem: leaderboard ranks are typically reported as point estimates without uncertainty. The existing rank-CI machinery is sound and the paper correctly credits prior work; the main contribution is an empirical decomposition of uncertainty sources on a widely used benchmark. The public code and the use of the PromptEval correctness matrix are concrete strengths. If the main claim holds, the paper would justify a simple, useful practice: attach subject-level rank intervals to top-model claims. However, the quantitative conclusion that subject variability dominates prompt variability is not rigorously established, and the subject-level intervals rely on an exchangeability assumption that is stated but not tested. These issues currently limit the strength of the empirical claim rather than invalidating the methodology.

major comments (2)
  1. [§3.1, Sensitivity to Prompt Variants] The paper claims that subject-level variability dominates prompt-level variability, but this rests on comparing Figure 2 (n=57 subjects) with Figure 3 (n=100 prompt variants). Different sample sizes change both the degrees of freedom and the power of the pairwise tests, so the wider intervals in Figure 2 may reflect sample size rather than subject heterogeneity. The sentence that random subsampling of prompt variants gave 'consistent results' is not accompanied by any reported figure, table, or quantitative summary. Please report the subsampled prompt-level rank CIs at n=57, or another matched-sample comparison; without this, the dominance claim is not supported by the evidence shown.
  2. [§3.1, Eq. (7)–(8)] The subject-level rank CIs in Figure 2 are interpreted as inference to a population of subjects, and Eq. (7)–(8) define the estimand as an expectation over a subject population. This requires the 57 MMLU subjects to be exchangeable draws from that population. MMLU subjects are hand-picked knowledge domains, and Figure 4 itself shows that model differences are domain-dependent (clinical knowledge vs. virology). If subject-level differences are clustered by field (STEM, humanities, social sciences), the i.i.d.-subject model is misspecified, the effective sample size is smaller than 57, and the reported interval widths are miscalibrated for any real 'new subject' population. The qualitative finding that rankings vary across the observed subjects is robust from the raw per-subject accuracies, but the quantitative widths and the prescription that subject-level intervals should accompany top-m
minor comments (5)
  1. [§3.1] Typos and unclear wording: 'prompt variants are considered fixed and averaged are averaged within each subject' and 'the variability between questions exceed that between prompts' should be corrected. Also, the sentence beginning 'To determine whether the number of units...' is vague about what 'consistent results' means; a reference to an appendix figure would help.
  2. [Appendix A.2, Algorithms 1 and 2] The comments in the algorithms refer to 'c j' (e.g., 'Count models significantly worse than c j'), but c j is not defined. Presumably this is model m_j. Please fix the notation.
  3. [§4 and Appendix B.4] The non-informative-subject filtering rule ('14 or 15 out of 15 models rank in the top 3') is ad hoc and the alpha-split (0.3α / 0.7α) is presented without justification. Since the filtering changes the benchmark-level rank CIs, a sensitivity analysis with respect to the threshold and split proportion would strengthen the claim.
  4. [Appendix B.3] The indifference margin δ=0.02 is arbitrary. The text says the margin can be chosen empirically with sample splitting, but no demonstration of that procedure is given. Consider reporting results for a range of δ values.
  5. [Various] The Wilcoxon robustness check in Appendix B.1 is reported only for the benchmark-level analyses, not for the subject-level analysis in Figure 2, which is the main source of the paper's headline claim. A Wilcoxon version of Figure 2 would make the normality assumption concern less salient.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: rank CI construction is explicit and data-driven; the empirical variability conclusions are observations, not artifacts.

full rationale

The paper's central derivation constructs rank CIs by counting rejected one-sided pairwise tests (Eqs. 3-4), with p-values from paired t-statistics on observed MMLU differences at three unit levels (Eqs. 5-10). Every step is explicit and data-driven: no fitted parameter is renamed as a prediction, and no estimand is defined in terms of the rank intervals it produces. The claim that ranking variability across MMLU subjects is substantial is read directly from the data (Figures 2 and 4), not manufactured by the interval-construction procedure. The self-citations to Neuhof & Benjamini (2024) for simultaneous rank CIs (Theorem A.4, Algorithm 2) cite a published, parameter-free theorem with stated assumptions and are not load-bearing for the main empirical claim, which uses Holm (2013) marginal rank CIs; thus they do not constitute circularity under the stated rules. The exchangeability assumption that MMLU subjects are a sample from a subject population (Section 3.1, Eq. 8) is a substantive modeling assumption and a generalizability limitation, but it is not circular: the reported intervals are not forced to equal the inputs by construction, and the observed subject-level variability is an empirical finding independent of the exchangeability assumption needed for extrapolation.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The paper is an application of an existing rank-CI framework; it introduces no new entities. The central claim rests mainly on the subject-population exchangeability assumption, plus standard t-test assumptions and the cited validity theorems. Free parameters appear mostly in the appendix demonstrations and do not materially affect the main subject-vs-prompt comparison.

free parameters (4)
  • Indifference margin δ = 0.02
    Used in Appendix B.3 as an example of practical equivalence; chosen ad hoc for demonstration, not motivated by domain data or a formal power analysis.
  • Non-informative subject filtering threshold = 14 or 15 out of 15 models ranking in top 3
    Defined in Appendix B.4; a data-derived criterion used to remove 27 of 57 subjects before recomputing benchmark intervals. The threshold is arbitrary.
  • Alpha split proportion for filtering = 0.3α / 0.7α
    Chosen in Appendix B.4 to split miscoverage between the filtering step and the final benchmark CIs; the 0.3/0.7 split is arbitrary and not derived from any optimality principle.
  • Number of sampled questions per subject = 100
    The authors randomly sample 100 questions per subject to balance the design (§3). This is a hand-chosen design parameter; the original dataset contains at least 100 questions per subject, and this choice affects variance estimates.
assumptions (4)
  • domain assumption MMLU subjects are a random sample from a population of subjects.
    Invoked in §3.1 Subject Heterogeneity: 'A better approach is to treat the subjects as a sample from a subject population.' The subject-level t-test and resulting rank CIs extrapolate to 'new subjects' only under this exchangeability assumption; if false, coverage claims for the subject-units analysis do not hold.
  • domain assumption Paired differences are approximately normally distributed (via CLT).
    Used to justify paired t-test p-values in §2.2 and §3. The authors note in Appendix B.1 that normality may fail and suggest the Wilcoxon signed-rank test as a robust alternative; the central results are shown to be mostly insensitive to this choice.
  • standard math The rank CI construction theorems of Holm (2013) and Neuhof & Benjamini (2024) are valid.
    The paper invokes Theorem A.3 and Theorem A.4 without proof, citing prior work. Treating these theorems as background is standard practice.
  • domain assumption Independence across subjects and across questions/prompts within subjects.
    The paired t-test treats the unit-level differences as i.i.d. The paper does not test for dependence between subjects or within-subject correlations in the PromptEval data; correlated units could make the effective sample size smaller than the nominal n.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Quantifying Ranking Uncertainty in LLM Benchmarks." pith.science (2026). https://pith.science/paper/ZICVXITX

@misc{pith2026260716259,
  author       = {Pith},
  title        = {Pith review of: Quantifying Ranking Uncertainty in LLM Benchmarks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZICVXITX}},
  note         = {Machine review of arXiv:2607.16259}
}
read the original abstract

Pretrained models are typically ranked on multi-task leaderboards to assess their effectiveness across diverse tasks. Rank confidence intervals were recently introduced as a method to quantify the uncertainty in these rankings by aggregating pairwise hypothesis tests. In this work, we analyze the sources of uncertainty in the knowledge evaluation benchmark MMLU and show how hypothesis tests can be modified to account for their effects. We demonstrate that ranking variability across MMLU subjects is substantial and should be considered when comparing LLMs or identifying the top-performing models.

Figures

Figures reproduced from arXiv: 2607.16259 by the authors.

Figure 1
Figure 1. Scores with 95% CIs and rank CIs across all subjects, prompt variants, and questions. The models are clearly distinguish￾able, as indicated by non-overlapping CIs. However, this test only supports inference for the specific evaluation set, not for new or unseen data. Treating all d[jk,b,bc,l] as independent observations would generally un￾derestimate uncertainty, because it ignores dependencies within subjects, ques… view at source ↗
Figure 3
Figure 3. reveals a pattern similar to [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Subject-specific rank CIs with aggregation across ques￾tions (a) and prompts (b); the variability between questions exceed that between prompts. A model can be ranked clearly as the best for one subject (left), but not for another (right). one model is statistically superior to, or equivalent with, another. These two analyses use the same observed scores but ad￾dress different inferential questions. The first analys… view at source ↗
Figures from the paper (4 more)
Figure 5
Figure 5. Figure 5: Rank CIs with Wilcoxon as the paired test and different definitions of units; observations (a), subjects (b), and prompt variants (c). B.2. Simultaneous Rank CIs Here, we present the results of constructing simultaneous rank CIs for the same data as in [PITH_FULL_IMAG…
Figure 6
Figure 6. Figure 6: Subject-specific rank CIs with aggregation across questions (a) and prompts (b), with simultaneous coverage. B.3. Indifference Zone For this demonstration, we introduce an indifference margin δ = 0.02, representing a 2% difference in accuracy between models. Including …
Figure 7
Figure 7. Figure 7: Rank CIs with indifference zone of 2% between the accuracy of two models, and different definitions of units; observations (a), subjects (b), and prompt variants (c). B.4. Filter Non-informative Subjects In [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: Rank CIs based on a subset of 30 subjects, after filtering out subjects with almost no difference between models. 11 [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

14 extracted references · 1 canonical work pages

  1. [1]

    Sta- tistical multi-metric evaluation and visualization of llm system predictive performance.arXiv preprint arXiv:2501.18243,

    Ackerman, S., Farchi, E., Raz, O., and Toledo, A. Sta- tistical multi-metric evaluation and visualization of llm system predictive performance.arXiv preprint arXiv:2501.18243,

  2. [4]

    The simultaneous rank CIs in Figure 6 are wider than the marginal rank CIs, as they account for multiple comparisons by considering all NM(NM −1) hypotheses as a single family (see (Al Mohamad et al., 2021; Neuhof & Benjamini, 2024)). This adjustment ensures a coverage guarantee: with probability at least 1−α , the true ranks of all selected models are co...

  3. [8]

    Adding error bars to evals: A statistical ap- proach to language model evaluations.arXiv preprint arXiv:2411.00640,

    Miller, E. Adding error bars to evals: A statistical ap- proach to language model evaluations.arXiv preprint arXiv:2411.00640,

  4. [10]

    Uncertainty in ranking.arXiv preprint arXiv:2107.03459,

    Rising, J. Uncertainty in ranking.arXiv preprint arXiv:2107.03459,

  5. [11]

    and Soares, C

    6 Quantifying Ranking Uncertainty in LLM Benchmarks Valdeira, F. and Soares, C. Ranking with confidence for large scale comparison data. InProceedings of the 2025 SIAM International Conference on Data Mining (SDM), pp. 223–232. SIAM,

  6. [13]

    However, the distribution may not be approximately normal, or even symmetric

    By the Central Limit Theorem, this average is approximately normally distributed. However, the distribution may not be approximately normal, or even symmetric. In these cases, use a more robust test or a nonparametric test, such as the Wilcoxon signed-rank test (Wilcoxon, 1945), instead of the t-test. In Figure 5, we show benchmark-level rank CIs using th...

  7. [1945]

    Rank CIs definitions and methods A.1

    7 Quantifying Ranking Uncertainty in LLM Benchmarks A. Rank CIs definitions and methods A.1. Marginal and Simultaneous Coverage Our goal is to estimate the true ranks for allNM models. Let X be an n×NM observed aggregation of scores so thatX∼P µ. Here Pµ is a distribution family indexed by the vector of meansµ∈R N M. Let ([L1(X), U1(X)], . . . ,[LNM (X), ...

  8. [2005]

    and Xie, M.-g

    Chandra, O. and Xie, M.-g. Finite-sample valid rank con- fidence sets for a broad class of statistical and machine learning models.arXiv preprint arXiv:2512.00316,

Show all 14 references
  1. [2006]

    Ranking the scores of algorithms with confidence

    Foucart, A., Elskens, A., and Decaestecker, C. Ranking the scores of algorithms with confidence. InESANN 2025 proceedings, pp. 431–436,

  2. [2015]

    v067.i01

    doi: 10.18637/jss. v067.i01. URL https://www.jstatsoft.org/ index.php/jss/article/view/v067i01. Benjamini, Y . and Yekutieli, D. False discovery rate– adjusted multiple confidence intervals for selected pa- rameters.Journal of the American Statistical Association, 100(469):71–81,

  3. [2021]

    Al Mohamad, D., Goeman, J

    doi: 10.1214/21-EJS1847. Al Mohamad, D., Goeman, J. J., and van Zwet, E. W. Simul- taneous confidence intervals for ranks with application to ranking institutions.Biometrics, 78(1):238–247,

  4. [2022]

    Statistical uncertainty quantification for aggregate performance met- rics in machine learning benchmarks.arXiv preprint arXiv:2501.04234,

    Longjohn, R., Gopalan, G., and Casleton, E. Statistical uncertainty quantification for aggregate performance met- rics in machine learning benchmarks.arXiv preprint arXiv:2501.04234,

  5. [2024]

    doi: 10.1093/restud/rdad006

    ISSN 0034-6527. doi: 10.1093/restud/rdad006. Neuhof, B. and Benjamini, Y . Confident feature ranking. In International Conference on Artificial Intelligence and Statistics, pp. 1468–1476. PMLR,

  6. [2025]

    csranks: an r package for estimation and inference involving ranks.arXiv preprint arXiv:2401.15205,

    Chetverikov, D., Mogstad, M., Morgen, P., Romano, J., Shaikh, A., and Wilhelm, D. csranks: an r package for estimation and inference involving ranks.arXiv preprint arXiv:2401.15205,

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.