Pith. sign in

REVIEW 3 major objections 6 minor 2 cited by

Most of the LLM routing gap is real specialist advantage; 12–36% is single-draw noise no single-commit router can close.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 02:29 UTC pith:LWDWWHH4

load-bearing objection Solid theory + honest protocol paper: recoverability asymmetry is real and elementary; 12–36% noise shares are prospective open-pool localizations, not an audit of LLMRouterBench’s 20-pt gap. the 3 major comments →

arxiv 2607.03436 v2 pith:LWDWWHH4 submitted 2026-07-03 cs.LG

How Much of the Routing Gap Is Real? Decomposing the Router-to-Oracle Gap into Reproducible Specialist Advantage and Single-Draw Label Noise

classification cs.LG
keywords LLM routingmodel selectionoracle biasstochastic decodingbenchmark reliabilityrecoverability asymmetrysingle-draw noisemulti-sample oracle
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Routing benchmarks measure remaining opportunity as the gap between a learned router and a per-query oracle that picks the best model in hindsight. Under temperature sampling that oracle is built from one random correct/incorrect draw per model, so it credits lucky successes that no model can reproduce. This paper decomposes the expected oracle into a reproducible ceiling (the best model’s true solve probability) plus a non-negative noise floor. It proves that every single-commit router—deterministic or randomized—is capped at the reproducible ceiling, yet best-of-K sampling of that same model, at the oracle’s own budget, recovers the floor. On controlled open-model re-runs the noise share is a substantial minority of the gap (12% on saturated arithmetic, 36% on competition math, 13% on hard science), larger on thin-support queries, while the majority remains genuine recoverable specialist advantage. The work releases a multi-sample oracle protocol so benchmarks can report both ceilings instead of a single lucky draw.

Core claim

The expected single-draw oracle decomposes as O^exp = O^repro + Δ into reproducible single-commit headroom and a non-negative selection floor. That floor is closed by no single-commit router (for any coupling of the draws), yet is recovered by test-time best-of-K on the committed maximizer at matched budget. Empirically, single-draw noise is 12–36% of the reported router-to-oracle gap on open pools—majority recoverable specialist advantage remains.

What carries the argument

Recoverability asymmetry (Theorem 2) together with the exact algebraic gap split G = G_rec + G_noise (Theorem 1): selection is capped at max_m p_im for any joint law, while sampling the maximizer lifts through that cap; the noise term is non-negative by Jensen on the max and needs no cross-model independence.

Load-bearing premise

The argument treats a router as committing to one model and being scored only on that model’s reproducible success probability, with the choice made independently of the scored draw.

What would settle it

On a multi-sample re-generation, if best-of-K accuracy on the per-query best model fails to meet or exceed the matched-budget independent-pool single-draw oracle, or if the estimated noise share collapses to near zero across unsaturated thin-support strata, the asymmetry and the claimed minority noise share would be refuted.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper argues that under stochastic decoding the per-instance routing oracle used by RouterBench/LLMRouterBench is an upward-biased single-draw estimator, and decomposes the router-to-oracle gap exactly as G = G_rec + G_noise into recoverable specialist advantage and a non-negative single-commit selection floor Δ. The main theoretical claim is a recoverability asymmetry (Thm. 2): for any joint law of Bernoulli draws, every single-commit router (deterministic or randomized) is capped at O^repro = max_m p_im, so G_noise is a hard floor for that class, yet best-of-K on a maximizer at matched budget weakly dominates the independent-pool single-draw oracle. Empirically, on controlled open-weight re-generations (GSM8K, MATH-500, GPQA; 11 models, k=30), the noise share of the gap is localized at 12%/36%/13%, larger in thin support. The paper releases a multi-sample oracle protocol and correctness tensors.

Significance. If the framing is adopted, routing benchmarks would stop treating single-draw union/argmax oracles as reproducible ceilings, and reported headroom would shrink by a measurable minority (here 12–36%). The contribution is primarily methodological and structural rather than a new router: the decomposition (Thm. 1), the selection-vs-sampling asymmetry (Thm. 2), the Fréchet/FKG robustness results, and the released multi-sample protocol are concrete and usable. Strengths include detailed appendix proofs, explicit assumption scoping (A1–A7), a falsifiable best-of-K check, lineage/pool-cardinality controls, and clear distinction from concurrent debiasing work (Capability Frontier) and from deterministic evaluation artifacts. The theorems are elementary but correctly scoped and load-bearing for how the field interprets oracle gaps.

major comments (3)
  1. [Abstract; §VI.B; Table V; Table VII] Abstract and §VI.B: the headline noise shares (12%/36%/13%) are computed against the best-single baseline, which the paper itself notes maximizes G_noise/G; against a weaker TF-IDF k-NN router the MATH-500 share falls to 31% of a larger gap while the floor G_noise stays fixed. The abstract’s phrasing “12–36% of the reported router-to-oracle gap” should state that the share is baseline-dependent and that the invariant object is the floor O^exp − O^repro, not the percentage of any particular router’s gap.
  2. [Abstract; Thm. 2(b); Cor. 3; Lemma 4; §VI.F] Thm. 2(b) and Cor. 3 / Lemma 4: the abstract and intro state that Δ is “provably recovered by test-time sampling,” but full recovery of Δ is verifier-gated. On MATH-500 the verifier-free majority-vote ceiling falls below O^repro (0.824 < 0.837), so the sampling-recoverable mass is dominated by Δ_guess. The claim should be tightened in the abstract and contribution list to “recoverable by sampling up to a verifier-gated residual,” matching the paper’s own A6/Cor. 3 scope, so readers do not over-read the operational recovery.
  3. [§I; Assumption 1 (A4)–(A5); Thm. 2(a); §VII–VIII] §I and Assumption 1 (A4–A5): the hard floor is proved only for single-commit selection under fresh-draw scoring. Cascading and fusion are introduced as peer paradigms in the opening, yet the operational reading of “not recoverable by routing” does not apply to multi-call systems that leave the single-commit class. A short explicit statement in the introduction and conclusion that the floor is scoped to single-commit routers (and shrinks under ensembling/cascading) would prevent misapplication to the broader multi-LLM literature the paper cites.
minor comments (6)
  1. [Fig. 2; Fig. 4] Fig. 2 and the worked example (10 models, p=0.1) are effective; consider adding the corresponding O^agg / Δ_know–Δ_guess split so the figure matches the three-oracle diagram in Fig. 4.
  2. [Table III; §V] Table III provenance is useful; the RouterBench decoding “—” cells could note more explicitly that generative/LLM-judge subsets are out of exact-match scope so readers do not treat the secondary pool as a full corroboration.
  3. [Table II; Cor. 1; §VI.D–E] Notation table (Table II) lists S(K) as noise share of the oracle; the main text also uses G_noise/G for the gap share. A one-line reminder that these are distinct (oracle share vs gap share) would reduce confusion in §VI.D–E where both appear.
  4. [§V; Algorithm 1] In §V the protocol says k≥20 overall and k≥30 in thin support; realized runs use k=30 throughout. Stating the realized budget once in the methods box (Alg. 1 / Fig. 6) would avoid a minor inconsistency for reimplementers.
  5. [Table I; §II] Related-work Table I is dense but helpful; a footnote clarifying that “✓” for recoverability asymmetry is unique to this work (as claimed) would make the novelty row self-contained.
  6. [Abstract; §I] Minor typography: “onesingle” / “asingle” spacing artifacts appear in the abstract and early introduction in the preprint text; clean before camera-ready.

Circularity Check

0 steps flagged

No significant circularity: decomposition is an acknowledged algebraic identity; recoverability cap follows from stated single-commit scoring; empirical noise shares come from fresh multi-sample re-generation, not fitted free parameters or self-citation chains.

full rationale

The paper's load-bearing claims do not reduce to their inputs by construction in the sense the circularity rubric targets. Theorem 1(a) is explicitly labeled an algebraic identity (G = G_rec + G_noise once O^exp and O^repro are fixed); its substance is non-negativity of G_noise via Jensen/monotonicity of max under (A1) alone, which is elementary but independent of any fit. Theorem 2(a)'s selection cap (score_i(π) ≤ O^repro for any single-commit π) is the simplex maximum under the paper's own scoring convention (A4-linear + A5); that is a scoped definition of the single-commit class, not a prediction forced by rearranging data or by a self-citation uniqueness theorem. The sampling lift (Thm. 2(b)) is a standard best-of-n bound under (A1). Empirical noise shares (12%/36%/13% on GSM8K/MATH-500/GPQA) are prospective localizations from controlled open-pool re-generation at k=30, seed-aligned, with one-sided dependence-corrected bounds—not rearrangements of the published k=1 LLMRouterBench matrix (which the paper correctly treats as non-identifiable for O^repro). Concurrent debiasing work (Capability Frontier) is cited and distinguished rather than silently reused as proof. No uniqueness theorem is imported from overlapping authors; no free parameter is fitted then re-presented as a prediction; no ansatz is smuggled via self-citation. The single-commit scoring scope is a modeling choice (already flagged by the reader), not circularity. Score 0 is the honest finding.

Axiom & Free-Parameter Ledger

4 free parameters · 8 axioms · 2 invented entities

The theorems rest on standard probability (Bernoulli marginals, Jensen on max, Fréchet bounds) plus domain scoring conventions that define what “router” means. Empirical magnitudes further depend on design choices (k, T, open text-only pool, exact-match scope) that are free parameters of the localization, not of the asymmetry proof. No new physical entities; the oracles are definitional constructs.

free parameters (4)
  • sample budget k = k=30 (realized); protocol k≥20, thin-support k≥30
    k≥20 protocol / k=30 realized; noise-share estimate rises with k until ~20–30 (Table VII). Choice affects finite-sample localization of G_noise though not the population theorems.
  • decoding temperature T = T=0.2
    Fixed to audited benchmark T=0.2; T-sweep is secondary. Different T would estimate different p_im and a different noise floor.
  • reliability threshold τ = 0.5 / 0.9
    Auxiliary O^thr(τ) reported at τ∈{0.5,0.9}; not load-bearing for main G_noise but a free analysis knob.
  • open-pool composition and N = 11 models; N=500 GSM8K/MATH-500, 198 GPQA
    11 open text-only models, N=500/500/198 stratified queries; noise share is explicitly pool-conditional (Table VI). Lineage dedup vs full pool changes reported shares.
axioms (8)
  • domain assumption A1: per-(query,model) draws are i.i.d. Bernoulli(p_im) at fixed T>0
    Invoked for sampling lift, concentration, and best-of-n formulas; gated empirically by over-dispersion checks.
  • domain assumption A2: optional cross-model independence for closed-form product O^exp,⊥
    Dropped for selection cap and sign of Δ; used only for sharp product form and as FKG upper envelope.
  • domain assumption A3: recorded benchmark label is one such draw
    Matches LLMRouterBench released matrices (one generation per cell at T=0.2).
  • domain assumption A4/A4-linear: router commits to one model and is scored by that model’s p (linear extension for randomized π)
    Defines the single-commit class for which Δ is a hard floor (Thm. 2(a), Lem. 2).
  • domain assumption A5: router choice independent of the scored draw (fresh-draw scoring)
    Makes reported accuracy equal score_i(π); claimed verified on LLMRouterBench train/test protocol.
  • domain assumption A6: deliverable recovery of full Δ needs a deploy-time verifier; else only O^agg
    Scopes “recovered by sampling” (Cor. 3); confirmed empirically when majority vote stays below O^repro on MATH-500.
  • domain assumption A7: re-generation records seed-aligned K-tuples across models
    Required for unbiased seed-aligned estimator of true O^exp without independence.
  • standard math Standard facts: max is convex/monotone; Fréchet bounds on unions; Hoeffding/sub-Gaussian maximal inequalities
    Used throughout Props. 1–3 and Thms. 1–4; not novel.
invented entities (2)
  • Reproducible oracle O^repro and single-commit selection floor Δ independent evidence
    purpose: Split expected single-draw oracle into committable ceiling vs residual that selection cannot close
    Definitional constructs from {p_im}, not new physical objects; independent handle is multi-sample estimation and best-of-K tests.
  • Verifier-free aggregation ceiling O^agg and Δ_know / Δ_guess split independent evidence
    purpose: Scope how much of Δ sampling can reclaim without a ground-truth verifier
    Operational refinement of recoverability; tested via majority-vote vs union on the re-generated tensors.

pith-pipeline@v1.1.0-grok45 · 54874 in / 4404 out tokens · 39332 ms · 2026-07-12T02:29:00.612842+00:00 · methodology

0 comments
read the original abstract

On real open-model pools, 12--36% of the reported router-to-oracle gap is single-draw label noise that no single-commit router can capture, while the majority is genuine, recoverable specialist advantage; this work proves why (a recoverability asymmetry) and releases a protocol to measure it. Routing among large language models (LLMs) trades cost for quality, motivated by the gap between learned routers and a per-instance oracle. But under stochastic decoding that oracle is a single Bernoulli draw, not a reproducible property. We recast the question structurally: the expected oracle decomposes as $O^{\exp}=O^{\mathrm{repro}}+\Delta$, into reproducible single-commit headroom $O^{\mathrm{repro}}$ and a non-negative single-commit selection floor $\Delta$. Our main result is a recoverability asymmetry: this floor is closed by no single-commit router (deterministic or randomized), yet is provably recovered by test-time sampling: best-of-$K$ on the committed model, at the oracle's own budget, dominates the independent-pool single-draw oracle. This cap needs no cross-model independence, pinning "not recoverable" to single-commit selection, not to information. The floor's magnitude is a prospective, conservative localization, not an audit: LLMRouterBench (33 models, 391,645 instances) builds its oracle as a per-query union of single $T=0.2$ draws, so its 20-point gap is by construction a union of stochastic draws; since $O^{\mathrm{repro}}$ is non-identifiable at $k=1$, we re-estimate by fresh $k\ge20$ resampling under one-sided, dependence-corrected bounds. Across three controlled open-model re-generations (arithmetic, competition math, and non-math science), single-draw noise is a substantial minority of the gap, larger on unsaturated benchmarks and approaching half on the hardest queries. We release a multi-sample oracle protocol that routing benchmarks can adopt.

Figures

Figures reproduced from arXiv: 2607.03436 by Teng-Ruei Chen.

Figure 1
Figure 1. Figure 1: The paper as a closed loop. Assumptions (A1–A7) yield the main theorems by deduction—the recoverability asymmetry (Thm. 2) and the exact decomposition G = Grec + Gnoise (Thm. 1), shaded. The method only re-samples (no new router); the experiment localizes the magnitude of Gnoise and runs the falsifiable best-of-K check. The dashed feedback path revises only the assumptions and localized magnitude, not the … view at source ↗
Figure 2
Figure 2. Figure 2: A worked example. With 10 models each correct with probability p = 0.1, the single-draw oracle reaches 1 − 0.9 10 ≈ 65%, but only the 10% reproducible mass (grey) is reachable by a router that commits to one model; the remaining 55% (blue) is single-draw label noise no router can capture. of every router’s gap is single-draw noise not recoverable by single-commit routing—though recoverable by test￾time sam… view at source ↗
Figure 3
Figure 3. Figure 3: The recoverability asymmetry (Thm. 2), the paper’s main result. On the selection axis every single-commit router—deterministic, randomized, or any mixture—is capped at the reproducible ceiling Orepro (the hatched mass above it is unreachable), and this cap needs no cross-model independence. On the sampling axis the same per-query budget, spent as best-of-K on the committed model, provably lifts through the… view at source ↗
Figure 4
Figure 4. Figure 4: Three oracles on a single query, in the proven order [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: With identical per-model success probability [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: The method in four steps: sample each model [PITH_FULL_IMAGE:figures/full_fig_p009_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Gap composition (Thm. 1) on an identical [PITH_FULL_IMAGE:figures/full_fig_p011_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Single-draw noise share by support stratum (Cor. 2), all three [PITH_FULL_IMAGE:figures/full_fig_p012_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Single-draw noise share by pool composition ( [PITH_FULL_IMAGE:figures/full_fig_p012_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models

    cs.LG 2026-07 conditional novelty 6.5

    An online resample-or-reroute policy that allocates each unit of a per-query budget by estimated marginal correctness per unit cost attains a better cost–quality Pareto front than single-route, cascade, and best-of-K ...

  2. Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models

    cs.LG 2026-07 conditional novelty 6.0

    A greedy online policy that allocates a per-query budget between resampling the committed LLM and rerouting to an alternative achieves favorable cost-quality trade-offs, with gains concentrated on heterogeneous model ...

Reference graph

Works this paper leans on

44 extracted references · 14 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Best arm identification in multi-armed bandits,

    J.-Y . Audibert, S. Bubeck, and R. Munos, “Best arm identification in multi-armed bandits,” inProc. Conf. on Learning Theory (COLT), 2010, pp. 41–53

  2. [2]

    On the complexity of best- arm identification in multi-armed bandit models,

    E. Kaufmann, O. Cappé, and A. Garivier, “On the complexity of best- arm identification in multi-armed bandit models,”Journal of Machine Learning Research, vol. 17, no. 1, pp. 1–42, 2016

  3. [3]

    Mixture- of-agents enhances large language model capabilities,

    J. Wang, J. Wang, B. Athiwaratkun, C. Zhang, and J. Zou, “Mixture- of-agents enhances large language model capabilities,” inProc. ICLR, 2025

  4. [4]

    RouteLLM: Learning to route LLMs from preference data,

    I. Ong, A. Almahairi, V . Wu, W.-L. Chiang, T. Wu, J. E. Gonzalez, M. W. Kadous, and I. Stoica, “RouteLLM: Learning to route LLMs from preference data,” inProc. ICLR, 2025

  5. [5]

    A unified approach to routing and cascading for LLMs,

    J. Dekoninck, M. Baader, and M. Vechev, “A unified approach to routing and cascading for LLMs,” inProc. ICML, 2025, arXiv:2410.10347, 2024

  6. [6]

    Dynamic model routing and cas- cading for efficient LLM inference: A survey,

    Y . Moslem and J. D. Kelleher, “Dynamic model routing and cas- cading for efficient LLM inference: A survey,”arXiv preprint arXiv:2603.04445, 2026

  7. [7]

    RouterBench: A benchmark for multi-LLM routing system,

    Q. J. Huet al., “RouterBench: A benchmark for multi-LLM routing system,”arXiv preprint arXiv:2403.12031, 2024

  8. [8]

    LLMRouterBench: A massive benchmark and unified framework for LLM routing,

    H. Li, Y . Zhang, Z. Guo, C. Wang, S. Tang, Q. Zhang, Y . Chen, B. Qi, P. Ye, L. Bai, Z. Wang, and S. Hu, “LLMRouterBench: A massive benchmark and unified framework for LLM routing,”arXiv preprint arXiv:2601.07206, 2026

  9. [9]

    Route-and-reason: Scaling LLM reasoning with a reinforced model router,

    C. Shaoet al., “Route-and-reason: Scaling LLM reasoning with a reinforced model router,”arXiv preprint arXiv:2506.05901, 2025

  10. [10]

    Best-of-∞: Asymp- totic performance of test-time LLM ensembling,

    J. Komiyama, D. Oba, and M. Oyamada, “Best-of-∞: Asymp- totic performance of test-time LLM ensembling,”arXiv preprint arXiv:2509.21091, 2025

  11. [11]

    A flexible defense against the winner’s curse,

    T. Zrnic and W. Fithian, “A flexible defense against the winner’s curse,” The Annals of Statistics, 2025

  12. [12]

    The greatest of a finite set of random variables,

    C. E. Clark, “The greatest of a finite set of random variables,”Operations Research, vol. 9, no. 2, pp. 145–162, 1961

  13. [13]

    Don’t pass@k: A bayesian framework for large language model evaluation,

    M. Hariri, A. Samandar, M. Hinczewski, and V . Chaudhary, “Don’t pass@k: A bayesian framework for large language model evaluation,” arXiv preprint arXiv:2510.04265, 2025

  14. [14]

    The capability frontier: Benchmarks miss 82% of model performance,

    B. Fowler, R. Smith, D. T. Graviet, W. Myers, J. Greaves, N. F. Oozeer, A. García, P. Quirke, A. Abdullah, F. Barez, and S. K. Upadhyay, “The capability frontier: Benchmarks miss 82% of model performance,”arXiv preprint arXiv:2606.26836, 2026

  15. [15]

    Unsolvability ceiling in multi-LLM rout- ing: An empirical study of evaluation artifacts,

    S. Garg and A. Sagtani, “Unsolvability ceiling in multi-LLM rout- ing: An empirical study of evaluation artifacts,”arXiv preprint arXiv:2605.07395, 2026

  16. [16]

    FrugalGPT: How to use large language models while reducing cost and improving performance,

    L. Chen, M. Zaharia, and J. Zou, “FrugalGPT: How to use large language models while reducing cost and improving performance,”Transactions on Machine Learning Research (TMLR), 2024

  17. [17]

    Hybrid LLM: Cost-efficient and quality-aware query routing,

    D. Ding, A. Mallick, C. Wang, R. Sim, S. Mukherjee, V . Rühle, L. V . S. Lakshmanan, and A. H. Awadallah, “Hybrid LLM: Cost-efficient and quality-aware query routing,” inProc. ICLR, 2024

  18. [18]

    Are more LLM calls all you need? towards the scaling properties of compound AI systems,

    L. Chen, J. Q. Davis, B. Hanin, P. Bailis, I. Stoica, M. Zaharia, and J. Zou, “Are more LLM calls all you need? towards the scaling properties of compound AI systems,” inAdvances in Neural Information Processing Systems (NeurIPS), 2024

  19. [19]

    Rethinking mixture-of-agents: Is mixing different large language models beneficial?

    W. Li, Y . Lin, M. Xia, and C. Jin, “Rethinking mixture-of-agents: Is mixing different large language models beneficial?” inProc. ICML, 2025, arXiv:2502.00674

  20. [20]

    RouterEval: A comprehensive benchmark for routing LLMs to explore model-level scaling up in LLMs,

    Z. Huang, G. Ling, Y . Lin, Y . Chen, S. Zhong, H. Wu, and L. Lin, “RouterEval: A comprehensive benchmark for routing LLMs to explore model-level scaling up in LLMs,”arXiv preprint arXiv:2503.10657, 2025

  21. [21]

    Within-model vs. between-prompt vari- ability in large language models for creative tasks,

    J. Haase, J. Gonnermann-Müller, P. H. P. Hanel, N. Leins, T. Kosch, J. Mendling, and S. Pokutta, “Within-model vs. between-prompt vari- ability in large language models for creative tasks,”arXiv preprint arXiv:2601.21339, 2026

  22. [22]

    Measuring all the noises of LLM evals,

    S. Wang, “Measuring all the noises of LLM evals,”arXiv preprint arXiv:2512.21326, 2025

  23. [23]

    On randomness in agentic evals,

    B. H. Bjarnason, A. Silva, and M. Monperrus, “On randomness in agentic evals,”arXiv preprint arXiv:2602.07150, 2026

  24. [24]

    The good, the bad, and the greedy: Evaluation of LLMs should not ignore non-determinism,

    Y . Song, G. Wang, S. Li, and B. Y . Lin, “The good, the bad, and the greedy: Evaluation of LLMs should not ignore non-determinism,” in Proc. NAACL, 2025, pp. 4195–4206

  25. [25]

    Quantifying variance in evaluation benchmarks,

    L. Madaan, A. K. Singh, R. Schaeffer, A. Poulton, S. Koyejo, P. Stene- torp, S. Narang, and D. Hupkes, “Quantifying variance in evaluation benchmarks,”arXiv preprint arXiv:2406.10229, 2024

  26. [26]

    Reporting score distributions makes a difference: Performance study of LSTM-networks for sequence tagging,

    N. Reimers and I. Gurevych, “Reporting score distributions makes a difference: Performance study of LSTM-networks for sequence tagging,” inProc. EMNLP. Association for Computational Linguistics, 2017, pp. 338–348

  27. [27]

    Inference on winners,

    I. Andrews, T. Kitagawa, and A. McCloskey, “Inference on winners,” The Quarterly Journal of Economics, vol. 139, no. 1, pp. 305–358, 2024, first circulated as NBER Working Paper 25456, 2019

  28. [28]

    The optimizer’s curse: Skepticism and postdecision surprise in decision analysis,

    J. E. Smith and R. L. Winkler, “The optimizer’s curse: Skepticism and postdecision surprise in decision analysis,”Management Science, vol. 52, no. 3, pp. 311–322, 2006

  29. [29]

    Tweedie’s formula and selection bias,

    B. Efron, “Tweedie’s formula and selection bias,”Journal of the Amer- ican Statistical Association, vol. 106, no. 496, pp. 1602–1614, 2011

  30. [30]

    Valid post- selection inference,

    R. Berk, L. Brown, A. Buja, K. Zhang, and L. Zhao, “Valid post- selection inference,”The Annals of Statistics, vol. 41, no. 2, pp. 802–837, 2013

  31. [31]

    Double Q-learning,

    H. van Hasselt, “Double Q-learning,” inAdvances in Neural Information Processing Systems (NeurIPS), vol. 23, 2010, pp. 2613–2621, analyzes the upward bias of the maximum over noisy value estimates

  32. [32]

    When routing collapses: On the degenerate convergence of LLM routers,

    G. Lai and H.-J. Ye, “When routing collapses: On the degenerate convergence of LLM routers,”arXiv preprint arXiv:2602.03478, 2026

  33. [33]

    Can large language models always solve easy problems if they can solve harder ones?

    Z. Yang, Y . Zhang, T. Liu, J. Yang, J. Lin, C. Zhou, and Z. Sui, “Can large language models always solve easy problems if they can solve harder ones?” inProc. EMNLP, 2024, pp. 1531–1555

  34. [34]

    Expected reward prediction, with applications to model routing,

    K. Hasanaliyev, S. Alberti, J. Hamer, D. Rajagopal, K. Robinson, J. Snoek, V . Veitch, and A. N. D’Amour, “Expected reward prediction, with applications to model routing,”arXiv preprint arXiv:2603.20217, 2026

  35. [35]

    LLMs encode their failures: Predicting success from pre-generation activations,

    W. Lugoloobi, T. Foster, W. Bankes, and C. Russell, “LLMs encode their failures: Predicting success from pre-generation activations,”arXiv preprint arXiv:2602.09924, 2026

  36. [36]

    Large language monkeys: Scaling inference compute with repeated sampling,

    B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V . Le, C. Ré, and A. Mirhoseini, “Large language monkeys: Scaling inference compute with repeated sampling,”arXiv preprint arXiv:2407.21787, 2024

  37. [37]

    Self-consistency improves chain of thought reasoning in language models,

    X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdh- ery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” inProc. ICLR, 2023, arXiv:2203.11171

  38. [38]

    The limits of inference scaling through resampling,

    B. Stroebl, S. Kapoor, and A. Narayanan, “The limits of inference scaling through resampling,” inProc. ICLR, 2026

  39. [39]

    Let’s verify step by step,

    H. Lightman, V . Kosaraju, Y . Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe, “Let’s verify step by step,” inProc. ICLR, 2024

  40. [40]

    When does combining language models help? a co-failure ceiling on routing, voting, and mixture-of-agents across 67 frontier models,

    J. Chen, “When does combining language models help? a co-failure ceiling on routing, voting, and mixture-of-agents across 67 frontier models,”arXiv preprint arXiv:2606.27288, 2026

  41. [41]

    Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy,

    L. I. Kuncheva and C. J. Whitaker, “Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy,”Machine Learning, vol. 51, no. 2, pp. 181–207, 2003

  42. [42]

    Neural network ensembles, cross validation, and active learning,

    A. Krogh and J. Vedelsby, “Neural network ensembles, cross validation, and active learning,” inAdvances in Neural Information Processing Systems (NIPS). MIT Press, 1994, pp. 231–238. PREPRINT, JULY 2026 29

  43. [43]

    GPQA: A graduate-level google-proof Q&A benchmark,

    D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y . Pang, J. Dirani, J. Michael, and S. R. Bowman, “GPQA: A graduate-level google-proof Q&A benchmark,” inFirst Conference on Language Modeling (COLM), 2024, arXiv:2311.12022

  44. [44]

    Devroye, L

    L. Devroye, L. Györfi, and G. Lugosi,A Probabilistic Theory of Pattern Recognition, ser. Stochastic Modelling and Applied Probability. New York: Springer, 1996, vol. 31, see Thm. 2.2: the plug-in (argmax) rule has excess risk at most twice theL 1 estimation error of the conditional probabilities