REVIEW 3 major objections 6 minor 2 cited by
Most of the LLM routing gap is real specialist advantage; 12–36% is single-draw noise no single-commit router can close.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 02:29 UTC pith:LWDWWHH4
load-bearing objection Solid theory + honest protocol paper: recoverability asymmetry is real and elementary; 12–36% noise shares are prospective open-pool localizations, not an audit of LLMRouterBench’s 20-pt gap. the 3 major comments →
How Much of the Routing Gap Is Real? Decomposing the Router-to-Oracle Gap into Reproducible Specialist Advantage and Single-Draw Label Noise
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The expected single-draw oracle decomposes as O^exp = O^repro + Δ into reproducible single-commit headroom and a non-negative selection floor. That floor is closed by no single-commit router (for any coupling of the draws), yet is recovered by test-time best-of-K on the committed maximizer at matched budget. Empirically, single-draw noise is 12–36% of the reported router-to-oracle gap on open pools—majority recoverable specialist advantage remains.
What carries the argument
Recoverability asymmetry (Theorem 2) together with the exact algebraic gap split G = G_rec + G_noise (Theorem 1): selection is capped at max_m p_im for any joint law, while sampling the maximizer lifts through that cap; the noise term is non-negative by Jensen on the max and needs no cross-model independence.
Load-bearing premise
The argument treats a router as committing to one model and being scored only on that model’s reproducible success probability, with the choice made independently of the scored draw.
What would settle it
On a multi-sample re-generation, if best-of-K accuracy on the per-query best model fails to meet or exceed the matched-budget independent-pool single-draw oracle, or if the estimated noise share collapses to near zero across unsaturated thin-support strata, the asymmetry and the claimed minority noise share would be refuted.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that under stochastic decoding the per-instance routing oracle used by RouterBench/LLMRouterBench is an upward-biased single-draw estimator, and decomposes the router-to-oracle gap exactly as G = G_rec + G_noise into recoverable specialist advantage and a non-negative single-commit selection floor Δ. The main theoretical claim is a recoverability asymmetry (Thm. 2): for any joint law of Bernoulli draws, every single-commit router (deterministic or randomized) is capped at O^repro = max_m p_im, so G_noise is a hard floor for that class, yet best-of-K on a maximizer at matched budget weakly dominates the independent-pool single-draw oracle. Empirically, on controlled open-weight re-generations (GSM8K, MATH-500, GPQA; 11 models, k=30), the noise share of the gap is localized at 12%/36%/13%, larger in thin support. The paper releases a multi-sample oracle protocol and correctness tensors.
Significance. If the framing is adopted, routing benchmarks would stop treating single-draw union/argmax oracles as reproducible ceilings, and reported headroom would shrink by a measurable minority (here 12–36%). The contribution is primarily methodological and structural rather than a new router: the decomposition (Thm. 1), the selection-vs-sampling asymmetry (Thm. 2), the Fréchet/FKG robustness results, and the released multi-sample protocol are concrete and usable. Strengths include detailed appendix proofs, explicit assumption scoping (A1–A7), a falsifiable best-of-K check, lineage/pool-cardinality controls, and clear distinction from concurrent debiasing work (Capability Frontier) and from deterministic evaluation artifacts. The theorems are elementary but correctly scoped and load-bearing for how the field interprets oracle gaps.
major comments (3)
- [Abstract; §VI.B; Table V; Table VII] Abstract and §VI.B: the headline noise shares (12%/36%/13%) are computed against the best-single baseline, which the paper itself notes maximizes G_noise/G; against a weaker TF-IDF k-NN router the MATH-500 share falls to 31% of a larger gap while the floor G_noise stays fixed. The abstract’s phrasing “12–36% of the reported router-to-oracle gap” should state that the share is baseline-dependent and that the invariant object is the floor O^exp − O^repro, not the percentage of any particular router’s gap.
- [Abstract; Thm. 2(b); Cor. 3; Lemma 4; §VI.F] Thm. 2(b) and Cor. 3 / Lemma 4: the abstract and intro state that Δ is “provably recovered by test-time sampling,” but full recovery of Δ is verifier-gated. On MATH-500 the verifier-free majority-vote ceiling falls below O^repro (0.824 < 0.837), so the sampling-recoverable mass is dominated by Δ_guess. The claim should be tightened in the abstract and contribution list to “recoverable by sampling up to a verifier-gated residual,” matching the paper’s own A6/Cor. 3 scope, so readers do not over-read the operational recovery.
- [§I; Assumption 1 (A4)–(A5); Thm. 2(a); §VII–VIII] §I and Assumption 1 (A4–A5): the hard floor is proved only for single-commit selection under fresh-draw scoring. Cascading and fusion are introduced as peer paradigms in the opening, yet the operational reading of “not recoverable by routing” does not apply to multi-call systems that leave the single-commit class. A short explicit statement in the introduction and conclusion that the floor is scoped to single-commit routers (and shrinks under ensembling/cascading) would prevent misapplication to the broader multi-LLM literature the paper cites.
minor comments (6)
- [Fig. 2; Fig. 4] Fig. 2 and the worked example (10 models, p=0.1) are effective; consider adding the corresponding O^agg / Δ_know–Δ_guess split so the figure matches the three-oracle diagram in Fig. 4.
- [Table III; §V] Table III provenance is useful; the RouterBench decoding “—” cells could note more explicitly that generative/LLM-judge subsets are out of exact-match scope so readers do not treat the secondary pool as a full corroboration.
- [Table II; Cor. 1; §VI.D–E] Notation table (Table II) lists S(K) as noise share of the oracle; the main text also uses G_noise/G for the gap share. A one-line reminder that these are distinct (oracle share vs gap share) would reduce confusion in §VI.D–E where both appear.
- [§V; Algorithm 1] In §V the protocol says k≥20 overall and k≥30 in thin support; realized runs use k=30 throughout. Stating the realized budget once in the methods box (Alg. 1 / Fig. 6) would avoid a minor inconsistency for reimplementers.
- [Table I; §II] Related-work Table I is dense but helpful; a footnote clarifying that “✓” for recoverability asymmetry is unique to this work (as claimed) would make the novelty row self-contained.
- [Abstract; §I] Minor typography: “onesingle” / “asingle” spacing artifacts appear in the abstract and early introduction in the preprint text; clean before camera-ready.
Circularity Check
No significant circularity: decomposition is an acknowledged algebraic identity; recoverability cap follows from stated single-commit scoring; empirical noise shares come from fresh multi-sample re-generation, not fitted free parameters or self-citation chains.
full rationale
The paper's load-bearing claims do not reduce to their inputs by construction in the sense the circularity rubric targets. Theorem 1(a) is explicitly labeled an algebraic identity (G = G_rec + G_noise once O^exp and O^repro are fixed); its substance is non-negativity of G_noise via Jensen/monotonicity of max under (A1) alone, which is elementary but independent of any fit. Theorem 2(a)'s selection cap (score_i(π) ≤ O^repro for any single-commit π) is the simplex maximum under the paper's own scoring convention (A4-linear + A5); that is a scoped definition of the single-commit class, not a prediction forced by rearranging data or by a self-citation uniqueness theorem. The sampling lift (Thm. 2(b)) is a standard best-of-n bound under (A1). Empirical noise shares (12%/36%/13% on GSM8K/MATH-500/GPQA) are prospective localizations from controlled open-pool re-generation at k=30, seed-aligned, with one-sided dependence-corrected bounds—not rearrangements of the published k=1 LLMRouterBench matrix (which the paper correctly treats as non-identifiable for O^repro). Concurrent debiasing work (Capability Frontier) is cited and distinguished rather than silently reused as proof. No uniqueness theorem is imported from overlapping authors; no free parameter is fitted then re-presented as a prediction; no ansatz is smuggled via self-citation. The single-commit scoring scope is a modeling choice (already flagged by the reader), not circularity. Score 0 is the honest finding.
Axiom & Free-Parameter Ledger
free parameters (4)
- sample budget k =
k=30 (realized); protocol k≥20, thin-support k≥30
- decoding temperature T =
T=0.2
- reliability threshold τ =
0.5 / 0.9
- open-pool composition and N =
11 models; N=500 GSM8K/MATH-500, 198 GPQA
axioms (8)
- domain assumption A1: per-(query,model) draws are i.i.d. Bernoulli(p_im) at fixed T>0
- domain assumption A2: optional cross-model independence for closed-form product O^exp,⊥
- domain assumption A3: recorded benchmark label is one such draw
- domain assumption A4/A4-linear: router commits to one model and is scored by that model’s p (linear extension for randomized π)
- domain assumption A5: router choice independent of the scored draw (fresh-draw scoring)
- domain assumption A6: deliverable recovery of full Δ needs a deploy-time verifier; else only O^agg
- domain assumption A7: re-generation records seed-aligned K-tuples across models
- standard math Standard facts: max is convex/monotone; Fréchet bounds on unions; Hoeffding/sub-Gaussian maximal inequalities
invented entities (2)
-
Reproducible oracle O^repro and single-commit selection floor Δ
independent evidence
-
Verifier-free aggregation ceiling O^agg and Δ_know / Δ_guess split
independent evidence
read the original abstract
On real open-model pools, 12--36% of the reported router-to-oracle gap is single-draw label noise that no single-commit router can capture, while the majority is genuine, recoverable specialist advantage; this work proves why (a recoverability asymmetry) and releases a protocol to measure it. Routing among large language models (LLMs) trades cost for quality, motivated by the gap between learned routers and a per-instance oracle. But under stochastic decoding that oracle is a single Bernoulli draw, not a reproducible property. We recast the question structurally: the expected oracle decomposes as $O^{\exp}=O^{\mathrm{repro}}+\Delta$, into reproducible single-commit headroom $O^{\mathrm{repro}}$ and a non-negative single-commit selection floor $\Delta$. Our main result is a recoverability asymmetry: this floor is closed by no single-commit router (deterministic or randomized), yet is provably recovered by test-time sampling: best-of-$K$ on the committed model, at the oracle's own budget, dominates the independent-pool single-draw oracle. This cap needs no cross-model independence, pinning "not recoverable" to single-commit selection, not to information. The floor's magnitude is a prospective, conservative localization, not an audit: LLMRouterBench (33 models, 391,645 instances) builds its oracle as a per-query union of single $T=0.2$ draws, so its 20-point gap is by construction a union of stochastic draws; since $O^{\mathrm{repro}}$ is non-identifiable at $k=1$, we re-estimate by fresh $k\ge20$ resampling under one-sided, dependence-corrected bounds. Across three controlled open-model re-generations (arithmetic, competition math, and non-math science), single-draw noise is a substantial minority of the gap, larger on unsaturated benchmarks and approaching half on the hardest queries. We release a multi-sample oracle protocol that routing benchmarks can adopt.
Figures
Forward citations
Cited by 2 Pith papers
-
Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models
An online resample-or-reroute policy that allocates each unit of a per-query budget by estimated marginal correctness per unit cost attains a better cost–quality Pareto front than single-route, cascade, and best-of-K ...
-
Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models
A greedy online policy that allocates a per-query budget between resampling the committed LLM and rerouting to an alternative achieves favorable cost-quality trade-offs, with gains concentrated on heterogeneous model ...
Reference graph
Works this paper leans on
-
[1]
Best arm identification in multi-armed bandits,
J.-Y . Audibert, S. Bubeck, and R. Munos, “Best arm identification in multi-armed bandits,” inProc. Conf. on Learning Theory (COLT), 2010, pp. 41–53
2010
-
[2]
On the complexity of best- arm identification in multi-armed bandit models,
E. Kaufmann, O. Cappé, and A. Garivier, “On the complexity of best- arm identification in multi-armed bandit models,”Journal of Machine Learning Research, vol. 17, no. 1, pp. 1–42, 2016
2016
-
[3]
Mixture- of-agents enhances large language model capabilities,
J. Wang, J. Wang, B. Athiwaratkun, C. Zhang, and J. Zou, “Mixture- of-agents enhances large language model capabilities,” inProc. ICLR, 2025
2025
-
[4]
RouteLLM: Learning to route LLMs from preference data,
I. Ong, A. Almahairi, V . Wu, W.-L. Chiang, T. Wu, J. E. Gonzalez, M. W. Kadous, and I. Stoica, “RouteLLM: Learning to route LLMs from preference data,” inProc. ICLR, 2025
2025
-
[5]
A unified approach to routing and cascading for LLMs,
J. Dekoninck, M. Baader, and M. Vechev, “A unified approach to routing and cascading for LLMs,” inProc. ICML, 2025, arXiv:2410.10347, 2024
Pith/arXiv arXiv 2025
-
[6]
Dynamic model routing and cas- cading for efficient LLM inference: A survey,
Y . Moslem and J. D. Kelleher, “Dynamic model routing and cas- cading for efficient LLM inference: A survey,”arXiv preprint arXiv:2603.04445, 2026
Pith/arXiv arXiv 2026
-
[7]
RouterBench: A benchmark for multi-LLM routing system,
Q. J. Huet al., “RouterBench: A benchmark for multi-LLM routing system,”arXiv preprint arXiv:2403.12031, 2024
Pith/arXiv arXiv 2024
-
[8]
LLMRouterBench: A massive benchmark and unified framework for LLM routing,
H. Li, Y . Zhang, Z. Guo, C. Wang, S. Tang, Q. Zhang, Y . Chen, B. Qi, P. Ye, L. Bai, Z. Wang, and S. Hu, “LLMRouterBench: A massive benchmark and unified framework for LLM routing,”arXiv preprint arXiv:2601.07206, 2026
arXiv 2026
-
[9]
Route-and-reason: Scaling LLM reasoning with a reinforced model router,
C. Shaoet al., “Route-and-reason: Scaling LLM reasoning with a reinforced model router,”arXiv preprint arXiv:2506.05901, 2025
arXiv 2025
-
[10]
Best-of-∞: Asymp- totic performance of test-time LLM ensembling,
J. Komiyama, D. Oba, and M. Oyamada, “Best-of-∞: Asymp- totic performance of test-time LLM ensembling,”arXiv preprint arXiv:2509.21091, 2025
arXiv 2025
-
[11]
A flexible defense against the winner’s curse,
T. Zrnic and W. Fithian, “A flexible defense against the winner’s curse,” The Annals of Statistics, 2025
2025
-
[12]
The greatest of a finite set of random variables,
C. E. Clark, “The greatest of a finite set of random variables,”Operations Research, vol. 9, no. 2, pp. 145–162, 1961
1961
-
[13]
Don’t pass@k: A bayesian framework for large language model evaluation,
M. Hariri, A. Samandar, M. Hinczewski, and V . Chaudhary, “Don’t pass@k: A bayesian framework for large language model evaluation,” arXiv preprint arXiv:2510.04265, 2025
Pith/arXiv arXiv 2025
-
[14]
The capability frontier: Benchmarks miss 82% of model performance,
B. Fowler, R. Smith, D. T. Graviet, W. Myers, J. Greaves, N. F. Oozeer, A. García, P. Quirke, A. Abdullah, F. Barez, and S. K. Upadhyay, “The capability frontier: Benchmarks miss 82% of model performance,”arXiv preprint arXiv:2606.26836, 2026
Pith/arXiv arXiv 2026
-
[15]
Unsolvability ceiling in multi-LLM rout- ing: An empirical study of evaluation artifacts,
S. Garg and A. Sagtani, “Unsolvability ceiling in multi-LLM rout- ing: An empirical study of evaluation artifacts,”arXiv preprint arXiv:2605.07395, 2026
Pith/arXiv arXiv 2026
-
[16]
FrugalGPT: How to use large language models while reducing cost and improving performance,
L. Chen, M. Zaharia, and J. Zou, “FrugalGPT: How to use large language models while reducing cost and improving performance,”Transactions on Machine Learning Research (TMLR), 2024
2024
-
[17]
Hybrid LLM: Cost-efficient and quality-aware query routing,
D. Ding, A. Mallick, C. Wang, R. Sim, S. Mukherjee, V . Rühle, L. V . S. Lakshmanan, and A. H. Awadallah, “Hybrid LLM: Cost-efficient and quality-aware query routing,” inProc. ICLR, 2024
2024
-
[18]
Are more LLM calls all you need? towards the scaling properties of compound AI systems,
L. Chen, J. Q. Davis, B. Hanin, P. Bailis, I. Stoica, M. Zaharia, and J. Zou, “Are more LLM calls all you need? towards the scaling properties of compound AI systems,” inAdvances in Neural Information Processing Systems (NeurIPS), 2024
2024
-
[19]
Rethinking mixture-of-agents: Is mixing different large language models beneficial?
W. Li, Y . Lin, M. Xia, and C. Jin, “Rethinking mixture-of-agents: Is mixing different large language models beneficial?” inProc. ICML, 2025, arXiv:2502.00674
Pith/arXiv arXiv 2025
-
[20]
RouterEval: A comprehensive benchmark for routing LLMs to explore model-level scaling up in LLMs,
Z. Huang, G. Ling, Y . Lin, Y . Chen, S. Zhong, H. Wu, and L. Lin, “RouterEval: A comprehensive benchmark for routing LLMs to explore model-level scaling up in LLMs,”arXiv preprint arXiv:2503.10657, 2025
Pith/arXiv arXiv 2025
-
[21]
Within-model vs. between-prompt vari- ability in large language models for creative tasks,
J. Haase, J. Gonnermann-Müller, P. H. P. Hanel, N. Leins, T. Kosch, J. Mendling, and S. Pokutta, “Within-model vs. between-prompt vari- ability in large language models for creative tasks,”arXiv preprint arXiv:2601.21339, 2026
arXiv 2026
-
[22]
Measuring all the noises of LLM evals,
S. Wang, “Measuring all the noises of LLM evals,”arXiv preprint arXiv:2512.21326, 2025
arXiv 2025
-
[23]
On randomness in agentic evals,
B. H. Bjarnason, A. Silva, and M. Monperrus, “On randomness in agentic evals,”arXiv preprint arXiv:2602.07150, 2026
arXiv 2026
-
[24]
The good, the bad, and the greedy: Evaluation of LLMs should not ignore non-determinism,
Y . Song, G. Wang, S. Li, and B. Y . Lin, “The good, the bad, and the greedy: Evaluation of LLMs should not ignore non-determinism,” in Proc. NAACL, 2025, pp. 4195–4206
2025
-
[25]
Quantifying variance in evaluation benchmarks,
L. Madaan, A. K. Singh, R. Schaeffer, A. Poulton, S. Koyejo, P. Stene- torp, S. Narang, and D. Hupkes, “Quantifying variance in evaluation benchmarks,”arXiv preprint arXiv:2406.10229, 2024
Pith/arXiv arXiv 2024
-
[26]
Reporting score distributions makes a difference: Performance study of LSTM-networks for sequence tagging,
N. Reimers and I. Gurevych, “Reporting score distributions makes a difference: Performance study of LSTM-networks for sequence tagging,” inProc. EMNLP. Association for Computational Linguistics, 2017, pp. 338–348
2017
-
[27]
Inference on winners,
I. Andrews, T. Kitagawa, and A. McCloskey, “Inference on winners,” The Quarterly Journal of Economics, vol. 139, no. 1, pp. 305–358, 2024, first circulated as NBER Working Paper 25456, 2019
2024
-
[28]
The optimizer’s curse: Skepticism and postdecision surprise in decision analysis,
J. E. Smith and R. L. Winkler, “The optimizer’s curse: Skepticism and postdecision surprise in decision analysis,”Management Science, vol. 52, no. 3, pp. 311–322, 2006
2006
-
[29]
Tweedie’s formula and selection bias,
B. Efron, “Tweedie’s formula and selection bias,”Journal of the Amer- ican Statistical Association, vol. 106, no. 496, pp. 1602–1614, 2011
2011
-
[30]
Valid post- selection inference,
R. Berk, L. Brown, A. Buja, K. Zhang, and L. Zhao, “Valid post- selection inference,”The Annals of Statistics, vol. 41, no. 2, pp. 802–837, 2013
2013
-
[31]
Double Q-learning,
H. van Hasselt, “Double Q-learning,” inAdvances in Neural Information Processing Systems (NeurIPS), vol. 23, 2010, pp. 2613–2621, analyzes the upward bias of the maximum over noisy value estimates
2010
-
[32]
When routing collapses: On the degenerate convergence of LLM routers,
G. Lai and H.-J. Ye, “When routing collapses: On the degenerate convergence of LLM routers,”arXiv preprint arXiv:2602.03478, 2026
arXiv 2026
-
[33]
Can large language models always solve easy problems if they can solve harder ones?
Z. Yang, Y . Zhang, T. Liu, J. Yang, J. Lin, C. Zhou, and Z. Sui, “Can large language models always solve easy problems if they can solve harder ones?” inProc. EMNLP, 2024, pp. 1531–1555
2024
-
[34]
Expected reward prediction, with applications to model routing,
K. Hasanaliyev, S. Alberti, J. Hamer, D. Rajagopal, K. Robinson, J. Snoek, V . Veitch, and A. N. D’Amour, “Expected reward prediction, with applications to model routing,”arXiv preprint arXiv:2603.20217, 2026
arXiv 2026
-
[35]
LLMs encode their failures: Predicting success from pre-generation activations,
W. Lugoloobi, T. Foster, W. Bankes, and C. Russell, “LLMs encode their failures: Predicting success from pre-generation activations,”arXiv preprint arXiv:2602.09924, 2026
Pith/arXiv arXiv 2026
-
[36]
Large language monkeys: Scaling inference compute with repeated sampling,
B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V . Le, C. Ré, and A. Mirhoseini, “Large language monkeys: Scaling inference compute with repeated sampling,”arXiv preprint arXiv:2407.21787, 2024
Pith/arXiv arXiv 2024
-
[37]
Self-consistency improves chain of thought reasoning in language models,
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdh- ery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” inProc. ICLR, 2023, arXiv:2203.11171
Pith/arXiv arXiv 2023
-
[38]
The limits of inference scaling through resampling,
B. Stroebl, S. Kapoor, and A. Narayanan, “The limits of inference scaling through resampling,” inProc. ICLR, 2026
2026
-
[39]
Let’s verify step by step,
H. Lightman, V . Kosaraju, Y . Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe, “Let’s verify step by step,” inProc. ICLR, 2024
2024
-
[40]
J. Chen, “When does combining language models help? a co-failure ceiling on routing, voting, and mixture-of-agents across 67 frontier models,”arXiv preprint arXiv:2606.27288, 2026
Pith/arXiv arXiv 2026
-
[41]
Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy,
L. I. Kuncheva and C. J. Whitaker, “Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy,”Machine Learning, vol. 51, no. 2, pp. 181–207, 2003
2003
-
[42]
Neural network ensembles, cross validation, and active learning,
A. Krogh and J. Vedelsby, “Neural network ensembles, cross validation, and active learning,” inAdvances in Neural Information Processing Systems (NIPS). MIT Press, 1994, pp. 231–238. PREPRINT, JULY 2026 29
1994
-
[43]
GPQA: A graduate-level google-proof Q&A benchmark,
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y . Pang, J. Dirani, J. Michael, and S. R. Bowman, “GPQA: A graduate-level google-proof Q&A benchmark,” inFirst Conference on Language Modeling (COLM), 2024, arXiv:2311.12022
Pith/arXiv arXiv 2024
-
[44]
Devroye, L
L. Devroye, L. Györfi, and G. Lugosi,A Probabilistic Theory of Pattern Recognition, ser. Stochastic Modelling and Applied Probability. New York: Springer, 1996, vol. 31, see Thm. 2.2: the plug-in (argmax) rule has excess risk at most twice theL 1 estimation error of the conditional probabilities
1996
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.