{"id":"97a6fda6-e339-4a41-8a87-b0ccda32d6fa","arxiv_id":"2606.05169","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Public LLM leaderboards have effective dimension ~3–5, so the geometric blind spot between models with identical scores exceeds runner-up gaps by ~100× and makes top rankings structurally unreliable.","lead":"LLM leaderboards only probe about 3–5 independent capability directions, so top-model rankings can flip under unobserved dimensions. The paper bounds that structural blind spot and shows it dwarfs ordinary score noise on real public suites.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The headline structural-indeterminacy claim rests on a width/linearization model that the paper itself places near its regime boundary on the frontier where the claim is made.","rationale":"The reader correctly isolates the width/linearization assumption as the softest load-bearing step for the central empirical claim. The multi-leaderboard d_eff pattern, half-split swaps, greedy coverage, and the Gardner-rate architecture are real contributions and do not collapse if the width residual is larger than hoped; what weakens is the specific stereological interpretation that the geometric blind spot (not ordinary projection/held-out variance) is what makes top rankings indeterminate by two orders of magnitude. The paper already flags the regime boundary and reports the frontier R^{2} drop, so the concern is internal, not external. A single residual-vs-Δ_2 audit on the Extended frontier would settle whether the 100\times ratios survive the authors’ own diagnostic. That keeps the verdict CONDITIONAL rather than moving it to REJECT or ACCEPT: the contribution is accept-shaped if the residual check passes or the claims are scoped to the linearizable core; it needs tightening if residuals are comparable to Δ_2 on the frontier. LiveBench CI and D-range issues are secondary relative to this mapping step.","tokens_in":51591,"tokens_out":859,"duration_ms":14074,"concrete_test":"On the Extended frontier, recompute per-benchmark quadratic-vs-linear R^{2} gaps and the support-function reconstruction residual (App. I H.15) in the same z-score units as Table 2’s Δ_2 and δ_vis_H. If for any of the 12 benchmarks (especially TruthfulQA or MATH Lvl 5) the residual or R^{2} gap is within a small constant of Δ_2 ≈ 0.072, recompute Table 2 after folding that residual into ε (or drop the width model for those columns). If the revised δ_vis_H / Δ_2 falls below ~10\times or loses dominance over bootstrap noise, the structural-indeterminacy headline weakens on the regime where it is claimed.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim is that on frontier leaderboards the visible Hausdorff radius δ_vis_H ≈ ε + C R m^{-1/(d_eff-1)} (or the practitioner form 2 R ω_emp) exceeds the runner-up gap Δ_2 by ~100\times, so top rankings are structurally indeterminate. That radius is a Hausdorff distance between convex bodies whose widths agree; it is only a bound on score-indistinguishability after Proposition 1 maps each benchmark to a width of the population convex hull, with residual η·diam^{2} absorbed into ε. The paper reports that on the frontier (top 50%, n=148) median linear R^{2} falls from 0.984 to 0.876 and the minimum to 0.710 (MATH Lvl 5), with TruthfulQA at 0.769 and the highest η; the explicit diagnostic is that the model is at its regime boundary when the quadratic-vs-linear R^{2} gap becomes comparable to the standardized Δ_2 (Setup; App. I H.15). The same frontier is where Δ_2 is tiny (0.072–0.17) and the 100\times ratios are reported (Table 2). If the residual is not negligible relative to Δ_2, the geometric radius no longer cleanly dominates the observed gap, and the half-split swap rates, while real, become evidence of ordinary held-out variance rather than of the stereological bound. Convexity is called conservative, but the rate and the practitioner formula still assume the width representation.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper develops a stereological theory of LLM benchmark coverage: benchmarks are treated as width-like measurements of convex capability profiles, and for a suite with effective dimensionality d_eff the visible Hausdorff distance between profiles consistent with the same scores is bounded by ε + C R m^{-1/(d_eff-1)} (Theorem 2), with a matching Lipschitz lower bound and a smooth C^{1,α}/C² extension. Empirically, three leaderboard families have frontier d_eff ∈ [2.86, 4.80]; the structural radius is reported to exceed runner-up gaps by ~10² and statistical noise by 52–127×; half-split experiments swap top-1 in 92% of trials; a submodular greedy algorithm with the Nemhauser guarantee finds a stable 4-benchmark core and 7/12 for 90% coverage with strong temporal transfer; counterfactual removal/addition correlates with eigenstructure (ρ = −0.69, +0.38). As a second contribution, the authors claim a resolution of Gardner’s Problem 1.5 for C² support functions via optimal recovery on S^{D−1}, with rate Θ(R/(κ m^{2/(D−1)})).","tokens_in":51973,"tokens_out":1949,"duration_ms":29311,"significance":"If the stereological bound and its empirical calibration hold, the paper supplies a quantitative, geometry-based account of why public LLM rankings are structurally underdetermined even with infinite data per benchmark—distinct from item noise or decoding variance—and a practical coverage algorithm with classical approximation guarantees. Strengths that should be credited: multi-suite empirics (Open LLM v2, extended 12-bench, LiveBench, Arena categories, Epoch AI) with bootstrap CIs, MP/permutation eigenvalue checks, and half-split swap counts; full proof appendices for Theorems 1–4 and the Gardner-rate argument via n-widths/Jackson–Bernstein; explicit regime diagnostics for linearization (Prop. 1, App. I H.15); and a reproducibility path (pip install -e .; reproduce.py). The independent geometric-tomography contribution, if correctly scoped, is of interest beyond ML. The work is therefore potentially high-impact for evaluation methodology, provided the width-model application on the competitive frontier is tightened.","major_comments":[{"comment":"Setup / Prop. 1 and Table 2: The headline claim that the structural blind spot exceeds Δ₂ by two orders of magnitude (e.g. 123× on OLLM v2 frontier, 385× on Extended) is load-bearing and rests on the width representation after linearization. The manuscript itself reports that on the frontier (top 50%, n=148) median linear R² falls from 0.984 to 0.876 and the minimum to 0.710 (MATH Lvl 5), with TruthfulQA at 0.769 and highest η, and states that the model is at its regime boundary when the quadratic-vs-linear R² gap is comparable to standardized Δ₂ (gap ≤0.067 vs Δ₂≈0.072 on Extended). The practitioner formula δ_vis_H≈2R ω_emp drops ε; if residual η·diam² is O(Δ₂), the geometric covering term no longer cleanly dominates the observed gap as a pure stereological effect. Please either (i) recompute Table 2 and the 52–127× noise comparisons with ε including the measured linearization residual","section":"Setup; Prop. 1; Table 2; App. I H.15"},{"comment":"Theorem 2 vs ranking claims (worked example, Corollary 1): δ_vis_H is a Hausdorff distance between convex bodies whose widths agree within ε on m directions in V_eff. The leap to “any ranking claim between models separated by less than 21.2 standardized units is structurally indeterminate” and to rank-equivalence classes (App. H.47) requires an explicit map from Hausdorff distance on profiles to aggregate score gaps under the actual aggregator. Half-split swap rates (92% top-1, mean 2.83/5 top-5) are strong and largely model-independent evidence of ranking instability; they should be cleanly separated from the stereological bound so that empirical instability does not silently substitute for the geometric theorem when the width residual is non-negligible.","section":"§4 Theorem 2; Worked example; Corollary 1; App. H.21"},{"comment":"Theorem 3 / App. G (Gardner’s Problem 1.5): The abstract and introduction claim a resolution of Gardner’s Problem 1.5 for C² support functions. Gardner’s original problem is stated for X-rays of planar convex bodies; the manuscript proves a minimax rate for width measurements and then argues X-ray parity for origin-symmetric bodies (Thm. 20 / App. G). That is a substantial and interesting contribution, but “resolution of Problem 1.5 as Gardner stated it” should be scoped precisely: state what is proved for widths, what is proved or conjectured for X-rays, and what remains open (exact constants, non-symmetric case, adaptive X-rays). Overclaiming the classical problem will invite specialist pushback and distract from the LLM results.","section":"Abstract; Theorem 3; App. E–G"},{"comment":"Corollary 1 / App. G.2–G.3: The χ² top-1 formula depends on ambient D, which three estimators place in a wide range [6,184]; the paper correctly notes robustness of P(swap) because Δ₂ is small, but the calibration table shows the bound is loose at small r (14–17 pp) and prior-sensitive (isotropic 0.79 vs empirical 0.29 in simulation). Present the χ² bound strictly as a rejection threshold / optimistic isotropic case (Prop. 2), not as a calibrated probability, and avoid headline intervals [0.38,0.49] without stating the prior and D assumptions in the same sentence.","section":"Corollary 1; Prop. 2; App. H G.1–G.4"}],"minor_comments":[{"comment":"LiveBench frontier n=19 is very small; H.39 notes d_eff=4.74 lies above the bootstrap CI upper bound and Horn finds 0 signal eigenvalues. Soften LiveBench as a third independent confirmation or report only full-slice d_eff=2.63 with CI.","section":"Table 1; App. I H.39"},{"comment":"NeurIPS checklist answers are filled [Yes] with “Justification: Not applicable,” which is inconsistent with the guidelines text. Replace with brief real justifications or N/A where appropriate.","section":"NeurIPS Paper Checklist"},{"comment":"Figure 3 uses private-use Unicode glyphs that render as boxes in the manuscript text; replace with standard PDF text or vector labels.","section":"Figure 3"},{"comment":"Notation: m is both “visible benchmarks” and generic sample size in covering bounds; d_eff is non-integer while covering is on S^{⌈d_eff⌉−1}—Remark 2 helps, but a single consistent convention in Theorem 2 would reduce confusion.","section":"§2 Notation; Theorem 2; Remark 2"},{"comment":"Related work: BenchScope [Sha and Zhao, 2026] is concurrent; keep the diagnostic/credit split clear and ensure arXiv identifiers/dates are stable at camera-ready.","section":"§1 Related work; §7"},{"comment":"H.12 reports geometric vs statistical radius 8.95× while the abstract/main text say 52–127×; reconcile the two definitions (πR/k vs 2R ω_emp / bootstrap radius) in one place.","section":"Table 2; App. I H.12"}],"recommendation":"major_revision","confidential_remarks":"The empirical core (low frontier d_eff, half-split rank instability, greedy coverage with transfer) is solid and would stand even if the stereological 100× ratios were softened. The main risk is overselling the geometric bound on the exact slice where linearization is weakest, plus a slightly expansive claim on Gardner 1.5. With the four major points addressed, this is a strong methods paper for a top ML venue; without them, specialist referees in convex geometry or evaluation will fixate on the regime-boundary issue. Fit is appropriate for cs.LG / evaluation methodology tracks."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful core is not “low d_eff” (already known and concurrent with BenchScope). It is the Lipschitz/smooth Hausdorff rate specialized to benchmark suites, the Nemhauser greedy with a stable 4-benchmark core and 93–97% quarterly transfer, the multi-prior swap envelope, and a serious secondary claim resolving Gardner 1.5 for C² support functions via n-widths on S^{D-1}. Those pieces are new and cleanly separated from the shared diagnostic.\n\nEmpirics are the paper’s strength. Three leaderboards plus Arena categories and Epoch AI, bootstrap CIs, MP/permutation checks, half-split swaps (92% top-1 flip), and counterfactual ρ = −0.69 / +0.38. Code and public data are there. The practitioner formula and rank-equivalence-class recommendation are immediately usable. Submodularity of tr(P_T Σ) is classical and correctly applied; the appendices give full proof architecture for Theorems 1–4 and the Gardner rate.\n\nThe stress-test lands, but only partially. Proposition 1 maps benchmarks to widths after linearization; residual η·diam² is absorbed into ε. On the same frontier where they report δ_vis_H / Δ_2 ~ 100×, median linear R² falls to ~0.876 (min 0.710) and TruthfulQA is worse. The authors state the diagnostic themselves: when the quadratic gap approaches Δ_2 the width model is at its regime boundary. So the geometric radius still dominates statistical noise, but the clean “structurally indeterminate by two orders of magnitude” claim is softer than the abstract. Half-split swaps remain real evidence of ranking fragility; they just do not uniquely prove the stereological bound. LiveBench n=19 CI anomaly, unidentified D, and standardization-sensitive ratios are secondary and mostly disclosed.\n\nWho it is for: people who design or report leaderboards, and anyone who wants a geometric language for evaluation coverage. It deserves a serious referee. I would engage, cite the coverage algorithm and the multi-suite d_eff numbers, and treat the 100× claim with the authors’ own regime caveat. Send to peer review; ask for tighter language on the width-model boundary and precise Gardner scope (widths, C²).","headline":"Solid multi-suite empirics and a usable coverage algorithm; the stereological bound is real geometry but the headline 100× indeterminacy sits on a width model the authors themselves flag as near its frontier limit.","tokens_in":52661,"tokens_out":649,"would_cite":true,"duration_ms":8664,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["52A20","62H25","68T50"],"pacs":[],"model":"grok-4.5","headline":"LLM leaderboard scores leave a geometric blind spot so large that top rankings are often structurally indeterminate.","keywords":["LLM evaluation","benchmark coverage","effective dimensionality","stereology","Hausdorff distance","submodular selection","Gardner problem","ranking reliability"],"falsifier":"On a new frontier leaderboard, compute d_eff and the empirical covering radius; if the resulting visible indistinguishability radius is smaller than the top-pair score gap, or if random half-splits of the suite almost never reverse the top ranking, the central claim that structural blind spots dominate is falsified for that population.","tokens_in":52378,"feed_emoji":"📐","tokens_out":772,"duration_ms":6533,"temperature":0.7,"pith_summary":"This paper treats public LLM leaderboard scores as one-dimensional projections of unknown multi-dimensional capability profiles, the same way X-rays project a body. Using stereology and convex-geometry tools, it proves that when a suite of benchmarks only probes a few independent directions, many different capability profiles remain consistent with the same scores. On three independent competitive leaderboards the effective dimensionality is only about three to five, so the structural gap between models that look identical on the suite dwarfs both the runner-up score gap and ordinary statistical noise by one to two orders of magnitude. Half-split experiments confirm the practical consequence: randomly holding out half the benchmarks swaps the top-ranked model in most trials. The paper then supplies a greedy algorithm that selects a small stable core of benchmarks covering most of the observed variance, and as a pure-math side result resolves a 1995 open problem on how stably widths recover a smooth convex body. A sympathetic reader cares because the work turns a familiar unease about leaderboard rankings into a quantitative, falsifiable geometric claim with a concrete remedy.","feed_headline":"LLM top rankings hide a geometric blind spot 50–100× noise","feed_subtitle":"Three leaderboards show d_eff ≈ 3–5; half-splits swap the #1 model in 92% of trials","key_machinery":"The indistinguishability bound (Theorem 2): after modelling each benchmark as a width (support-function) measurement of an origin-centred convex capability body, the Hausdorff distance that remains invisible to m directions in the effective subspace is controlled by the covering radius on the sphere of dimension d_eff−1, yielding the rate m^{−1/(d_eff−1)} with a matching Lipschitz lower bound.","core_discovery":"For any benchmark suite whose effective dimensionality is d_eff, the visible Hausdorff distance between two convex capability profiles that produce identical scores is at most ε plus a term that shrinks only as m to the power −1/(d_eff−1). Empirically d_eff lies in [2.86, 4.80] on three frontier leaderboards, so this structural blind spot exceeds the observed runner-up gap by roughly two orders of magnitude and dominates statistical noise by 52–127×; half-split trials swap the top model in 92 percent of draws.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["LLM rankings hide geometric blind spot 50-100x larger than noise","d_eff of 3-5 leaves Hausdorff gaps two orders past runner-up scores","Half-split trials flip top LLM in 92% of draws on frontier benches","Structural coverage hole dwarfs score gaps and noise by 52-127x","Four-benchmark core covers 90%; half-splits reshuffle top-5 often"],"cache_read_input_tokens":49280,"weakest_assumption_plain":"Each benchmark must be well approximated by a linear width measurement of a convex capability profile; when models become very similar or a task is highly non-linear, that linearisation residual can grow as large as ordinary score gaps and the geometric bound no longer cleanly separates structure from noise.","fun_headline_variants_meta":{"raw":{"variants":["LLM rankings hide geometric blind spot 50-100x larger than noise","d_eff of 3-5 leaves Hausdorff gaps two orders past runner-up scores","Half-split trials flip top LLM in 92% of draws on frontier benches","Structural coverage hole dwarfs score gaps and noise by 52-127x","Four-benchmark core covers 90%; half-splits reshuffle top-5 often"]},"model":"grok-4.5","effort":"low","cost_usd":0.003672,"raw_usage":{"total_tokens":1330,"prompt_tokens":981,"num_sources_used":0,"completion_tokens":90,"cost_in_usd_ticks":36720000,"prompt_tokens_details":{"text_tokens":981,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":259,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":981,"tokens_out":90,"duration_ms":3489,"temperature":1.0,"reasoning_tokens":259,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T20:42:17.322234+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a new frontier leaderboard, compute d_eff and the empirical covering radius; if the resulting visible indistinguishability radius is smaller than the top-pair score gap, or if random half-splits of the suite almost never reverse the top ranking, the central claim that structural blind spots dominate is falsified for that population.","supporting_citations":[],"review_version":1}