Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

This paper constructs anytime-valid confidence sequences for an LLM's benchmark accuracy, proving the RIPr-based construction is growth-optimal among valid e-values and identifying prediction mismatch and sampling spikiness as the two facto

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 18:02 UTC pith:TIWN2SGF

load-bearing objection New problem setup and honest empirical work, but the method in the experiments is an acknowledged heuristic CS, so the central anytime-valid claim is not yet established for what is actually run. the 3 major comments →

arxiv 2607.17409 v1 pith:TIWN2SGF submitted 2026-07-19 stat.ML cs.LGstat.ME

Efficient Sequential Evaluation of Large Language Models

classification stat.ML cs.LGstat.ME MSC 62L1262F2560G42
keywords confidence sequenceanytime-valid inferenceactive queryingLLM evaluationreverse information projectiontesting-by-bettinge-valuesadaptive sampling
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Sequential evaluation of a large language model on a fixed question bank is treated as an active statistical inference problem with a time-uniform coverage guarantee. The paper constructs confidence sequences by inverting test supermartingales and studies two conditional e-value constructions, reverse information projection (RIPr) and testing-by-betting, under adaptive querying. In the oracle setting — true question-level correctness probabilities known — the RIPr construction is shown to be growth-optimal among valid e-values for any fixed query rule. In practice the true probabilities are replaced by question-level predictions learned from historical data, which preserves anytime-valid coverage; the shrinkage analysis identifies accumulated prediction mismatch and spikiness of the querying distribution as the two factors that slow interval narrowing. Experiments across synthetic benchmarks show that no single querying rule dominates, and the simplest rule, uniform sampling, is often competitive.

Core claim

The central claim is that a model-free level-α confidence sequence for the average correctness of a new LLM on a fixed question set can be constructed even when questions are chosen adaptively: at every stopping time the realized benchmark accuracy lies in the confidence set with probability at least 1−α, regardless of whether the question-level predictions are accurate. The proof routes through conditional e-values and supermartingales: for each candidate mean m, an e-value is built from the density ratio between the predicted distribution and the reverse information projection onto the null class of vectors with mean m (RIPr), or from an AIPW-type estimator combined with a predictable bet

What carries the argument

Test supermartingales and conditional e-values for each candidate mean m, with the confidence sequence defined as the set of m whose e-process has not crossed 1/α. Two e-value constructions carry the argument: the RIPr e-value, the density ratio of the (predicted) observation distribution to its reverse information projection onto the null simplex {p: mean = m}, which is the least-favorable null closest under KL divergence and is growth-optimal; and the betting e-value, 1 + λ(Φ − m), where Φ is an AIPW-style estimator and λ is constrained to an interval that depends on the minimum query probability h. A growth-oriented query rule selects the sampling distribution maximizing the minimum one-s

Load-bearing premise

The load-bearing premise is that the practical grid-based confidence sequence inherits the validity of the underlying e-values, yet the paper explicitly states this heuristic 'is not in general a valid level-α CS,' leaving the experimental procedure without a proven anytime-valid guarantee.

What would settle it

Run the practical grid/hull algorithm with plug-in predictions on a benchmark with a fixed binary correctness vector, stop at the adaptive time when the interval width first falls below a target, and check coverage of the realized accuracy across many repetitions; if the observed miss rate exceeds the nominal α for a reasonable prediction model, the claimed anytime-validity for the practical procedure is falsified.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • An evaluator can query questions one at a time, watch the confidence interval shrink, and stop at an adaptive stopping time while retaining time-uniform coverage.
  • The RIPr construction provides a benchmark for query design: among valid e-values it achieves maximal expected log-growth, so the active-query problem reduces to choosing the sampling distribution for that e-value.
  • Prediction mismatch and query-distribution spikiness are the two controllable factors that slow shrinkage; mixture rules that refine predictions and keep the sampling distribution spread out mitigate both.
  • The testing-by-betting approach is less sensitive to prediction error than RIPr but becomes conservative when the query rule assigns very low probability to some questions.
  • Uniform sampling, despite making no use of predictions or adaptation, is a competitive baseline and sometimes outperforms sophisticated rules.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The paper's validity proofs cover the idealized procedure, while the grid/hull version used in all experiments is acknowledged in Appendix A.1 not to be a valid confidence sequence in general; the empirical stopping-time comparisons therefore rest on a heuristic whose coverage is only approximately checked.
  • The theory models correctness as independent Bernoulli variables, but the oracle experiments condition on a fixed binary correctness vector; under that degenerate law the KL/MSE mismatch terms in the width theorems are undefined or infinite, so the shrinkage-rate theorems as stated do not directly describe the evaluated protocol.
  • The observed uniform-sampling competitiveness suggests that when predictions are poor or the benchmark is near-homogeneous, the gains from active querying are small; a testable extension is to compare the proposed rules under real benchmark conditions with human-labeled correctness rather than synthetic factor-model draws, where prediction quality varies.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies sequential evaluation of a new LLM's average correctness on a fixed question bank, using historical predictions from prior LLMs. It constructs confidence sequences (CSs) by inverting test supermartingales built from two families of conditional e-values: a reverse-information-projection (RIPr) likelihood ratio and a testing-by-betting AIPW-type factor. Under an oracle with known question-level probabilities p*, it proves conditional e-value validity and RIPr growth optimality; it then replaces p* by a predictable prediction p̂ and argues validity is preserved. The authors propose a maxmin growth-oriented querying rule, derive 'approximated width' bounds involving KL/MSE prediction mismatch and the minimum query probability, design mixture querying rules, and compare them in synthetic experiments. A central empirical finding is that uniform sampling is often competitive.

Significance. The sound components are genuinely useful: Propositions 4.1 and 4.2 establish simple conditional e-value constructions under predictable adaptive querying, and the plug-in observation in Section 5.1 that any predictable prediction preserves validity is correct. The oracle optimality of the RIPr construction is a known theorem and is reproduced appropriately. These pieces provide a clean framework for anytime-valid active benchmarking. However, the paper's headline claim—an implemented model-free anytime-valid CS with guaranteed shrinkage—is not established: the algorithm actually used in all experiments is explicitly a heuristic (Appendix A.1), and the width theorems are informal, contain uncontrolled terms, and do not match the evaluated betting implementation. The paper is therefore best read as a promising framework with interesting experiments, not as a finished method with proven guarantees.

major comments (3)
  1. [Appendix A.1, Algorithm 1] This appendix states that the algorithm used in every experiment is 'not in general a valid level-α CS'. Line 11 replaces the valid continuum set {m: W_t(m)<1/α} by the interval [L_t,U_t] formed from the hull of surviving grid points. Since W_t is not shown to be unimodal and no discretization or mesh bound is given, this interval can exclude the true θ* even when W_t(θ*)<1/α, for instance if all surviving grid points lie on one side of θ*. Consequently the anytime-valid coverage asserted in the abstract and Section 1 is not established for the evaluated algorithm, and the stopping-time results in Figures 1–7 concern this heuristic interval. The paper must either prove a validity-preserving discretization/enlargement, or change the experiments to use a demonstrably valid algorithm.
  2. [Section 5.2, Theorems 5.1/5.2; Remarks C.2/C.4] The shrinkage-rate results are explicitly informal. The formal versions in Appendices C.3/C.4 still contain the uncontrolled variance terms σ_RIPr and σ_bet in the numerator, and Remarks C.2 and C.4 state that the bounds 'may be too coarse' and 'will revise this issue'. Moreover, Theorem C.3 assumes λ_t(m)∈1/2Λ_h(m), whereas the implemented Betting-UpdateQueryPmf optimizes over the full Λ_h(m). Thus the betting width theorem does not describe the method evaluated in Figures 2–7. The qualitative conclusions about prediction mismatch and spikiness slowing shrinkage are plausible, but they are not rigorous consequences of the stated theorems. The theorems should either be tightened, or relabeled as diagnostic heuristics with assumptions matching the implementation.
  3. [Appendix A.2] The theoretical setting and Theorems 5.1/5.2 assume independent Z_i∼Ber(p*_i) with p*∈(0,1)^N, and the KL/MSE mismatch terms are defined with respect to that law. The experimental protocol instead conditions on a fixed binary correctness vector A^{(b)} and defines the target as its empirical mean. Under the degenerate law of A, the divergence D(Q_{p*,q_s}||Q_{p̂_{s-1},q_s}) and the MSE term used in the proofs are infinite or undefined, so Theorems 5.1/5.2 do not cover the evaluated fixed-benchmark protocol. The coverage-extension argument for boundary A is reasonable for the conditional e-value property, but the shrinkage-rate analysis does not transfer. The authors should either run experiments with independent Bernoulli draws per replication or extend the width theory to degenerate benchmark vectors.
minor comments (5)
  1. [Appendix A.4] The theory curves in Figures 8 and 9 are fitted to the empirical widths by choosing a calibration constant c that minimizes squared error. Without a clear caveat, these figures may be read as parameter-free validation of the predicted rates. Please state explicitly that c is fitted and that the comparison is qualitative.
  2. [Figures 4, 6, 7] Coverage curves overlap heavily, and the practical setting uses only B=8 ground-truth realizations and R=2 querying runs, i.e., 16 repetitions. The 95% binomial uncertainty interval is roughly ±0.21, so 'consistent with 0.95 up to Monte Carlo error' is not informative. Use separate panels and report confidence bands or more repetitions.
  3. [Theorems 5.1/5.2] The main-text statements mix big-O notation with random sets, e.g., |C_t∩M_K|∈O(...). Please state the high-probability event explicitly, as in the formal versions C.1/C.3.
  4. [General] Minor wording and notation issues: 'with probability at least probability 1−δ', 'the proof of Theorem 5.1 is proved', inconsistent 'Mk' vs 'M_K', and remarks in C.2/C.4 that refer to planned revisions should be cleaned before resubmission.
  5. [Section 5.1] The update equations for the prediction p̂_t are omitted and deferred to [2]. Since the prediction update is part of the proposed method, please include at least the explicit pseudo-code or a precise equation reference to [2] for reproducibility.

Circularity Check

0 steps flagged

No significant circularity: the CS construction and plug-in extension are self-contained, and the paper's own stated limitations are not circular steps.

full rationale

The central derivation is not circular. Propositions 4.1 and 4.2 establish conditional e-value validity for any predictable reference vector, and Section 5.1 correctly observes that replacing p* with an F_{t-1}-measurable prediction preserves the proof; the proof of Proposition 4.1 explicitly states that p* only needs to be predictable and in (0,1). The RIPr growth optimality is quoted from external references [21,22] and also proved for the proposition, while the self-citations ([8],[12]) are contextual rather than load-bearing. The shrinkage theorems (5.1, 5.2) are explicitly labeled as diagnostic bounds in Remark 5.3, and Remarks C.2 and C.4 acknowledge they are loose; the calibration constant in Appendix A.4 is for plotting only and does not enter the derivation. Appendix A.1 explicitly states that the grid/hull algorithm 'is not in general a valid level-α CS'—this is a candid correctness limitation of the experiments, not a circular argument. No fitted parameter is renamed as a prediction, and no claimed prediction is forced by construction.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 0 invented entities

The central validity result depends only on standard e-value theory and predictability of the plug-in predictions; no new entity or physical constant is introduced. The efficiency story rests on heuristic criteria, hand-set hyperparameters, and loose diagnostic bounds that the authors explicitly defer revising.

free parameters (6)
  • Grid size K = 200 (oracle), 400 (practical)
    CS and stopping times are computed only on this finite grid; K is chosen by hand and controls the discretization error of the heuristic CS hull.
  • Query-smoothing weight ζ = 0.05
    Used in RIPr/Betting-UpdateQueryPmf to guarantee min_i q_t(i) ≥ ζ/N (lower bound h); hand-set; also ζ=1 gives uniform.
  • Mixture schedule hyperparameters = ρ_init=0.75, a=1.5, β_target=0.7, l=10, m_0=0.7
    Define the time-scheduled and data-dependent mixture querying rules; authors state the benefit of the growth component 'has not been theoretically analyzed.'
  • EG/bisection iteration counts = R_EG=25, B_bet=40; η=1/(2√i) or 0.5/√i
    Approximate the maxmin growth rule; no convergence guarantee to the argmax in (1).
  • Theory-curve calibration constant c = fit via argmin Σ(avg_width − c·raw_bound)²
    Appendix A.4 scales theoretical width bounds to match empirical curves; a post-hoc fit for visualization.
  • UniformPreference tolerance δ = predefined (details in code)
    Used in data-dependent mixture UniformPreference score; hyperparameter whose value is deferred to unlinked code.
axioms (6)
  • standard math Standard e-value machinery: Ville's inequality, conditional e-value products form test supermartingales, Freedman's inequality.
    Invoked in Section 3 and Lemmas B.4–B.6 without proof; unproved background.
  • domain assumption Z_i ~ Ber(p*_i) independently across questions i∈[N], with p*_i∈(0,1).
    Section 1 model for the correctness of the new LLM; used in oracle analysis and in the width theorems.
  • domain assumption Historical responses are fully observed, and the Bayesian factor model of [2] yields predictable predictions p-hat_t ∈ (0,1)^N after a Laplace update.
    Section 5.1 replaces p* with p-hat_t; update equations are omitted and referenced to [2].
  • domain assumption Experiments condition on a fixed realized binary vector A~(Ber(p*)) with target θ_A, while oracle querying uses p*.
    Appendix A.2 fixed-benchmark protocol; validity is argued via e-value validity for all p∈P_m, but the width theorems' KL/MSE terms are undefined for degenerate A.
  • ad hoc to paper The maxmin growth criterion over current CS endpoints is a suitable surrogate for CS-width shrinkage.
    Eq. (1); the paper states no single querying rule is optimal for all test supermartingales and provides no theorem linking this criterion to minimal CS width.
  • ad hoc to paper Regularity conditions for Theorems 5.1/5.2: bounded log-e-values and an 'oracle regularized' bet that stays in ½Λ_h(m).
    Stated informally in Section 5.2; the authors call the bounds too coarse and announce revisions (Remarks C.2, C.4).

pith-pipeline@v1.3.0-alltime-deepseek · 31282 in / 21469 out tokens · 196449 ms · 2026-08-01T18:02:28.385152+00:00 · methodology

0 comments
read the original abstract

We study the problem of sequentially evaluating a new large language model (LLM) on a fixed question set using historical performance data from prior LLMs. Our goal is to construct a confidence sequence (CS) for the model's capability on this question set and to design active querying rules that shrink the CS width as quickly as possible. For CS construction, we invert a family of test supermartingales and focus on two representative approaches: a reverse information projection (RIPr)-based approach and a testing-by-betting-based approach. We first study these approaches under an oracle setting, and demonstrate the oracle optimality of the RIPr-based construction. We then propose a growth-oriented querying rule that aims to maximize the worst-case one-step expected log-increment over the endpoints of the current CS. In practice, we build these test supermartingales and the querying rule on predictions of question-level correctness learned from historical data. We then analyze the shrinkage behavior of the resulting CSs and identify two key factors that slow the shrinkage rate of CSs: accumulated prediction mismatch and the spikiness of the querying distribution. Finally, motivated by this analysis, we propose several mixture querying rules that combine growth-oriented querying, prediction refinement, and uniform exploration, trying to mitigate the effects that slow the shrinkage rate. We provide experiments comparing different querying rules for the RIPr-based and testing-by-betting-based CSs across several synthetic testing datasets. Interestingly, we observe that the simplest querying rule, uniform sampling, can sometimes outperform more adaptive querying rules for both methods.

Figures

Figures reproduced from arXiv: 2607.17409 by Chia-Yu Hsu, Shubhanshu Shekhar.

Figure 1
Figure 1. Figure 1: Average stopping time versus the prescribed accuracy levels [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Average stopping time versus ϵ for RIPr-based methods across five synthetic regimes: train like, which is close to the training distribution; mean shift mismatch and mean shift mismatch high accuracy, which have shifted overall accuracy levels; out of family model mismatch, which exhibits structural mis￾match from the training family; and normal regime, which is a handcrafted heterogeneous setting with a s… view at source ↗
Figure 3
Figure 3. Figure 3: Average stopping time versus ϵ for testing-by-betting-based methods across five synthetic regimes described in [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Oracle setting: anytime coverage rate for each target accuracy [PITH_FULL_IMAGE:figures/full_fig_p019_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Average CS width over time. It is obtained in the same experiments conducted in Figure 1. [PITH_FULL_IMAGE:figures/full_fig_p019_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Practical setting: anytime coverage rate for each target accuracy [PITH_FULL_IMAGE:figures/full_fig_p023_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Practical setting: anytime coverage rate for each target accuracy [PITH_FULL_IMAGE:figures/full_fig_p024_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Plug-in RIPr-based CS: Average CS width over time for different values of [PITH_FULL_IMAGE:figures/full_fig_p025_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Plug-in testing-by-betting-based CS: Average CS width over time for different values of [PITH_FULL_IMAGE:figures/full_fig_p025_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. BayesAME: Bayesian Active Model Evaluation

    cs.LG 2026-07 conditional novelty 6.0

    A sequential Bayesian method automatically grows a coreset until performance estimate and uncertainty stabilize, outperforming adapted baselines and showing active selection beats random when reference signals are rich.

Reference graph

Works this paper leans on

30 extracted references · 16 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Evaluating large language models: A comprehensive survey,

    Z. Guo, R. Jin, C. Liu, Y. Huang, D. Shi, Supryadi, L. Yu, Y. Liu, J. Li, B. Xiong, and D. Xiong, “Evaluating large language models: A comprehensive survey,” 2023. [Online]. Available: https://arxiv.org/abs/2310.19736

  2. [2]

    Efficient evaluation of llm performance with statistical guarantees,

    S. Wu, Y. Nair, and E. J. Cand` es, “Efficient evaluation of llm performance with statistical guarantees,”

  3. [3]

    Active statistical inference,

    T. Zrnic and E. J. Cand` es, “Active statistical inference,” 2026. [Online]. Available: https: //arxiv.org/abs/2403.03208

  4. [4]

    Confidence sequences for mean, variance, and median,

    D. A. Darling and H. Robbins, “Confidence sequences for mean, variance, and median,”Proceedings of the National Academy of Sciences of the United States of America, vol. 58, no. 1, pp. 66–68, July 1967. [Online]. Available: https://doi.org/10.1073/pnas.58.1.66

  5. [5]

    Estimating means of bounded random variables by betting,

    I. Waudby-Smith and A. Ramdas, “Estimating means of bounded random variables by betting,” Journal of the Royal Statistical Society: Series B (Statistical Methodology), vol. 86, no. 1, pp. 1–27, February 2024. [Online]. Available: https://doi.org/10.1093/jrsssb/qkad009

  6. [6]

    Time-uniform, nonparametric, nonasymptotic confidence sequences,

    S. R. Howard, A. Ramdas, J. McAuliffe, and J. Sekhon, “Time-uniform, nonparametric, nonasymptotic confidence sequences,”Annals of Statistics, vol. 49, no. 2, pp. 1055–1080, April 2021. [Online]. Available: https://doi.org/10.1214/20-AOS1991

  7. [7]

    Time-uniform confidence spheres for means of random vectors,

    B. Chugg, H. Wang, and A. Ramdas, “Time-uniform confidence spheres for means of random vectors,”

  8. [8]

    On the near-optimality of betting confidence sets for bounded means,

    S. Shekhar and A. Ramdas, “On the near-optimality of betting confidence sets for bounded means,”

  9. [9]

    Safe testing,

    P. Gr¨ unwald, R. de Heide, and W. Koolen, “Safe testing,” 2023. [Online]. Available: https://arxiv.org/abs/1906.07801

  10. [10]

    Game-theoretic statistics and safe anytime-valid inference,

    A. Ramdas, P. Gr¨ unwald, V. Vovk, and G. Shafer, “Game-theoretic statistics and safe anytime-valid inference,” 2023. [Online]. Available: https://arxiv.org/abs/2210.01948

  11. [11]

    Exact anytime-valid confidence intervals for contingency tables and beyond,

    R. Turner and P. Gr¨ unwald, “Exact anytime-valid confidence intervals for contingency tables and beyond,” 2022. [Online]. Available: https://arxiv.org/abs/2203.09785

  12. [12]

    Risk-limiting financial audits via weighted sampling without replacement,

    S. Shekhar, Z. Xu, Z. Lipton, P. Liang, and A. Ramdas, “Risk-limiting financial audits via weighted sampling without replacement,” inProceedings of the Thirty-Ninth Conference on Uncertainty in Artificial Intelligence, ser. Proceedings of Machine Learning Research, R. J. Evans and I. Shpitser, Eds., vol. 216. PMLR, 31 Jul–04 Aug 2023, pp. 1932–1941. [Onli...

  13. [13]

    Design-based confidence sequences: A general approach to risk mitigation in online experimentation,

    D. W. Ham, I. Bojinov, M. Lindon, and M. Tingley, “Design-based confidence sequences: A general approach to risk mitigation in online experimentation,” 2023. [Online]. Available: https://arxiv.org/abs/2210.08639

  14. [14]

    Semiparametric efficient inference in adaptive experiments,

    T. Cook, A. Mishler, and A. Ramdas, “Semiparametric efficient inference in adaptive experiments,”

  15. [15]

    Anytime-valid off-policy inference for contextual bandits,

    I. Waudby-Smith, L. Wu, A. Ramdas, N. Karampatziakis, and P. Mineiro, “Anytime-valid off-policy inference for contextual bandits,” 2024. [Online]. Available: https://arxiv.org/abs/2210.10768

  16. [16]

    Valid best-model identification for llm evaluation via low-rank factorization,

    E. Tolochinsky, Y. Tenzer, and Y. Romano, “Valid best-model identification for llm evaluation via low-rank factorization,” 2026. [Online]. Available: https://arxiv.org/abs/2605.10405

  17. [17]

    Prediction-powered inference,

    A. N. Angelopoulos, S. Bates, C. Fannjiang, M. I. Jordan, and T. Zrnic, “Prediction-powered inference,” 2023. [Online]. Available: https://arxiv.org/abs/2301.09633

  18. [18]

    Revisiting active sequential prediction-powered mean estimation,

    M.-E. Sfyraki and J.-K. Wang, “Revisiting active sequential prediction-powered mean estimation,”

  19. [19]

    A new interpretation of information rate,

    J. L. Kelly, “A new interpretation of information rate,”the bell system technical journal, vol. 35, no. 4, pp. 917–926, 1956. [Online]. Available: https://www.princeton.edu/ ∼wbialek/rome/refs/kelly 56.pdf

  20. [20]

    Optimal gambling systems for favorable games,

    L. Breiman, “Optimal gambling systems for favorable games,”Kelly Capital Growth Investment Criterion, The: Theory And Practice, vol. 3, p. 47, 2011. [Online]. Available: https: //projecteuclid.org/proceedings/berkeley-symposium-on-mathematical-statistics-and-probability/ Proceedings-of-the-Fourth-Berkeley-Symposium-on-Mathematical-Statistics-and/Chapter/ ...

  21. [21]

    Hypothesis testing with e-values,

    A. Ramdas and R. Wang, “Hypothesis testing with e-values,”Foundations and Trends®in Statistics, vol. 1, no. 1-2, p. 1–390, 2025. [Online]. Available: http://dx.doi.org/10.1561/3600000002

  22. [22]

    The numeraire e-variable and reverse information projection,

    M. Larsson, A. Ramdas, and J. Ruf, “The numeraire e-variable and reverse information projection,”

  23. [23]

    Available: https://arxiv.org/abs/2604.18569 11

    [Online]. Available: https://arxiv.org/abs/2604.18569 11

  24. [24]

    Collabeval: Statistically efficient collaborative model evaluation via matrix completion,

    A. Fisch, D. Deutsch, J. Maynez, A. Agarwal, J. Berant, W. Cohen, A. Globerson, and J. Eisenstein, “Collabeval: Statistically efficient collaborative model evaluation via matrix completion,” 2026. [Online]. Available: https://arxiv.org/abs/2607.05046 12 A Experiment details and supplemental results A.1 A heuristic algorithm of the CS We present aheuristic...

  25. [28]

    Available: https://arxiv.org/abs/2402.18810

    [Online]. Available: https://arxiv.org/abs/2402.18810

  26. [29]

    On tail probabilities for martingales,

    D. A. Freedman, “On tail probabilities for martingales,”the Annals of Probability, pp. 100–118, 1975

  27. [2023]

    [Online]. Available: https://arxiv.org/abs/2310.01547 10 0.06 0.08 0.10 0.12 0.14 epsilon 150 200 250 300 350 400 450 500Average stopping time mean_shift_mismatch 0.06 0.08 0.10 0.12 0.14 epsilon 300 350 400 450 500Average stopping time mean_shift_mismatch_high_accuracy 0.06 0.08 0.10 0.12 0.14 epsilon 420 440 460 480 500Average stopping time normal_regim...

  28. [2024]

    Available: https://arxiv.org/abs/2311.18274

    [Online]. Available: https://arxiv.org/abs/2311.18274

  29. [2025]

    Available: https://arxiv.org/abs/2311.08168

    [Online]. Available: https://arxiv.org/abs/2311.08168

  30. [2026]

    Available: https://arxiv.org/abs/2601.20251

    [Online]. Available: https://arxiv.org/abs/2601.20251