Pith. sign in

REVIEW 2 major objections 4 minor 15 references

Supermartingales for One-Sided Tests: Sufficient Monotone Likelihood Ratios are Sufficient

T0 review · 2 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read One-sided t-tests are anytime-valid after all.

desk verdict A short, correct paper that resolves the Wang–Ramdas open question by proving a sufficient-statistic/MLR condition, with honest examples and a useful counterexample. read the letter →

arxiv 2502.04208 v2 pith:CKIL56GK submitted 2025-02-06 math.ST stat.TH

classification math.STstat.TH MSC 62L1062F0362B05
keywords testsupermartingalee-processmonotonelikelihoodratiosufficientstatisticanytime-validinferencet-testone-sidedhypothesis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proves a general theorem: if at every sample size $n$ a one-dimensional sufficient statistic $T_n$ for the coarsened model has the monotone likelihood ratio property, then the likelihood ratio process $(p^{U^n}_{\delta_+})_n$ is a test supermartingale for the one-sided null $H_{\delta \leq \delta_0}$. Applied to the scale-invariant $t$-test, this answers a previously open question in the affirmative: the $t$-likelihood ratio process is an e-process for the one-sided null, so sequential one-sided $t$-tests remain valid under optional stopping even when the true mean is negative. The same argument covers the location-invariant $\chi^2$-test, linear regression with nuisance covariates, and a label-agnostic Bernoulli test. The proof is short once the right object is examined: the conditional density of the sufficient statistic $T_n$, not of the raw outcome $U_n$.

What carries the argument

The load-bearing machinery is Lemma 3, which shows that sufficiency forces the conditional likelihood ratio of the raw outcome to equal the conditional likelihood ratio of the sufficient statistic, $p^{U^n}_{\delta_+}(u^n \mid u^{n-1}) = p^{T_n}_{\delta_+}(t_n(u^n) \mid u^{n-1})$, and that the latter is non-decreasing in $t$ under the monotone likelihood ratio property. This lets the proof apply a known fixed-sample fact pointwise, conditional on the past: a monotone likelihood ratio likelihood ratio has expectation at most $1$ under every $\delta \leq \delta_0$. The resulting past-conditional e-variables multiply into a supermartingale. For the $t$-test, the sufficiency of the $t$-statistic and the monotone likelihood ratio property of noncentral $t$ densities supply the two ingredients.

What would settle it

Construct a coarsened model that has a one-dimensional sufficient statistic with the monotone likelihood ratio property at every $n$, simulate data from any $\delta < \delta_0$, and estimate $\mathbb{E}_\delta[p^{U^n}_{\delta_+}(U^n)]$ for increasing $n$; the theorem predicts each value is at most $1$, so any estimate reliably above $1$ would refute the central claim.

Watch

Extended reading notes

Core claim

Theorem 4 is the central claim: let $T_n = t_n(U^n)$ be a sufficient statistic for the coarsened model $\{P^{U^n}_{\delta} : \delta \in \Delta\}$, and suppose $T_n$ satisfies the monotone likelihood ratio property for every $n$. Then the process formed by multiplying the conditional likelihood ratios $p^{T_i}_{\delta_+}(T_i \mid U^{i-1})$ for $i = 1, \ldots, n$ is identical to the likelihood ratio process $(p^{U^n}_{\delta_+})_n$, and both are test supermartingales relative to the one-sided null $H_{\delta \leq \delta_0}$. For the anytime-valid $t$-test, the theorem resolves the open question about the one-sided null: because the $t$-statistic is sufficient for the maximal invariants and noncentral $t$ densities have monotone likelihood ratios, the $t$-likelihood ratio process is a supermartingale for $H_{\delta \leq \delta_0}$. The fixed-sample e-variable result then extends to the whole process by conditioning on the past and multiplying.

Load-bearing premise

The argument assumes that at every sample size $n$ the coarsened model admits a one-dimensional sufficient statistic $T_n$ whose likelihood ratio is monotone in $T_n$; if no such sufficient statistic exists, the paper does not establish the supermartingale property.

Editorial extensions

If this is right

  • The anytime-valid $t$-test for the one-sided null $H_{\delta \leq \delta_0}$ is a genuine e-process: practitioners may stop at arbitrary data-dependent times and retain type-I error control.
  • The same conclusion holds when a prior $W$ on the alternative replaces the point mass at $\delta_+$, giving log-optimal anytime-valid e-values for the one-sided null.
  • The location-invariant $\chi^2$-test and the linear-regression likelihood ratio with nuisance covariates are supermartingales under their respective one-sided nulls.
  • The argument is not a blank cheque: a zero-mean sub-Gaussian (symmetric Bernoulli) example shows the $t$-test likelihood ratio can have expectation greater than $1$ at fixed $n$, so the sufficient-statistic MLR condition is doing real work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • For the designer of new anytime-valid tests, the proof suggests a practical recipe: find a one-dimensional sufficient statistic for the coarsened invariants and verify its MLR property, instead of trying to prove monotonicity of the raw conditional likelihood ratio, which often fails.
  • The argument likely transfers to any group-invariant testing problem whose maximal invariants admit a one-dimensional sufficient statistic with MLR; spherical and elliptical families are natural candidates for new e-processes.
  • A direct extension would replace the one-dimensional statistic by a vector sufficient statistic with a suitable stochastic ordering; success there would apply to multivariate analogues of the $t$-test.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This paper studies sequential testing of one-sided composite null hypotheses of the form H_{δ≤δ0} using likelihood ratio processes built from coarsened data (e.g., maximal invariants). The main result, Lemma 3 and Theorem 4, states that if for every sample size n there exists a one-dimensional sufficient statistic T_n for the coarsened model satisfying the monotone likelihood ratio (MLR) property, then the likelihood ratio process (p^{U^n}_{δ+}) is a test supermartingale relative to H_{δ≤δ0}. The key step shows that the conditional likelihood ratio of the next coarsened outcome given the past is equal to a conditional likelihood ratio of the sufficient statistic, which is increasing in the sufficient statistic by the MLR assumption. Applications are developed for the scale-invariant t-test (answering an open question of Wang and Ramdas 2024), a χ²-test for variances with translation nuisance, sequential linear regression with nuisance covariates, and a label-agnostic Bernoulli test. An appendix demonstrates that the theorem's conditions are not vacuous by giving a counterexample where the t-likelihood ratio fails to be an e-variable under Rademacher data.

Significance. The paper gives a simple, general, and checkable sufficient condition for converting one-sided parametric likelihood ratio tests into anytime-valid e-processes. The t-test application is particularly valuable, as it resolves a previously open question about whether the scale-invariant t-likelihood ratio process controls the one-sided composite null. The proof is elementary and transparent, and the authors are explicit about the condition being load-bearing. The applications cover several canonical settings, and each verifies the required MLR and sufficiency assumptions. The counterexample in Appendix C strengthens the paper by delimiting the scope of the theorem. The paper is honest about the simplicity of the proof and gives proper credit to prior work.

major comments (2)
  1. [Section 2, Theorem 4] The statement of Theorem 4 defines the two processes as (∏_{i=1}^n p^{T_i}_δ(T_i | U^{i-1}))_{n∈N} and (p^{U^n}_δ)_{n∈N}, but the supermartingale claim relative to H_{δ≤δ0} requires these processes to be built from an alternative δ+ ≥ δ0, not from the generic parameter δ. As written, the theorem asserts that the likelihood ratio process (p^{U^n}_δ) is a test supermartingale for H_{δ≤δ0}, which is false in general: for δ ≤ δ0, this process is typically not a supermartingale under the null distributions. The proof of Theorem 4 and all applications correctly use δ+; in particular, the product appearing after (8) has the same subscript error. Please correct the subscripts in the theorem statement and proof.
  2. [Section 1, Definition 1] The definition of the monotone likelihood ratio property is ambiguous because the notation p^T_{δ+}(t) was introduced as the density of P^T_{δ+} relative to the fixed baseline P_{δ0} of the testing problem. If the phrase 'for all δ0, δ+ ∈ Δ with δ0 ≤ δ+' is read with that fixed baseline, the condition is not the standard pairwise MLR property, and it would not generally imply the stochastic dominance used in the proof of Proposition 2 in Appendix A. Please restate Definition 1 as the usual pairwise MLR property (for all δ1 ≤ δ2, the density ratio f_{δ2}(t)/f_{δ1}(t) is increasing in t), or clarify that the ratios are taken with respect to the δ0 appearing in the quantifier.
minor comments (4)
  1. [Section 2, proof of Theorem 4] The product in the sentence after (8) uses p^{T_i}_δ(· | U^{i-1}) instead of p^{T_i}_{δ+}(· | U^{i-1}); this is the same subscript typo as in the theorem statement.
  2. [Section 1 and Lemma 3] The notation p^{U^n}_{δ+}(u^n|u^{n-1}) denotes the conditional density of the last outcome U_n given the past; using p^{U_n}_{δ+}(u_n|u^{n-1}) would avoid confusion with the density of the full vector U^n.
  3. [Appendix C] The Taylor expansion E_{Rad}[M^δ_n] = 1 + (n-1)/6 δ^4 + O(δ^6) is stated without derivation; a short justification would make the counterexample reproducible.
  4. [References] The reference to 'Wang (2024) Personal communication' for the numerical observation that the conditional likelihood ratio is not monotone in the t-test is not verifiable; the authors should either provide the details in the paper or cite a public version of this work.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the supermartingale theorem follows directly from sufficiency and the MLR property, with all load-bearing ingredients either proved in the paper or drawn from independent external results.

full rationale

The paper's central claim is a conditional theorem: if each T_n is sufficient and satisfies the MLR property, then the likelihood-ratio process is a test supermartingale for the one-sided null. The proof is self-contained. Lemma 3 derives p^{U^n}_{δ+}(u^n|u^{n-1}) = p^{T_n}_{δ+}(t_n(u^n)|u^{n-1}) directly from the sufficiency factorization (4), so no equality is transformed into the conclusion by construction. Theorem 4 invokes Proposition 2, which is stated as following from Grünwald et al. (2024) but is proved in Appendix A via Lehmann–Romano stochastic dominance and the cancellation identity h(δ0) = 1, so the self-citation is not load-bearing. The t-test application rests on external facts: the noncentral t MLR property (Kruskal, 1954) and sufficiency of the t-statistic, which the paper verifies directly and also attributes to Pérez-Ortiz et al. (2024). The 'No Free For All' appendix demonstrates non-triviality by exhibiting a sub-Gaussian distribution under which the same process is not an e-variable, confirming that the theorem is not a tautology. No parameter is fitted and later called a prediction; the sufficient-statistic assumption is explicitly stated and verified in each example rather than engineered to force the conclusion.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper contributes a theorem whose cost is the MLR condition on a sufficient statistic; no free parameters or invented entities are introduced.

assumptions (4)
  • standard math The family of distributions Pδ and the coarsenings U^n admit regular conditional densities and are mutually absolutely continuous.
    Technical condition stated in the first paragraph to ensure likelihood ratios exist; standard in martingale hypothesis testing.
  • domain assumption For each n, there exists a sufficient statistic T_n for the model {Pδ: δ∈Δ} such that T_n has the MLR property (Def. 1).
    The central condition of Theorem 4; it is verified for each application (t-test via Kruskal 1954, chi-square directly, Bernoulli by algebra).
  • standard math The t-statistic is sufficient for the coarsened data U^n and follows a noncentral t distribution (for the Gaussian scale model).
    Used in Example 2; the sufficiency is from Pérez-Ortiz et al. (2024, Example 1) and the distribution is classical.
  • standard math The noncentral t family has the monotone likelihood ratio property (Kruskal 1954).
    External classical theorem used to satisfy the MLR requirement in the t-test example.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Supermartingales for One-Sided Tests: Sufficient Monotone Likelihood Ratios are Sufficient." pith.science (2026). https://pith.science/paper/CKIL56GK

@misc{pith2026250204208,
  author       = {Pith},
  title        = {Pith review of: Supermartingales for One-Sided Tests: Sufficient Monotone Likelihood Ratios are Sufficient},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CKIL56GK}},
  note         = {Machine review of arXiv:2502.04208}
}
abstract

The t-statistic is a widely-used scale-invariant statistic for testing the null hypothesis that the mean is zero. Martingale methods enable sequential testing with the t-statistic at every sample size, while controlling the probability of falsely rejecting the null. For one-sided sequential tests, which reject when the t-statistic is too positive, a natural question is whether they also control false rejection when the true mean is negative. We prove that this is the case using monotone likelihood ratios and sufficient statistics. We develop applications to the scale-invariant t-test, the location-invariant $\chi^2$-test and sequential linear regression with nuisance covariates.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

15 extracted references · 13 canonical work pages

  1. [1]

    Bhowmik and M

    J. Bhowmik and M. King. Maximal invariant likelihood based testing of semi-linear models. Statistical Papers, 48: 0 357--–383, 2007

  2. [2]

    L. D. Brown, I. M. Johnstone, and K. B. MacGibbon. Variation diminishing transformations: a direct approach to total positivity and its statistical applications. Journal of the American Statistical Association, 76 0 (376): 0 824--832, 1981

  3. [3]

    P. D. Gr\"unwald, R. de Heide, and W. M. Koolen. Safe testing. Journal of the Royal Statistical Society Series B: Statistical Methodology, 86 0 (5): 0 1091--1128, 2024. With Discussion

  4. [4]

    W. M. Koolen and P. Gr \"u nwald. Log-optimal anytime-valid e-values. International Journal of Approximate Reasoning, 2021. Festschrift for G. Shafer's 75th Birthday

  5. [5]

    W. Kruskal. The Monotonicity of the Ratio of Two Noncentral t-Density Functions . The Annals of Mathematical Statistics, 25 0 (1): 0 162--165, 1954

  6. [6]

    Larsson, A

    M. Larsson, A. Ramdas, and J. Ruf. The numeraire e-variable and reverse information projection. arXiv preprint arXiv:2402.18810, 2024

  7. [7]

    Lehmann and J

    E. Lehmann and J. P. Romano. Testing statistical hypotheses, volume 3. Springer, 1986

  8. [8]

    Lindon, D

    M. Lindon, D. W. Ham, M. Tingley, and I. Bojinov. Anytime-valid linear models and regression adjusted causal inference in randomized experiments, 2024

Show all 15 references
  1. [9]

    M. F. P \'e rez-Ortiz, T. Lardy, R. de Heide, and P. Gr \"u nwald. E-statistics, group invariance and anytime valid testing. The Annals of Statistics, 52 0 (4): 0 1410--1432, 2024

  2. [10]

    Ramdas, P

    A. Ramdas, P. Gr\"unwald, V. Vovk, and G. Shafer. Game-theoretic statistics and safe anytime-valid inference. Statist. Sci., 38 0 (4): 0 576--601, 2023. ISSN 0883-4237. doi:10.1214/23-STS894

  3. [11]

    Turner, A

    R. Turner, A. Ly, and P. Gr \"u nwald. Generic e-variables for exact sequential k-sample tests that allow for optional stopping. Statistical Planning and Inference, 230: 0 106116, 2024

  4. [12]

    Schure Ter ter Schure, M

    J. Schure Ter ter Schure, M. Pérez-Ortiz, A. Ly, and P. Gr \"u nwald. The anytime-valid logrank test: Error control under continuous monitoring with unlimited horizon. New England Journal of Statistics in Data Science, 2 0 (2): 0 190--214, 2024

  5. [13]

    H. Wang. Personal communication, 2024

  6. [14]

    Wang and A

    H. Wang and A. Ramdas. Anytime-valid t-tests and confidence sequences for G aussian means with unknown variance, 2024

  7. [15]

    Williams

    D. Williams. Probability with Martingales. Cambridge Mathematical Textbooks, 1991

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.