REVIEW 2 major objections 5 minor 33 references
Real-time Program Evaluation using Anytime-valid Rank Tests
T0 review · 2 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Rank tests let you monitor treatment effects with exact error control
desk verdict The anytime-valid rank-test core is solid and worth engaging, but the SCM bridge in Proposition 2 is missing an assumption on the training period and needs a fix before the application is justified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central objects are the sequential rank $R_t$, the position of the newest treatment estimate among all previous estimates, and the reduced sequential rank $\tilde{R}_t$, its position among the $T_0$ pre-treatment estimates only. Their role is to reduce the highly composite exchangeability null to a simple probabilistic statement: under the null, $R_t$ is uniform on $\{1,\ldots,t\}$, while $\tilde{R}_t$ has the categorical distribution with cell probabilities $q_i^t=(1+\#\{s<t:\tilde{R}_s=i\})/t$. This reduction makes it possible to write down closed-form sequential e-values, whose running product is a test martingale; Ville's inequality converts that martingale into an anytime-valid p-value. The theorem structure allows any non-negative test statistic that is conditionally independent of the current rank, and for a simple alternative the log-optimal statistic is proportional to the conditional density of the rank, which is what the Gaussian and plug-in versions implement.
What would settle it
Run the Section 6.1 difference-in-differences simulation under the null with independent normal errors and the stopping rule “reject at the first time $p_t\le\alpha$”; the fraction of simulated paths that ever reject must not exceed $\alpha$ at any horizon. A second run with AR(1) errors should show the same procedure rejecting on up to 20–26% of paths with block size one, which would locate the failure in the exchangeability premise rather than in the rank construction.
Extended reading notes
Core claim
The paper's central claim is that anytime-valid p-values for the null hypothesis that the treatment-effect estimators $\hat{\tau}_t$ are exchangeable can be built from reduced information. Under the exchangeability null, the sequential rank $R_t$ of the newest treatment estimate among all previous estimates is uniform on its support, while the reduced sequential rank $\tilde{R}_t$, which ranks the newest post-treatment estimate among the $T_0$ pre-treatment estimates only, follows a categorical distribution with probabilities $q_i^t=(1+\#\{s<t:\tilde{R}_s=i\})/t$. Theorems 1 and 2 show that for any non-negative test statistic that is independent of the current rank given the past, the normalized ratio is a sequential e-value; choosing the statistic proportional to the conditional density under a simple alternative gives the log-optimal test martingale. The running product $W_t$ is a test martingale, and $p_t=1/W_t$ satisfies the anytime-validity bound in Eq. (3), so rejecting at the first time $p_t\le\alpha$, or at any later data-dependent time, controls the Type-I error exactly in finite samples. The paper also claims that for post-exchangeable alternatives the reduced-rank construction dominates the sequential-rank plug-in construction in expected log-growth (Theorem 3), and that in the interactive fixed-effects model both difference-in-differences and synthetic control estimators are exchangeable under the null under the conditions of Propositions 1 and 2.
Load-bearing premise
The whole construction rests on the treatment-effect estimates being exchangeable over the blank and post-treatment periods when the treatment has no effect; if serial dependence or imperfect factor-loading reconstruction breaks that, the anytime-valid size guarantee no longer holds, and the paper's own simulations show rejection rates can climb to 0.20–0.26 with block size one.
Editorial extensions
If this is right
- Under the exact exchangeability conditions of Propositions 1 and 2, decision makers can stop at the first time the anytime-valid p-value falls below $\alpha$, or continue past any pre-planned $T$, and still keep the probability of ever rejecting under the null at most $\alpha$.
- Repeatedly applying a fixed-$T$ permutation test is shown to inflate size to about 20% in the DiD simulation, while the anytime-valid tests remain below $\alpha$ for up to 1000 post-treatment observations.
- For the discounted-utility criterion in Eq. (27), the anytime-valid Gaussian reduced-rank test is preferred over every fixed-$T$ test for discount factors $\delta\ge 0.8$ in the stylized DiD setting.
- When the post-treatment data are exchangeable among themselves, the reduced-rank construction has weakly higher expected log-growth than the sequential-rank plug-in construction (Theorem 3), and the adaptive mixture over effect sizes recovers the growth rate of the best candidate up to a $\log k/t$ regret bound.
- In the interactive fixed-effects model, both DiD and SCM estimators are exchangeable under the null under the stated assumptions, so the same anytime-valid machinery applies to both estimators.
Reading between the lines
- A natural extension is to use the same rank reduction on other exchangeable test statistics from program evaluation, such as event-study coefficients or conformal prediction residuals, whenever a batch of pre-treatment values is available at the start of monitoring.
- The paper's size distortions under serial dependence suggest the anytime-valid p-value itself can serve as a continuous diagnostic for non-exchangeability: a run of low p-values under a supposedly null treatment is evidence that the estimator's blank-period distribution is not representative of the post-treatment period.
- Because the reduced-rank construction intentionally discards the ordering of post-treatment observations, it points to a design trade-off: choosing which statistic to monitor is also choosing which alternatives the test can detect, so monitoring several coarsened statistics simultaneously could protect against both mean shifts and dynamic effects at the cost of a small regret bound.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops anytime-valid tests for program evaluation by testing the null hypothesis that treatment-effect estimators are exchangeable over pre- and post-treatment periods. The main construction converts the estimators into sequential ranks or reduced sequential ranks and builds sequential e-values from arbitrary non-negative test statistics (Theorems 1 and 2), yielding p-processes with finite-sample type-I error control at all data-dependent stopping times via Ville's inequality. Theorem 3 shows that, under a post-exchangeable alternative, the reduced-rank plug-in statistic dominates the sequential-rank plug-in statistic in expected log-growth. The methodology is illustrated for difference-in-differences and synthetic control in an interactive fixed-effects model, with simulations comparing anytime-valid tests to fixed-T permutation tests under exact and violated exchangeability, including a block-based remedy for serial dependence.
Significance. If the results hold, the paper makes a useful contribution: it supplies finite-sample anytime-valid inference for a class of program-evaluation settings where the number of post-treatment observations need not be pre-specified, and it exploits the pre-treatment batch through reduced sequential ranks. The core e-value proofs are clean and correct: Theorems 1 and 2 are valid conditional e-value constructions, and the Jensen argument in Theorem 3 is sound. The simulations are informative and honestly document size distortions under serial dependence. The main caveat is that the synthetic-control bridge from 'no treatment effect' to 'exchangeable estimators' is not justified as stated, and the abstract's 'optimal' claim is stronger than what is proved.
major comments (2)
- [Section 5.2.2, Proposition 2] The proposition imposes assumptions only on t in B union {T0+1,...,T}, but the SCM weights are estimated on the training set E = {1,...,T0}\B. Nothing in the stated conditions prevents the training period from being dependent with the evaluation period, so the weight vector can be correlated with some evaluation errors but not others. Concretely, take r=0, no covariates, one control unit, so btau_t = eps_1t - w eps_2t with w a function of training data; let E contain t=T0, set eps_1,T0 = eps_1,T0+1 and eps_2,T0 = eps_2,T0+1, and let the remaining evaluation errors be i.i.d. N(0,1). The proposition's evaluation-block conditions then hold, yet w is correlated with eps_T0+1 and not with eps_T0+2, so (btau_T0+1, btau_T0+2) is not exchangeable. Proposition 2 therefore does not establish the claimed exchangeability bridge for SCM. An additional condition analogous to Proposition 1's full exchangeability over t=1,...,T, or an explicit independence restriction on E relative to B union post-treatment periods, is needed.
- [Abstract and Section 1.1] The abstract's phrase 'optimal finite-sample valid sequential tests' and the text's claim of 'a p-value with a type of optimal shrinkage rate' overstate what Theorems 1 and 2 prove. Log-optimality is established only for a fully specified simple alternative with known conditional density (St proportional to gt). For the composite alternatives used in Section 4.1 and the adaptive mixtures of Section 4.2 and Appendix B, no optimality theorem is proved; Theorem 3 gives only a dominance relation between two plug-in statistics. The claims should be qualified, for example as 'log-optimal under a correctly specified simple alternative'.
minor comments (5)
- [Section 4.1, footnote 3 and Appendix A.3] The proof of Theorem 3 relies on the fact that the normalization constant in the plug-in e-value can be set to 1, but this is stated only in a footnote. Move this point into the main text or add a cross-reference in the proof of Theorem 3 to avoid confusion.
- [Section 4.2 and Appendix C] The appendix says the reduced-rank Gaussian statistic conditions on the pre-treatment outcomes, while the main text emphasizes that inference may only depend on the coarsened rank filtration. The reason for marginalizing over the pre-treatment outcomes (to maintain measurability with respect to the rank filtration) should be explained more prominently in the main text.
- [Section 6.1.3, Eq. (27)] The discounted utility E[U_S] uses P[H0 rejected by test S at t' <= t] without specifying the probability measure; it should be stated that this probability is evaluated under the alternative.
- [Abstract and Table 2] The abstract's claim that the methods 'control size even under mild exchangeability violations' should be calibrated against the B=1 results in Table 2, where rejection rates reach 0.20-0.26 under serial dependence. The text should clarify that this robustness applies to mild violations or relies on the block structure.
- [Appendix A.1 and A.2] There are small typographical issues in the proofs, including 'demoninator' for 'denominator', and the convention 0/0=1 should be stated before the first use in Theorem 1 rather than after it.
Circularity Check
No significant circularity: the anytime-valid rank e-values are proved from exchangeability via Ville's inequality, and the SCM proposition's assumption gap is a correctness issue, not a circular reduction.
full rationale
I walked the derivation chain. The central validity claim, Eq. (3), follows from Ville's inequality applied to the test martingale W_t = product of sequential e-values. Theorems 1 and 2 prove the e-value property for arbitrary predictable non-negative test statistics using only the rank-uniform null (10) or the reduced-rank categorical null (13)-(14); no parameter is fitted to make the p-values valid, and the Gaussian effect size affects power but not validity. The reduction from 'no treatment effect' to exchangeability is an explicit assumption (Eq. (1), Section 1.2, and Section 5), not a conclusion smuggled into the premise. Proposition 1 spells out the DiD conditions under which this reduction holds. Proposition 2 (Section 5.2.2) attempts the same for SCM but omits assumptions on the training period E; as stated, the implication from no treatment effect to exchangeable SCM estimates is not justified, as the skeptic's example shows. This is an assumption gap / correctness risk, not circularity: the conclusion is not definitionally equal to the premise, and the missing condition is not an output of the paper's own construction. The self-citations (Koning 2024a,b) are contextual and non-load-bearing; the exchangeability-to-ranks connection is traced to Vovk et al. (2003), and the proofs here are self-contained. The simulations report honest size and power properties rather than claiming a fitted quantity as a prediction. Therefore no circular step meets the quoted-evidence bar, and the circularity score is 0.
Assumptions & free parameters
free parameters (4)
- Gaussian alternative effect size
- Kernel bandwidth for plug-in density
- Block size B =
B=1,3 in simulations
- Candidate effect sizes C for adaptive mixture =
{0.25, 0.5, 1, 2, 4}
assumptions (5)
- standard math Ville's inequality for nonnegative supermartingales
- domain assumption Exchangeability of treatment effect estimators under H0
- domain assumption The IFE model (21) with conditions in Proposition 2 (exchangeable (lambda_t,theta_t), independent and exchangeable epsilon_t, and adequate reconstruction of factor loadings by SCM weights)
- standard math Sequential ranks are sufficient for testing exchangeability under the coarsened filtration
- domain assumption Conditional density of alternative is known or learnable for log-optimality
Cite this review
Pith. "Pith review of Real-time Program Evaluation using Anytime-valid Rank Tests." pith.science (2026). https://pith.science/paper/AOKFZLOA
@misc{pith2026250421595,
author = {Pith},
title = {Pith review of: Real-time Program Evaluation using Anytime-valid Rank Tests},
year = {2026},
howpublished = {\url{https://pith.science/paper/AOKFZLOA}},
note = {Machine review of arXiv:2504.21595}
}
read the original abstract
Counterfactual mean estimators such as difference-in-differences and synthetic control have grown into workhorse tools for program evaluation. Inference for these estimators is well-developed in settings where all post-treatment data is available at the time of analysis. However, in settings where data arrives sequentially, these tests do not permit real-time inference, as they require a pre-specified sample size T. We introduce real-time inference for program evaluation through anytime-valid rank tests. Our methodology relies on interpreting the absence of a treatment effect as exchangeability of the treatment estimates. We then convert these treatment estimates into sequential ranks, and construct optimal finite-sample valid sequential tests for exchangeability. We illustrate our methods in the context of difference-in-differences and synthetic control. In simulations, they control size even under mild exchangeability violations. While our methods suffer slight power loss at T, they allow for early rejection (before T) and preserve the ability to reject later (after T).
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Abadie, A., Diamond, A., and Hainmueller, J. (2010). Synthetic control methods for comparative case studies: Estimating the effect of california’s tobacco control program. Journal of the American statistical Association , 105(490):493--505
work page 2010
-
[2]
Abadie, A. and Gardeazabal, J. (2003). The economic costs of conflict: A case study of the basque country. American economic review , 93(1):113--132
work page 2003
-
[3]
Abadie, A. and Zhao, J. (2021). Synthetic controls for experimental design. arXiv preprint arXiv:2108.02196
arXiv 2021
-
[4]
Bai, J. (2009). Panel data models with interactive fixed effects. Econometrica , 77(4):1229--1279
2009
-
[5]
Bates, S., Cand \`e s, E., Lei, L., Romano, Y., and Sesia, M. (2023). Testing for outliers with conformal p-values. The Annals of Statistics , 51(1):149--178
2023
-
[6]
Chernozhukov, V., W \"u thrich, K., and Yinchu, Z. (2018). Exact and robust conformal inference methods for predictive machine learning with dependent data. In Conference On learning theory , pages 732--749. PMLR
2018
-
[7]
Chernozhukov, V., W \"u thrich, K., and Zhu, Y. (2021). An exact and robust conformal inference method for counterfactual and synthetic controls. Journal of the American Statistical Association , 116(536):1849--1864
work page 2021
-
[8]
Doudchenko, N., Gilinson, D., Taylor, S., and Wernerfelt, N. (2019). Designing experiments with synthetic controls. Unpublished working paper
work page 2019
Show all 33 references
-
[9]
Doudchenko, N., Khosravi, K., Pouget-Abadie, J., Lahaie, S., Lubin, M., Mirrokni, V., Spiess, J., et al. (2021). Synthetic design: An optimization approach to experimental design with synthetic controls. Advances in Neural Information Processing Systems , 34:8691--8701
2021
-
[10]
Fedorova, V., Gammerman, A., Nouretdinov, I., and Vovk, V. (2012). Plug-in martingales for testing exchangeability on-line. In Proceedings of the 29th International Coference on International Conference on Machine Learning , pages 923--930
2012
-
[11]
and Ramdas, A
Fischer, L. and Ramdas, A. (2024). Improving the (approximate) sequential probability ratio test by avoiding overshoot. arXiv preprint arXiv:2410.16076
2024 arXiv
-
[12]
and Ramdas, A
Fischer, L. and Ramdas, A. (2025). Sequential monte carlo testing by betting. Journal of the Royal Statistical Society: Series B (Statistical Methodology) . Advance online publication
2025
-
[13]
Gr \"u nwald, P., de Heide, R., and Koolen, W. (2024). Safe testing. Journal of the Royal Statistical Society Series B: Statistical Methodology , 86(5):1091--1128
2024
-
[14]
W., Bojinov, I., Lindon, M., and Tingley, M
Ham, D. W., Bojinov, I., Lindon, M., and Tingley, M. (2022). Design-based confidence sequences for anytime-valid causal inference. arXiv preprint arXiv:2210.08639 , 11
2022 arXiv
-
[15]
and Law, M
Henzi, A. and Law, M. (2024). A rank-based sequential test of independence. Biometrika , 111(4):1169--1186
2024
-
[16]
R., Ramdas, A., McAuliffe, J., and Sekhon, J
Howard, S. R., Ramdas, A., McAuliffe, J., and Sekhon, J. (2021). Time-uniform, nonparametric, nonasymptotic confidence sequences. The Annals of Statistics , 49(2):1055--1080
2021
-
[17]
Kelly, J. L. (1956). A new interpretation of information rate. The Bell System Technical Journal , 35(4):917--926
1956
-
[18]
Koning, N. W. (2024a). Measuring evidence with a continuous test. arXiv preprint arXiv:2409.05654
2024 arXiv
-
[19]
Koning, N. W. (2024b). Post-hoc and anytime valid permutation and group invariance testing. arXiv preprint arXiv:2310.01153
2024
-
[20]
Koolen, W. M. and Gr \"u nwald, P. (2022). Log-optimal anytime-valid e-values. International Journal of Approximate Reasoning , 141:69--82
2022
-
[21]
Larsson, M., Ramdas, A., and Ruf, J. (2024). The numeraire e-variable and reverse information projection. arXiv preprint arXiv:2402.18810
2024 arXiv
-
[22]
Z., Sinha, M., Addanki, R., Ramdas, A., Garg, M., and Swaminathan, V
Maharaj, A., Sinha, R., Arbour, D., Waudby-Smith, I., Liu, S. Z., Sinha, M., Addanki, R., Ramdas, A., Garg, M., and Swaminathan, V. (2023). Anytime-valid confidence sequences in an enterprise a/b testing platform. In Companion Proceedings of the ACM Web Conference 2023 , pages...
2023
-
[23]
Nagler, T. (2018). Asymptotic analysis of the jittering kernel density estimator. Mathematical Methods of Statistics , 27:32--46
2018
-
[24]
Ramdas, A., Gr \"u nwald, P., Vovk, V., and Shafer, G. (2023). Game-theoretic statistics and safe anytime-valid inference. Statistical Science , 38(4):576--601
2023
-
[25]
and Manole, T
Ramdas, A. and Manole, T. (2023). Randomized and exchangeable improvements of markov's, chebyshev's and chernoff's inequalities. arXiv preprint arXiv:2304.02611
2023 arXiv
-
[26]
Ramdas, A., Ruf, J., Larsson, M., and Koolen, W. (2020). Admissible anytime-valid sequential inference must rely on nonnegative martingales. arXiv preprint arXiv:2009.03167
2020 arXiv
-
[27]
Ramdas, A., Ruf, J., Larsson, M., and Koolen, W. M. (2022). Testing exchangeability: Fork-convexity, supermartingales and e-processes. International Journal of Approximate Reasoning , 141:83--109
2022
-
[28]
Shafer, G., Shen, A., Vereshchagin, N., and Vovk, V. (2011). Test martingales, bayes factors and p-values. Statistical Science , pages 84--101
2011
-
[29]
Ville, J. (1939). Étude critique de la notion de collectif . PhD thesis, University of Paris
1939
-
[30]
Vovk, V. (2021). Testing randomness online. Statistical Science , 36(4):595--611
2021
-
[31]
Vovk, V. (2023). The power of forgetting in statistical hypothesis testing. In Conformal and Probabilistic Prediction with Applications , pages 347--366. PMLR
2023
-
[32]
Vovk, V., Nouretdinov, I., and Gammerman, A. (2003). Testing exchangeability on-line. In Proceedings of the 20th international conference on machine learning (ICML-03) , pages 768--775
2003
-
[33]
and Wang, R
Vovk, V. and Wang, R. (2021). E-values: Calibration, combination and applications. The Annals of Statistics , 49(3):1736--1754
2021
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.