Pith. sign in

REVIEW 2 major objections 5 minor 33 references

Real-time Program Evaluation using Anytime-valid Rank Tests

T0 review · 2 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Rank tests let you monitor treatment effects with exact error control

desk verdict The anytime-valid rank-test core is solid and worth engaging, but the SCM bridge in Proposition 2 is missing an assumption on the training period and needs a fix before the application is justified. read the letter →

arxiv 2504.21595 v1 pith:AOKFZLOA submitted 2025-04-30 econ.EM stat.ME

classification econ.EMstat.ME MSC 62L1062G1062F0362P20
keywords anytime-validinferencesequentialranksreducedexchangeabilitye-valuestestmartingalesdifference-in-differencessyntheticcontrol
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that program evaluation can be done in real time without pre-specifying how many post-treatment observations will be collected. The trick is to interpret “no treatment effect” as exchangeability of the treatment-effect estimators over the blank and post-treatment periods, and to convert those estimators into sequential ranks. From the ranks the paper constructs e-values whose running product is a test martingale, and the reciprocal of that martingale is an anytime-valid p-value: under the null, the probability that the p-value ever drops below $\alpha$ is at most $\alpha$, at every data-dependent stopping time. If this is right, difference-in-differences and synthetic control analyses can be monitored continuously, rejecting early when evidence is strong or continuing past a conventional end date without losing the Type-I error guarantee.

What carries the argument

The central objects are the sequential rank $R_t$, the position of the newest treatment estimate among all previous estimates, and the reduced sequential rank $\tilde{R}_t$, its position among the $T_0$ pre-treatment estimates only. Their role is to reduce the highly composite exchangeability null to a simple probabilistic statement: under the null, $R_t$ is uniform on $\{1,\ldots,t\}$, while $\tilde{R}_t$ has the categorical distribution with cell probabilities $q_i^t=(1+\#\{s<t:\tilde{R}_s=i\})/t$. This reduction makes it possible to write down closed-form sequential e-values, whose running product is a test martingale; Ville's inequality converts that martingale into an anytime-valid p-value. The theorem structure allows any non-negative test statistic that is conditionally independent of the current rank, and for a simple alternative the log-optimal statistic is proportional to the conditional density of the rank, which is what the Gaussian and plug-in versions implement.

What would settle it

Run the Section 6.1 difference-in-differences simulation under the null with independent normal errors and the stopping rule “reject at the first time $p_t\le\alpha$”; the fraction of simulated paths that ever reject must not exceed $\alpha$ at any horizon. A second run with AR(1) errors should show the same procedure rejecting on up to 20–26% of paths with block size one, which would locate the failure in the exchangeability premise rather than in the rank construction.

Watch

Extended reading notes

Core claim

The paper's central claim is that anytime-valid p-values for the null hypothesis that the treatment-effect estimators $\hat{\tau}_t$ are exchangeable can be built from reduced information. Under the exchangeability null, the sequential rank $R_t$ of the newest treatment estimate among all previous estimates is uniform on its support, while the reduced sequential rank $\tilde{R}_t$, which ranks the newest post-treatment estimate among the $T_0$ pre-treatment estimates only, follows a categorical distribution with probabilities $q_i^t=(1+\#\{s<t:\tilde{R}_s=i\})/t$. Theorems 1 and 2 show that for any non-negative test statistic that is independent of the current rank given the past, the normalized ratio is a sequential e-value; choosing the statistic proportional to the conditional density under a simple alternative gives the log-optimal test martingale. The running product $W_t$ is a test martingale, and $p_t=1/W_t$ satisfies the anytime-validity bound in Eq. (3), so rejecting at the first time $p_t\le\alpha$, or at any later data-dependent time, controls the Type-I error exactly in finite samples. The paper also claims that for post-exchangeable alternatives the reduced-rank construction dominates the sequential-rank plug-in construction in expected log-growth (Theorem 3), and that in the interactive fixed-effects model both difference-in-differences and synthetic control estimators are exchangeable under the null under the conditions of Propositions 1 and 2.

Load-bearing premise

The whole construction rests on the treatment-effect estimates being exchangeable over the blank and post-treatment periods when the treatment has no effect; if serial dependence or imperfect factor-loading reconstruction breaks that, the anytime-valid size guarantee no longer holds, and the paper's own simulations show rejection rates can climb to 0.20–0.26 with block size one.

Editorial extensions

If this is right

  • Under the exact exchangeability conditions of Propositions 1 and 2, decision makers can stop at the first time the anytime-valid p-value falls below $\alpha$, or continue past any pre-planned $T$, and still keep the probability of ever rejecting under the null at most $\alpha$.
  • Repeatedly applying a fixed-$T$ permutation test is shown to inflate size to about 20% in the DiD simulation, while the anytime-valid tests remain below $\alpha$ for up to 1000 post-treatment observations.
  • For the discounted-utility criterion in Eq. (27), the anytime-valid Gaussian reduced-rank test is preferred over every fixed-$T$ test for discount factors $\delta\ge 0.8$ in the stylized DiD setting.
  • When the post-treatment data are exchangeable among themselves, the reduced-rank construction has weakly higher expected log-growth than the sequential-rank plug-in construction (Theorem 3), and the adaptive mixture over effect sizes recovers the growth rate of the best candidate up to a $\log k/t$ regret bound.
  • In the interactive fixed-effects model, both DiD and SCM estimators are exchangeable under the null under the stated assumptions, so the same anytime-valid machinery applies to both estimators.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension is to use the same rank reduction on other exchangeable test statistics from program evaluation, such as event-study coefficients or conformal prediction residuals, whenever a batch of pre-treatment values is available at the start of monitoring.
  • The paper's size distortions under serial dependence suggest the anytime-valid p-value itself can serve as a continuous diagnostic for non-exchangeability: a run of low p-values under a supposedly null treatment is evidence that the estimator's blank-period distribution is not representative of the post-treatment period.
  • Because the reduced-rank construction intentionally discards the ordering of post-treatment observations, it points to a design trade-off: choosing which statistic to monitor is also choosing which alternatives the test can detect, so monitoring several coarsened statistics simultaneously could protect against both mean shifts and dynamic effects at the cost of a small regret bound.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper develops anytime-valid tests for program evaluation by testing the null hypothesis that treatment-effect estimators are exchangeable over pre- and post-treatment periods. The main construction converts the estimators into sequential ranks or reduced sequential ranks and builds sequential e-values from arbitrary non-negative test statistics (Theorems 1 and 2), yielding p-processes with finite-sample type-I error control at all data-dependent stopping times via Ville's inequality. Theorem 3 shows that, under a post-exchangeable alternative, the reduced-rank plug-in statistic dominates the sequential-rank plug-in statistic in expected log-growth. The methodology is illustrated for difference-in-differences and synthetic control in an interactive fixed-effects model, with simulations comparing anytime-valid tests to fixed-T permutation tests under exact and violated exchangeability, including a block-based remedy for serial dependence.

Significance. If the results hold, the paper makes a useful contribution: it supplies finite-sample anytime-valid inference for a class of program-evaluation settings where the number of post-treatment observations need not be pre-specified, and it exploits the pre-treatment batch through reduced sequential ranks. The core e-value proofs are clean and correct: Theorems 1 and 2 are valid conditional e-value constructions, and the Jensen argument in Theorem 3 is sound. The simulations are informative and honestly document size distortions under serial dependence. The main caveat is that the synthetic-control bridge from 'no treatment effect' to 'exchangeable estimators' is not justified as stated, and the abstract's 'optimal' claim is stronger than what is proved.

major comments (2)
  1. [Section 5.2.2, Proposition 2] The proposition imposes assumptions only on t in B union {T0+1,...,T}, but the SCM weights are estimated on the training set E = {1,...,T0}\B. Nothing in the stated conditions prevents the training period from being dependent with the evaluation period, so the weight vector can be correlated with some evaluation errors but not others. Concretely, take r=0, no covariates, one control unit, so btau_t = eps_1t - w eps_2t with w a function of training data; let E contain t=T0, set eps_1,T0 = eps_1,T0+1 and eps_2,T0 = eps_2,T0+1, and let the remaining evaluation errors be i.i.d. N(0,1). The proposition's evaluation-block conditions then hold, yet w is correlated with eps_T0+1 and not with eps_T0+2, so (btau_T0+1, btau_T0+2) is not exchangeable. Proposition 2 therefore does not establish the claimed exchangeability bridge for SCM. An additional condition analogous to Proposition 1's full exchangeability over t=1,...,T, or an explicit independence restriction on E relative to B union post-treatment periods, is needed.
  2. [Abstract and Section 1.1] The abstract's phrase 'optimal finite-sample valid sequential tests' and the text's claim of 'a p-value with a type of optimal shrinkage rate' overstate what Theorems 1 and 2 prove. Log-optimality is established only for a fully specified simple alternative with known conditional density (St proportional to gt). For the composite alternatives used in Section 4.1 and the adaptive mixtures of Section 4.2 and Appendix B, no optimality theorem is proved; Theorem 3 gives only a dominance relation between two plug-in statistics. The claims should be qualified, for example as 'log-optimal under a correctly specified simple alternative'.
minor comments (5)
  1. [Section 4.1, footnote 3 and Appendix A.3] The proof of Theorem 3 relies on the fact that the normalization constant in the plug-in e-value can be set to 1, but this is stated only in a footnote. Move this point into the main text or add a cross-reference in the proof of Theorem 3 to avoid confusion.
  2. [Section 4.2 and Appendix C] The appendix says the reduced-rank Gaussian statistic conditions on the pre-treatment outcomes, while the main text emphasizes that inference may only depend on the coarsened rank filtration. The reason for marginalizing over the pre-treatment outcomes (to maintain measurability with respect to the rank filtration) should be explained more prominently in the main text.
  3. [Section 6.1.3, Eq. (27)] The discounted utility E[U_S] uses P[H0 rejected by test S at t' <= t] without specifying the probability measure; it should be stated that this probability is evaluated under the alternative.
  4. [Abstract and Table 2] The abstract's claim that the methods 'control size even under mild exchangeability violations' should be calibrated against the B=1 results in Table 2, where rejection rates reach 0.20-0.26 under serial dependence. The text should clarify that this robustness applies to mild violations or relies on the block structure.
  5. [Appendix A.1 and A.2] There are small typographical issues in the proofs, including 'demoninator' for 'denominator', and the convention 0/0=1 should be stated before the first use in Theorem 1 rather than after it.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the anytime-valid rank e-values are proved from exchangeability via Ville's inequality, and the SCM proposition's assumption gap is a correctness issue, not a circular reduction.

full rationale

I walked the derivation chain. The central validity claim, Eq. (3), follows from Ville's inequality applied to the test martingale W_t = product of sequential e-values. Theorems 1 and 2 prove the e-value property for arbitrary predictable non-negative test statistics using only the rank-uniform null (10) or the reduced-rank categorical null (13)-(14); no parameter is fitted to make the p-values valid, and the Gaussian effect size affects power but not validity. The reduction from 'no treatment effect' to exchangeability is an explicit assumption (Eq. (1), Section 1.2, and Section 5), not a conclusion smuggled into the premise. Proposition 1 spells out the DiD conditions under which this reduction holds. Proposition 2 (Section 5.2.2) attempts the same for SCM but omits assumptions on the training period E; as stated, the implication from no treatment effect to exchangeable SCM estimates is not justified, as the skeptic's example shows. This is an assumption gap / correctness risk, not circularity: the conclusion is not definitionally equal to the premise, and the missing condition is not an output of the paper's own construction. The self-citations (Koning 2024a,b) are contextual and non-load-bearing; the exchangeability-to-ranks connection is traced to Vovk et al. (2003), and the proofs here are self-contained. The simulations report honest size and power properties rather than claiming a fitted quantity as a prediction. Therefore no circular step meets the quoted-evidence bar, and the circularity score is 0.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

No new physical entities are introduced. The free parameters are user-specified inputs to the testing procedure, not fitted to make the derivation work. The axioms are standard probabilistic tools plus domain assumptions inherited from the interactive fixed-effects model and the exchangeability-based causal inference framework.

free parameters (4)
  • Gaussian alternative effect size
    User-specified standardized mean shift (mu_post - mu_pre)/sigma used in the Gaussian alternative; appears in Section 4.2 and Algorithm 1. Power depends on it but validity does not.
  • Kernel bandwidth for plug-in density
    The plug-in test statistic (18) uses a kernel density estimate; the bandwidth is not specified in the paper, so it is an implicit free choice affecting power.
  • Block size B = B=1,3 in simulations
    Block structure in Section 5.3 is imposed to mitigate serial dependence; B is chosen by the user and affects the set of stopping times.
  • Candidate effect sizes C for adaptive mixture = {0.25, 0.5, 1, 2, 4}
    In Section 6.2.2, the adaptive mixture averages over a finite set of Gaussian alternatives; the set is chosen by the user.
assumptions (5)
  • standard math Ville's inequality for nonnegative supermartingales
    Used in Section 2.2 to convert test martingales into anytime-valid p-values; a standard result in probability theory.
  • domain assumption Exchangeability of treatment effect estimators under H0
    The paper assumes btau_t are exchangeable over blank and post-treatment periods under no treatment effect; derived from IFE assumptions in Propositions 1 and 2 but inherently untestable in practice.
  • domain assumption The IFE model (21) with conditions in Proposition 2 (exchangeable (lambda_t,theta_t), independent and exchangeable epsilon_t, and adequate reconstruction of factor loadings by SCM weights)
    Used in Section 5 to justify exchangeability of SCM estimators; the factor loading reconstruction is asymptotic (Abadie et al. 2010).
  • standard math Sequential ranks are sufficient for testing exchangeability under the coarsened filtration
    Uses the result that non-trivial test martingales require coarsening to ranks (Vovk 2021; Ramdas et al. 2022), cited in Section 3.1.
  • domain assumption Conditional density of alternative is known or learnable for log-optimality
    Log-optimality in Theorems 1 and 2 requires a specified alternative distribution; in composite settings this is approximated by plug-in or mixtures.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Real-time Program Evaluation using Anytime-valid Rank Tests." pith.science (2026). https://pith.science/paper/AOKFZLOA

@misc{pith2026250421595,
  author       = {Pith},
  title        = {Pith review of: Real-time Program Evaluation using Anytime-valid Rank Tests},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AOKFZLOA}},
  note         = {Machine review of arXiv:2504.21595}
}
read the original abstract

Counterfactual mean estimators such as difference-in-differences and synthetic control have grown into workhorse tools for program evaluation. Inference for these estimators is well-developed in settings where all post-treatment data is available at the time of analysis. However, in settings where data arrives sequentially, these tests do not permit real-time inference, as they require a pre-specified sample size T. We introduce real-time inference for program evaluation through anytime-valid rank tests. Our methodology relies on interpreting the absence of a treatment effect as exchangeability of the treatment estimates. We then convert these treatment estimates into sequential ranks, and construct optimal finite-sample valid sequential tests for exchangeability. We illustrate our methods in the context of difference-in-differences and synthetic control. In simulations, they control size even under mild exchangeability violations. While our methods suffer slight power loss at T, they allow for early rejection (before T) and preserve the ability to reject later (after T).

Figures

Figures reproduced from arXiv: 2504.21595 by the authors.

Figure 1
Figure 1. Illustration of the null distribution of the sequential rank [PITH_FULL_IMAGE:figures/full_fig_p018_1.png] view at source ↗
Figure 2
Figure 2. Rejection rates based on the DiD estimator under [PITH_FULL_IMAGE:figures/full_fig_p028_2.png] view at source ↗
Figure 3
Figure 3. Differences in power between the anytime-valid test with Gaussian alternative [PITH_FULL_IMAGE:figures/full_fig_p030_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Regions of discount factors δ in which the fixed-T test performed at different levels of T achieves higher discounted utility, as defined in (27), than anytime-valid tests. Simulation settings are the same as in [PITH_FULL_IMAGE:figures/full_fig_p031_4.png]
Figure 5
Figure 5. Figure 5: Rejection rates based on the SCM estimator under [PITH_FULL_IMAGE:figures/full_fig_p035_5.png]
Figure 6
Figure 6. Figure 6: Regions of discount factors δ in which the fixed-T test at different levels of T attains higher discounted utility than the anytime-valid tests. Test is based on the SCM estimator under H1 : τt = 1 + t−T0 15 . We consider the fixed-T test, the adaptive mixture of reduc…
Figure 7
Figure 7. Figure 7: Rejection rates based on the SCM estimator under [PITH_FULL_IMAGE:figures/full_fig_p054_7.png]
Figure 8
Figure 8. Figure 8: Regions of discount factors δ in which the fixed-T test at different levels of T attains higher discounted utility than the anytime-valid tests. The simulation settings are the same as in Figure 7b. though they note that more general Lp norms can be employed. For one-s…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

33 extracted references · 15 canonical work pages

  1. [1]

    Abadie, A., Diamond, A., and Hainmueller, J. (2010). Synthetic control methods for comparative case studies: Estimating the effect of california’s tobacco control program. Journal of the American statistical Association , 105(490):493--505

  2. [2]

    and Gardeazabal, J

    Abadie, A. and Gardeazabal, J. (2003). The economic costs of conflict: A case study of the basque country. American economic review , 93(1):113--132

  3. [3]

    and Zhao, J

    Abadie, A. and Zhao, J. (2021). Synthetic controls for experimental design. arXiv preprint arXiv:2108.02196

  4. [4]

    Bai, J. (2009). Panel data models with interactive fixed effects. Econometrica , 77(4):1229--1279

  5. [5]

    Bates, S., Cand \`e s, E., Lei, L., Romano, Y., and Sesia, M. (2023). Testing for outliers with conformal p-values. The Annals of Statistics , 51(1):149--178

  6. [6]

    Chernozhukov, V., W \"u thrich, K., and Yinchu, Z. (2018). Exact and robust conformal inference methods for predictive machine learning with dependent data. In Conference On learning theory , pages 732--749. PMLR

  7. [7]

    Chernozhukov, V., W \"u thrich, K., and Zhu, Y. (2021). An exact and robust conformal inference method for counterfactual and synthetic controls. Journal of the American Statistical Association , 116(536):1849--1864

  8. [8]

    Doudchenko, N., Gilinson, D., Taylor, S., and Wernerfelt, N. (2019). Designing experiments with synthetic controls. Unpublished working paper

Show all 33 references
  1. [9]

    Doudchenko, N., Khosravi, K., Pouget-Abadie, J., Lahaie, S., Lubin, M., Mirrokni, V., Spiess, J., et al. (2021). Synthetic design: An optimization approach to experimental design with synthetic controls. Advances in Neural Information Processing Systems , 34:8691--8701

  2. [10]

    Fedorova, V., Gammerman, A., Nouretdinov, I., and Vovk, V. (2012). Plug-in martingales for testing exchangeability on-line. In Proceedings of the 29th International Coference on International Conference on Machine Learning , pages 923--930

  3. [11]

    and Ramdas, A

    Fischer, L. and Ramdas, A. (2024). Improving the (approximate) sequential probability ratio test by avoiding overshoot. arXiv preprint arXiv:2410.16076

  4. [12]

    and Ramdas, A

    Fischer, L. and Ramdas, A. (2025). Sequential monte carlo testing by betting. Journal of the Royal Statistical Society: Series B (Statistical Methodology) . Advance online publication

  5. [13]

    Gr \"u nwald, P., de Heide, R., and Koolen, W. (2024). Safe testing. Journal of the Royal Statistical Society Series B: Statistical Methodology , 86(5):1091--1128

  6. [14]

    W., Bojinov, I., Lindon, M., and Tingley, M

    Ham, D. W., Bojinov, I., Lindon, M., and Tingley, M. (2022). Design-based confidence sequences for anytime-valid causal inference. arXiv preprint arXiv:2210.08639 , 11

  7. [15]

    and Law, M

    Henzi, A. and Law, M. (2024). A rank-based sequential test of independence. Biometrika , 111(4):1169--1186

  8. [16]

    R., Ramdas, A., McAuliffe, J., and Sekhon, J

    Howard, S. R., Ramdas, A., McAuliffe, J., and Sekhon, J. (2021). Time-uniform, nonparametric, nonasymptotic confidence sequences. The Annals of Statistics , 49(2):1055--1080

  9. [17]

    Kelly, J. L. (1956). A new interpretation of information rate. The Bell System Technical Journal , 35(4):917--926

  10. [18]

    Koning, N. W. (2024a). Measuring evidence with a continuous test. arXiv preprint arXiv:2409.05654

  11. [19]

    Koning, N. W. (2024b). Post-hoc and anytime valid permutation and group invariance testing. arXiv preprint arXiv:2310.01153

  12. [20]

    Koolen, W. M. and Gr \"u nwald, P. (2022). Log-optimal anytime-valid e-values. International Journal of Approximate Reasoning , 141:69--82

  13. [21]

    Larsson, M., Ramdas, A., and Ruf, J. (2024). The numeraire e-variable and reverse information projection. arXiv preprint arXiv:2402.18810

  14. [22]

    Z., Sinha, M., Addanki, R., Ramdas, A., Garg, M., and Swaminathan, V

    Maharaj, A., Sinha, R., Arbour, D., Waudby-Smith, I., Liu, S. Z., Sinha, M., Addanki, R., Ramdas, A., Garg, M., and Swaminathan, V. (2023). Anytime-valid confidence sequences in an enterprise a/b testing platform. In Companion Proceedings of the ACM Web Conference 2023 , pages...

  15. [23]

    Nagler, T. (2018). Asymptotic analysis of the jittering kernel density estimator. Mathematical Methods of Statistics , 27:32--46

  16. [24]

    Ramdas, A., Gr \"u nwald, P., Vovk, V., and Shafer, G. (2023). Game-theoretic statistics and safe anytime-valid inference. Statistical Science , 38(4):576--601

  17. [25]

    and Manole, T

    Ramdas, A. and Manole, T. (2023). Randomized and exchangeable improvements of markov's, chebyshev's and chernoff's inequalities. arXiv preprint arXiv:2304.02611

  18. [26]

    Ramdas, A., Ruf, J., Larsson, M., and Koolen, W. (2020). Admissible anytime-valid sequential inference must rely on nonnegative martingales. arXiv preprint arXiv:2009.03167

  19. [27]

    Ramdas, A., Ruf, J., Larsson, M., and Koolen, W. M. (2022). Testing exchangeability: Fork-convexity, supermartingales and e-processes. International Journal of Approximate Reasoning , 141:83--109

  20. [28]

    Shafer, G., Shen, A., Vereshchagin, N., and Vovk, V. (2011). Test martingales, bayes factors and p-values. Statistical Science , pages 84--101

  21. [29]

    Ville, J. (1939). Étude critique de la notion de collectif . PhD thesis, University of Paris

  22. [30]

    Vovk, V. (2021). Testing randomness online. Statistical Science , 36(4):595--611

  23. [31]

    Vovk, V. (2023). The power of forgetting in statistical hypothesis testing. In Conformal and Probabilistic Prediction with Applications , pages 347--366. PMLR

  24. [32]

    Vovk, V., Nouretdinov, I., and Gammerman, A. (2003). Testing exchangeability on-line. In Proceedings of the 20th international conference on machine learning (ICML-03) , pages 768--775

  25. [33]

    and Wang, R

    Vovk, V. and Wang, R. (2021). E-values: Calibration, combination and applications. The Annals of Statistics , 49(3):1736--1754

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.