Pith. sign in

REVIEW 2 major objections 4 minor 41 references

Doubly Robust Estimation of Causal Excursion Effects in Micro-Randomized Trials with Missing Longitudinal Outcomes

T0 review · 2 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper proposes the first doubly robust estimator for causal excursion effects in micro-randomized trials when longitudinal outcomes are missing at random, consistent when either the missingness model or the outcome regression is…

desk verdict Solid doubly robust estimator for Δ=1 CEE with missing outcomes, but the general-Δ claim overreaches and needs a corrected outcome regression. read the letter →

arxiv 2411.10620 v1 pith:UERHYNFK submitted 2024-11-15 stat.ME

classification stat.ME MSC 62D2062F1262G05
keywords causalexcursioneffectmicro-randomizedtrialmissingatrandomdoublyrobustestimationmobilehealthlongitudinaloutcomestwo-stage
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Micro-randomized trials randomize treatments many times per person to optimize mobile health interventions, and the causal excursion effect (CEE) measures how a treatment changes a subsequent outcome under a policy of interest. Missing outcomes—missed self-reports, sensors not worn—are common and can bias CEE estimates. This paper proposes a two-stage estimator for the CEE that is doubly robust: it remains consistent and asymptotically normal if either the missingness model or the outcome regression model is correctly specified. The method covers identity and log link functions, hence continuous, binary, and count outcomes, and allows the nuisance models to be fitted by flexible nonparametric or machine-learning methods. If the claim is right, MRT analyses no longer need to bet on a single correct missingness model.

What carries the argument

The key identity is the augmented missing-data score $\tilde U_t$ in (3.3), which combines the stabilized treatment weight $W_t$, the excursion weight $W_{t,\Delta}$ (a product of likelihood-ratio terms for future treatments under policy $\pi$ versus the randomization policy $p$), and the augmentation term $E\{U_t(\beta,\mu_t,\tilde p_t)\mid H_t,A_t\}$. The augmentation is what converts inverse probability weighting into double robustness: it has mean zero under a correct missingness model and exactly cancels the bias from a wrong missingness model when the outcome regression is correct. Theorems 1 and 2 then provide the asymptotic distribution, with Theorem 2 relying on the product-rate condition $\|\hat e-e^\star\|\,\|\hat\mu-\mu^\star\|=o_p(n^{-1/2})$ for data-adaptive nuisance estimators.

What would settle it

Evaluate the estimating equation at $\Delta=2$ with a future-treatment policy $\pi\ne p$, generating data where $E(W_{t,2}Y_{t,2}\mid H_t,A_t)\ne E(Y_{t,2}\mid H_t,A_t)$, then misspecify the outcome regression while keeping the missingness model correct; if $\hat\beta$ shows bias or confidence intervals under-cover, the stated double-robustness property does not hold as written for delayed effects.

Watch

Extended reading notes

Core claim

At the center is the doubling of protection provided by equation (3.3). The proposed estimator weights the CEE score by the inverse probability of observing the outcome, then subtracts the conditional expectation of that weighted score given history and treatment. Because the subtracted term corrects the misspecified nuisance in one direction, the score stays unbiased when the missingness model is correct even if the outcome regression is wrong, and stays unbiased when the outcome regression is correct even if the missingness model is wrong. Theorems 1 and 2 state this as consistency and asymptotic normality with parametric and nonparametric nuisance estimation, respectively, under missing at random and positivity. Simulations with continuous outcomes confirm low bias and near-nominal coverage when either one nuisance model is misspecified, and the HeartSteps analysis shows the estimator in a real 9.4% missingness setting.

Load-bearing premise

The central guarantee rests on missingness being explainable by history and treatment, and on the outcome regression targeting the plain conditional mean even when estimating delayed effects under a different future policy; if either fails, the double-protection property can collapse.

Editorial extensions

If this is right

  • Analyses of MRTs with missing-at-random outcomes can replace complete-case analysis or ad hoc imputation with an estimator that is consistent when either the missingness model or the outcome regression is correct.
  • Because the same estimating equations handle identity and log links, the method covers continuous, binary, and count outcomes without separate developments.
  • Stage-1 nuisance fits can be flexible: with nonparametric estimators, the estimator remains $\sqrt{n}$-consistent and asymptotically normal under the product-rate condition $\|\hat e-e^\star\|\,\|\hat\mu-\mu^\star\|=o_p(n^{-1/2})$.
  • Sandwich variance estimators from Theorems 1 and 2 provide Wald confidence intervals with near-nominal coverage in the simulation settings, including when exactly one nuisance model is misspecified.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • When the horizon $\Delta>1$ and the future policy $\pi$ differs from the randomization probabilities, the double-robustness argument as written appears to require the excursion-weighted regression $E(W_{t,\Delta}Y_{t,\Delta}\mid H_t,A_t)$ inside the augmentation, while the paper defines the outcome regression as $E(Y_{t,\Delta}\mid H_t,A_t)$; the two coincide only when $W_{t,\Delta}=1$, so the del
  • It would be informative to stress-test the estimator in a $\Delta=2$, $\pi\ne p$ simulation where the outcome regression is misspecified and the missingness model is correct: visible bias would confirm that the excursion-weighted target is the necessary augmentation.
  • The augmentation term subtracts the conditional mean of the weighted score, so even when missingness is low the method can reduce variance relative to plain inverse-probability weighting; this gain should grow as the missingness rate increases.
  • Because the estimator requires missing at random, a natural next step would be to replace the augmentation with a shadow-variable construction to allow missing-not-at-random mechanisms, preserving the same two-stage structure.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper proposes a two-stage, doubly robust estimator for the causal excursion effect (CEE) in micro-randomized trials when longitudinal outcomes are missing at random. It defines a stabilized inverse-probability-weighted estimating function and augments it with an outcome-regression term in the style of Robins, Rotnitzky, and Zhao. The authors claim that the estimator is consistent and asymptotically normal if either the missingness model or the outcome regression model is correctly specified, for identity or log links, for both parametric and nonparametric nuisance estimation. Simulation results for Δ=1 with continuous outcomes support the claim under correctly specified or one-sided misspecification models, and an application to HeartSteps illustrates the method. The central claim is double robustness for general excursion windows Δ≥1 and arbitrary future treatment regimes π, but the formal development in the main text only verifies the required algebra for the fully observed, unweighted outcome regression, which agrees with the excursion-weighted regression only when Δ=1 or π=p.

Significance. If the general-Δ double-robustness claim is correct, this would be a useful contribution: it would provide a principled alternative to ad hoc imputation for missing outcomes in MRTs, with a simple two-stage construction and explicit asymptotic variance estimators. The paper's Delta=1 results appear internally consistent, the estimating equation algebra for the augmented missing-data term is standard, and the simulation study is carefully designed with multiple generating functions and implementations. The availability of replication code and the inclusion of a real-data comparison against common imputation approaches are strengths. However, the manuscript's headline claim is broader than what is actually proved, because the outcome regression is not defined as the excursion-weighted regression required by the identification result for Δ>1 and non-π=p regimes.

major comments (2)
  1. [Section 3.1 and Eq. (3.2)-(3.5)] The double-robustness claim for general Δ>1 is not supported by the definitions given. Identification (3.2) expresses the CEE as a contrast of q_t(h,a)=E(W_{t,Δ}Y_{t,Δ}|H_t=h,A_t=a), but Section 3.1 defines the outcome regression as μ*_t(h,a)=E(Y_{t,Δ}|H_t=h,A_t=a), and this unweighted μ is the object estimated in Algorithm 1 and used in the augmenting term of (3.4)-(3.5). These two regressions coincide only when W_{t,Δ}=1, i.e. Δ=1 or π=p. For Δ>1 with a regime such as π=0, the augmentation algebra requires the conditional moment E{Ũ_t(β*, μ*)|H_t,A_t}=0 to hold at the true weighted regression q_t; with μ=m the displayed estimating function contains terms of the form q_a − p f^Tβ − (1−p)m_1 − p m_0, which do not generally vanish. As written, the estimator is therefore likely inconsistent for Δ>1 and π≠p even when the unweighted outcome regression is correctly specified. The manuscript should either restrict all claims to Δ=1, or redefine μ as E(W_{t,Δ}Y_{t,Δ}|H_t,A_t) (or the π-regime conditional mean) and reprove Theorem 1 and Theorem 2 under that definition.
  2. [Section 5 and Section 6] The numerical evidence does not exercise the general-Δ claim. Section 5 explicitly states that all simulations focus on Δ=1, and the HeartSteps application uses the 30-minute proximal outcome, which also corresponds to Δ=1. There is therefore no simulation or data analysis that checks consistency, coverage, or double robustness for Δ>1 with π≠p. Given the definitional issue in the previous comment, the empirical work cannot distinguish between the claimed general theorem and a Δ=1-only result. At minimum, the abstract and theorems should be restated to the Δ=1 setting, or the manuscript should add a simulation with Δ>1 and a non-π=p regime.
minor comments (4)
  1. [Abstract] The phrase 'identify or log link' should be 'identity or log link'.
  2. [Theorem 1] In the definition of Φ(θ, p̃), the second and third blocks are both written as U_e(γ_e); the third block should presumably be the estimating function for the outcome regression nuisance parameters, U_μ(γ_μ). This typo makes the displayed sandwich variance formula ambiguous.
  3. [Algorithm 1, Stage 1] Stage 1 says to fit Pr(A_t=1|S_t, I_t=1), while Section 3.1 denotes the corresponding nuisance function as p̃_t(S_t); the notation is consistent in substance but the algorithm should also state that this is the stabilized numerator probability, not a model for the true randomization probability.
  4. [Section 5, Figures 1 and 2] The four-panel figures compress three generating functions and four sample sizes into small panels; labeling each panel with the implementation (A-D) directly in the plot, rather than only in Table 2 and the figure legend, would improve readability.

Circularity Check

0 steps flagged · score 1.0 of 10

No construction-level circularity: the DR estimator is derived from standard semiparametric theory and validated against simulated truth and external HeartSteps data; minor self-citations are building blocks, and the only substantive concern (a Delta>1 definitional mismatch) is a correctness gap, not a circular reduction.

full rationale

The derivation chain is not circular. The target beta_star is defined by the CEE model in (3.1), the fully observed estimating function U_t is an existing CEE score from Boruvka/Qian/Cheng, and the missing-data version U_tilde_t is obtained by the standard augmented inverse-probability-weighting construction of Robins et al. (1994), as shown in (3.3). The double-robustness claim (Assumption 5, Theorems 1 and 2) is a theorem about a moment condition, not a statement that beta is a renamed nuisance fit. Stage 1 nuisance fits are inputs, not the target, and the estimator is evaluated against a simulated truth and an external HeartSteps dataset, so it is not a fitted-input-called-prediction setup. The citations to Cheng et al. (2023) and Qian et al. (2021) are self-citations, but they supply building-block estimating functions for the fully observed case; the present paper's contribution, the missing-data augmentation and rate double-robustness theory, is independent of those citations and is not justified by them alone. I therefore find no circular step. One non-circular correctness risk should be flagged: Section 3.1 defines mu_star_t(h,a) := E(Y_t,Delta | H_t = h, A_t = a), while identification (3.2) requires E(W_t,Delta Y_t,Delta | H_t, A_t). These coincide only when W_t,Delta = 1, i.e., Delta = 1 or pi = p. For Delta > 1 with pi != p, the unweighted outcome regression is not the regression needed by the weighted estimating function, so the general-Delta double-robustness claim is unsubstantiated. This is a definitional mismatch and a correctness gap, not a circular reduction; the paper's own simulations and application use Delta = 1 only.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The estimator introduces no free constants; the target beta is estimated from data, and the nuisance models are estimated from data. The central assumptions are MAR, correct specification of at least one nuisance model, convergence of nuisance fits, and rate double robustness for nonparametric nuisance. No new physical or statistical entities are posited.

assumptions (5)
  • domain assumption Assumption 1: SUTVA, positivity of treatment randomization, and sequential ignorability.
    Identifies the causal excursion effect from observed data; positivity and sequential ignorability are satisfied by MRT design, SUTVA rules out interference.
  • domain assumption Assumption 2: positivity for missingness and missing at random, Y_t,Delta independent of R_t,Delta given H_t,A_t.
    Unverifiable in practice; needed for inverse-probability weighting of observed outcomes.
  • domain assumption Assumptions 3 and 5: nuisance estimators converge in L2 and at least one of e or mu is correctly specified.
    This is the double-robustness premise; the theorem requires either e'=e* or mu'=mu*.
  • domain assumption Theorem 2 product-rate condition: ||e_hat-e*|| ||mu_hat-mu*|| = op(n^{-1/2}).
    Standard rate double robustness for nonparametrically fitted nuisance functions.
  • standard math Assumption 6: compact parameter space, bounded support, invertible Jacobian, Donsker classes, propensity bounded away from 0 and 1.
    Regularity conditions for asymptotic normality; cross-fitting is noted as an alternative to the Donsker assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Doubly Robust Estimation of Causal Excursion Effects in Micro-Randomized Trials with Missing Longitudinal Outcomes." pith.science (2026). https://pith.science/paper/UERHYNFK

@misc{pith2026241110620,
  author       = {Pith},
  title        = {Pith review of: Doubly Robust Estimation of Causal Excursion Effects in Micro-Randomized Trials with Missing Longitudinal Outcomes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UERHYNFK}},
  note         = {Machine review of arXiv:2411.10620}
}
read the original abstract

Micro-randomized trials (MRTs) are increasingly utilized for optimizing mobile health interventions, with the causal excursion effect (CEE) as a central quantity for evaluating interventions under policies that deviate from the experimental policy. However, MRT often contains missing data due to reasons such as missed self-reports or participants not wearing sensors, which can bias CEE estimation. In this paper, we propose a two-stage, doubly robust estimator for CEE in MRTs when longitudinal outcomes are missing at random, accommodating continuous, binary, and count outcomes. Our two-stage approach allows for both parametric and nonparametric modeling options for two nuisance parameters: the missingness model and the outcome regression. We demonstrate that our estimator is doubly robust, achieving consistency and asymptotic normality if either the missingness or the outcome regression model is correctly specified. Simulation studies further validate the estimator's desirable finite-sample performance. We apply the method to HeartSteps, an MRT for developing mobile health interventions that promote physical activity.

Figures

Figures reproduced from arXiv: 2411.10620 by the authors.

Figure 1
Figure 1. Bias, mean squared error, and coverage probability of βˆ 0 in simulation. 16 [PITH_FULL_IMAGE:figures/full_fig_p016_1.png] view at source ↗
Figure 2
Figure 2. Bias, mean squared error, and coverage probability of βˆ 1 in simulation. 17 [PITH_FULL_IMAGE:figures/full_fig_p017_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

41 extracted references · 21 canonical work pages

  1. [1]

    and Robins, J

    Bang, H. and Robins, J. M. (2005). Doubly robust estimation in missing data and causal inference models. Biometrics , 61(4):962--973

  2. [2]

    W., Williamson, E., et al

    Bell, L., Garnett, C., Bao, Y., Cheng, Z., Qian, T., Perski, O., Potts, H. W., Williamson, E., et al. (2023). How notifications affect engagement with a behavior change app: Results from a micro-randomized trial. JMIR mHealth and uHealth , 11(1):e38342

  3. [3]

    Boruvka, A., Almirall, D., Witkiewitz, K., and Murphy, S. A. (2018). Assessing time-varying causal effect moderation in mobile health. Journal of the American Statistical Association , 113(523):1112--1121

  4. [4]

    A., and Davidian, M

    Cao, W., Tsiatis, A. A., and Davidian, M. (2009). Improving efficiency and robustness of the doubly robust estimator for a population mean with incomplete data. Biometrika , 96(3):723--734

  5. [5]

    Cheng, Z., Bell, L., and Qian, T. (2023). Efficient and globally robust causal excursion effect estimation. arXiv preprint arXiv:2311.16529

  6. [6]

    Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018). Double/debiased machine learning for treatment and structural parameters

  7. [7]

    Dempsey, W., Liao, P., Klasnja, P., Nahum-Shani, I., and Murphy, S. A. (2015). Randomised trials for the fitbit generation. Significance , 12(6):20--23

  8. [8]

    Dempsey, W., Liao, P., Kumar, S., and Murphy, S. A. (2020). The stratified micro-randomized trial design: sample size considerations for testing nested causal effects of time-varying treatments. The annals of applied statistics , 14(2):661

Show all 41 references
  1. [9]

    D \' az, I., Carone, M., and van der Laan, M. J. (2016). Second-order inference for the mean of a variable missing at random. The international journal of biostatistics , 12(1):333--349

  2. [10]

    and Gijbels, I

    Fan, J. and Gijbels, I. (1996). Local polynomial modelling and its applications: monographs on statistics and applied probability 66 , volume 66. CRC Press

  3. [11]

    Horowitz, J. L. (2009). Semiparametric and nonparametric methods in econometrics , volume 12. Springer

  4. [12]

    Hudgens, M. G. and Halloran, M. E. (2008). Toward causal inference with interference. Journal of the American Statistical Association , 103(482):832--842

  5. [13]

    Kennedy, E. H. (2016). Semiparametric theory and empirical processes in causal inference. Statistical causal inferences and their applications in public health research , pages 141--167

  6. [14]

    Kennedy, E. H. (2022). Semiparametric doubly robust targeted double machine learning: a review. arXiv preprint arXiv:2203.06469

  7. [15]

    B., Shiffman, S., Boruvka, A., Almirall, D., Tewari, A., and Murphy, S

    Klasnja, P., Hekler, E. B., Shiffman, S., Boruvka, A., Almirall, D., Tewari, A., and Murphy, S. A. (2015). Microrandomized trials: An experimental design for developing just-in-time adaptive interventions. Health Psychology , 34(S):1220

  8. [16]

    J., Lee, A., Hall, K., Luers, B., Hekler, E

    Klasnja, P., Smith, S., Seewald, N. J., Lee, A., Hall, K., Luers, B., Hekler, E. B., and Murphy, S. A. (2019). Efficacy of contextually tailored suggestions for physical activity: a micro-randomized optimization trial of heartsteps. Annals of Behavioral Medicine , 53(6):573--582

  9. [17]

    K., Lessler, J., and Stuart, E

    Lee, B. K., Lessler, J., and Stuart, E. A. (2011). Weight trimming and propensity score weighting. PloS one , 6(3):e18174

  10. [18]

    Lok, J. J. (2024). How estimating nuisance parameters can reduce the variance (with consistent variance estimation). Statistics in Medicine , 43(23):4456--4480

  11. [19]

    J., and Geng, Z

    Miao, W., Liu, L., Li, Y., Tchetgen Tchetgen, E. J., and Geng, Z. (2024). Identification and semiparametric efficiency theory of nonignorable missing data with a shadow variable. ACM/JMS Journal of Data Science , 1(2):1--23

  12. [20]

    Murphy, S. A. (2003). Optimal dynamic treatment regimes. Journal of the Royal Statistical Society Series B: Statistical Methodology , 65(2):331--355

  13. [21]

    E., Collins, L

    Qian, T., Walton, A. E., Collins, L. M., Klasnja, P., Lanza, S. T., Nahum-Shani, I., Rabbi, M., Russell, M. A., Walton, M. A., Yoo, H., et al. (2022). The microrandomized trial for developing digital interventions: Experimental design and data analysis considerations. Psycholo...

  14. [22]

    Qian, T., Xiang, S., and Cheng, Z. (2023). MRTAnalysis: Primary and Secondary Analyses for Micro-Randomized Trials . R package version 0.1.2

  15. [23]

    Qian, T., Yoo, H., Klasnja, P., Almirall, D., and Murphy, S. A. (2021). Estimating time-varying causal excursion effects in mobile health with binary outcomes. Biometrika , 108(3):507--527

  16. [24]

    Robins, J. (1986). A new approach to causal inference in mortality studies with a sustained exposure period—application to control of the healthy worker survivor effect. Mathematical modelling , 7(9-12):1393--1512

  17. [25]

    Robins, J. M. (2004). Optimal structural nested models for optimal sequential decisions. In Proceedings of the Second Seattle Symposium in Biostatistics: analysis of correlated data , pages 189--326. Springer

  18. [26]

    M., Rotnitzky, A., and Zhao, L

    Robins, J. M., Rotnitzky, A., and Zhao, L. P. (1994). Estimation of regression coefficients when some regressors are not always observed. Journal of the American statistical Association , 89(427):846--866

  19. [27]

    Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of educational Psychology , 66(5):688

  20. [28]

    P., and Carroll, R

    Ruppert, D., Wand, M. P., and Carroll, R. J. (2003). Semiparametric regression . Cambridge University Press

  21. [29]

    O., Rotnitzky, A., and Robins, J

    Scharfstein, D. O., Rotnitzky, A., and Robins, J. M. (1999). Adjusting for nonignorable drop-out using semiparametric nonresponse models. Journal of the American Statistical Association , 94(448):1096--1120

  22. [30]

    J., Smith, S

    Seewald, N. J., Smith, S. N., Lee, A. J., Klasnja, P., and Murphy, S. A. (2019). Practical considerations for data collection and management in mobile health micro-randomized trials. Statistics in biosciences , 11:355--370

  23. [31]

    and Dempsey, W

    Shi, J. and Dempsey, W. (2023). A meta-learning method for estimation of causal excursion effects to assess time-varying moderation. arXiv preprint arXiv:2306.16297

  24. [32]

    Shi, J., Wu, Z., and Dempsey, W. (2023). Assessing time-varying causal effect moderation in the presence of cluster-level treatment effect heterogeneity and interference. Biometrika , 110(3):645--662

  25. [33]

    Smucler, E., Rotnitzky, A., and Robins, J. M. (2019). A unifying approach for doubly-robust l_1 regularized estimation of causal contrasts. arXiv preprint arXiv:1904.03737

  26. [34]

    Tan, Z. (2010). Bounded, efficient and doubly robust estimation with inverse weighting. Biometrika , 97(3):661--682

  27. [35]

    Tsiatis, A. A. (2006). Semiparametric theory and missing data , volume 4. Springer

  28. [36]

    van der Laan, M. (2017). A generally efficient targeted minimum loss based estimator based on the highly adaptive lasso. The international journal of biostatistics , 13(2)

  29. [37]

    J., Rose, S., Zheng, W., and van der Laan, M

    van der Laan, M. J., Rose, S., Zheng, W., and van der Laan, M. J. (2011). Cross-validated targeted minimum-loss-based estimation. Targeted learning: causal inference for observational and experimental data , pages 459--474

  30. [38]

    and Vansteelandt, S

    Vermeulen, K. and Vansteelandt, S. (2015). Bias-reduced doubly robust estimation. Journal of the American Statistical Association , 110(511):1024--1036

  31. [39]

    E., and Thall, P

    Wang, L., Rotnitzky, A., Lin, X., Millikan, R. E., and Thall, P. F. (2012). Evaluation of viable dynamic treatment regimes in a sequentially randomized trial of advanced prostate cancer. Journal of the American Statistical Association , 107(498):493--508

  32. [40]

    Wood, S. N. (2017). Generalized additive models: an introduction with R . CRC press

  33. [41]

    Yang, S., Wang, L., and Ding, P. (2019). Causal inference with confounders missing not at random. Biometrika , 106(4):875--888

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.