{"id":"ff553131-c358-4e1d-a0ed-401fdfcaa205","arxiv_id":"2411.10620","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A two-stage doubly robust estimator for causal excursion effects in micro-randomized trials keeps consistency when either the missingness model or the outcome regression is correctly specified.","lead":"This paper builds a two-stage doubly robust estimator for causal excursion effects in micro-randomized trials when outcomes are missing at random, consistent if either the missingness or the outcome model is right. It matters because mobile health trials routinely lose outcome data, and common imputation or complete-case analyses can bias intervention effect estimates.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 3.1 defines the outcome regression as E(Y_t,Δ|H_t,A_t), but identification (3.2) and the W_t,Δ-weighted estimating equations require E(W_t,ΔY_t,Δ|H_t,A_t); for Δ>1 and π≠p these differ, so the general-Δ double-robustness claim is unproven and likely false as stated.","rationale":"The reader's weakest assumption and my independent reading point to the same issue: the outcome regression used in the estimating equations is not the W_t,Δ-weighted regression required by the identified CEE for Δ>1 and π≠p. This is not a disagreement with a scientific consensus; it is an internal definitional inconsistency in the paper's general-Δ claims. If the concern lands, the proposed estimator is not doubly robust for Δ>1 with a non-p future regime, and the stated contribution needs either a corrected μ_t=E(W_t,ΔY_t,Δ|H_t,A_t) with corresponding changes to (3.4)–(3.5) and the proofs, or an explicit restriction of the claims to Δ=1 or π=p. For Δ=1, the estimator reduces to a standard augmented inverse-probability-weighted CEE estimator, and the simulations and application support it. The paper also assumes MAR (Assumption 2(ii)) and positivity, which are explicit assumptions rather than hidden gaps. Since the general-Δ claim is central to the abstract but the Δ=1 contribution remains solid, conditional acceptance is the appropriate recommendation, matching the reader's verdict.","tokens_in":13374,"tokens_out":11092,"duration_ms":110446,"concrete_test":"Re-run the Section 5 simulation with Δ=2 and π_t=0 for all t, using oracle nuisance functions: e_t=e⋆_t and μ_t=m_t=E(Y_t,2|H_t,A_t). If Algorithm 1 with (3.4) yields bias that does not vanish as n=50,100,200,400, the general-Δ claim fails; if the bias vanishes, the concern is wrong. Independently, inspect the supplementary proof of Theorems 1 and 2 for the step where E(W_t,ΔY_t,Δ|H_t,A_t) is replaced by E(Y_t,Δ|H_t,A_t); if that replacement occurs without W_t,Δ=1, the proof has a gap.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The load-bearing issue is a definitional mismatch in the outcome regression. Section 3.1 sets μ⋆_t(h,a)=E(Y_t,Δ|H_t=h,A_t=a), and this μ enters the estimating function U_t through ε in (3.4)/(3.5). But the identification result (3.2) expresses the CEE as a contrast of q_t(h,a):=E(W_t,ΔY_t,Δ|H_t=h,A_t=a) — equivalently the π-regime mean E_π(Y_t,Δ|H_t,A_t) — not of the unweighted mean m_t(h,a)=E(Y_t,Δ|H_t,A_t). The two coincide only when W_t,Δ=1, i.e. Δ=1 or π=p. For Δ>1 with a non-p future regime (e.g. π_t=0, the Qian et al. excursion), W_t,Δ is a nontrivial likelihood ratio depending on future treatments and covariates that are also correlated with Y_t,Δ, so q≠m. In the double-robustness algebra, the augmentation term must satisfy E{U_t(β⋆,μ⋆,p̃)|H_t,A_t}=0 when μ⋆ is the correct regression. With μ=m, this conditional moment is generally nonzero because it contains terms such as q_a − p f^Tβ − (1−p)m_1 − p m_0 rather than the q-based version; hence even the fully observed version of U_t is not unbiased for β⋆ when Δ>1 and π≠p. The paper's simulations (Section 5) and application (Section 6) only use Δ=1, and the supplementary proof was not independently checked. Thus the abstract's general claim of double robustness for CEE with missing longitudinal outcomes is unsubstantiated for Δ>1, and as stated the algorithm is likely inconsistent in that regime.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-stage, doubly robust estimator for the causal excursion effect (CEE) in micro-randomized trials when longitudinal outcomes are missing at random. It defines a stabilized inverse-probability-weighted estimating function and augments it with an outcome-regression term in the style of Robins, Rotnitzky, and Zhao. The authors claim that the estimator is consistent and asymptotically normal if either the missingness model or the outcome regression model is correctly specified, for identity or log links, for both parametric and nonparametric nuisance estimation. Simulation results for Δ=1 with continuous outcomes support the claim under correctly specified or one-sided misspecification models, and an application to HeartSteps illustrates the method. The central claim is double robustness for general excursion windows Δ≥1 and arbitrary future treatment regimes π, but the formal development in the main text only verifies the required algebra for the fully observed, unweighted outcome regression, which agrees with the excursion-weighted regression only when Δ=1 or π=p.","tokens_in":13775,"tokens_out":2865,"duration_ms":35549,"significance":"If the general-Δ double-robustness claim is correct, this would be a useful contribution: it would provide a principled alternative to ad hoc imputation for missing outcomes in MRTs, with a simple two-stage construction and explicit asymptotic variance estimators. The paper's Delta=1 results appear internally consistent, the estimating equation algebra for the augmented missing-data term is standard, and the simulation study is carefully designed with multiple generating functions and implementations. The availability of replication code and the inclusion of a real-data comparison against common imputation approaches are strengths. However, the manuscript's headline claim is broader than what is actually proved, because the outcome regression is not defined as the excursion-weighted regression required by the identification result for Δ>1 and non-π=p regimes.","major_comments":[{"comment":"The double-robustness claim for general Δ>1 is not supported by the definitions given. Identification (3.2) expresses the CEE as a contrast of q_t(h,a)=E(W_{t,Δ}Y_{t,Δ}|H_t=h,A_t=a), but Section 3.1 defines the outcome regression as μ*_t(h,a)=E(Y_{t,Δ}|H_t=h,A_t=a), and this unweighted μ is the object estimated in Algorithm 1 and used in the augmenting term of (3.4)-(3.5). These two regressions coincide only when W_{t,Δ}=1, i.e. Δ=1 or π=p. For Δ>1 with a regime such as π=0, the augmentation algebra requires the conditional moment E{Ũ_t(β*, μ*)|H_t,A_t}=0 to hold at the true weighted regression q_t; with μ=m the displayed estimating function contains terms of the form q_a − p f^Tβ − (1−p)m_1 − p m_0, which do not generally vanish. As written, the estimator is therefore likely inconsistent for Δ>1 and π≠p even when the unweighted outcome regression is correctly specified. The manuscript should either restrict all claims to Δ=1, or redefine μ as E(W_{t,Δ}Y_{t,Δ}|H_t,A_t) (or the π-regime conditional mean) and reprove Theorem 1 and Theorem 2 under that definition.","section":"Section 3.1 and Eq. (3.2)-(3.5)"},{"comment":"The numerical evidence does not exercise the general-Δ claim. Section 5 explicitly states that all simulations focus on Δ=1, and the HeartSteps application uses the 30-minute proximal outcome, which also corresponds to Δ=1. There is therefore no simulation or data analysis that checks consistency, coverage, or double robustness for Δ>1 with π≠p. Given the definitional issue in the previous comment, the empirical work cannot distinguish between the claimed general theorem and a Δ=1-only result. At minimum, the abstract and theorems should be restated to the Δ=1 setting, or the manuscript should add a simulation with Δ>1 and a non-π=p regime.","section":"Section 5 and Section 6"}],"minor_comments":[{"comment":"The phrase 'identify or log link' should be 'identity or log link'.","section":"Abstract"},{"comment":"In the definition of Φ(θ, p̃), the second and third blocks are both written as U_e(γ_e); the third block should presumably be the estimating function for the outcome regression nuisance parameters, U_μ(γ_μ). This typo makes the displayed sandwich variance formula ambiguous.","section":"Theorem 1"},{"comment":"Stage 1 says to fit Pr(A_t=1|S_t, I_t=1), while Section 3.1 denotes the corresponding nuisance function as p̃_t(S_t); the notation is consistent in substance but the algorithm should also state that this is the stabilized numerator probability, not a model for the true randomization probability.","section":"Algorithm 1, Stage 1"},{"comment":"The four-panel figures compress three generating functions and four sample sizes into small panels; labeling each panel with the implementation (A-D) directly in the plot, rather than only in Table 2 and the figure legend, would improve readability.","section":"Section 5, Figures 1 and 2"}],"recommendation":"major_revision","confidential_remarks":"The core Δ=1 construction appears sound and the simulations are persuasive for that case. However, the abstract and theorems make an unqualified claim for general Δ, and the definitional mismatch in Section 3.1 appears to make the general-Δ claim false as stated. This is not a merely cosmetic issue; it affects the central contribution. I would like the authors to either sharply restrict the paper to Δ=1, or provide a corrected excursion-weighted outcome regression and a full proof for Δ>1. If they choose the latter, they should also add Δ>1 simulations with π≠p. The manuscript is otherwise well-written and the topic is suitable for the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper delivers a useful contribution for the common case Δ=1: a two-stage doubly robust estimator for causal excursion effects in MRTs with longitudinal outcomes missing at random. The construction is a clean adaptation of standard AIPW ideas to the CEE setting, the simulations support consistency and coverage when one nuisance model is correct, and the HeartSteps application is a reasonable illustration. The writing is clear, and the authors correctly position this against the existing IPW-only approach of Shi and Dempsey.\n\nThe stress-test concern is real and worth taking seriously. Section 3.1 defines μ as E(Y|H,A), but identification (3.2) and the estimating equations require the excursion-weighted regression E(WY|H,A). For Δ>1 with π≠p, these differ, and the double-robustness algebra falls apart. The abstract's general claim of double robustness is therefore unproven and likely false as stated. The fact that the theory and simulations only exercise Δ=1 is not incidental; the authors appeal to Cheng et al.'s result for fully observed outcomes that holds for Δ=1, but they do not carry that restriction into their own abstract or theorems.\n\nThe fix is straightforward conceptually: either restrict the claims to Δ=1 or π=p, or redefine the outcome regression as the W-weighted expectation. The latter would make the method genuinely general but requires reworking the proofs and simulations. As it stands, the paper is a solid contribution for the immediate-outcome setting, which is the most common in practice, but the overclaim will mislead readers if not corrected.\n\nI also note the MAR assumption is strong but standard and explicitly acknowledged. The positivity assumptions are reasonable for MRTs. The self-citation pattern is unremarkable; the cited prior work is directly relevant.\n\nThis deserves a serious referee. The referee should push the authors to clarify the Δ>1 case, but the paper is publishable in revised form, even if only for Δ=1.","headline":"Solid doubly robust estimator for Δ=1 CEE with missing outcomes, but the general-Δ claim overreaches and needs a corrected outcome regression.","tokens_in":14281,"tokens_out":3264,"would_cite":true,"duration_ms":37349,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D20","62F12","62G05"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes the first doubly robust estimator for causal excursion effects in micro-randomized trials when longitudinal outcomes are missing at random, consistent when either the missingness model or the outcome regression is…","keywords":["causal excursion effect","micro-randomized trial","missing at random","doubly robust estimation","mobile health","longitudinal outcomes","two-stage estimation"],"falsifier":"Evaluate the estimating equation at $\\Delta=2$ with a future-treatment policy $\\pi\\ne p$, generating data where $E(W_{t,2}Y_{t,2}\\mid H_t,A_t)\\ne E(Y_{t,2}\\mid H_t,A_t)$, then misspecify the outcome regression while keeping the missingness model correct; if $\\hat\\beta$ shows bias or confidence intervals under-cover, the stated double-robustness property does not hold as written for delayed effects.","tokens_in":13166,"feed_emoji":"📲","tokens_out":9839,"duration_ms":87660,"temperature":0.7,"pith_summary":"Micro-randomized trials randomize treatments many times per person to optimize mobile health interventions, and the causal excursion effect (CEE) measures how a treatment changes a subsequent outcome under a policy of interest. Missing outcomes—missed self-reports, sensors not worn—are common and can bias CEE estimates. This paper proposes a two-stage estimator for the CEE that is doubly robust: it remains consistent and asymptotically normal if either the missingness model or the outcome regression model is correctly specified. The method covers identity and log link functions, hence continuous, binary, and count outcomes, and allows the nuisance models to be fitted by flexible nonparametric or machine-learning methods. If the claim is right, MRT analyses no longer need to bet on a single correct missingness model.","feed_headline":"One correct model suffices for missing-data MRT estimates","feed_subtitle":"Two-stage estimator stays consistent when either the missingness model or the outcome regression is right.","key_machinery":"The key identity is the augmented missing-data score $\\tilde U_t$ in (3.3), which combines the stabilized treatment weight $W_t$, the excursion weight $W_{t,\\Delta}$ (a product of likelihood-ratio terms for future treatments under policy $\\pi$ versus the randomization policy $p$), and the augmentation term $E\\{U_t(\\beta,\\mu_t,\\tilde p_t)\\mid H_t,A_t\\}$. The augmentation is what converts inverse probability weighting into double robustness: it has mean zero under a correct missingness model and exactly cancels the bias from a wrong missingness model when the outcome regression is correct. Theorems 1 and 2 then provide the asymptotic distribution, with Theorem 2 relying on the product-rate condition $\\|\\hat e-e^\\star\\|\\,\\|\\hat\\mu-\\mu^\\star\\|=o_p(n^{-1/2})$ for data-adaptive nuisance estimators.","core_discovery":"At the center is the doubling of protection provided by equation (3.3). The proposed estimator weights the CEE score by the inverse probability of observing the outcome, then subtracts the conditional expectation of that weighted score given history and treatment. Because the subtracted term corrects the misspecified nuisance in one direction, the score stays unbiased when the missingness model is correct even if the outcome regression is wrong, and stays unbiased when the outcome regression is correct even if the missingness model is wrong. Theorems 1 and 2 state this as consistency and asymptotic normality with parametric and nonparametric nuisance estimation, respectively, under missing at random and positivity. Simulations with continuous outcomes confirm low bias and near-nominal coverage when either one nuisance model is misspecified, and the HeartSteps analysis shows the estimator in a real 9.4% missingness setting.","pith_inferences":["When the horizon $\\Delta>1$ and the future policy $\\pi$ differs from the randomization probabilities, the double-robustness argument as written appears to require the excursion-weighted regression $E(W_{t,\\Delta}Y_{t,\\Delta}\\mid H_t,A_t)$ inside the augmentation, while the paper defines the outcome regression as $E(Y_{t,\\Delta}\\mid H_t,A_t)$; the two coincide only when $W_{t,\\Delta}=1$, so the del","It would be informative to stress-test the estimator in a $\\Delta=2$, $\\pi\\ne p$ simulation where the outcome regression is misspecified and the missingness model is correct: visible bias would confirm that the excursion-weighted target is the necessary augmentation.","The augmentation term subtracts the conditional mean of the weighted score, so even when missingness is low the method can reduce variance relative to plain inverse-probability weighting; this gain should grow as the missingness rate increases.","Because the estimator requires missing at random, a natural next step would be to replace the augmentation with a shadow-variable construction to allow missing-not-at-random mechanisms, preserving the same two-stage structure."],"forward_implications":["Analyses of MRTs with missing-at-random outcomes can replace complete-case analysis or ad hoc imputation with an estimator that is consistent when either the missingness model or the outcome regression is correct.","Because the same estimating equations handle identity and log links, the method covers continuous, binary, and count outcomes without separate developments.","Stage-1 nuisance fits can be flexible: with nonparametric estimators, the estimator remains $\\sqrt{n}$-consistent and asymptotically normal under the product-rate condition $\\|\\hat e-e^\\star\\|\\,\\|\\hat\\mu-\\mu^\\star\\|=o_p(n^{-1/2})$.","Sandwich variance estimators from Theorems 1 and 2 provide Wald confidence intervals with near-nominal coverage in the simulation settings, including when exactly one nuisance model is misspecified."],"supporting_citations":[{"why":"Introduces the micro-randomized trial design that the causal excursion effect and this estimator target.","marker":"[Klasnja et al., 2015]"},{"why":"Documents missed self-reports and sensor non-wear as the missingness mechanisms that motivate Assumption 2.","marker":"[Seewald et al., 2019]"},{"why":"Provides the semiparametric inverse-probability plus augmentation structure that equation (3.3) adapts to the MRT setting.","marker":"[Robins et al., 1994]"},{"why":"Establishes doubly robust estimation for missing data and causal inference, the template for requiring only one of the missingness or outcome model to be correct.","marker":"[Bang and Robins, 2005]"},{"why":"Defines the weighted and centered least squares estimator for CEE with fully observed outcomes, which the identity-link version generalizes.","marker":"[Boruvka et al., 2018]"},{"why":"Shows the fully-observed two-stage CEE estimator is globally robust for $\\Delta=1$, the baseline the proposed method extends to missing outcomes.","marker":"[Cheng et al., 2023]"},{"why":"Defines the marginal excursion effect for binary outcomes under log link, the source of the log-link estimating equation.","marker":"[Qian et al., 2021]"},{"why":"Supplies the product-rate condition and cross-fitting framework used to allow nonparametric nuisance estimators in Theorem 2.","marker":"[Chernozhukov et al., 2018]"},{"why":"Provides the HeartSteps MRT data used to illustrate the estimator in Section 6.","marker":"[Klasnja et al., 2019]"}],"fun_headline_variants":["Double robustness: Either model works for MRT missing data","Missing MRT outcomes? One correct model suffices","Causal effects in MRTs with missing data: doubly robust","One right model fixes missing-data bias in MRTs","Doubly robust CEE estimates survive missing outcomes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central guarantee rests on missingness being explainable by history and treatment, and on the outcome regression targeting the plain conditional mean even when estimating delayed effects under a different future policy; if either fails, the double-protection property can collapse.","fun_headline_variants_meta":{"raw":{"variants":["Double robustness: Either model works for MRT missing data","Missing MRT outcomes? One correct model suffices","Causal effects in MRTs with missing data: doubly robust","One right model fixes missing-data bias in MRTs","Doubly robust CEE estimates survive missing outcomes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000593,"raw_usage":{"total_tokens":2746,"prompt_tokens":882,"completion_tokens":1864,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":1784}},"tokens_in":498,"tokens_out":1864,"duration_ms":12561,"temperature":1.0,"reasoning_tokens":1784,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:31:17.507202+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Evaluate the estimating equation at $\\Delta=2$ with a future-treatment policy $\\pi\\ne p$, generating data where $E(W_{t,2}Y_{t,2}\\mid H_t,A_t)\\ne E(Y_{t,2}\\mid H_t,A_t)$, then misspecify the outcome regression while keeping the missingness model correct; if $\\hat\\beta$ shows bias or confidence intervals under-cover, the stated double-robustness property does not hold as written for delayed effects.","supporting_citations":[{"cited_title":"B., Shiffman, S., Boruvka, A., Almirall, D., Tewari, A., and Murphy, S","cited_arxiv_id":null,"evidence_quote":"Introduces the micro-randomized trial design that the causal excursion effect and this estimator target."},{"cited_title":"J., Smith, S","cited_arxiv_id":null,"evidence_quote":"Documents missed self-reports and sensor non-wear as the missingness mechanisms that motivate Assumption 2."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the marginal excursion effect for binary outcomes under log link, the source of the log-link estimating equation."},{"cited_title":"J., Lee, A., Hall, K., Luers, B., Hekler, E","cited_arxiv_id":null,"evidence_quote":"Provides the HeartSteps MRT data used to illustrate the estimator in Section 6."}],"review_version":1}