Pith. sign in

REVIEW 3 major objections 2 minor 28 references

Discussion of "Causal and counterfactual views of missing data models" by Razieh Nabi, Rohit Bhattacharya, Ilya Shpitser, & James M. Robins

T0 review · 3 major / 2 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Under the permutation missingness model, the mean of a partially missing variable can be estimated at the parametric rate with local efficiency, and binary outcomes reduce to a posterior odds calculation.

desk verdict A useful discussion with a real new influence function for the permutation MNAR model, but the efficiency claim is under-supported as written. read the letter →

arxiv 2506.13025 v1 pith:S3ZSCZID submitted 2025-06-16 stat.ME

classification stat.ME MSC 62D1062G0562G20
keywords missingnotatrandompermutationmissingnessmodelinfluencefunctionsemiparametricefficiencyone-stepestimationcounterfactualindependencesingleworldinterventiongraphs
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This discussion argues that the causal and counterfactual framework of the paper under discussion can be pushed in two directions: borrowing additional causal identification tools for missing-not-at-random problems, and constructing efficient estimators for target functionals rather than the full data law. Its central concrete result concerns the permutation missingness model: the mean $\psi = E(Y(1))$ of the partially missing variable is identified by a ratio of conditional expectations, and the piece $\theta = E(Y(1) \mid R_1 = 0)$ satisfies a von Mises expansion with an explicit influence function. That expansion yields a one-step estimator that is $\sqrt{n}$-consistent and locally efficient, so a researcher who trusts the model's independence restrictions does not need to estimate the entire observed-data distribution. When $Y$ is binary, the estimand simplifies to a posterior odds calculation, and the influence function becomes a simple product of odds and density ratios.

What carries the argument

The load-bearing object is the permutation missingness model's pair of nonparametric independence restrictions: $R_1 \perp\!\!\perp Y(1) \mid X(1)$ and $R_2 \perp\!\!\perp (Y(1), X(1)) \mid Y, R_1$. These restrictions identify the full data law and, via Bayes' rule, imply the ratio formula for $\theta$. The analytical engine is the von Mises expansion of the functional $\theta = E(\beta(X)/\alpha(X) \mid R_1 = 0, R_2 = 1)$, whose influence function is derived in Lemma 1; the remainder term is explicitly displayed, making the second-order rate transparent. In the binary case the key simplification is the odds identity $\xi(X) = \lambda(X) \times \text{odds}(Y = 1 \mid R_1 = R_2 = 1)$, where $\lambda$ is a density ratio, so the mean functional becomes the posterior probability of $Y = 1$ from prior odds times likelihood ratio, standardized to the $R_1 = 0, R_2 = 1$ group.

What would settle it

Generate data from a model that satisfies both independence restrictions, compute the proposed one-step estimator with correctly specified nuisance functions on many samples, and check that the standardized estimator is approximately standard normal; a systematic deviation would disprove the influence function.

Watch

Extended reading notes

Core claim

Under the permutation missingness model, the discussion establishes that the target mean $\psi = E(Y(1))$ is identified from observed data $(X, R, Y)$ by $\psi = P(R_1 = 1)E(Y \mid R_1 = 1) + P(R_1 = 0)\theta$, where $\theta$ is a ratio of conditional expectations involving $\zeta(Y) = P(R_2 = 1 \mid R_1 = 1, Y)$. Proposition 1 states this identification, and Lemma 1 shows that $\theta$ has the von Mises expansion $\theta(P) - \theta(P) = \int \varphi(o; P) \, d(P - P)(o) + R_\theta(P; P)$ with the displayed influence function $\varphi(O; P)$; this gives the local asymptotic minimax lower bound. For binary $Y$, $\theta = E\{\xi(X)/(\rho + \xi(X)) \mid R_1 = 0, R_2 = 1\}$, where $\xi$ is a conditional odds and $\rho$ an odds ratio, interpreted as a posterior odds combining prior information from the $R_1 = 1$ group with a likelihood ratio from the doubly observed group. Corollary 2 provides the corresponding influence function in the binary case, and the resulting one-step estimator is displayed.

Load-bearing premise

The permutation model assumes that $R_1$ is unrelated to the unobserved outcome given the first covariates, and that $R_2$ is unrelated to both unobserved variables once $Y$ and $R_1$ are known; if either is false, the identifying formula and the efficient estimator are invalid.

Editorial extensions

If this is right

  • Under the permutation model, estimating $\psi$ does not require estimating the joint distribution of $(X(1), Y(1))$; the closed-form ratio in Proposition 1 suffices.
  • The influence function in Lemma 1 gives the efficiency bound for $\theta$, so the one-step estimator attains local asymptotic minimax optimality.
  • With binary $Y$, the estimand has a direct posterior odds reading, making the missing-not-at-random adjustment transparent to practitioners.
  • Replacing the unknown odds ratio $\rho$ by its plug-in estimate preserves $\sqrt{n}$-consistency and asymptotic normality, adding only an extra asymptotically linear term.
  • SWIGs (m-SWIGs) can deliver counterfactual independence conditions such as $A(1) \perp\!\!\perp (Y^{a(1)}, R^{a(1)}) \mid X$ that are not directly visible from m-DAG d-separation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same ratio-of-expectations structure suggests that double/debiased machine learning with flexible nuisance estimators could be applied directly; the discussion only sketches the one-step estimator.
  • The posterior odds interpretation implies the estimand is monotone in the prior odds and in the density ratio $\lambda(x)$, which could be used to construct sensitivity bounds under partial violations of the independence assumptions.
  • The explicit remainder in the von Mises expansion shows products of second-order nuisance errors, the structure required for rate double robustness; the discussion does not develop this point.
  • A natural extension is to models with more than two missingness indicators, where the same permutation-style independence restrictions would yield recursive ratio formulas; the paper does not state this.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. This discussion paper responds to Nabi et al. (2022) in three directions: (i) it surveys causal-inference tools (instrumental variables, shadow variables, negative controls) that could transfer to missing-data problems; (ii) it introduces the idea of m-SWIGs for simultaneous interventions on treatment and missingness, illustrated on a missing-exposure example; and (iii) it derives identification and estimation results for the permutation missingness model of Robins (1997), which is a primary example in Nabi et al. (2022). Specifically, Proposition 1 gives an identifying expression for ψ = E(Y(1)) in terms of observed-data functionals, Lemma 1 claims a von Mises expansion and influence function for θ = E(Y(1) | R1=0) in the general (non-binary) outcome case, and Corollary 2 gives a simplified influence function and a one-step estimator in the binary-outcome case with known odds ratio ρ. The paper argues that the proposed estimator is sqrt(n)-consistent and asymptotically efficient when ρ is known, and that plugging in a sqrt(n)-consistent estimator of ρ introduces only an extra asymptotically linear term.

Significance. If the efficiency results are correct, the paper provides the first practical semiparametric efficient estimator for a mean functional in the permutation missingness model, a non-standard MNAR model where estimation theory is otherwise underdeveloped. The m-SWIG discussion and the survey of transferable causal tools are useful conceptual contributions that will likely stimulate further work. The identification result in Proposition 1 is derived in detail and is a useful simplification of the full-law identification of Nabi et al. (2022). However, the central new technical contribution—the influence function and asymptotic-efficiency claim—is not fully established: Lemma 1's proof is omitted, and the extension to unknown ρ is only sketched. The efficiency claim therefore rests on unverified calculations rather than a complete theoretical argument. This is a support gap rather than a demonstrated error, and the paper is clearly written and well organized.

major comments (3)
  1. [Section 4, Lemma 1] Lemma 1 is the load-bearing step for the paper's efficiency claim, but its proof is omitted: the manuscript says 'We omit details, but the result follows from calculations similar to those discussed for example in Section 4 of Kennedy (2022).' The influence function φ(O;P) must satisfy the pathwise derivative identity, and the remainder Rθ(P;P̄) must be o_P(n^{-1/2}) under suitable conditions, but neither is verified. As written, a reader cannot check the correctness of the displayed φ. Please provide a full proof (or at least a detailed derivation) of the von Mises expansion, including the mean-zero property of φ and a bound on the remainder that would make the one-step estimator asymptotically linear.
  2. [Section 4, Corollary 2 and the one-step estimator] The influence function in Corollary 2 is derived under the assumption that ρ is known, but the proposed estimator then plugs in an estimated ρ̂, with the statement that 'the resulting estimator of θ will just have an extra asymptotically linear term.' No theorem is stated that gives conditions under which the one-step estimator with estimated ρ is asymptotically linear with a known influence function, nor is the extra term characterized. To support the efficiency claim, the authors should state a formal theorem that includes the regularity conditions (e.g., consistency and rate conditions on ξ̂ and ϖ̂, Donsker or empirical-process conditions, and the structure of the extra term from estimating ρ).
  3. [Appendix, proof of Corollary 2] The appendix derivation uses 'the influence function for ξ(x) when X is discrete.' The main results are presented for arbitrary X, including continuous covariates in the HIV example. For continuous X, the influence function for the conditional odds ξ(x) involves nonparametric estimation over a continuum, and the displayed remainder and one-step estimator require additional conditions (e.g., smoothness, rate conditions, or sample-splitting) that are not provided. This gap limits the generality of the efficiency claim to discrete X unless the authors supply the appropriate continuous-X theory.
minor comments (2)
  1. [Section 4, Proposition 1] In the displayed expression for ψ, the term θ is used before it is formally defined; please define θ = E(Y(1) | R1=0) immediately before the proposition to avoid confusion.
  2. [Section 4, Corollary 1] The interpretation of the posterior odds result is helpful, but the notation ζ(Y) is reused from Lemma 1 for a different quantity; consider a distinct symbol (e.g., ρ0(Y)) for the odds ratio to avoid ambiguity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the identification and influence-function claims are either inherited from external prior work or newly stated with explicit remainders, not derived from their own conclusions.

full rationale

I walked the derivation chain from the permutation-model assumptions through Proposition 1, Lemma 1, and Corollary 2. The identifying expression for ψ is inherited from the full-law identification of Nabi et al. (2022), not defined in terms of ψ, and Proposition 1 re-derives it by integration and Bayes' rule. The influence function in Lemma 1 is stated as a new mathematical claim with the remainder explicitly displayed; the proof is omitted and deferred to 'calculations similar to' Kennedy (2022), but this is a support gap, not a circular reduction, since the expansion is not true by construction and no fitted parameter is relabeled as a prediction. Corollary 2 likewise derives a specialized influence function from elementary differentiation and a stated discrete-X influence function for ξ(x). The paper's self-citations (Kennedy 2022 for method; Levis et al. 2024, 2025 for context) are not load-bearing uniqueness claims or definitions of the target. I therefore find no circular step; the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new free parameters and no new entities. It relies on the permutation model's independence assumptions, causal model semantics for SWIGs, consistency relations, and unstated regularity conditions. The absence of fitted parameters is a point in its favor; the unstated regularity conditions are a minor gap.

assumptions (4)
  • domain assumption Permutation missingness model: R1 is independent of Y(1) given X(1), and R2 is independent of (Y(1), X(1)) given Y and R1
    Stated at the start of Section 4; all identification and efficiency results depend on these independence restrictions inherited from Robins (1997) and Nabi et al. (2022).
  • domain assumption Nonparametric structural equation model or FFRCISTG independence semantics, so d-separation in m-DAGs and m-SWIGs implies counterfactual independence
    Invoked in Section 3 when deriving A(1) independent of (Y(a(1)), R(a(1))) given X from the SWIG in Figure 1b.
  • domain assumption Consistency and deterministic relations, e.g., A = R A(1) + (1 - R) '?'
    Used in Section 3 to connect observed, missingness, and counterfactual variables; the authors note such deterministic relationships may complicate d-separation in m-SWIGs.
  • domain assumption Regularity conditions for von Mises expansions and second-order remainders, including existence of densities and nuisance estimators converging at appropriate rates
    Implicit in Lemma 1 and Corollary 2; not stated explicitly in the paper, yet required for the influence function to represent a local asymptotic minimax lower bound.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Discussion of "Causal and counterfactual views of missing data models" by Razieh Nabi, Rohit Bhattacharya, Ilya Shpitser, & James M. Robins." pith.science (2026). https://pith.science/paper/S3ZSCZID

@misc{pith2026250613025,
  author       = {Pith},
  title        = {Pith review of: Discussion of "Causal and counterfactual views of missing data models" by Razieh Nabi, Rohit Bhattacharya, Ilya Shpitser, & James M. Robins},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/S3ZSCZID}},
  note         = {Machine review of arXiv:2506.13025}
}
read the original abstract

We congratulate Nabi et al. (2022) on their impressive and insightful paper, which illustrates the benefits of using causal/counterfactual perspectives and tools in missing data problems. This paper represents an important approach to missing-not-at-random (MNAR) problems, exploiting nonparametric independence restrictions for identification, as opposed to parametric/semiparametric models, or resorting to sensitivity analysis. Crucially, the authors represent these restrictions with missing data directed acyclic graphs (m-DAGs), which can be useful to determine identification in complex and interesting MNAR models. In this discussion we consider: (i) how/whether other tools from causal inference could be useful in missing data problems, (ii) problems that combine both missing data and causal inference together, and (iii) some work on estimation in one of the authors' example MNAR models.

Figures

Figures reproduced from arXiv: 2506.13025 by the authors.

Figure 1
Figure 1. m-DAG and m-SWIG for missing point exposure [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Directed acyclic graph for the permutation missingness model. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

28 extracted references · 22 canonical work pages

  1. [1]

    and Pearl, J

    Balke, A. and Pearl, J. (1994), Probabilistic evaluation of counterfactual queries, in Proceedings of the 12th Conference on Artificial Intelligence-Volume 1, Menlo Park, CA: MIT Press, pp. 230--237

  2. [2]

    --- (1997), Bounds on treatment effects from studies with imperfect compliance, Journal of the American Statistical Association, 92, 1171--1176

  3. [3]

    (2024), Semiparametric proximal causal inference, Journal of the American Statistical Association, 119, 1348--1359

    Cui, Y., Pu, H., Shi, X., Miao, W., and Tchetgen Tchetgen, E. (2024), Semiparametric proximal causal inference, Journal of the American Statistical Association, 119, 1348--1359

  4. [4]

    Frangakis, C. E. and Rubin, D. B. (2002), Principal stratification in causal inference, Biometrics, 58, 21--29

  5. [5]

    Heckman, J. J. (1979), Sample selection bias as a specification error, Econometrica: Journal of the econometric society, 153--161

  6. [6]

    Hern \'a n, M. A. and Robins, J. M. (2020), Causal inference: what if, Boca Raton: Chapman & Hill/CRC

  7. [7]

    Imbens, G. W. and Angrist, J. D. (1994), Identification and Estimation of Local Average Treatment Effects, Econometrica, 62, 467--475

  8. [8]

    Kennedy, E. H. (2020), Efficient nonparametric causal inference with missing exposure information, The International Journal of Biostatistics, 16

Show all 28 references
  1. [9]

    --- (2022), Semiparametric doubly robust targeted double machine learning: a review, arXiv preprint arXiv:2203.06469

  2. [10]

    W., Bonvini, M., Zeng, Z., Keele, L., and Kennedy, E

    Levis, A. W., Bonvini, M., Zeng, Z., Keele, L., and Kennedy, E. H. (2025), Covariate-assisted bounds on causal effects with instrumental variables, Journal of the Royal Statistical Society: Series B (Statistical Methodology), qkaf028

  3. [11]

    W., Kennedy, E

    Levis, A. W., Kennedy, E. H., and Keele, L. (2024), Nonparametric identification and efficient estimation of causal effects with instrumental variables, arXiv preprint arXiv:2402.09332

  4. [12]

    Li, W., Miao, W., and Tchetgen Tchetgen, E. (2023), Non-parametric inference about mean functionals of non-ignorable non-response data without identifying the joint distribution, Journal of the Royal Statistical Society Series B: Statistical Methodology, 85, 913--935

  5. [13]

    Miao, W., Geng, Z., and Tchetgen Tchetgen, E. J. (2018), Identifying causal effects with proxy variables of an unmeasured confounder, Biometrika, 105, 987--993

  6. [14]

    T., and Geng, Z

    Miao, W., Liu, L., Tchetgen, E. T., and Geng, Z. (2015), Identification, doubly robust estimation, and semiparametric efficiency theory of nonignorable missing data with a shadow variable, arXiv preprint arXiv:1509.02556

  7. [15]

    and Tchetgen Tchetgen, E

    Miao, W. and Tchetgen Tchetgen, E. J. (2016), On varieties of doubly robust estimators under missingness not at random with a shadow variable, Biometrika, 103, 475--482

  8. [16]

    (2022), Causal and counterfactual views of missing data models, arXiv preprint arXiv:2210.05558

    Nabi, R., Bhattacharya, R., Shpitser, I., and Robins, J. (2022), Causal and counterfactual views of missing data models, arXiv preprint arXiv:2210.05558

  9. [17]

    B., and Tchetgen Tchetgen, E

    Park, C., Richardson, D. B., and Tchetgen Tchetgen, E. J. (2024), Single proxy control, Biometrics, 80, ujae027

  10. [18]

    (2009), Causality, Cambridge U niversity P ress

    Pearl, J. (2009), Causality, Cambridge U niversity P ress

  11. [19]

    Richardson, T. S. and Robins, J. M. (2013), Single world intervention graphs (SWIGs): A unification of the counterfactual and graphical approaches to causality, Center for the Statistics and the Social Sciences, University of Washington Series. Working Paper, 128, 2013

  12. [20]

    Robins, J. M. (1997), Non-response models for the analysis of non-monotone non-ignorable missing data, Statistics in medicine, 16, 21--37

  13. [21]

    Robins, J. M. and Richardson, T. S. (2010), Alternative graphical causal models and the identification of direct effects, Tech. rep., Center for Statistics and the Social Sciences, University of Washington

  14. [22]

    S., and Robins, J

    Shpitser, I., Richardson, T. S., and Robins, J. M. (2022), Multivariate counterfactual systems and causal graphical models, in Probabilistic and causal inference: The works of Judea Pearl, New York, NY: Association for Computing Machinery, pp. 813--852

  15. [23]

    Sun, B., Liu, L., Miao, W., Wirth, K., Robins, J., and Tchetgen, E. J. T. (2018), Semiparametric estimation with data missing not at random using an instrumental variable, Statistica Sinica, 28, 1965

  16. [24]

    A., Hern \'a n, M

    Swanson, S. A., Hern \'a n, M. A., Miller, M., Robins, J. M., and Richardson, T. S. (2018), Partial identification of the average treatment effect using instrumental variables: review of methods for binary instruments, treatments, and outcomes, Journal of the American Statisti...

  17. [25]

    (2014), The control outcome calibration approach for causal inference with unobserved confounding, American journal of epidemiology, 179, 633--640

    Tchetgen Tchetgen, E. (2014), The control outcome calibration approach for causal inference with unobserved confounding, American journal of epidemiology, 179, 633--640

  18. [26]

    Tchetgen Tchetgen, E. J. and Wirth, K. E. (2017), A general instrumental variable framework for regression analysis with outcome missing not at random, Biometrics, 73, 1123--1131

  19. [27]

    and Tchetgen Tchetgen, E

    Wang, L. and Tchetgen Tchetgen, E. (2018), Bounded, efficient and multiply robust estimation of average treatment effects using instrumental variables, Journal of the Royal Statistical Society Series B: Statistical Methodology, 80, 531--550

  20. [28]

    (2012), Doubly robust estimators of causal exposure effects with missing data in the outcome, exposure or a confounder, Statistics in Medicine, 31, 4382--4400

    Williamson, E., Forbes, A., and Wolfe, R. (2012), Doubly robust estimators of causal exposure effects with missing data in the outcome, exposure or a confounder, Statistics in Medicine, 31, 4382--4400

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.