REVIEW 3 major objections 5 minor 40 references
Fortified Proximal Causal Inference with Many Invalid Proxies
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper proves that the ATE is nonparametrically identified when at least $\gamma$ of $K$ candidate treatment confounding proxies are valid, without knowing which ones, via a fortified confounding bridge function and a strengthened…
desk verdict A genuine extension of PCI to invalid treatment proxies, with honest assumptions and a real but nonfatal soft spot in the fortified completeness condition. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the fortified subspace $H_\gamma$ of functions of $(Z,A,X)$ whose conditional expectation given any subvector $Z_{-\nu}$ of the candidate proxies with $|\nu| \ge \gamma$ is zero; it encodes the union of all possible validity sets without naming one. Solving the moment equation over $H_\gamma$ produces the fortified outcome confounding bridge $h^*$, and the fortified completeness condition (Assumption 5)—completeness of the smaller space $H_\gamma + L^2(Z_{-\nu^*},X)$—forces $h^*$ to satisfy the conditional-independence property that yields the ATE formula. For the weighting route, the fortified treatment confounding bridge $q^*$ with $(-1)^{1-A}q^* \in H_\gamma$ plays the role of an inverse propensity score and yields a second identification via Theorem 2.2. Both constructions reduce to the conventional proximal causal inference bridge functions when $\gamma = K$.
What would settle it
Construct a data-generating process satisfying Assumption 3 with $K=2$ and $\gamma=1$, for example discrete $U$ with $Z_1$ valid and $Z_2$ invalid, chosen so that the matrix in Remark 2.3 is rank-deficient, making the fortified completeness assumption fail; then solve (2.4) numerically in large samples and compare $E[h^*(W,1,X) - h^*(W,0,X)]$ to the true ATE. If the two differ materially, the identification theorem cannot hold without the fortified completeness condition.
Extended reading notes
Core claim
The paper's central claim is that knowing how many treatment confounding proxies are valid—not which ones—suffices for proximal causal inference. Under the assumption that at least $\gamma$ of $K$ candidates satisfy the proxy conditions, it defines the fortified subspace $H_\gamma$, the set of square-integrable functions of $(Z,A,X)$ whose conditional expectation given any subvector $Z_{-\nu}$ with $|\nu| \ge \gamma$ is zero, and shows that any solution $h^*$ of the moment condition $E[d(Z,A,X)\{Y - h^*(W,A,X)\}] = 0$ for all $d \in H_\gamma$ identifies the ATE as $\tau^* = E[h^*(W,1,X) - h^*(W,0,X)]$, provided a fortified completeness condition holds. A parallel construction using a fortified treatment confounding bridge function $q^*$ gives the weighting identification $\tau^* = E[q^*(Z,1,X)AY - q^*(Z,0,X)(1-A)Y]$. These results hold for any $\gamma$, including values below a majority of the candidates, and they avoid any model-selection step that would need to identify the valid proxies. The companion estimation theory proves consistency and asymptotic normality of the fortified proximal multiply robust estimator under the union of three working models and local semiparametric efficiency at their intersection.
Load-bearing premise
The load-bearing premise is the fortified completeness assumption: the unobserved confounder must be nonparametrically recoverable from the proxies even after the analysis deliberately works with the smaller fortified subspace $H_\gamma$ plus functions of the possibly invalid proxies, a condition strictly stronger than the completeness needed when the valid proxies are known and untestable because it involves the unobserved $U$.
Editorial extensions
If this is right
- At $\gamma = K$ the fortified bridges coincide with the conventional proximal bridge functions, so standard proximal causal inference is a special case of the new framework.
- No proxy-validity model selection is ever needed: the analyst only states a lower bound $\gamma$ and the estimating procedure works with all $K$ candidates jointly.
- Because $\gamma$ can be smaller than $K/2$, the approach tolerates a minority of valid proxies, unlike majority-rule or plurality-rule methods for invalid instruments.
- Sensitivity analysis is built in: the identified parameter must be constant across all $\gamma$ that are valid lower bounds, so changing $\gamma$ and inspecting stability offers a diagnostic for the assumption that at least $\gamma$ proxies are valid.
- The fortified proximal multiply robust (fPMR) estimator remains consistent if any one of its three working models is correctly specified, and it attains the local semiparametric efficiency bound at the intersection submodel.
Reading between the lines
- A natural extension the paper leaves implicit: the same fortified construction applies symmetrically to invalid outcome confounding proxies by swapping the roles of $W$ and $Z$, which would make the robustness two-sided rather than one-sided.
- The identification result suggests a concrete falsification check beyond the paper's simulations: compute the fortified estimator at two different values of $\gamma$; if the true valid count is at least the larger value, both estimates must agree within sampling variability, so systematic divergence directly indicates that the assumed lower bound is too high.
- Since $H_\gamma$ is defined by conditional expectations over all $\gamma$-subsets, the construction grows combinatorially in $K$; for large proxy sets the alternating conditional expectations (ACE) implementation will likely need dimension reduction or screening, a regime the paper does not analyze.
- Setting $\gamma=1$ makes the method depend on a much stronger completeness condition; reporting estimates across the whole range of $\gamma$ effectively converts an untestable completeness assumption into a sensitivity curve over the analyst's confidence in the proxies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a proximal causal inference framework for settings where, among K candidate treatment confounding proxies, at least gamma are valid (Assumption 3) but the identity of valid proxies is unknown. It introduces fortified outcome and treatment confounding bridge functions, defined through moment restrictions over the subspace H_gamma in (2.3), and proves nonparametric identification of the ATE by Theorems 2.1 and 2.2. A locally semiparametric efficient influence function is derived in Theorem 3.1, and a multiply robust estimator fPMR is proposed in Theorem 3.2. The methods are evaluated by simulations and applied to the SUPPORT right heart catheterization data. The identification proofs in the appendix are internally consistent, and the simulation results match the theory under correct specification, but the central identification claim rests on strengthened completeness conditions whose practical scope is not fully resolved.
Significance. If the strengthened completeness conditions hold, the paper offers a meaningful advance over standard PCI: it removes the need to know which treatment confounding proxies are valid, avoids a model-selection step, and provides a multiply robust estimator with local efficiency. The paper is also commendably explicit that Assumptions 5 and 7 are strictly stronger than the standard completeness Assumption 2, and it explains the trade-off between the analyst's lower bound gamma and the strength of the required completeness. The proofs are detailed and the simulations, including misspecification scenarios, support the multiply robustness claim in the correctly specified cases. However, the practical estimator used in the simulations and data analysis (the ACE implementation with bootstrap inference) is not covered by the asymptotic theorems as written, and the completeness assumptions are not verified or given nontrivial sufficient conditions.
major comments (3)
- [Section 2.2.1, Assumption 5 and Theorem 2.1] Assumption 5 is the load-bearing premise of Theorem 2.1 and is strictly stronger than the standard completeness condition Assumption 2(i), as the paper acknowledges. The strength depends on the analyst's chosen gamma: for 1 ≤ gamma < |nu*|, H_gamma is a proper subspace of H_{|nu*|}, so Assumption 5 requires completeness of a smaller space than would be needed if the true valid set were known. Since nu* and U are unobserved, Assumption 5 cannot be checked from data. The proof of Theorem 2.1 uses Assumption 5 precisely to conclude that the residual E{Y - h* - l* | U, A, Z_{-nu*}, X} equals zero; if Assumption 5 fails, different solutions to (2.4) can yield different values of E[h*(W,1,X) - h*(W,0,X)], so tau* is not identified. The paper should provide either a finite-dimensional example satisfying Assumptions 3 and 4 and the standard completeness Assumption 2(i) but violating Assumption 5, or a nontrivial sufficient condition for Assumption 5 in terms of the observed data distribution. Without one of these, the practical scope of the central identification claim remains unclear.
- [Section 4 and Remark 3.1] The estimator actually implemented in the simulations and data analysis is not the estimator covered by Theorem 3.2. The theorem is stated for an abstract mapping d((-)1^{1-A}q(^t); â) under regularity conditions relegated to the supplementary material, whereas the implementation uses the ACE algorithm d-dagger fitted by iterative linear regression models, and Remark 3.1 explicitly states that this implementation lacks closed-form expressions and therefore uses the nonparametric bootstrap for inference. No theorem establishes consistency or asymptotic normality of this practical ACE-based estimator, nor does the paper account for the first-stage estimation of â and the working models in equation (3.6). Consequently, the coverage probabilities reported in Tables 1 and 2 are not consequences of Theorem 3.2. The authors should either extend the asymptotic theory to the practical ACE implementation or explicitly present that implementation as heuristic and restrict the formal inferential claims to estimators based on the closed-form mapping d-dagger.
- [Section 3.1, Theorem 3.1] The semiparametric efficiency result is stated for the submodel M_eff,gamma, where Assumption 8 holds and h* + l* and q* are uniquely defined. Assumption 8 is itself a strong surjectivity condition, so the claimed local efficiency bound does not apply to the whole model M_gamma. The paper's use of the qualifier 'local' is accurate, but the text in Section 3.1 repeatedly refers to 'the semiparametric local efficiency bound of M_gamma' without consistently reminding the reader that the bound is obtained only at a submodel. This should be clarified, and the role of Assumption 8 in the efficiency claim should be stated explicitly in the main text rather than only in the proof.
minor comments (5)
- [Section 2.2, equation (2.3)] The equivalence between the two displayed characterizations of H_gamma is stated to follow by the law of iterative expectations, but the argument is not shown; a one-sentence proof or pointer to the supplementary material would help readers see why conditioning on Z_{-nu} for all subsets nu of size gamma suffices for all subsets of size at least gamma.
- [Theorem 3.1, equation (3.1)] The influence function uses l*(Z,X), whereas in Theorem 2.1 and Proposition 2.2 the function l* is introduced as l*(Z_{-nu*},X). The argument of l* should be kept consistent to avoid confusion about the domain of this nuisance function.
- [Section 4, Tables 1 and 2] The text states that 500 bootstrap samples were used for standard errors, but it does not explicitly say whether the same number was used in the misspecification scenarios of Table 2. Please state the bootstrap size for each table.
- [Section 5, Table 3] The sentence 'our estimates ... are more conservative in magnitude' is ambiguous; the authors should specify that the fPIPW and fPMR estimates are closer to zero than the conventional proximal estimates, rather than merely 'more conservative.'
- [Introduction and Section 2.1] The term 'disconnected proxies' is used when citing Kummerfeld et al. (2024) but is not defined; a brief definition or reference to the definition would improve readability.
Circularity Check
No significant circularity: the fortified-bridge identification theorems are proved from the stated assumptions, and self-citations are background, not load-bearing.
full rationale
The central identification claims are derived, not assumed. Theorem 2.1's proof (Appendix B.1) starts from the fortified moment restriction (2.4) and, using the proxy conditional independences in Definition 2.1 plus the fortified completeness Assumption 5, establishes the residual property (2.5); the ATE formula then follows by standard potential-outcome algebra with the ℓ* term canceling. No step defines h* or Hγ in terms of τ*, and no fitted parameter is relabeled as a prediction. Theorem 2.2 and the multiply robust estimator Theorem 3.2 are likewise proved from the stated models. The citations to Miao et al. (2018) and Cui et al. (2024) are used only to state the conventional ('oracle') baseline Proposition 2.1 and to position the framework; the fortified results have self-contained proofs in the appendix, so these citations are not load-bearing. The assumptions most likely to fail, such as fortified completeness (Assumption 5), are stated explicitly as identification conditions, not as consequences of the target result. Simulations and the RHC application are evaluations, not predictions derived from the estimator's own construction.
Assumptions & free parameters
free parameters (2)
- gamma (lower bound on number of valid treatment proxies) =
user-specified (simulation: gamma=1; data analysis: gamma in {2,4,6,8})
- Nuisance parameters (b, r, t) of working models =
estimated from data (simulation values in Section 4 and Appendix A.4)
assumptions (8)
- domain assumption Assumption 3: |nu*| >= gamma, at least gamma of K candidate treatment confounding proxies are valid.
- domain assumption Assumption 4: existence of a fortified outcome bridge function h* satisfying E[d(Z,A,X){Y-h*(W,A,X)}]=0 for all d in H_gamma.
- domain assumption Assumption 5: fortified completeness for the outcome bridge.
- domain assumption Assumptions 6 and 7: existence of a fortified treatment bridge q* and corresponding completeness condition.
- domain assumption Definition 2.1(i)-(iii): latent ignorability, positivity, and proxy exclusion restrictions for the true valid set nu*, including validity of W as an outcome proxy.
- standard math Assumption 1: consistency Y = Y(a) when A=a.
- domain assumption Assumption 8: surjectivity of the projection operator Pi_gamma onto H_gamma.
- standard math Regularity conditions from Robins et al. (1994) for asymptotic normality of the estimating equations.
Cite this review
Pith. "Pith review of Fortified Proximal Causal Inference with Many Invalid Proxies." pith.science (2026). https://pith.science/paper/DQAVPYVC
@misc{pith2026250613152,
author = {Pith},
title = {Pith review of: Fortified Proximal Causal Inference with Many Invalid Proxies},
year = {2026},
howpublished = {\url{https://pith.science/paper/DQAVPYVC}},
note = {Machine review of arXiv:2506.13152}
}
abstract
Causal inference from observational data often relies on the assumption of no unmeasured confounding, an assumption frequently violated in practice due to unobserved or poorly measured covariates. Proximal causal inference (PCI) offers a promising framework for addressing unmeasured confounding using a pair of outcome and treatment confounding proxies. However, existing PCI methods typically assume all specified proxies are valid, which may be unrealistic and is untestable without extra assumptions. In this paper, we develop a semiparametric approach for a many-proxy PCI setting that accommodates potentially invalid treatment confounding proxies. We introduce a new class of fortified confounding bridge functions and establish nonparametric identification of the population average treatment effect (ATE) under the assumption that at least $\gamma$ out of $K$ candidate treatment confounding proxies are valid, for any $\gamma \leq K$ set by the analyst without requiring knowledge of which proxies are valid. We establish a local semiparametric efficiency bound and develop a class of multiply robust, locally efficient estimators for the ATE. These estimators are thus simultaneously robust to invalid treatment confounding proxies and model misspecification of nuisance parameters. The proposed methods are evaluated through simulation and applied to assess the effect of right heart catheterization in critically ill patients.
Figures
Reference graph
Works this paper leans on
-
[1]
P. J. Bickel and K. A. Doksum. Mathematical Statistics: Basic Ideas and Selected Topics. Vol. 1. Chapman and Hall/CRC, 2015
work page 2015
-
[2]
P. J. Bickel, C. A. J. Klaassen, Y. Ritov, and J. A. Wellner. Efficient and Adaptive Estimation for Semiparametric Models. New York: Springer-Verlag, 1993
work page 1993
-
[3]
L. Breiman and J. H. Friedman. Estimating optimal transformations for multiple regression and correlation. Journal of the American Statistical Association, 80 0 (391): 0 580--598, 1985
work page 1985
-
[4]
M. Carrasco, J.-P. Florens, and E. Renault. Linear inverse problems in structural econometrics estimation based on spectral decomposition and regularization. Handbook of Econometrics, 6: 0 5633--5751, 2007
work page 2007
-
[5]
A. F. Connors, T. Speroff, N. V. Dawson, C. Thomas, F. E. Harrell, D. Wagner, N. Desbiens, L. Goldman, A. W. Wu, R. M. Califf, W. J. Fulkerson, H. Vidaillet, S. Broste, P. Bellamy, J. Lynn, and W. A. Knaus. The effectiveness of right heart catheterization in the initial care of critically iii patients. Journal of the American Medical Association, 276 0 (1...
work page 1996
-
[6]
J. B. Conway. A Course in Functional Analysis. Springer, 1990
work page 1990
-
[7]
Y. Cui, H. Pu, X. Shi, W. Miao, and E. J. Tchetgen Tchetgen. Semiparametric proximal causal inference. Journal of the American Statistical Association, 119 0 (546): 0 1348--1359, 2024
work page 2024
-
[8]
A. E. Ghassami, A. Ying, I. Shpitser, and E. J. Tchetgen Tchetgen. Minimax kernel machine learning for a class of doubly robust functionals with application to proximal causal inference. In International Conference on Artificial Intelligence and Statistics, pages 7210--7239. PMLR, 2022
work page 2022
Show all 40 references
-
[9]
Z. Guo, H. Kang, T. T. Cai, and D. S. Small. Confidence intervals for causal effects with invalid instruments by using two-stage hard thresholding with voting. Journal of the Royal Statistical Society Series B: Statistical Methodology, 80 0 (4): 0 793--815, 2018
2018
-
[10]
C. Han. Detecting invalid instruments using l_1 -gmm. Economics Letters, 101 0 (3): 0 285--287, 2008
2008
-
[11]
Hirano and G
K. Hirano and G. W. Imbens. Estimation of causal effects using propensity score weighting: An application to data on right heart catheterization. Health Services and Outcomes Research Methodology, 2: 0 259--278, 2001
2001
-
[12]
P. W. Holland. Causal inference, path analysis and recursive structural equations models. Sociological Methodology, 18: 0 449--484, 1988
1988
-
[13]
Hu and S
Y. Hu and S. M. Schennach. Instrumental variable treatment of nonclassical measurement error models. Econometrica, 76 0 (1): 0 195--216, 2008
2008
-
[14]
Imbens, N
G. Imbens, N. Kallus, X. Mao, and Y. Wang. Long-term causal inference under persistent confounding via data combination. Journal of the Royal Statistical Society Series B: Statistical Methodology, 87 0 (2): 0 362--388, 2025
2025
-
[15]
L. A. Jackson, M. L. Jackson, J. C. Nelson, K. M. Neuzil, and N. S. Weiss. Evidence of bias in estimates of influenza vaccine effectiveness in seniors. International Journal of Epidemiology, 35 0 (2): 0 337--344, 2006
2006
-
[16]
Kallus, X
N. Kallus, X. Mao, and M. Uehara. Causal inference under unmeasured confounding with negative controls: A minimax learning approach. arXiv preprint arXiv:2103.14029, 2021
2021 arXiv
-
[17]
H. Kang, A. Zhang, T. T. Cai, and D. S. Small. Instrumental variables estimation with some invalid instruments and its application to mendelian randomization. Journal of the American Statistical Association, 111 0 (513): 0 132--144, 2016
2016
-
[18]
J. D. Y. Kang and J. L. Schafer. Demystifying double robustness: A comparison of alternative strategies for estimating a population mean from incomplete data. Statistical Science, 22 0 (4): 0 523--539, 2007
2007
-
[19]
R. Kress. Linear Integral Equations. Springer, 1989
1989
-
[20]
Kummerfeld, J
E. Kummerfeld, J. Lim, and X. Shi. Data-driven automated negative control estimation (dance): Search for, validation of, and causal inference with negative controls. Journal of Machine Learning Research, 25 0 (229): 0 1--35, 2024
2024
-
[21]
Lipsitch, E
M. Lipsitch, E. J. Tchetgen Tchetgen, and T. Cohen. Negative controls: a tool for detecting confounding and bias in observational studies. Epidemiology, 21 0 (3): 0 383--388, 2010
2010
-
[22]
Mastouri, Y
A. Mastouri, Y. Zhu, L. Gultchin, A. Korba, R. Silva, M. Kusner, A. Gretton, and K. Muandet. Proximal causal learning with kernels: Two-stage estimation and moment restriction. In International Conference on Machine Learning, pages 7512--7523. PMLR, 2021
2021
-
[23]
W. Miao, Z. Geng, and E. J. Tchetgen Tchetgen. Identifying causal effects with proxy variables of an unmeasured confounder. Biometrika, 105 0 (4): 0 987--993, 2018
2018
-
[24]
W. Miao, X. Shi, Y. Li, and E. J. Tchetgen Tchetgen. A confounding bridge approach for double negative control inference on causal effects. Statistical Theory and Related Fields, pages 1--12, 2024
2024
-
[25]
W. K. Newey and J. L. Powell. Instrumental variable estimation of nonparametric models. Econometrica, 71 0 (5): 0 1565--1578, 2003
2003
-
[26]
J. Neyman. On the application of probability theory to agricultural experiments. essay on principles. section 9. Statistical Science, 5 0 (4): 0 465--472, 1923. Trans. Dorota M. Dabrowska and Terence P. Speed (1990)
1990
-
[27]
O'Sullivan
F. O'Sullivan. A statistical perspective on ill-posed inverse problems. Statistical Science, pages 502--518, 1986
1986
-
[28]
J. M. Robins. A new approach to causal inference in mortality studies with a sustained exposure period—application to control of the healthy worker survivor effect. Mathematical Modelling, 7 0 (9-12): 0 1393--1512, 1986
1986
-
[29]
J. M. Robins, A. Rotnitzky, and L. P. Zhao. Estimation of regression coefficients when some regressors are not always observed. Journal of the American Statistical Association, 89 0 (427): 0 846--866, 1994
1994
-
[30]
D. B. Rubin. Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66 0 (5): 0 688, 1974
1974
-
[31]
X. Shi, W. Miao, J. C. Nelson, and E. J. Tchetgen Tchetgen. Multiply robust causal inference with double-negative control adjustment for categorical unmeasured confounding. Journal of the Royal Statistical Society Series B: Statistical Methodology, 82 0 (2): 0 521--540, 2020 a
2020
-
[32]
X. Shi, W. Miao, and E. J. Tchetgen Tchetgen. A selective review of negative control methods in epidemiology. Current Epidemiology Reports, 7: 0 190--202, 2020 b
2020
-
[33]
R. Singh. Kernel methods for unobserved confounding: Negative controls, proxies, and instruments. arXiv preprint arXiv:2012.10315, 2020
2012 arXiv
-
[34]
D. S. Small. Sensitivity analysis for instrumental variables regression with overidentifying restrictions. Journal of the American Statistical Association, 102 0 (479): 0 1049--1058, 2007
2007
-
[35]
Sofer, D
T. Sofer, D. B. Richardson, E. Colicino, J. Schwartz, and E. J. Tchetgen Tchetgen Tchetgen. On negative outcome control of unobserved confounding as a generalization of difference-in-differences. Statistical Science, 31 0 (3): 0 348, 2016
2016
-
[36]
B. Sun, Z. Liu, and E. J. Tchetgen Tchetgen. Semiparametric efficient g-estimation with invalid instrumental variables. Biometrika, 110 0 (4): 0 953--971, 2023
2023
-
[37]
E. J. Tchetgen Tchetgen, J. M. Robins, and A. Rotnitzky. On doubly robust estimation in a semiparametric odds ratio model. Biometrika, 97 0 (1): 0 171--180, 2010
2010
-
[38]
E. J. Tchetgen Tchetgen, A. Ying, Y. Cui, X. Shi, and W. Miao. An Introduction to Proximal Causal Inference . Statistical Science, 39 0 (3): 0 375--390, 2024
2024
-
[39]
Vansteelandt, T
S. Vansteelandt, T. J. VanderWeele, E. J. Tchetgen Tchetgen, and J. M. Robins. Multiply robust inference for statistical interactions. Journal of the American Statistical Association, 103 0 (484): 0 1693--1704, 2008
2008
-
[40]
A. Ying, W. Miao, X. Shi, and E. J. Tchetgen Tchetgen. Proximal causal inference for complex longitudinal studies. Journal of the Royal Statistical Society Series B: Statistical Methodology, 85 0 (3): 0 684--704, 2023
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.