REVIEW 1 major objections 5 minor 1 cited by
Adding covariates to bounds: What is the question?
T0 review · 1 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read The sharpness of covariate-averaged causal bounds reduces to uniform sharpness of the conditional bounds, and in the all-binary IV graph where S is independent of Z and a confounded parent of X and Y, the averaged Balke-Pearl bounds…
desk verdict A useful formal distinction for covariate-conditional sharpness, undermined by an overgeneralized equivalence that fails for continuous covariates. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the definition of uniform sharpness: a pointwise valid set of conditional bounds $\{L_s:s\in S\}$ is uniformly sharp if for every observed joint distribution $P_O$ and every $\varepsilon>0$ there exists a single underlying distribution $P_V$ compatible with the model whose conditional estimands $\theta_s$ all lie within $\varepsilon$ of the corresponding $L_s$. The covariate-averaged bound $\bar L=\int L_s\,dP_S(s)$ is the mechanism that connects this global property to the marginal estimand $\theta=\int\theta_s\,dP_S(s)$. The proofs use the Balke-Pearl bounds, the sharp bounds on a binary causal risk difference under an instrumental variable, as the stratum-specific bounds in the IV setting, and a symbolic linear-programming computation of covariate-optimal bounds to compare against the averaged bounds.
What would settle it
Find a continuous covariate $S$ and any model in the Figure 5e graph for which the covariate-averaged Balke-Pearl bound equals the covariate-optimal bound for every observed distribution, yet no single compatible full distribution attains all conditional bounds within $\varepsilon$ for arbitrarily small $\varepsilon$; that would disprove Proposition 2 as printed. Concretely, simulate with $S$ continuous and check whether the only-if direction fails when stratum probabilities are zero or when the supremum over distributions of the average gap is zero but the infimum over distributions of the maximum stratum gap is positive.
Extended reading notes
Core claim
The central claim is Proposition 2: the covariate-averaged lower bound $\bar L(P_O)=\int L_s(P_{O|s})\,dP_S(s)$ is sharp for $\theta$ under $G$ if and only if the set of conditional bounds $\{L_s: s\in S\}$ is uniformly sharp for the set of conditional estimands $\{\theta_s: s\in S\}$. The paper also proves Proposition 4: in the all-binary IV model $G'$ of Figure 5e, where $S$ is independent of $Z$ and a confounded parent of $X$ and $Y$, the Balke-Pearl bounds conditioned on $S$ are uniformly sharp, so the covariate-averaged lower bound is valid and sharp for $\theta$. Together these claims identify the exact sense in which covariate averaging can be optimal: it inherits sharpness exactly when one distribution can realize all conditional bounds simultaneously, and in the binary IV setting this happens in the one graph where $S\perp\!\!\perp Z$ and $S$ confounds $X$ and $Y$.
Load-bearing premise
The equivalence between sharpness of the averaged bound and uniform sharpness of the conditional bounds requires that the covariate space is finite or atomic with positive probability in every stratum, a regularity condition the paper states only as a proof deferred to the supplement.
Editorial extensions
If this is right
- If Proposition 2 holds, checking whether an averaged bound is sharp reduces to checking uniform sharpness of the conditional bounds; pointwise sharpness alone is not enough.
- In the all-binary IV graph where $S$ is independent of $Z$ and is a confounded parent of $X$ and $Y$, covariate-averaged Balke-Pearl bounds coincide with the covariate-optimal bounds, so no information is lost by averaging.
- When $S$ is an additional instrument or an unconfounded mediator, the covariate-averaged bounds are not sharp, and in the mediator case the full covariate-optimal model can even point-identify the effect.
- Covariate-averaged bounds are never wider than covariate-marginal bounds when $S\perp\!\!\perp Z$, but can be wider than the marginal bounds when relevance of $Z$ is broken through $S$.
- Averaging can still narrow valid bounds without being sharp, so width improvement and sharpness are separate claims that should be reported distinctly.
Reading between the lines
- We infer that Proposition 2's only-if direction needs a regularity condition on $S$: the equivalence likely holds for finite or atomic covariate spaces, and for continuous $S$ a sequence of distributions could make the average gap vanish without a single distribution making all conditional gaps small.
- A testable extension would be to approximate a continuous covariate by finer stratifications in the Figure 5e model and check whether the averaged bounds converge to the covariate-optimal bounds only under an added uniformity condition on the strata.
- The paper's examples suggest a practical caution the authors leave implicit: covariates that are predictive of $X$ and $Y$ should not automatically be averaged over; the graph hosting $S$ decides whether averaging helps, and in some graphs it hurts.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper formalizes the distinction between pointwise and uniform sharpness for covariate-conditional bounds, proves general properties of covariate-averaged bounds, and applies these concepts to the binary instrumental-variable setting with covariates. The main positive result is Proposition 4, which states that in the DAG of Figure 5e—where S is independent of Z and is a confounded parent of X and Y—the covariate-conditional Balke-Pearl bounds are uniformly sharp and the covariate-averaged bound is sharp. The paper also gives two negative examples, in which covariate averaging is not sharp, and a simulation study over six DAGs comparing covariate-marginal, covariate-averaged, and covariate-optimal bounds.
Significance. If the results hold, the paper makes a useful conceptual contribution by carefully separating pointwise from uniform sharpness, a distinction that is often blurred in the partial-identification literature. The definitions are clear, the running examples are instructive, and the explicit algebraic comparison with causaloptim for Proposition 4 provides a concrete case where covariate averaging attains the optimal width. The reproducible simulation code is a further strength. However, the paper's most general theorem, Proposition 2, is false as stated for continuous covariate spaces, so the general framework needs a regularity condition or a restriction to finite/atomic S before the paper's broader claims can be accepted.
major comments (1)
- [Section 4, Proposition 2] A second, related point: because Proposition 2 is invoked in the narrative that connects covariate-averaged sharpness to uniform sharpness, the authors should state clearly in the main text whether the result is intended for all covariate spaces or only for finite S. The examples and Proposition 4 use binary S, so they are not affected by the counterexample, but the general formulation in Section 4 and the abstract's 'general conditions' claim must be revised.
minor comments (5)
- [Section 5.2] In the displayed definition of A|s, the third term is written as 'p10.0 + p01.0 − p00.1 − p01.1s'; the first three summands appear to be missing the s subscript and should presumably be p10.0s + p01.0s − p00.1s − p01.1s. As printed, the expression mixes conditional and unconditional probabilities and makes the subsequent algebra impossible to follow.
- [Section 4.1] The sentence '¯L(PO) = P (X = 0, Y= 1) − P (X = 1, Y= 0)' has a sign error: the preceding line defines L(PO) = −P(X=0,Y=1) − P(X=1,Y=0), and the later covariate-optimal expression confirms that the averaged bound should be the negative sum, not the difference.
- [Section 7] In the concluding paragraph, 'unformly sharp' should be 'uniformly sharp'.
- [Section 6] The simulation section does not report Monte Carlo standard errors for the proportions in Table 1; with 10^5 replicates the error is likely small, but a brief statement would be helpful.
- [Section 5.3.2] The phrase 'This example needs less explanation' is informal for a journal article and could be replaced by a more neutral transition.
Circularity Check
No significant circularity: central results are proven from model assumptions; self-citations to causaloptim/Sachs et al. are independent computational/theorem support.
full rationale
The paper's derivation chain is self-contained. Proposition 2 is a general equivalence whose proof is deferred to the supplement; whether or not it requires a finite/atomic S regularity condition, it is a mathematical claim, not a definitional reduction. Proposition 4 is established by showing, algebraically in the supplement, that the covariate-averaged Balke-Pearl bound equals the covariate-optimal bound Lco computed by causaloptim. The use of causaloptim and Sachs et al. (2023) to certify sharpness is a self-citation, but it is not circular: Sachs et al. is an externally published, code-reproduced theorem with stated assumptions that do not include the covariate-averaging result, and the equality Lbar = Lco is checked algebraically rather than assumed. The examples in Section 5.3 use the same software to exhibit non-sharpness, which is again a computation, not a fitted prediction. No free parameter is fitted and later called a prediction; no estimand is defined in terms of the bound it is supposed to constrain; no uniqueness theorem from the authors' prior work is invoked to forbid alternatives. The only notable caveat is the possible failure of Proposition 2's only-if direction for continuous S without extra compactness or positivity conditions, but that is a correctness or regularity concern, not circularity.
Assumptions & free parameters
assumptions (6)
- domain assumption Consistency and no interference for counterfactuals (Section 2)
- domain assumption NPSEM/DAG encoding of the causal model for each setting (Figures 3-5)
- domain assumption IV assumptions 1-2: Z ⊥⊥ {Y(z),X(z)}|S and Y(x,z)=Y(x), plus positivity (Section 5.1)
- standard math Sharpness of causaloptim bounds proven in Sachs et al. 2023
- domain assumption S ⊥⊥ Z and binary S in Proposition 4 (Figure 5e)
- ad hoc to paper Unstated regularity: S finite/atomic with positive stratum probabilities for the only-if direction of Proposition 2
Cite this review
Pith. "Pith review of Adding covariates to bounds: What is the question?." pith.science (2026). https://pith.science/paper/26G4OVWZ
@misc{pith2026250203156,
author = {Pith},
title = {Pith review of: Adding covariates to bounds: What is the question?},
year = {2026},
howpublished = {\url{https://pith.science/paper/26G4OVWZ}},
note = {Machine review of arXiv:2502.03156}
}
read the original abstract
Symbolic nonparametric bounds for partial identification of causal effects now have a long history in the causal literature. Sharp bounds, bounds that use all available information to make the range of values as narrow as possible, are often the goal. For this reason, many publications have focused on deriving sharp bounds, but the concept of sharp bounds is nuanced and can be misleading. In settings with ancillary covariates, the situation becomes more complex. We provide clear definitions for pointwise and uniform sharpness of covariate-conditional bounds, that we then use to prove some general and some specific to the IV setting results about the relationship between these two concepts. As we demonstrate, general conditions are much more difficult to determine and thus, we urge authors to be clear when including ancillary covariates in bounds via conditioning about the setting of interest and the assumptions made.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
Partial identification via conditional linear programs: estimation and policy learning
Two debiased estimators, one based on linear programming solutions and one on entropic smoothing, provide asymptotic confidence intervals for covariate-dependent partial identification bounds and support policy learning.
Reference graph
Works this paper leans on
-
[1]
Bounds on treatment effects from studies with imperfect compliance
Alexander Balke and Judea Pearl. Bounds on treatment effects from studies with imperfect compliance. Journal of the American Statistical Association, 92 0 (439): 0 1171--1176, 1997. doi:10.1080/01621459.1997.10474074. URL https://doi.org/10.1080/01621459.1997.10474074
-
[2]
Nonparametric bounds on causal effects from partial compliance data
Alexander Balke and Judea Pearl. Nonparametric bounds on causal effects from partial compliance data. 2011. URL https://api.semanticscholar.org/CorpusID:142574577
work page 2011
-
[3]
Non-parametric bounds on treatment effects with non-compliance by covariate adjustment
Zhihong Cai, Manabu Kuroki, and Tosiya Sato. Non-parametric bounds on treatment effects with non-compliance by covariate adjustment. Statistics in Medicine, 26 0 (16): 0 3188--3204, 2007. doi:https://doi.org/10.1002/sim.2766. URL https://onlinelibrary.wiley.com/doi/abs/10.1002/sim.2766
-
[4]
Vanessa Didelez, Sha Meng, and Nuala A. Sheehan. Assumptions of IV Methods for Observational Epidemiology . Statistical Science, 25 0 (1): 0 22 -- 40, 2010. doi:10.1214/09-STS316. URL https://doi.org/10.1214/09-STS316
-
[5]
Causal bounds for outcome-dependent sampling in observational studies
Erin E Gabriel, Michael C Sachs, and Arvid Sj \"o lander. Causal bounds for outcome-dependent sampling in observational studies. Journal of the American Statistical Association, 117 0 (538): 0 939--950, 2022
work page 2022
-
[6]
Gustav Jonzon, Michael C. Sachs, and Erin E. Gabriel. Accessible computation of tight symbolic bounds on causal effects using an intuitive graphical interface. R JOURNAL, 15 0 (4): 0 53--68, DEC 2023. ISSN 2073-4859
work page 2023
-
[7]
Covariate-assisted bounds on causal effects with instrumental variables
Alexander W Levis, Matteo Bonvini, Zhenghao Zeng, Luke Keele, and Edward H Kennedy. Covariate-assisted bounds on causal effects with instrumental variables. arXiv preprint arXiv:2301.12106, 2023
arXiv 2023
-
[8]
Nonparametric bounds on treatment effects
Charles F Manski. Nonparametric bounds on treatment effects. The American Economic Review, 80 0 (2): 0 319--323, 1990
work page 1990
Show all 13 references
-
[9]
Causality
Judea Pearl. Causality. Cambridge University Press, New York, 2000
2000
-
[10]
Ramsahai
Roland R. Ramsahai. Causal bounds and instruments. In Proceedings of the Twenty-Third Conference on Uncertainty in Artificial Intelligence, UAI'07, page 310–317, Arlington, Virginia, USA, 2007. AUAI Press. ISBN 0974903930
2007
-
[11]
J.M. Robins. The analysis of randomized and non-randomized AIDS treatment trials using a new approach to causal inference in longitudinal studies. In L. Sechrest, H. Freeman, and A. Mulley, editors, Health service research methodology: a focus on AIDS, pages 113--159. US Publi...
1989
-
[12]
Sachs, Gustav Jonzon, Arvid Sjolander, and Erin E
Michael C. Sachs, Gustav Jonzon, Arvid Sjolander, and Erin E. Gabriel. A general method for deriving tight symbolic bounds on causal effects. JOURNAL OF COMPUTATIONAL AND GRAPHICAL STATISTICS, 32 0 (2): 0 567--576, APR 3 2023. ISSN 1061-8600. doi:10.1080/10618600.2022.2071905
2023
-
[13]
Swanson, Miguel A
Sonja A. Swanson, Miguel A. Hernán, Matthew Miller, James M. Robins, and Thomas S. Richardson. Partial identification of the average treatment effect using instrumental variables: Review of methods for binary instruments, treatments, and outcomes. Journal of the American Stati...
2018
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.