Pith. sign in

REVIEW 3 major objections 4 minor 25 references

Using causal diagrams to assess parallel trends in difference-in-differences studies

T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A causal diagram can tell you when parallel trends is untenable.

desk verdict The paper's main graphical rejection criteria rest on a false lemma; the supplied counterexample is valid, so Condition 1 is invalid and the central contribution collapses. read the letter →

arxiv 2505.03526 v1 pith:MHBPS4BK submitted 2025-05-06 stat.ME

classification stat.ME MSC 62D20
keywords difference-in-differencesparalleltrendscausaldiagramsDAGunmeasuredconfoundinglinearfaithfulnessminimallysufficientsetsadditivehomogeneous
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Difference-in-differences (DID) delivers causal estimates only when parallel trends holds, and researchers have had little guidance for judging that assumption ahead of time. This paper shows that a causal diagram can provide such guidance once one adds linear faithfulness: every pair of variables the graph connects has nonzero covariance. Under that assumption, parallel trends is incompatible with pre-treatment outcomes affecting treatment (together with unmeasured confounding of treatment and the post-treatment outcome), and with pre- and post-treatment outcomes having distinct minimally sufficient adjustment sets. The paper also argues that pre-treatment outcomes affecting post-treatment outcomes is a serious warning sign, though this is proven only in restricted semiparametric models. If all three warning signs are absent, the remaining content of parallel trends is exactly additive homogeneous confounding: a common confounder set must have constant additive association with the untreated potential outcome across the two periods.

What carries the argument

The machine that does the work is linear faithfulness (Assumption 1): whenever the graph does not d-separate (block) two variables, their conditional covariance is nonzero. This lets the authors convert 'the graph says this association should be present' into 'parallel trends would force this association to be absent.' The supporting machinery consists of directed single-world intervention graphs (SWIGs) for reading counterfactual independencies, minimally sufficient adjustment sets for the effect of treatment on each outcome, and Lemma 2, which states that parallel trends plus a common sufficient set M implies additive homogeneous confounding, E($Y1^{0}$-$Y0^{0}$|M)=E($Y1^{0}$-$Y0^{0}$).

What would settle it

Simulate a two-period linear structural equation model that exactly follows the graph with disjoint minimally sufficient sets (a U3 affecting A and Y0, a U4 affecting A and Y1, no Y0-to-A or Y0-to-Y1 edge), and choose coefficients so that E($Y1^{0}$-$Y0^{0}$|A)=E($Y1^{0}$-$Y0^{0}$) exactly. Then check every d-connected pair in the graph for nonzero covariance; a coefficient vector that satisfies both parallel trends and nonzero covariances would refute the claim that Condition 2 plus linear faithfulness rejects parallel trends.

Watch

Extended reading notes

Core claim

The central claim is that parallel trends, although scale-dependent, can be assessed with a scale-independent graph if linear faithfulness holds. In that setting, adopting parallel trends forces conditional mean equalities that the graph contradicts whenever (i) the pre-treatment outcome Y0 directly affects treatment A while unmeasured confounding connects A to the post-treatment outcome $Y1^{0}$, or (ii) the minimally sufficient adjustment sets for Y0 and $Y1^{0}$ differ, as when separate unmeasured confounders affect treatment with only one outcome. The paper further argues, without a full proof in the general nonparametric model, that (iii) an arrow from Y0 to $Y1^{0}$ should be regarded as suspect, because in partially linear and additively separable models parallel trends can hold with and without that arrow only through exact cancellation. When none of these features appears, the maximal graph compatible with parallel trends consists of a common confounder for both outcomes and a separate source of correlation between Y0 and Y1, and the assumption reduces to additive homogeneous confounding.

Load-bearing premise

The load-bearing premise is linear faithfulness: any two variables connected by the graph must have nonzero covariance, so exact cancellation of associations is ruled out; if such cancellations occur naturally, a graph can contain all three warning features while parallel trends still holds.

Editorial extensions

If this is right

  • Researchers can reject parallel trends before estimation when their causal diagram shows pre-treatment outcomes influencing treatment while unmeasured confounding between treatment and the post-treatment outcome remains.
  • A graph whose minimally sufficient adjustment sets differ between the pre- and post-treatment outcomes is incompatible with parallel trends under linear faithfulness; such graphs should steer analysts toward other designs.
  • An arrow from the pre-treatment to the post-treatment outcome should be treated as a warning flag, since in reasonable semiparametric models it makes parallel trends depend on exact coincidence.
  • Even a graph with none of the three features does not verify parallel trends; it only narrows the required justification to additive homogeneous confounding.
  • These results extend earlier warnings, which were confined to linear structural equation models, to nonparametric structural models and graphs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The checklist suggests a sensitivity analysis: for a graph that violates Conditions 1 or 2, one could quantify how large a linear-faithfulness violation would have to be for parallel trends to survive, turning the rejection into a graded warning.
  • The paper's logic implies that empirically observed parallel pre-trends cannot rescue a graph with disjoint sufficient sets; if the graph is right, the pre-trends must be a coincidence, which is a testable prediction when many similar policy evaluations are analyzed together.
  • Condition 3's conjecture could be probed by constructing a fully nonparametric model in which h(Y0) enters nonlinearly and checking whether parallel trends forces h to be uncorrelated with the common confounder set, a strict condition the paper demonstrates only in separable models.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper aims to provide graphical guidance for deciding whether the parallel trends assumption underlying difference-in-differences is plausible. Under a linear faithfulness assumption and a nonparametric structural equation model with unmeasured common causes, the authors claim that parallel trends implies three conditions: (1) no direct effect of the pre-treatment outcome Y0 on treatment A in the presence of unmeasured confounding; (2) the pre- and post-treatment outcomes have common minimally sufficient adjustment sets; and (3) no direct effect of Y0 on the untreated potential outcome Y1^0. They further claim that, absent these features, parallel trends is equivalent to an 'additive homogeneous confounding' condition with respect to a common sufficient set. The paper applies this framework to Medicaid expansion and insurance coverage. The writing is clear and the authors are transparent about the heuristic status of Condition 3, but the central lemmas used to derive Conditions 1 and 2 are false as stated.

Significance. If the central results were correct, the paper would provide a useful bridge between causal diagrams and the scale-dependent parallel trends assumption, allowing applied researchers to use substantive graphical knowledge to assess a key DID assumption. The paper also contributes a useful explicit statement of linear faithfulness and an application to a realistic policy question, and it includes reproducible R code for illustrative minimal sufficient set calculations. However, the main rejection criteria rest on incorrect implications; the counterexample below satisfies the paper's own assumptions and shows that Condition 1 and the additive homogeneous confounding necessity are invalid. The contribution as stated therefore does not stand, although the broader goal of connecting DID assumptions to graphical structure remains valuable.

major comments (3)
  1. [§3.7, Lemma 1; §4.2, Eq. (8)] Lemma 1 is false as stated. Let U, εY0, eA, e1 be independent mean-zero normal variables with variances 1, 2, 1, 1, and define Y0 = U + εY0, A = 1{2U + εY0 + eA > 0}, Y1^0 = 2U + e1. Put D = Y1^0 - Y0^0 and L = 2U + εY0 + eA. Then (D, L) is jointly normal with Cov(D, L) = 2Var(U) - Var(εY0) = 0, so D⊥⊥L; since A is a function of L and the independent eA, D⊥⊥A, so parallel trends holds. Yet pa(A) includes U and Y0, and E(D | U, Y0) = 2U - Y0 is not constant, so the conclusion E(D | pa(A)) = E(D) of Lemma 1 fails. This model has the Y0→A arrow and the open confounded path A←U→Y1^0 that Condition 1 declares incompatible with parallel trends, and it satisfies linear faithfulness: the d-connected vertices in this graph have nonzero conditional covariances, for example Cov(A, Y1^0 | Y0) > 0. Equation (8) is exactly the invalid step: parallel trends gives mean independence of D from A, not from the parents of A. Consequently Condition 1 is not a valid necessary condition.
  2. [§4.1, Lemma 2; Appendix A; Remark 1] Lemma 2 and the claimed necessity of additive homogeneous confounding are also false. In the same model, M = {U, Y0} is a common sufficient set for Y0 and Y1^0: conditional on M, A depends only on eA, which is independent of Y0 and of Y1^0, so E(Yt^0 | A, M) = E(Yt^0 | M) for t = 0, 1. Parallel trends holds, but E(D | M) = E(D | U, Y0) = 2U - Y0, which is not constant, contradicting Eq. (6). The proof in Appendix A is invalid: the equality E{π(M)E(D|M)} = 0 is obtained only for the actual propensity score π(M) = E[A | M], but the proof then replaces π(M) by indicator functions of {E(D|M) ≥ 0} and {E(D|M) ≤ 0}, as if parallel trends held for every propensity score. That inference is not licensed. Since Lemma 3 and Condition 2 are derived from Lemma 2, the common-minimally-sufficient-set criterion is unsupported.
  3. [§4.4, Condition 3; §4.5 summary] The treatment of Condition 3 is explicitly conditional and does not support the summary claim in §4.5 that 'no arrow from Y0 to Y1' is one of the conditions implied by parallel trends under the paper's assumptions. The proof in §4.4 only shows that, in an additively separable model, parallel trends cannot hold in both G0 and G1 without violating linear faithfulness; it does not establish that parallel trends is impossible in G1, nor does it quantify the 'strongly questioned' claim for general nonparametric models. The extension to the nonparametric setting is a conjecture. The authors are candid about this limitation, but the abstract and Section 4.5 present Condition 3 as part of the operative checklist, which exceeds what is proven.
minor comments (4)
  1. [References and Section 3.7] The citation to Ghanem et al. is inconsistent: Section 3.7 cites Lemma F.3 as Ghanem et al. (2024), while the reference list and Section 2 identify the paper as Ghanem et al. (2022).
  2. [§3.5, Figure 1] The status of the edges among U1, U2, and U3 is described only informally; the text says U1, U2, and U3 can impact U4 but leaves their mutual relationships otherwise unspecified. A clearer statement of which edges are definitely present versus unknown would help the reader interpret the partially directed SWIG.
  3. [Appendix D] The reliance on dagitty output for the minimal sufficient sets of Figure 4, with the comment that showing the result analytically is complex, leaves the reader without a verifiable argument for a claim that is used in the main text. A proof or a more detailed derivation would strengthen the paper.
  4. [Throughout] Several minor language issues remain: 'canonical' is misspelled in the caption of Figure 1, Appendix B contains the phrase 'a colliders', and Assumption 2 is phrased in a way that is close to tautological ('are either not all positive or not all negative').

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the derivation chain is anchored in an external lemma and self-contained proofs.

full rationale

The paper's central results are conditional logical implications, not empirical predictions. Parallel trends (Definition 1) is taken as a given assumption, and the paper derives graphical conditions under which parallel trends can be rejected or supported. The load-bearing Lemma 1 is explicitly imported from Ghanem et al. (2022, Lemma F.3), an external source, and the paper states that its Lemma 2 proof 'builds directly on Ghanem et al. (2022)'s proof of Lemma 1.' This is legitimate external support rather than circular self-citation. Lemma 3 and the additive homogeneous confounding condition are derived in the appendices with explicit arguments; the sufficiency direction is proved directly from the definition of a sufficient set and equation (6). No fitted parameter is later relabeled as a prediction, no quantity is defined in terms of the quantity it is supposed to establish, and no load-bearing claim is justified only by a citation to the authors' own prior work. The paper even acknowledges the scale-dependence of parallel trends and explicitly refrains from claiming the graph alone implies parallel trends, instead identifying additive homogeneous confounding as a separate, extra-graphical condition that must be justified. The skeptical counterexample concerning Lemma 1, if valid, would be a mathematical error in an imported lemma, not a circularity; per the hard rules, lack of correctness is not itself evidence of circularity. The analysis is therefore self-contained apart from its stated external foundation in Ghanem et al., and no circular step can be exhibited.

Assumptions & free parameters 0 free parameters · 5 assumptions · 1 invented entities

The analysis relies on standard causal inference assumptions (Markov, faithfulness, no anticipation) and a mild regularity condition. The unmeasured variables U1-U4 are a modeling device, not entities with independent evidence. No free parameters are fitted.

assumptions (5)
  • domain assumption Causal Markov assumption: each node is independent of its non-descendants given its parents.
    Stated in Section 3.2; used to connect d-separation to conditional independence, which is essential for all SWIG-based arguments.
  • domain assumption Linear faithfulness: d-connected variables have nonzero covariance.
    Assumption 1, Section 3.3. This is the key extra-graphical input that lets the paper move from graph structure to rejection of parallel trends. It is explicitly acknowledged as possibly violated by exact cancellations.
  • domain assumption No anticipation and causal consistency.
    Section 3.1, standard DID assumptions needed to identify the ATT and to connect observed outcomes to potential outcomes.
  • domain assumption Regularity condition of varying conditional trends (Assumption 2).
    Introduced in Appendix A to prove Lemma 2. It requires that conditional trends given M are not all positive or all negative. This is a mild but not automatic condition.
  • standard math Existence of a common sufficient set for both outcomes.
    Proven as Proposition 1 in Appendix B using d-separation rules and results from Shpitser et al. (2012). Ensures that the additive homogeneous confounding condition is not vacuous.
invented entities (1)
  • Unmeasured common-cause categories U1-U4
    purpose: Represent all possible unmeasured common causes of A, Y0 and Y1 in the nonparametric structural equation model.
    These are generic latent variables used to structure the graphical analysis, not new physical or causal entities with independent empirical handles. They are defined by which observed variables they affect.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Using causal diagrams to assess parallel trends in difference-in-differences studies." pith.science (2026). https://pith.science/paper/MHBPS4BK

@misc{pith2026250503526,
  author       = {Pith},
  title        = {Pith review of: Using causal diagrams to assess parallel trends in difference-in-differences studies},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MHBPS4BK}},
  note         = {Machine review of arXiv:2505.03526}
}
read the original abstract

Difference-in-differences (DID) is popular because it can allow for unmeasured confounding when the key assumption of parallel trends holds. However, there exists little guidance on how to decide a priori whether this assumption is reasonable. We attempt to develop such guidance by considering the relationship between a causal diagram and the parallel trends assumption. This is challenging because parallel trends is scale-dependent and causal diagrams are generally scale-independent. We develop conditions under which, given a nonparametric causal diagram, one can reject or fail to reject parallel trends. In particular, we adopt a linear faithfulness assumption, which states that all graphically connected variables are correlated, and which is often reasonable in practice. We show that parallel trends can be rejected if either (i) the treatment is affected by pre-treatment outcomes, or (ii) there exist unmeasured confounders for the effect of treatment on pre-treatment outcomes that are not confounders for the post-treatment outcome, or vice versa (more precisely, the two outcomes possess distinct minimally sufficient sets). We also argue that parallel trends should be strongly questioned if (iii) the pre-treatment outcomes affect the post-treatment outcomes (though the two can be correlated) since there exist reasonable semiparametric models in which such an effect violates parallel trends. When (i-iii) are absent, a necessary and sufficient condition for parallel trends is that the association between the common set of confounders and the potential outcomes is constant on an additive scale, pre- and post-treatment. These conditions are similar to, but more general than, those previously derived in linear structural equations models. We discuss our approach in the context of the effect of Medicaid expansion under the U.S. Affordable Care Act on health insurance coverage rates.

Figures

Figures reproduced from arXiv: 2505.03526 by the authors.

Figure 1
Figure 1. Partially directed SWIG (left) and corresponding structu [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. A directed SWIG in which Y0 and Y1 contain the disjoint sufficient adjustment sets {U1, U3} and {U1, U4}, respectively, and which are a each subset of the common adjustment set {U1, U3, U4}. M0 = {U1, U3} sufficient for Y0 but not Y1. In other words, the graph implies that if we omit U4 from M, we still have a sufficient set for Y0, but we no longer have a sufficient set for Y1. However, from Lemma 3 we have that pa… view at source ↗
Figure 3
Figure 3. Simplified SWIG representing G1 which, from (9), implies αE(Y˙ 0 0 |M) = 0, a violation of linear faithfulness. If parallel trends cannot hold in both G0 and G1, which model should one prefer? We argue that one should prefer G0 for reasons we will now explain. For parallel trends to hold in G1 requires special balancing of associations, which is evident from equation (10) as well. In a slight abuse of notation, conc… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: A partially directed SWIG in which parallel trends would be rej [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]
Figure 5
Figure 5. Figure 5: Maximal SWIG consistent with conditions 1, 2, and 3; when a [PITH_FULL_IMAGE:figures/full_fig_p018_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

25 extracted references · 21 canonical work pages

  1. [1]

    A. Abadie. Semiparametric difference-in-differences estimators. The review of economic studies, 72 0 (1): 0 1--19, 2005

  2. [2]

    Andersen

    H. Andersen. When to expect violations of causal faithfulness and why it matters. Philosophy of Science, 80 0 (5): 0 672--683, 2013

  3. [3]

    J. D. Angrist and J.-S. Pischke. Mostly harmless econometrics: An empiricist's companion. Princeton university press, 2009

  4. [4]

    Ashenfelter

    O. Ashenfelter. Estimating the effect of training programs on earnings. The Review of Economics and Statistics, pages 47--57, 1978

  5. [5]

    Callaway

    B. Callaway. Difference-in-differences for policy evaluation. Handbook of Labor, Human Resources and Population Economics, pages 1--61, 2023

  6. [6]

    Cunningham

    S. Cunningham. Causal inference: The mixtape. Yale university press, 2021

  7. [7]

    I. J. Dahabreh and M. A. Hern \'a n. Extending inferences from a randomized trial to a target population. European journal of epidemiology, 34: 0 719--722, 2019

  8. [8]

    Ghanem, P

    D. Ghanem, P. H. Sant'Anna, and K. W \"u thrich. Selection and parallel trends. arXiv preprint arXiv:2203.09001, 2022

Show all 25 references
  1. [9]

    Greenland, J

    S. Greenland, J. Pearl, and J. M. Robins. Causal diagrams for epidemiologic research. Epidemiology, 10 0 (1): 0 37--48, 1999

  2. [10]

    Kim and P

    Y. Kim and P. M. Steiner. Gain scores revisited: A graphical models perspective. Sociological Methods & Research, 50 0 (3): 0 1353--1375, 2021

  3. [11]

    T. L. Lash, M. P. Fox, R. F. MacLehose, G. Maldonado, L. C. McCandless, and S. Greenland. Good practices for quantitative bias analysis. International journal of epidemiology, 43 0 (6): 0 1969--1985, 2014

  4. [12]

    Lechner et al

    M. Lechner et al. The estimation of causal effects by difference-in-difference methods. Foundations and Trends in Econometrics , 4 0 (3): 0 165--224, 2011

  5. [13]

    J. Pearl. Causality. Cambridge university press, 2009

  6. [14]

    T. S. Richardson and J. M. Robins. Single world intervention graphs (swigs): A unification of the counterfactual and graphical approaches to causality. Center for the Statistics and the Social Sciences, University of Washington Series. Working Paper, 128 0 (30): 0 2013, 2013

  7. [15]

    P. R. Rosenbaum and D. B. Rubin. The central role of the propensity score in observational studies for causal effects. Biometrika, 70 0 (1): 0 41--55, 1983

  8. [16]

    Shpitser, T

    I. Shpitser, T. VanderWeele, and J. M. Robins. On the validity of covariate adjustment for estimating causal effects. arXiv preprint arXiv:1203.3515, 2012

  9. [17]

    Shrier and R

    I. Shrier and R. W. Platt. Reducing bias through directed acyclic graphs. BMC medical research methodology, 8: 0 1--15, 2008

  10. [18]

    Sofer, D

    T. Sofer, D. B. Richardson, E. Colicino, J. Schwartz, and E. J. T. Tchetgen. On negative outcome control of unobserved confounding as a generalization of difference-in-differences. Statistical science: a review journal of the Institute of Mathematical Statistics, 31 0 (3): 0 348, 2016

  11. [19]

    Spirtes, C

    P. Spirtes, C. Glymour, and R. Scheines. Causation, prediction, and search. MIT press, 2001

  12. [20]

    D. Steel. Homogeneity, selection, and the faithfulness condition. Minds and Machines, 16: 0 303--317, 2006

  13. [21]

    A. M. Weber, M. J. van der Laan, and M. L. Petersen. Assumption trade-offs when choosing identification strategies for pre-post treatment effect estimation: an illustration of a community-based intervention in madagascar. Journal of causal inference, 3 0 (1): 0 109--130, 2015

  14. [22]

    J. M. Wooldridge. Econometric Analysis of Cross Section and Panel Data. MIT Press, 2010

  15. [23]

    J. M. Wooldridge. Two-way fixed effects, the two-way mundlak regression, and difference-in-differences estimators. Available at SSRN 3906345, 2021

  16. [24]

    Zeldow and L

    B. Zeldow and L. A. Hatfield. Confounding and regression adjustment in difference-in-differences studies. Health services research, 56 0 (5): 0 932--941, 2021

  17. [25]

    Zhang, C

    C. Zhang, C. Cinelli, B. Chen, and J. Pearl. Exploiting equality constraints in causal inference. In International Conference on Artificial Intelligence and Statistics, pages 1630--1638. PMLR, 2021

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.