REVIEW 3 major objections 4 minor 44 references
Partial Homogeneity in Staggered Difference-in-Differences
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A Dirichlet process mixture over cohort-time effects turns the choice between fully flexible and fully pooled staggered difference-in-differences into a partition-selection problem, yielding unbiased variance reductions of 26–52% when…
desk verdict A useful and honest method for pooling CATT cells when effects are clumpy, but the abstract oversells the regime of applicability and the paper leaves the key formal theory for later. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the Dirichlet Process mixture prior on the $K$ CATT parameters, which induces a random partition through ties and is marginalised to a Chinese Restaurant Process prior over partitions. A collapsed Gibbs sampler draws cluster assignments from the marginal likelihood with group effects integrated out, yielding posterior means, credible intervals, and co-clustering probabilities. The paper also proves that with the error variance fixed and a pairwise prior charging each separated pair, the joint MAP is exactly an $\ell_0$-penalized regression, $RSS(\mathcal{P}) + \lambda \cdot c(\mathcal{P})$, connecting the Bayesian procedure to homogeneity pursuit.
What would settle it
Extend the paper's Table 3 grid at $\delta=6$, $m^*=9$ to much larger $N$: if the variance ratio of Bayes-PH and $\ell_0$-PH to flexible TWFE stays above 1 even as the adjusted Rand index approaches 1, then the claim that separation alone determines the efficiency gain is false; if it drops below 1, the paper's regime boundary is a finite-sample artifact of its 2,000-unit panels.
Extended reading notes
Core claim
Under partial homogeneity—some cohort-time effects exactly equal, others distinct—the paper establishes that grouping the effects according to the true partition makes the restricted estimator unbiased and more efficient than the fully flexible estimator, group by group, by the Gauss–Markov theorem. The unknown partition is recovered by a Dirichlet Process mixture via collapsed Gibbs sampling; the posterior mean and the $\ell_0$-penalized MAP both target the same restricted estimator, differing in how they handle selection uncertainty. The central empirical claim is that in the clumpy-effects regime the feasible estimators approach the oracle variance reduction while avoiding the pooled estimator's bias, and that posterior averaging over partitions delivers honest coverage where conditional-on-selection intervals can under-cover.
Load-bearing premise
The method's promised gains require that some true cohort-time effects are exactly equal to each other and that the distinct effect levels are separated by roughly six or more standard errors of the flexible estimates, so that the partition can actually be recovered; without that separation the variance advantage reverses.
Editorial extensions
If this is right
- When the true CATT partition has well-separated groups, both the $\ell_0$-PH and Bayes-PH estimators reduce the sampling variance of cohort-time effects by 26–52% relative to flexible TWFE while remaining essentially unbiased.
- Bayesian credible intervals that average over the unknown partition reach near-nominal coverage (about 0.92–0.94) in the motivating regime, while plug-in intervals that condition on a selected partition can under-cover badly when separation is low.
- The method degrades gracefully outside its target regime: at low separation ($\delta=3$) or with a fine partition ($m^*=9$ at $\delta=6$), neither feasible estimator beats flexible TWFE on variance, so the gains are specific to clumpy, separable heterogeneity.
- The same procedure can act as a formal specification test, as shown by its two applications: it recovers a partially homogeneous structure where effects are genuinely heterogeneous and correctly reports that full pooling is adequate where they are not.
- The fixed-variance MAP under a pairwise partition prior is exactly an $\ell_0$-penalized regression, so the Bayesian formulation is connected to the homogeneity-pursuit literature through an explicit penalty on separated CATT pairs.
Reading between the lines
- Going beyond the paper, the efficiency gains likely extend to any panel setting where coefficients are believed to cluster into exact levels, not just staggered DiD; the partition-selection formulation is generic.
- The paper leaves formal consistency theory for future work; a natural testable extension is to derive the separation threshold analytically as a function of $N$ and the per-cell effective sample size, rather than leaving it as a simulation-calibrated heuristic.
- The applications show that the exact cross-cell covariance is essential: the paper notes a diagonal approximation understates posterior variance of aggregates by about a factor of 3.5, so practitioners should supply a joint first-stage covariance rather than cell-wise standard errors.
- Because the method can also identify cells that lack a clean comparison through their group's effect, it could be combined with limited-overlap designs, though the paper cautions that such identification rests entirely on the assumed grouping.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper frames the specification choice in staggered difference-in-differences as a partition-selection problem on the cohort-time cells: instead of estimating every CATT freely or pooling all CATTs into one TWFE coefficient, it proposes recovering a partial-homogeneity structure in which some CATTs are exactly equal. The main estimator is a Dirichlet Process mixture prior on the CATTs with a collapsed Gibbs sampler; a fixed-variance MAP version is shown to coincide with an l0-penalized regression. The paper reports a calibrated simulation in which the feasible estimators cut sampling variance by 26–52% relative to the flexible estimator in clumpy-effect regimes, with near-nominal credible-interval coverage, plus two empirical applications. It also explicitly discloses regimes where the variance advantage reverses (low separation, large number of true groups) and defers formal consistency theory to future work.
Significance. If the central claims were fully established, the paper would be a useful addition to the staggered-DiD toolkit: it addresses a real specification problem and connects a Bayesian partition model to the homogeneity-pursuit literature. The strengths are substantial: Proposition 1 is a clean Gauss–Markov argument, Proposition 2 gives an exact MAP-to-l0 connection with a clearly stated pairwise prior, the simulation design includes negative controls (delta=3, m*=9, m*=18) and does not hide them, Remark 1 correctly insists on the exact cross-cell covariance, and Appendix E validates the sampler against exact enumeration for K=7. The core difficulty is that the abstract and conclusion state the efficiency and coverage results in a stronger form than the paper's own tables support, and the missing consistency theory is load-bearing for the stated condition.
major comments (3)
- [Abstract; §4.2, Table 2 (m*=9, delta=6)] The abstract's proviso 'provided the distinct effects are separated enough to be recovered' is not sufficient, by the paper's own results. At m*=9, delta=6, the adjusted Rand index for l0-PH is 0.91 and for Bayes-PH is 0.77, so the true partition is largely recovered under the paper's ARI metric, yet the variance ratios are 1.26 and 1.63, meaning both feasible estimators are less precise than fully flexible TWFE. Thus separation plus successful partition recovery is not enough to deliver the promised variance gain; the number of true groups m* (and the K/m selection overhead) also matters. The abstract, Section 1, and Section 6 should be revised to state the condition as 'a small number of well-separated equal-effect groups,' and the paper should either characterize the selection overhead as a function of K and m* or explicitly label the 26–52% gain as a small-m* phenomenon.
- [Abstract; §4.3, Table 5] The abstract claims that 'the posterior delivers near-nominal confidence-interval coverage by averaging over the unknown partition,' but Table 5 reports Bayes-PH CATT coverage of 0.81 at delta=3, m*=6 and 0.79–0.80 at m*=18 across delta values. These are not near-nominal 95% coverages. The coverage claim is therefore only valid in the separable, moderate-m* cells, and the abstract and Section 6 should be qualified accordingly, either by stating the hard-regime degradation explicitly or by describing it as graceful decline rather than near-nominal performance.
- [§3.2, Proposition 3; §6] The paper's only theoretical support for the separation threshold is the heuristic calculation in Proposition 3, in which a fixed lambda leaves a constant over-splitting probability because the correct-merge cost is O_p(sigma^2), and the paper explicitly defers a formal consistency theorem. This missing theory is load-bearing rather than a routine extension: Table 2 shows that recovery (ARI) and variance gains can diverge sharply, so a formal result would need to characterize the joint dependence on delta and on K/m, not just on delta. I am not asking for the theorem in this revision, but the paper's claims about what separation 'provides' should be restated as simulation-based and heuristic until that characterization exists.
minor comments (4)
- [Eq. (35)] The l0 penalty term is written as a sum over g≠g' of 1[tau_gt ≠ tau_g't], which is not well defined because t is not indexed in the summation; it should sum over pairs of cohort-time cells, e.g. (g,t) and (g',t').
- [§5.2.1] The sentence 'Because the design matrix of Equation 43 is block-diagonal, the estimation of each hat_tau_g is perfectly separable' conflicts with the immediately following sentence that control states recur across stacks and the event estimates are not independent; please rephrase to say that the point estimates are separable under the working design while the sampling covariance is not diagonal.
- [§4.1 vs. Appendix E] Section 4.1 says the DP sampler runs for 150 Gibbs sweeps with a 50-sweep burn-in, but Appendix E refers to a '100-draw setting used in the main experiments'; please reconcile the two numbers and, ideally, report convergence diagnostics at the actual settings used in the tables.
- [Throughout] There are minor OCR/reproduction artifacts in names, such as 'Bijani' for 'Bijani' and some broken inline math (e.g. the 'ci' fragment in Table 1); these should be cleaned in the final version.
Circularity Check
One by-construction l0-MAP identity; the central variance-reduction claims come from correctly specified simulations with disclosed failure cells, so circularity is minor and not load-bearing.
-
self definitional
[Section 3.2, Eq. (33) and Proposition 2]
"It is cleanest under a pairwise partition prior that charges each pair of CATTs assigned to different groups, Pr(P) ∝ exp(−λ0 c(P)), c(P)=... Proposition 2 (ℓ0-PH as a fixed-variance MAP). Fix σ², place an improper flat prior p(φ|P) ∝ 1 on the group effects, and adopt the pairwise prior (33). Then the joint maximum a posteriori estimator ... has partition component Q(P)=RSS(P)+λc(P), λ=2σ²λ0."
The claimed connection between the Bayesian model and ℓ0 homogeneity pursuit is produced by construction: the pairwise partition prior's log probability is defined to equal −λ0 times the number of cross-group pairs c(P), which is exactly the ℓ0 penalty appearing in Q(P). Exponentiating that penalty as the prior guarantees the MAP minimizes RSS plus the same penalty, so Proposition 2 is an algebraic identity rather than an independent equivalence. The paper itself acknowledges in Remark 2 that the CRP prior does not reduce to RSS, confirming that the ℓ0 reduction is an artifact of the specially chosen pairwise prior, not a consequence of the DP model.
full rationale
The paper's central efficiency claim is not circular: the 26–52% variance reductions are Monte Carlo outcomes from a correctly specified partial-homogeneity design, not quantities fitted to the data to manufacture the result. The estimators are evaluated with fixed tuning choices (BIC for ℓ0-PH, diffuse base measure and α=7 for Bayes-PH), and the paper honestly reports cells where the gain reverses (δ=3; m*=9 at δ=6; m*=18). The external applications use real data and pre-specified first-stage covariances, so they are not self-confirmatory. The only step that reduces by construction is the Proposition 2 identity between the fixed-variance MAP under the pairwise prior and the ℓ0 objective; because the prior is defined to be the exponentiated penalty, this is a definitional equivalence. It is explicitly labeled a 'connection' and is not load-bearing for the abstract's performance claims. The m*=9/δ=6 cell shows the abstract's 'separated enough to be recovered' proviso is insufficient, but that is a correctness and calibration concern about the strength of the claim, not circularity. Overall, the derivation chain is self-contained apart from the one minor by-construction identity.
Assumptions & free parameters
free parameters (7)
- DP concentration alpha =
7 in main simulation; sensitivity across 0.1-100 in applications
- Base measure variance sigma0^2 =
100
- Base mean mu0 =
0
- Inverse-gamma hyperparameters a0, b0 for error variance =
not specified numerically
- l0 penalty lambda =
chosen by BIC in simulation and applications
- Separation delta (simulation DGP parameter) =
6 in headline; also 3 and 12
- Number of true groups m* (simulation DGP parameter) =
swept {1,3,6,9,18}
assumptions (11)
- domain assumption Parallel trends (Assumption 3)
- domain assumption No treatment anticipation (Assumption 4)
- domain assumption Random sample and irreversible treatment (Assumptions 1, 2)
- domain assumption Homoskedastic normal errors in the main model
- standard math Gauss-Markov theorem
- standard math Woodbury identity and matrix determinant lemma
- standard math Dirichlet Process and Chinese Restaurant Process properties
- ad hoc to paper Exact equality partial homogeneity
- standard math Schwarz / Laplace approximation exact up to O(1)
- domain assumption First-stage asymptotic normality and consistent covariance
- domain assumption Within-event exchangeability for the randomization test
Cite this review
Pith. "Pith review of Partial Homogeneity in Staggered Difference-in-Differences." pith.science (2026). https://pith.science/paper/EF6R73RN
@misc{pith2026260808047,
author = {Pith},
title = {Pith review of: Partial Homogeneity in Staggered Difference-in-Differences},
year = {2026},
howpublished = {\url{https://pith.science/paper/EF6R73RN}},
note = {Machine review of arXiv:2608.08047}
}
abstract
In staggered difference-in-differences (DiD) designs, units enter treatment at different calendar times, so the treatment effect is not a single number but a set of Cohort-Average Treatment effects on the Treated (CATTs), one per cohort-time cell. Estimating every CATT as its own parameter, as the standard fully flexible estimator does, is unbiased but inefficient when some of these effects are in fact equal, whereas pooling them all into a single two-way fixed effects (TWFE) coefficient is efficient but, whenever the heterogeneity is genuine, biased for the individual effects. We frame the choice between these extremes as a partition-selection problem on the cohort-time cells and address it with a Dirichlet Process (DP) mixture prior on the CATTs. The model favors parsimonious groupings without fixing their number, and a collapsed Gibbs sampler delivers point estimates, credible intervals that marginalize the unknown partition, and co-clustering probabilities for every pair of CATTs. With the error variance held fixed and a pairwise penalty placed on the partition, a maximum a posteriori (MAP) partition reduces to an $\ell_0$-penalized regression, connecting the Bayesian formulation to the homogeneity-pursuit literature. In a calibrated simulation, the model cuts the sampling variance of the cohort-time effects by 26--52\% relative to the fully flexible estimator, without the pooled estimator's bias, provided the distinct effects are separated enough to be recovered, and the posterior delivers near-nominal confidence-interval coverage by averaging over the unknown partition. In two applications the method recovers a precision-improving partial-homogeneity structure in one, where the cohort-time effects are genuinely heterogeneous, and reports that full pooling is adequate in the other, where they are not.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Journal of Econometrics , year =
Goodman-Bacon, Andrew , title =. Journal of Econometrics , year =
-
[2]
Two-Way Fixed Effects Estimators with Heterogeneous Treatment Effects , journal =
de Chaisemartin, Cl\'ement and D'Haultf. Two-Way Fixed Effects Estimators with Heterogeneous Treatment Effects , journal =. 2020 , volume =
work page 2020
-
[3]
de Chaisemartin, Cl\'ement and D'Haultf. Two-Way Fixed Effects and Differences-in-Differences with Heterogeneous Treatment Effects: A Survey , journal =. 2023 , volume =
work page 2023
- [4]
-
[5]
Callaway, Brantly and Sant'Anna, Pedro H. C. , title =. Journal of Econometrics , year =
-
[6]
Journal of Econometrics , year =
Sun, Liyang and Abraham, Sarah , title =. Journal of Econometrics , year =
-
[7]
Review of Economic Studies , year =
Borusyak, Kirill and Jaravel, Xavier and Spiess, Jann , title =. Review of Economic Studies , year =
-
[8]
arXiv preprint arXiv:2207.05943 , year =
Gardner, John , title =. arXiv preprint arXiv:2207.05943 , year =
Show all 44 references
-
[9]
American Journal of Political Science , year =
Liu, Licheng and Wang, Ye and Xu, Yiqing , title =. American Journal of Political Science , year =
-
[10]
, title =
Wooldridge, Jeffrey M. , title =. Empirical Economics , year =
-
[11]
Su, Liangjun and Shi, Zhentao and Phillips, Peter C. B. , title =. Econometrica , year =
-
[12]
Econometrica , year =
Bonhomme, St\'ephane and Manresa, Elena , title =. Econometrica , year =
-
[13]
Journal of the American Statistical Association , year =
Ke, Yuan and Fan, Jianqing and Wu, Yichao , title =. Journal of the American Statistical Association , year =
-
[14]
Working paper, arXiv:2510.05454 , year =
Kwon, Soonwoo and Sun, Liyang , title =. Working paper, arXiv:2510.05454 , year =
-
[15]
and Koles\'ar, Michal , title =
Armstrong, Timothy B. and Koles\'ar, Michal , title =. Econometrica , year =
-
[16]
and Kline, Patrick and Sun, Liyang , title =
Armstrong, Timothy B. and Kline, Patrick and Sun, Liyang , title =. Forthcoming at Econometrica , year =
-
[17]
, title =
Ferguson, Thomas S. , title =. The Annals of Statistics , year =
-
[18]
, title =
Antoniak, Charles E. , title =. The Annals of Statistics , year =
-
[19]
and West, Mike , title =
Escobar, Michael D. and West, Mike , title =. Journal of the American Statistical Association , year =
-
[20]
Biometrika , year =
M\"uller, Peter and Erkanli, Alaattin and West, Mike , title =. Biometrika , year =
-
[21]
, title =
Neal, Radford M. , title =. Journal of Computational and Graphical Statistics , year =
-
[22]
, title =
M\"uller, Peter and Quintana, Fernando A. , title =. Statistical Science , year =
-
[23]
and Harrison, Matthew T
Miller, Jeffrey W. and Harrison, Matthew T. , title =. Journal of the American Statistical Association , year =
-
[24]
, title =
Hartigan, John A. , title =. Communications in Statistics - Theory and Methods , year =
-
[25]
IEEE Transactions on Automatic Control , year =
Akaike, Hirotugu , title =. IEEE Transactions on Automatic Control , year =
-
[26]
The Annals of Statistics , year =
Schwarz, Gideon , title =. The Annals of Statistics , year =
-
[27]
Journal of the American Statistical Association , year =
Fan, Jianqing and Li, Runze , title =. Journal of the American Statistical Association , year =
-
[28]
Journal of the Royal Statistical Society: Series B , year =
Tibshirani, Robert , title =. Journal of the Royal Statistical Society: Series B , year =
-
[29]
Journal of the Royal Statistical Society: Series B , year =
Tibshirani, Robert and Saunders, Michael and Rosset, Saharon and Zhu, Ji and Knight, Keith , title =. Journal of the Royal Statistical Society: Series B , year =
-
[30]
The Quarterly Journal of Economics , year =
Cengiz, Doruk and Dube, Arindrajit and Lindner, Attila and Zipperer, Ben , title =. The Quarterly Journal of Economics , year =
-
[31]
Bayesian Analysis , year =
Wade, Sara and Ghahramani, Zoubin , title =. Bayesian Analysis , year =
-
[32]
and Green, Peter J
Lau, John W. and Green, Peter J. , title =. Journal of Computational and Graphical Statistics , year =
-
[33]
, title =
Leeb, Hannes and P\"otscher, Benedikt M. , title =. Econometric Theory , year =
-
[34]
The Annals of Statistics , year =
Rinaldo, Alessandro and Wasserman, Larry and G'Sell, Max , title =. The Annals of Statistics , year =
-
[35]
arXiv:1410.2597 , year =
Fithian, William and Sun, Dennis and Taylor, Jonathan , title =. arXiv:1410.2597 , year =
-
[36]
Wang, Wuyi and Phillips, Peter C. B. and Su, Liangjun , title =. Journal of Applied Econometrics , year =
-
[37]
Quantitative Economics , year =
Lu, Xun and Su, Liangjun , title =. Quantitative Economics , year =
-
[38]
Roth, Jonathan and Sant'Anna, Pedro H. C. and Bilinski, Alyssa and Poe, John , title =. Journal of Econometrics , year =
-
[39]
The Review of Economic Studies , year =
Rambachan, Ashesh and Roth, Jonathan , title =. The Review of Economic Studies , year =
-
[40]
, title =
Barry, Daniel and Hartigan, John A. , title =. The Annals of Statistics , year =
-
[41]
The Annals of Statistics , year =
Bertsimas, Dimitris and King, Angela and Mazumder, Rahul , title =. The Annals of Statistics , year =
-
[42]
Journal of the American Statistical Association , year =
Shen, Xiaotong and Huang, Hsin-Cheng , title =. Journal of the American Statistical Association , year =
-
[43]
, title =
Hill, Jennifer L. , title =. Journal of Computational and Graphical Statistics , year =
-
[44]
, title =
Fisher, Walter D. , title =. Journal of the American Statistical Association , year =
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.