{"id":"cdc266fa-64be-49ca-99e0-f89de9300c9d","arxiv_id":"2606.17977","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A staggered difference-in-differences estimator that identifies treatment effects by extrapolating a pre-treatment polynomial gap, instead of requiring a flat gap.","lead":"This paper relaxes the standard \"parallel trends\" assumption in difference-in-differences, allowing the pre-existing gap between treated and untreated groups to follow a polynomial trend instead of staying flat. It shows how to estimate treatment effects in staggered rollouts under this weaker assumption and applies the method to Medicaid expansion, where standard parallel trends is rejected.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central identification is valid but rests entirely on the untestable post-treatment extrapolation in Assumption 4; the sequential pre-test does not certify it.","rationale":"I read the paper in good faith and find the central identification theorem mathematically correct. The weakest point is exactly the one the reader identified: Assumption 4 must hold in the post-treatment period, and this is untestable. The paper is transparent about this limitation, and it even recommends sensitivity bounds as a complement. The sequential order-selection procedure is a further concern, but it is secondary to the extrapolation assumption; even if post-selection inference were fully valid, the point estimates would still be conditional on the polynomial continuing after treatment. My proposed placebo test does not settle the actual post-treatment question, but it would test the credibility of the polynomial extrapolation in a setting where the truth is observed. Since the concern is inherent to the identifying assumption rather than an internal error, I do not recommend changing the reader's CONDITIONAL verdict; the paper should be published with the stated caveats and with the promised code and replication materials.","tokens_in":21183,"tokens_out":7296,"duration_ms":75194,"concrete_test":"Use the 16 never-treated states from the Medicaid application as a pseudo-treated panel. Assign each never-treated state a placebo expansion year in the middle of the sample (e.g., 2014), estimate the DD[p] polynomial using only pre-placebo gaps, and compare the extrapolated post-placebo counterfactual gap to the observed post-placebo gap for the pseudo-treated state versus the remaining never-treated states. Repeat for every never-treated state and for p=1,2,3. If extrapolation errors systematically exceed the bootstrap standard error or grow with the post-placebo horizon, the post-treatment Parallel[p] assumption is not credible even when the pre-treatment polynomial fit is good.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The point-identification result in Theorem 4.3 is internally sound: if Assumption 4 holds for all periods, including post-treatment, then the untreated gap is a polynomial of degree p−1 everywhere and the counterfactual is identified. The load-bearing step is precisely that post-treatment extrapolation. The paper's own sequential order-selection procedure (Section 6.2) uses only pre-treatment data to choose p and then reports the chosen order as 'not rejected' by the data (e.g., Table 8). But this only checks in-sample fit; it does not test whether the polynomial continues after treatment. Any divergence between the true untreated gap and the fitted polynomial — a slope change after treatment, a structural break, or higher-order curvature that appears only post-treatment — becomes bias in ATT(g,t) and in the aggregate θ. This is acknowledged in Remarks 7–9 and Section 9, but it means the practical contribution is not a fully assumption-free relaxation: it replaces the flat-gap extrapolation with a polynomial extrapolation of possibly high order, and the pre-treatment data cannot validate that projection. The Monte Carlo evidence is consistent with this: under DGP-1, the test selects p=2 or p=3 in 14.6% of draws and coverage falls to 89.8%. Thus the main risk is not a proof error but the credibility of the maintained extrapolation, which is structurally untestable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a hierarchy of identifying assumptions, Parallel[p], for staggered difference-in-differences. For a cohort with at least p pre-treatment periods, Parallel[p] requires that the p-th time difference of the untreated potential-outcome gap between that cohort and never-treated units is zero; the pre-treatment gap is then a polynomial of degree p-1 whose forward extrapolation identifies the counterfactual. The paper proves group-time ATT identification (Theorem 4.3), an aggregation theorem allowing cohort-specific orders (Theorem 4.4), and develops a sequential order-selection procedure with cluster-bootstrap inference. Monte Carlo experiments and a Medicaid expansion application illustrate the method. The identification argument is internally sound; the main weaknesses concern the credibility of the post-treatment extrapolation and the inferential claims attached to the order-selection step.","tokens_in":21543,"tokens_out":7017,"duration_ms":75799,"significance":"The paper gives a clean, self-contained proof that group-time ATTs are point-identified under a higher-order parallel-trends assumption, extending Mora and Reggio (2019) to staggered designs. Theorem 4.4, while not deep, addresses a real practical issue: cohorts with different pre-period lengths can be identified under different polynomial orders. The paper is also unusually transparent about the untestable nature of the extrapolation, which is a genuine strength. Its practical value, however, depends on the order-selection procedure and on the claimed post-selection coverage. These inferential foundations are not yet established, and some numerical claims are not supported by the reported tables.","major_comments":[{"comment":"The abstract and Section 7.2 state that post-selection bootstrap coverage is 'near-nominal'. Table 2 reports 89.8% for the flat-gap DGP and 89.2% for the borderline DGP, both based on B=99. These are not near 95% in any conventional sense. The text attributes the shortfall to finite-sample bootstrap behaviour at B=99, but no coverage results are reported at B=999, the value used in the application. The claim should either be supported by B=999 simulations or weakened substantially.","section":"§7.2, Table 2"},{"comment":"The order-selection statistic T_g(p) is explicitly admitted to be only a descriptive diagnostic because the chi-square approximation is invalid under the induced moving-average dependence of differenced outcomes. Yet the sequential algorithm in §6.2 uses exactly this statistic to select the order, and the resulting order is then treated as the identifying assumption in the main estimates and confidence intervals. No formal post-selection validity is established, and the paper's own Section 9 lists this as future work. As a result, the bootstrap CIs in Tables 2 and 7 do not have a proven repeated-sampling property after selection. The paper should either provide a formal uniform post-selection result or clearly label the whole pipeline as heuristic and remove the 'confirms near-nominal coverage' language.","section":"§6.1–6.2, Eq. (12)"},{"comment":"The identification result is conditional on the untestable extension of Parallel[p] to all post-treatment periods. The pre-treatment diagnostics in Section 6 cannot detect a post-treatment structural break or a change in the polynomial's coefficients after treatment; any such divergence enters directly as bias in ATT(g,t) and in the aggregate θ. The paper acknowledges this in Remarks 7–9, but it would be more accurate to state in the abstract and in Theorem 4.3 that point identification holds under this maintained extrapolation, not under an assumption 'the data do not reject'. The Appendix C simulations consider only smooth small curvature; a concrete robustness check allowing a post-treatment trend break of plausible magnitude would help calibrate the practical risk.","section":"§4.2, Assumption 4 / Proposition 4.2"},{"comment":"Under the true Parallel[3] DGP, the post-selection point estimate has bias -0.099, roughly 20% of the true ATT of 0.5, while the text emphasizes only that coverage is 93.8%. This bias is not discussed. It appears to reflect the sequential test stopping at p=2 in 59.2% of draws, so the reported estimates are generated under a misspecified order in a substantial share of replications. This undermines the practical claim that the selector recovers the correct order and should be either explained as an expected cost of the procedure or addressed by a different selection rule.","section":"§7.2, Table 2, True Parallel[3] row"}],"minor_comments":[{"comment":"Strategy III defines p_g = m_g - 1, which is not defined when m_g = 1. Clarify that the feasible set is p_g ∈ {1, ..., m_g} and that Strategy III should be capped at max(1, m_g - 1).","section":"§5.1, Strategy III"},{"comment":"The denominator dVar(Δ^p γ̂_{g,t}) is not defined. Specify the estimator used for this variance and note explicitly that it is not robust to the serial dependence induced by the difference operator, consistent with the descriptive interpretation.","section":"§6.1, Eq. (12)"},{"comment":"Coverage in Table 2 is reported for B=99 bootstrap draws, but the empirical application and the bootstrap stability table use B=999. Reporting coverage at B=999 would make the post-selection claim directly comparable to the application.","section":"Table 2"},{"comment":"The replication code and Stata command are described as 'will be available'; for a methods paper, providing the code at submission would substantially strengthen reproducibility.","section":"Appendix D"},{"comment":"The 'Pre-period implication of Parallel[2] not rejected' statement appears to rely on R² and visual diagnostics rather than on a formal test. Since the test statistic in Eq. (12) is described as descriptive, the language should be consistent and not imply a formal hypothesis-testing result.","section":"§8.2, Table 5"}],"recommendation":"major_revision","confidential_remarks":"The core identification argument is sound and the paper is a useful extension of higher-order parallel trends to staggered designs. The main barrier is not the mathematics but the inferential claims: the order-selection step lacks formal validity, and the Monte Carlo coverage numbers do not support 'near-nominal' as stated. The bias under the true Parallel[3] DGP also needs discussion. I would be willing to see a revised version that either supplies formal post-selection inference or recalibrates the language to describe the procedure as heuristic with sensitivity analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — you can read this one quickly. The core identification under Parallel[p] is Mora and Reggio's, and Egami and Yamauchi already extended it to staggered designs for p=1,2. What's actually new: arbitrary p in staggered settings, a heterogeneous-order aggregation theorem (Thm 4.4), and a practical sequential order-selection workflow with a Stata command. Thm 4.4 is simple — a weighted average of identified ATT(g,t) is identified — but it is a useful formalization that had not been written down, and it answers a real question when cohorts have different pre-period lengths.\n\nThe paper does several things well. The identification proofs are short and correct. The authors are unusually transparent: they state plainly that the post-treatment polynomial extrapolation is untestable, that T_g(p) is a descriptive diagnostic rather than a formal test, and that small cohorts (including a single-state cohort in the application) cannot support asymptotic inference. The Monte Carlo design is sensible, and the extrapolation-horizon appendix is a nice addition.\n\nThe soft spots are where the paper itself says they are. The load-bearing step is Assumption 4 applied to all periods: the pre-treatment gap polynomial continues unchanged after treatment. The sequential pre-test only checks in-sample fit; it cannot detect a post-treatment structural break or slope change. That is not a proof error — every DiD design has an extrapolation step — but the practical gain over Rambachan and Roth is a point estimate in exchange for a stronger maintained assumption. The Monte Carlo coverage is 89–94% at B=99, which is 'near-nominal' only if you squint; the paper should report B=999 in the main tables. There is no formal post-selection inference for the order choice, and the paper acknowledges it. These are addressable.\n\nI'd send this to a serious referee. The contribution is modest but real, the proofs are clean, and the weaknesses are openly acknowledged. Applied readers will use it. My main referee requests: formal or at least improved post-selection coverage results, B=999 Monte Carlo, and a sharper discussion of when polynomial extrapolation is credible relative to sensitivity bounds. Recommend: peer review, expect moderate revision.","headline":"A clean, honest extension of higher-order parallel trends to staggered DiD; the identification is valid, the new aggregation theorem is simple but useful, and the main risk is the untestable post-treatment extrapolation, which the paper acknowledges.","tokens_in":21999,"tokens_out":2782,"would_cite":true,"duration_ms":25733,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that group-time average treatment effects in staggered difference-in-differences designs remain point identified when the pre-treatment outcome gap follows a low-degree polynomial, by replacing the flat-gap parallel trends","keywords":["difference-in-differences","staggered adoption","parallel trends","higher-order parallelism","group-time ATT","polynomial counterfactual","order selection","Medicaid expansion"],"falsifier":"Use the never-treated units to test the extrapolation directly: assign a placebo adoption date g*, split the never-treated units into two random groups, fit the gap between them under Parallel[p] on periods before g*, and compare the projected polynomial to the observed gap after g*. A systematic, non-polynomial deviation larger than sampling noise would falsify the order-p extrapolation and therefore undermine the identifying assumption for treated cohorts.","tokens_in":21095,"feed_emoji":"📈","tokens_out":3882,"duration_ms":40849,"temperature":0.7,"pith_summary":"The paper asks what a researcher can do when the standard difference-in-differences assumption of parallel trends fails because the untreated outcome gap between treated and control groups is trending before treatment. It shows that point identification of treatment effects is still possible under a weaker hierarchy of assumptions called Parallel[p], which only requires that the p-th time difference of the untreated gap is equal across groups. Under this assumption, each cohort's untreated gap is a polynomial of degree p-1, and the post-treatment counterfactual is obtained by extrapolating that polynomial. The paper also proves that this works in staggered adoption settings, even when different cohorts are identified under different polynomial orders, and provides a data-driven procedure to select the order. This matters because it gives applied researchers a middle path between an assumption the data reject and giving up point identification in favor of bounds.","feed_headline":"Polynomial pre-trends restore point identification in staggered DiD","feed_subtitle":"Replacing 'flat gap' with 'same p-th order difference' keeps treatment effects identified and testable.","key_machinery":"The central object is the Parallel[p] assumption: E[Δ^p Y_it(∞) | G_i = g] = E[Δ^p Y_it(∞) | G_i = ∞] for all periods t, where Δ^p is the p-th time difference. This generalizes standard parallel trends (p=1) to allow the untreated gap to be a polynomial of degree p-1. The identification argument uses the finite-difference fact that a sequence whose p-th difference is zero is a polynomial of that degree, and the counterfactual is recovered by projecting that polynomial forward. The aggregation theorem handles cohort-heterogeneous orders through the observation that each cohort's ATT is the same structural parameter regardless of the identifying order.","core_discovery":"The central claim is that the group-time average treatment effect ATT(g,t) is identified whenever the p-th order difference of the never-treated potential outcome gap between cohort g and the never-treated group is zero for all periods. Lemma 4.1 shows this implies the pre-treatment gap is a polynomial of degree p-1; Proposition 4.2 shows the same polynomial identifies the counterfactual gap in every post-treatment period by extrapolation; Theorem 4.3 then writes ATT(g,t) as the observed post-treatment gap minus this polynomial counterfactual. Theorem 4.4 extends identification to weighted aggregates when each cohort is allowed to use its own feasible order, because each order identifies the","pith_inferences":["The extrapolation step is the Achilles' heel: the same polynomial that fits the pre-treatment gap is assumed to hold exactly after treatment. A natural extension would combine this polynomial projection with sensitivity bounds that allow deviations from the polynomial, indexed by a smoothness parameter, trading point identification for robustness.","The sequential order-selection procedure could be replaced or supplemented by an information-criterion approach (e.g., AIC or BIC on the pre-treatment gaps), which might yield better finite-sample coverage than sequential testing when pre-treatment series are short.","The R-squared diagnostic could serve as an informal smoothness measure: low R-squared at a given order flags non-polynomial dynamics, and a researcher could report the highest order at which R-squared remains high as a more objective choice than a single sequential test.","The aggregation theorem suggests an immediate testable extension: in applications where later-treated cohorts have longer pre-series, using cohort-specific orders should reduce variance without introducing bias, and the sensitivity of estimates to the choice of order assignment can be reported as a table."],"forward_implications":["Applied researchers whose pre-treatment event studies reject flat parallel trends can still report a point estimate, provided they are willing to assume the pre-treatment gap follows a low-degree polynomial that continues to hold post-treatment.","The pre-treatment implications of Parallel[p] are testable: the gap must lie on a degree p-1 polynomial, so over-identifying restrictions and pre-period R-squared provide diagnostics.","The sequential order-selection procedure chooses the lowest order not rejected by the data, giving a principled response to the pre-testing critique: estimation is conditioned on a selection rule but Monte Carlo evidence shows bootstrap coverage remains near nominal.","In staggered designs, cohorts with different pre-treatment lengths can be identified under different polynomial orders, and the resulting cohort-specific ATTs still aggregate into a single interpretable parameter.","The method recovers a positive and growing effect of Medicaid expansion on insurance coverage under Parallel[2], an assumption the pre-treatment data do not reject, whereas the flat-gap assumption is decisively rejected."],"fun_headline_variants":["Polynomial pre-trends identify ATT in staggered DiD","Beyond parallel trends: higher-order conditions for DiD","Staggered DiD: point ID via higher-order parallelism","Replace flat-gap with polynomial: DiD identification","Relaxing parallel trends: staggered DiD point ID"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that parallel[p] holds for every period, including all post-treatment periods, so the polynomial shape of the untreated gap fitted to pre-treatment data continues unchanged after treatment; this is untestable, and if the true untreated gap bends away from the polynomial after treatment, every DD[p] estimate is biased.","fun_headline_variants_meta":{"raw":{"variants":["Polynomial pre-trends identify ATT in staggered DiD","Beyond parallel trends: higher-order conditions for DiD","Staggered DiD: point ID via higher-order parallelism","Replace flat-gap with polynomial: DiD identification","Relaxing parallel trends: staggered DiD point ID"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001026,"raw_usage":{"total_tokens":4158,"prompt_tokens":736,"completion_tokens":3422,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":480,"completion_tokens_details":{"reasoning_tokens":3341}},"tokens_in":480,"tokens_out":3422,"duration_ms":26377,"temperature":1.0,"reasoning_tokens":3341,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T11:00:49.922710+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use the never-treated units to test the extrapolation directly: assign a placebo adoption date g*, split the never-treated units into two random groups, fit the gap between them under Parallel[p] on periods before g*, and compare the projected polynomial to the observed gap after g*. A systematic, non-polynomial deviation larger than sampling noise would falsify the order-p extrapolation and therefore undermine the identifying assumption for treated cohorts.","supporting_citations":[],"review_version":2}