{"id":"21425306-943b-4504-aa58-6ec849df7fe7","arxiv_id":"2601.12120","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":1,"one_line_summary":"IV estimates of an aggregate treatment's effect are generally not the aggregate's causal effect unless components affect the outcome proportionally or the intervention is tuned to the instrument.","lead":"This paper shows that when a treatment is an aggregate of unobserved components with different effects, standard instrumental-variable estimates do not generally equal a well-defined causal effect of that aggregate. It identifies two very restrictive conditions under which the IV estimand does match the aggregate causal effect, urging caution in applied IV studies of education, GDP, caloric intake, and similar aggregates.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the paper's conditional claim is sound; the surgicality limitation is acknowledged and does not undermine the conclusion.","rationale":"The paper's central argument is a well-defined conditional claim: under the explicit definition of the aggregate causal effect via a valid ACID (surgical, aggregation-restricted, value-independent), the IV estimand matches ACE only under proportional aggregation or instrument-tuned interventions. The mathematics is correct: the IV estimand (Eq. 3) and ACE (Eq. 15) are different linear combinations of component effects, and the constraints on ACID parameters leave the difference unbounded. The reader's weakest assumption about surgicality is a real limitation, but the paper explicitly acknowledges it (Section 3.3) and argues that weakening surgicality risks confounding. This does not invalidate the main result; it only clarifies the conditions under which the result applies. I see no internal inconsistency or unsupported derivation that would change the ACCEPT verdict. The simulations, while lacking error bars, are not necessary for the theoretical conclusions. I therefore find no load-bearing concern that would justify altering the reader's verdict.","tokens_in":20660,"tokens_out":19205,"duration_ms":199664,"concrete_test":"Verify the central mismatch analytically for a simple case: set k=2, α=(1,1), β=(1,2), δ=(1,2), so β_IV = 5/3. Let d=(d1, d2) range over all valid Gaussian ACIDs with d1+d2=1. Compute ACE = 2 - d1, which varies over all real numbers as d1 varies, while β_IV remains 5/3, demonstrating the unbounded difference. This confirms the paper's generic claim and shows that the unboundedness is not an artifact of a special parameter choice.","verdict_should_be":"UNCHANGED","load_bearing_attack":"No significant objection identified. The paper's central claim—that the IV estimand for an aggregate treatment coincides with the aggregate causal effect only under proportional aggregation or an instrument-tuned valid ACID—is mathematically sound under the explicit definition of ACE in Definition 1. The reader's concern that non-surgical interventions make ACE undefined is addressed in Section 3.3, where the authors acknowledge the limitation and justify surgicality as a standard desideratum to avoid confounding. This delimits scope but does not undermine the conditional claim: whenever an ACE is defined via a valid ACID, the IV estimand generically differs by an unbounded amount unless the two special conditions hold. The unboundedness follows from the freedom to choose ACID parameters d_j (subject to Σ α_j d_j = 1) independently of instrument effects δ_j, as shown in Eq. (3) and (15). The simulations are illustrative rather than load-bearing; the theoretical derivation is sufficient. The only minor issue is the lack of error bars in the simulation plots, which does not affect the core theoretical claims.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies instrumental variable (IV) estimation when the treatment A is an aggregate of k unobserved components A_1,...,A_k with possibly heterogeneous effects on the outcome. In the linear SCM of Section 2, the standard IV estimand is derived as β_IV = (Σ_j β_j δ_j)/(Σ_j α_j δ_j). The paper introduces the Aggregate-Constrained Component Intervention Distribution (ACID) to define the aggregate causal effect (ACE) as the effect of a unit intervention on A instantiated through the components, requiring surgicality, the aggregation restriction, and value independence. Under a Gaussian ACID the ACE is Σ_j β_j d_j with Σ_j α_j d_j = 1. The central result is that β_IV equals ACE only under two special conditions: proportional aggregation (β_j/α_j constant) or an 'instrument-tuned' ACID in which the component slopes d_j align with the normalized instrument effects δ_j. Otherwise, the difference can be arbitrarily large. Section 5 shows that, in the linear Gaussian case, the aggregate setting is observationally equivalent to a violation of the exclusion restriction, and demonstrates that the Sargan test cannot generally certify proportional aggregation. The paper concludes that IV estimates of aggregate treatments require strong, usually unjustified, assumptions.","tokens_in":20930,"tokens_out":3866,"duration_ms":41434,"significance":"If accepted, this paper makes a substantive contribution to the causal-inference literature on treatments with multiple versions and aggregation. It provides a clean formalization of the aggregate causal effect, a precise algebraic characterization of when IV identifies that effect, and a striking equivalence result connecting aggregation to exclusion-restriction violation in the Gaussian case. The derivation of β_IV and ACE is transparent and reproducible from the manuscript; the identification condition in Eq. (19) is a genuine constraint rather than a tautology. The Sargan-test analysis is a useful practical warning that passing an overidentification test does not establish proportional aggregation. The main message—that IV analyses with aggregate treatments need either proportional aggregation or an explicit, defensible ACID—is likely to influence applied practice. The principal limitation, acknowledged by the authors, is that the ACE is defined only under a valid ACID, and the surgicality requirement may not match real interventions on aggregates; however, this caveat is stated explicitly and does not undermine the conditional theoretical claim.","major_comments":[],"minor_comments":[{"comment":"The simulation results appear to be based on a single dataset per configuration. Reporting the number of replications and adding confidence intervals or standard-error bands would make the convergence claims visually and statistically supported. The same applies to Figure 4, where 100 datasets are used but no uncertainty measures are shown.","section":"§4.3, Figure 2"},{"comment":"The phrase 'the difference ... is unbounded' could be misread as a fixed-model property. What is established is that, across the class of models satisfying the aggregate setting but not the two sufficient conditions, the difference can be made arbitrarily large. Rewording to 'can be made arbitrarily large' or adding a formal statement would prevent overstatement.","section":"§5, first paragraph"},{"comment":"The instrument-tuned solution d_j = δ_j/(Σ α_j δ_j) is one particular solution to Eq. (19), but the text says 'the most straightforward to see.' It may be worth noting explicitly that other solutions exist and are equally contrived, so that the argument does not depend on this specific choice.","section":"§4.2, Eq. (20)"},{"comment":"The recommendation to use a large type-I error (0.5) for the Sargan test as a diagnostic tool is based on simulation evidence without formal justification. Adding a sentence that this is a heuristic diagnostic, not a significance test with calibrated error control, would improve the presentation.","section":"§5.1, Figure 4"},{"comment":"There are a few typographical and formatting issues: 'highights' in the introduction, 'Hernan' without the accent in the reference list is inconsistent with standard spelling, and some equations have slightly awkward inline notation. These do not affect the technical content.","section":"Throughout"}],"recommendation":"minor_revision","confidential_remarks":"I found no load-bearing technical error in the central derivation. The ACID assumptions are strong and contestable, but the authors are explicit about their scope and the conditional nature of the main theorem. The paper is a good fit for a statistical methods journal. With the addition of error bars and a few clarifications, I would be happy to see it published."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper deserves a serious referee. The core contribution is the explicit formalization of the aggregate-constrained component intervention distribution (ACID) and the derivation of two precise conditions under which the standard linear IV estimand equals the aggregate causal effect: proportional aggregation and instrument-tuned interventions. The negative result—that outside these conditions the difference between the IV estimand and the aggregate causal effect is unbounded—is what applied people will remember. It is correct, and it is genuinely new in this form. The existing literature on treatment variation invariance and aggregation bias does not make the IV identification conditions explicit, and Proposition 1 cleanly shows that aggregation is observationally equivalent to a violation of the exclusion restriction. That is a useful bridge to a familiar diagnostic.\n\nThe math in Sections 2–4 is straightforward linear algebra and checks out. The covariance matching in Appendix D is careful. The paper is also honest about what it does not do: the definition of the aggregate causal effect depends on a researcher-chosen ACID, and the surgicality assumption is strong. Both limitations are acknowledged in Section 3.3, and they delimit the scope rather than undermine the conditional claim. The simulations are illustrative rather than load-bearing; the lack of error bars is a minor presentation issue, not a substantive flaw. I would not hold that against the paper.\n\nThe Sargan test simulation is a nice practical addition. It confirms that with weak instruments, passing Sargan gives almost no support for proportional aggregation—useful caution for empirical work.\n\nWho should read this: applied IV users in economics and epidemiology, and methodologists working on treatment effect heterogeneity and multiple versions of treatment. The writing is clear enough that an applied reader can follow the main argument without grinding through every derivation.\n\nMy recommendation: send it to peer review. The theory is sound, the claims are scoped, and the practical implications are important. Minor revision would be appropriate—mostly to make the simulation presentation a bit more standard—but the core should be accepted.\n\nIn short: a solid paper, not a flashy one. Worth engaging with.","headline":"A clean, honest formalization of when IV estimands match aggregate causal effects, with the central conditional claim holding and limitations explicitly scoped.","tokens_in":21403,"tokens_out":1301,"would_cite":true,"duration_ms":39497,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D20","62P20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Instrumental-variable estimates of a treatment that is an aggregate of components with heterogeneous effects generally cannot be interpreted as the causal effect of the treatment; only proportional aggregation or instrument-tuned interventi","keywords":["instrumental variables","aggregate treatment","causal effect ambiguity","exclusion restriction","structural causal models","two-stage least squares","treatment variation invariance","ACID"],"falsifier":"Generate data from a linear Gaussian SCM with two components, α1=α2=1, β1=1, β2=2, δ1=δ2=1, and a confounder U affecting both components. Define the aggregate causal effect using the natural Gaussian ACID that sets components to their observational conditional distribution given A=a. The population IV estimand is (1+2)/(1+1)=1.5; the ACE computed from the conditional mean of components given A is (β1 cov(A1,A)+β2 cov(A2,A))/var(A), which depends on the confounder and instrument strengths and is generally not 1.5. A reader can pick parameter values and compute both quantities to see the mismatc","tokens_in":20557,"feed_emoji":"⚠️","tokens_out":4452,"duration_ms":44787,"temperature":0.7,"pith_summary":"This paper argues that when a treatment like education, GDP, or caloric intake is a sum of finer-grained components with different causal effects, the standard IV estimand is not, in general, the causal effect of the aggregate. The authors formalize the aggregate causal effect through a distribution over how the intervention is realized at the component level (the ACID) and prove that the IV estimand matches this effect only if either the components have proportional effects on the outcome or the intervention distribution is deliberately tuned to the instrument's component-wise effects. Outside these two scenarios, the gap between the IV estimate and the aggregate causal effect can be arbitrarily large. Because the aggregate setting is distributionally equivalent to a violation of the exclusion restriction in linear Gaussian models, standard diagnostics cannot detect the problem. The paper concludes that IV analyses with aggregate treatments require explicit, defensible justification of one of these two conditions before estimates can guide policy.","feed_headline":"IV estimates betray causal effects of aggregate treatments","feed_subtitle":"Only proportional aggregation or instrument-tuned interventions make the IV estimand match the aggregate causal effect","key_machinery":"The aggregate-constrained component intervention distribution (ACID) is the distribution over the unobserved components given an intervention on the aggregate. A 'valid' ACID must be surgical (each component set independently of its causes), respect the aggregation rule, and yield value-independent effects. The ACID carries the argument because the aggregate causal effect is defined as an expectation under it, and matching the IV estimand to that effect reduces to a constraint on the ACID's slope parameters d_j.","core_discovery":"The paper formalizes a linear structural causal model in which the treatment A is a weighted sum of unobserved components A_j, each with its own effect β_j on Y and its own response δ_j to the instrument I. The standard IV estimand is β_IV = (Σ_j β_j δ_j) / (Σ_j α_j δ_j). Defining the aggregate causal effect ACE(A,Y) as a unit intervention effect under a 'valid ACID', the paper shows ACE = Σ_j β_j d_j for Gaussian ACIDs, and that β_IV equals ACE only if the components satisfy proportional aggregation (β_j/α_j constant for all j) or the ACID's slopes d_j are aligned with the instrument effects δ_j (an 'instrument-tuned intervention'). Otherwise the difference between β_IV and ACE is unbounded","pith_inferences":["An editorial extension: the failure described is an identifiability failure at the population level, not a finite-sample artifact, so collecting more data or adding stronger instruments will not close the gap.","A practical diagnostic follows: if component-level data are available, one can directly estimate β_j and δ_j to check whether β_j/α_j are equal; if they are not, the IV estimate should be reported together with an explicit ACID assumption.","The equivalence to an exclusion violation suggests that sensitivity analyses for direct instrument effects could double as aggregation-robustness checks, provided the plausible magnitude of component heterogeneity can be bounded.","For Mendelian randomization, where instruments are genetic variants and treatments like BMI are aggregates, the paper implies that gene-exposure effects on components must be homogeneous or the ACID must be variant-specific—an unlikely condition across multiple genetic instruments."],"forward_implications":["In applied IV studies with aggregate treatments—education, GDP, caloric intake, BMI—the reported coefficient cannot be read as the causal effect of the treatment unless proportional aggregation is assumed or an instrument-tuned ACID is justified.","Under proportional aggregation, the IV estimand equals the aggregate effect regardless of how the intervention is instantiated; this is the only ACID-independent scenario.","When aggregation is not proportional, the linear Gaussian aggregate setting is distributionally equivalent to a violation of the exclusion restriction, so standard tests of instrument validity cannot detect the problem.","The Sargan test can offer evidence for proportional aggregation only when instruments are strong and the type-I error is set high; a non-rejection does not imply proportional aggregation.","Non-linear models do not rescue the interpretation: they lose value independence and make the aggregate causal effect still more sensitive to the ACID."],"fun_headline_variants":["Lost in aggregation: IV estimates miss the causal effect","Aggregate treatments make IV estimates ambiguous","IV estimates for aggregates need special conditions to be causal","IV causal effect for aggregates is not unique"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The aggregate causal effect is defined only under a 'valid ACID' whose surgicality requires that an intervention on the aggregate acts as a clean, independent intervention on each component, severing their links to the instrument and confounder; if real policy interventions on the aggregate do not have this surgical character, the aggregate causal effect is not well-defined and the paper's contrast does not apply.","fun_headline_variants_meta":{"raw":{"variants":["Lost in aggregation: IV estimates miss the causal effect","Aggregate treatments make IV estimates ambiguous","IV estimates for aggregates need special conditions to be causal","IV causal effect for aggregates is not unique"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001729,"raw_usage":{"total_tokens":6680,"prompt_tokens":761,"completion_tokens":5919,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":5870}},"tokens_in":505,"tokens_out":5919,"duration_ms":42816,"temperature":1.0,"reasoning_tokens":5870,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T06:19:01.656924+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate data from a linear Gaussian SCM with two components, α1=α2=1, β1=1, β2=2, δ1=δ2=1, and a confounder U affecting both components. Define the aggregate causal effect using the natural Gaussian ACID that sets components to their observational conditional distribution given A=a. The population IV estimand is (1+2)/(1+1)=1.5; the ACE computed from the conditional mean of components given A is (β1 cov(A1,A)+β2 cov(A2,A))/var(A), which depends on the confounder and instrument strengths and is generally not 1.5. A reader can pick parameter values and compute both quantities to see the mismatc","supporting_citations":[],"review_version":1}