{"id":"55c19268-237d-4ab0-8036-b23dc3c02b06","arxiv_id":"2505.10519","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"NURVA, a weaker replacement for SUTVA, supports within-experiment estimates but cannot support causal claims about interventions not actually assigned in the study.","lead":"This paper creates a general framework for design-based causal inference that allows for interference between units, and defines new statistical targets that remain well-defined even when the usual no-interference assumption fails. It shows that a weaker substitute for the standard assumption is enough for internal comparisons but not for generalizing results beyond the experiment, clarifying when the field's standard assumptions are truly needed.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The insufficiency of NURVA is conditional on an unformalized choice of design space Z; if Z is set to the support of the realized design, the NURVA/SUTVA distinction disappears.","rationale":"The paper's formal development is internally consistent and the examples correctly show that when Z is strictly larger than the support, NURVA does not identify AEEDs under alternative designs. The reader's weakest assumption captures this condition in terms of the inferential target; my concern is slightly broader, since it applies to the definition of SUTVA itself. Still, the paper explicitly frames its claims as applying to settings where inferential scope may go beyond the assigned interventions, so this is a boundary condition rather than an error. I would not change the accept verdict, but I would encourage the authors to state explicitly that Z is part of the estimand specification and that the NURVA/SUTVA distinction is relative to a researcher-chosen design space.","tokens_in":17651,"tokens_out":12111,"duration_ms":126606,"concrete_test":"Re-run the Table 1 household example with the design space redefined as Z = Supp(Z) = {(0,1),(1,0)}. Under Condition 5.2, SUTVA quantifies only over these two assignments; no unit has two support assignments with the same exposure, so SUTVA holds vacuously and is identical to NURVA. Compute the AEED under any design with support in this Z: it must be -1, so the contrast between NURVA and SUTVA and the nonidentification across designs vanishes. This demonstrates that the paper's central claim is contingent on the choice of Z, not on NURVA itself.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Condition 5.1 (NURVA) and Condition 5.2 (SUTVA) are identical except for the quantifier domain: Supp(Z) versus Z. The paper's central claim that NURVA is practically insufficient rests entirely on the availability of interventions in Z that lie outside Supp(Z). The framework, however, never formalizes how Z is to be chosen; Section 3.1 defines Z only as the image of a bijective map from a sample space of 'all conceivable experimental interventions,' and Section 5 notes that Z 'may include assignments outside the design.' If a researcher adopts the conventional design-based convention that the design is the actual randomization distribution and sets Z = Supp(Z), then NURVA and SUTVA coincide. The household voter-turnout example in Table 1 becomes a case where SUTVA holds vacuously on the support and the AEED is the same under every design with support in Z; the claimed failure of extrapolation in Table 2 does not arise. Thus the 'practical insufficiency of NURVA' is not an intrinsic property of the assumption but a consequence of the analyst's decision to broaden Z to include out-of-support assignments and to define inferential targets under those assignments. The paper is transparent about this condition in places, but it never provides a formal criterion (e.g., the closure of Z under the researcher's causal queries) that would make the central claim independent of an arbitrary modeling choice. This is the weakest link in the argument.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a design-based framework for causal inference and survey sampling that does not begin from SUTVA. It defines potential outcomes at the level of the full assignment vector, introduces exposure mappings, and then defines expected potential outcomes (EPO), average expected potential outcomes (AEPO), expected exposure differences (EED), and average expected exposure differences (AEED). These quantities are well-defined even under arbitrary interference. The paper contrasts a new assumption, NURVA (stability of potential outcomes on the support of the realized design), with SUTVA (stability on the entire design space), and argues that NURVA is practically insufficient for identifying substantively interesting quantities because it cannot support generalization to designs whose supports lie outside the realized design. The paper also provides estimation theory for AEPO and AEED, including Horvitz-Thompson type estimators, conservative variance estimators, covariate adjustment, and asymptotic results, with proofs in an appendix. Its central claims are illustrated with fully specified toy examples, including a household voter-turnout example and general-equilibrium job-training and campaign-ad examples.","tokens_in":1717,"tokens_out":2205,"duration_ms":100551,"significance":"If the central claim is accepted with appropriate qualifications, the paper makes a useful conceptual contribution: it isolates precisely what SUTVA adds over a design-specific stability assumption, and it offers a coherent set of estimands that remain meaningful under interference. The paper's strengths include its explicit formal definitions, the hand-checkable toy examples, the observation that SUTVA and NURVA are observationally equivalent for a given realized design, and the demonstration that HT-type estimators can be unbiased for the new targets without SUTVA. The estimation appendix is internally consistent and extends the Aronow and Samii (2017) framework. The main significance is foundational: it clarifies that SUTVA is a generalizing assumption rather than merely a regularity condition for design-based estimation.","major_comments":[{"comment":"The central claim of the paper, that NURVA is practically insufficient for identifying substantively interesting quantities, depends critically on how the design space Z is specified, and the paper never formalizes that choice. Condition 5.1 (NURVA) and Condition 5.2 (SUTVA) differ only in the quantifier domain: Supp(Z) versus Z. In Section 3.1, Z is defined only as the image of a bijective map from all conceivable experimental interventions, which is not a well-defined criterion, and Section 5 notes that Z may include assignments outside the design. If an analyst adopts the standard design-based convention that the design is the actual randomization distribution and sets Z = Supp(Z), then NURVA and SUTVA coincide, and the failure of extrapolation illustrated in Tables 1, 2, and 4 does not arise. In that case the household voter-turnout example satisfies SUTVA vacuously on the support, and the AEED is the same under every design whose support lies in Z. Thus the claimed practical insufficiency of NURVA is not an intrinsic property of the assumption; it is a consequence of the analyst's decision to broaden Z to include out-of-support interventions and to define inferential targets under those interventions. The paper is transparent about this in some passages, but the abstract and conclusion present the insufficiency as unconditional. The manuscript should state explicitly that the insufficiency claim is relative to a design space Z that is closed under the researcher's causal queries (for example, the union of supports of all designs under consideration), and should adjust the abstract and conclusion accordingly. Without such a qualification, the central claim is an artifact of an unformalized modeling choice.","section":"Sections 3.1, 5.1, 5.2, and 6"},{"comment":"The paper claims, in Section 7, that finite-sample unbiasedness can be attained for covariate-adjusted estimators even if NURVA does not hold. The proof in Appendix A.4, however, relies on the condition f_d(X_i, beta-hat) being independent of D_i, which is not stated in the main text. This is a substantive assumption: it fails when beta-hat is estimated from data that depend on the treatment or exposure assignment, which is the typical case in regression-adjusted estimation. The main-text claim as written is therefore stronger than what is proved. Please state the independence condition explicitly in Section 7, and discuss how it relates to standard practice, such as cross-fitting or using auxiliary data, so that readers do not infer that covariate adjustment is unbiased under NURVA without any additional condition.","section":"Section 7 and Appendix A.4"}],"minor_comments":[{"comment":"The sentence above Table 1 reads the AEPO under the treatment exposure is 1, y(1) = 0, and the AEPO under the control exposure is 0, y(0) = 1, which is internally inconsistent with the subsequent computation of AEED = -1. The text should say that the AEPO under treatment exposure is 0 and the AEPO under control exposure is 1.","section":"Section 5, toy example"},{"comment":"The notation for the design space Z and the random vector Z is confusing: the sentence We define the design space Z = Z(Omega), where Z is a bijective random vector Z : Omega to R^N uses the same symbol for both the design space and the vector. Consider using a different symbol for one of the two objects.","section":"Section 3.1"},{"comment":"The proof of Proposition 1 is only a sketch: it states that the result follows from plugging terms into the variance formula and using Chebyshev's inequality, but it does not show how Conditions A.1 through A.3 jointly control the variance. Since the appendix is the only place where the asymptotic claims are developed, a more detailed derivation would increase confidence in the consistency result.","section":"Appendix A.3"},{"comment":"The final sentence claims that inference on AEEDs is asymptotically equivalent to inference on the average direct effect if both statistical dependence in D and unmodeled interference are sufficiently local. This is an informal, unproven claim; either state the precise regularity conditions or flag it as a conjecture.","section":"Section 8"}],"recommendation":"major_revision","confidential_remarks":"The paper is potentially valuable for the design-based causal inference community, and the toy examples are carefully constructed and reproducible by hand. The main concern is the unformalized choice of the design space Z, which determines whether the central claim about NURVA's insufficiency holds. This concern is not a rejection of the framework; it is a request to make the paper's central claim precise and to align the abstract and conclusion with the conditionality that the body of the paper already acknowledges. The covariate-adjustment claim also needs a small but important correction. I see no evidence of circularity or missing citation issues in the manuscript itself."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, this is a conceptual paper and a reasonably good one. It formalizes a distinction that has been floating around: SUTVA is a claim about all feasible interventions, while NURVA only claims outcome stability across assignments in the support of the realized design. The new objects AEPO and AEED are natural design-dependent estimands that reduce to standard targets when SUTVA holds, and the toy examples (household voter turnout, job training, campaign ad) cleanly show why within-support stability does not buy cross-design generalizability.\n\nThe core contribution is real. The paper does not resolve an open quantitative question; it clarifies what SUTVA is doing, and that is worth something. The estimation appendix is mostly an adaptation of Aronow–Samii machinery, but the finite-sample unbiasedness of the Horvitz–Thompson estimator for the AEPO without any NURVA is stated and proved adequately, and the covariate-adjusted unbiasedness result is a nice small addition.\n\nSoft spots: the stress-test concern is fair but not fatal. The NURVA/SUTVA distinction hinges exactly on the choice of Z. If you set Z equal to the support of the design, the distinction collapses and the paper's 'practical insufficiency' argument evaporates. The paper is transparent that it is assuming a broader design space, but it never gives a formal criterion for how Z should be chosen; it just says Z 'may include assignments outside the design.' A skeptic can reasonably say the punchline is conditional on an analyst's decision. I do not think this is a hidden fatal flaw—external validity questions are precisely about interventions beyond the support—but the authors would serve readers better by stating the target design space as a primitive of the inferential question rather than a free parameter.\n\nThe asymptotic results are sketches: Proposition 1 relies on Conditions A.1–A.3 and the proof is one sentence. Fine for a foundations paper, but it would not carry a heavy applied load. The variance estimator is conservative and potentially very loose; they acknowledge this and cite sharper bounds. Citation patterns look honest—exposure mappings are credited to Aronow and Samii (2017), and the dependence on prior work is a borrowed definition, not a circular step.\n\nWho this is for: anyone working on interference, design-based inference, or external validity. It would make a good reading-group paper. I would send it to peer review; the central distinction is worth putting on record, and revisions can tighten the design-space choice and the asymptotic details.","headline":"A clear conceptual distinction between SUTVA and NURVA, with a real but conditional claim about external validity; the paper deserves a serious referee, though the design-space choice needs tightening.","tokens_in":18474,"tokens_out":2634,"would_cite":true,"duration_ms":26334,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D05","62K99"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that SUTVA's real work is not just ruling out interference within a study; it is what licenses generalizing causal claims from one design to another, and the weaker NURVA assumption cannot do that.","keywords":["design-based inference","SUTVA","NURVA","interference","exposure mapping","average expected exposure difference","survey sampling","causal inference"],"falsifier":"Ask whether NURVA alone determines the AEED under a design $Z'$ whose support strictly contains the support of the realized design $Z$. If two raw outcome schedules both satisfy NURVA under $Z$, are observationally identical under $Z$, and yet give different AEEDs under $Z'$, the paper's claim is demonstrated; a proof that such schedules cannot exist would refute it.","tokens_in":17469,"feed_emoji":"🎲","tokens_out":9953,"duration_ms":92891,"temperature":0.7,"pith_summary":"The paper tries to rebuild design-based causal inference and survey sampling without starting from Rubin's SUTVA. It allows arbitrary interference through a nonparametric model and defines new estimands, the average expected potential outcome (AEPO) and the average expected exposure difference (AEED), which reduce to conventional treatment effects when SUTVA holds. Its central move is to separate NURVA, which only requires stable potential outcomes on the support of the realized design, from SUTVA, which requires stability across all feasible interventions. The paper's negative result is that NURVA and SUTVA are observationally equivalent for a given design, but NURVA cannot support causal claims about interventions outside that design's support. A sympathetic reader should care because this pinpoints when a design licenses generalizable claims and when it only describes its own assignment mechanism.","feed_headline":"Without SUTVA, experiment effects do not leave the design","feed_subtitle":"Effects estimated under the weaker NURVA assumption need not apply to any other design.","key_machinery":"The load-bearing object is the exposure mapping $g_i:\\mathcal{Z}\\to\\mathbb{R}$, which assigns each unit an exposure for every intervention in the design space, together with the raw potential outcomes $y_i(z)$ for $z\\in\\mathcal{Z}$. NURVA (Condition 5.1) demands that when $g_i(z)=g_i(z')$ for two interventions in $\\mathrm{Supp}(Z)$, the outcomes agree; SUTVA (Condition 5.2) demands the same across all of $\\mathcal{Z}$. The machinery works by making the support of the design the boundary of what can be learned: NURVA can hold while outcomes under assignments outside the support vary freely, and that hidden variation is exactly what blocks the AEED from generalizing to other designs.","core_discovery":"On the paper's own terms, the discovery is that SUTVA, once NURVA is separated out, is a generalizing assumption rather than a mere no-interference/no-hidden-variations restriction. For any fixed design, NURVA and SUTVA are observationally equivalent, so all within-design estimation and inference proceeds as in the standard paradigm. But NURVA is design-dependent: it only requires stability of potential outcomes on the support of the realized design. Consequently, AEEDs identified under NURVA describe the contrast induced by using the realized design to place a unit in exposure $d$ rather than $d'$, not the contrast under an arbitrary intervention that assigns exposures differently. The paper's two-unit examples show the AEED can change sign when the design's support is extended to uniform treatment or uniform control even though NURVA holds under the realized design. The paper then reconstructs the standard paradigm, with SUTVA placed at the end as the assumption that carries causal conclusions beyond the realized design.","pith_inferences":["External-validity debates may be better framed around whether the target intervention lies in the support of the realized assignment mechanism than around effect heterogeneity, since NURVA cannot bridge a support gap.","A field experiment that compares AEEDs under two designs on the same population, one with a support nested in the other, could serve as a diagnostic: agreement would suggest SUTVA approximately holds, disagreement would reveal hidden variation that NURVA misses.","When generalization to a policy intervention is the goal, the design principle implied by the paper is to include the policy intervention in the design's support—for example, whole-community treatment arms—rather than relying on SUTVA to extrapolate from partial treatment."],"forward_implications":["For a fixed randomized or sampling design, replacing SUTVA with NURVA changes nothing observable: finite-sample unbiased Horvitz–Thompson estimation of AEPO/AEED, conservative variance estimation, and standard asymptotics all go through.","An AEED estimated under NURVA does not identify the corresponding contrast under a design with different support, such as treating everyone versus treating no one.","Claims that an experimental effect will persist under a scaled-up or differently targeted intervention require SUTVA or an equivalent assumption, and the realized design's data cannot certify them.","SUTVA's role shifts from a technical condition for unbiasedness to the assumption that converts design-specific exposure contrasts into treatment effects as usually understood.","A researcher can trim the population or redefine exposures to restore individual positivity, which keeps the AEPO and AEED well-defined in settings with interference."],"supporting_citations":[{"why":"Establishes the SUTVA-based causal inference framework that the paper reconstructs.","marker":"Rubin (1974)"},{"why":"Defines design-based inference in survey sampling and the finite-population, design-as-source-of-randomness view adopted throughout.","marker":"Särndal (1978)"},{"why":"Introduces exposure mappings and Horvitz–Thompson estimation under general interference, which the new estimands and estimators generalize.","marker":"Aronow and Samii (2017)"},{"why":"Defines individual average potential outcomes, the object the EPO extends to arbitrary assignment schemes.","marker":"Hudgens and Halloran (2008)"},{"why":"Provides the assignment-conditional unit-level treatment effect contrast and convergence results used to compare against the AEED.","marker":"Sävje, Aronow, and Hudgens (2021)"},{"why":"Analyzes causal estimands under misspecified exposure mappings, the setting of the paper's main examples.","marker":"Sävje (2024)"},{"why":"Supplies the notion of effective treatments that correctly specified exposure mappings are said to align with.","marker":"Manski (2013)"},{"why":"Gives the inverse-probability estimator that estimates the AEPO and AEED.","marker":"Horvitz and Thompson (1952)"}],"fun_headline_variants":["SUTVA is the generalizing assumption, not just no-interference","NURVA: inference works, but effects are design-bound","Without SUTVA, effects are confined to the design","Design-based effects need SUTVA to generalize","Why SUTVA matters: it makes effects portable"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's main negative claim applies only when the researcher wants conclusions about interventions whose possible assignments reach beyond those of the realized design; for conclusions strictly about the realized design, NURVA is enough and the claimed insufficiency does not arise.","fun_headline_variants_meta":{"raw":{"variants":["SUTVA is the generalizing assumption, not just no-interference","NURVA: inference works, but effects are design-bound","Without SUTVA, effects are confined to the design","Design-based effects need SUTVA to generalize","Why SUTVA matters: it makes effects portable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000561,"raw_usage":{"total_tokens":2640,"prompt_tokens":897,"completion_tokens":1743,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":1661}},"tokens_in":513,"tokens_out":1743,"duration_ms":16072,"temperature":1.0,"reasoning_tokens":1661,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:07:38.452963+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask whether NURVA alone determines the AEED under a design $Z'$ whose support strictly contains the support of the realized design $Z$. If two raw outcome schedules both satisfy NURVA under $Z$, are observationally identical under $Z$, and yet give different AEEDs under $Z'$, the paper's claim is demonstrated; a proof that such schedules cannot exist would refute it.","supporting_citations":[{"cited_title":"(10) 29 The Horvitz-Thompson estimator can be seen as a generalization of the sample mean","cited_arxiv_id":null,"evidence_quote":"Gives the inverse-probability estimator that estimates the AEPO and AEED."}],"review_version":1}