{"id":"8f6867e3-8081-4be9-996e-41f6240555b5","arxiv_id":"2505.01324","paper_version":8,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Design-based estimators can target mechanism-level causal estimands, averages over latent stochastic environments, when local dependence makes cross-sectional averaging ergodic.","lead":"Under sparse local dependence, a single randomized experiment can yield estimates that are consistent for causal effects defined as averages over a latent random environment, not just for one realized world. The paper provides conditions for consistency, normality, and variance estimation, but the variance estimator only works when individual effects are known, such as under a sharp null hypothesis.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Variance-consistency claim depends on knowing every unit-level effect θ_i(ỹ_i); Section 4 supplies ζ_i only under sharp nulls, so \"consistently estimable from a single experiment\" is not operational for τ_n.","rationale":"The reader's CONDITIONAL verdict is appropriately cautious, and my independent read converges on the same load-bearing premise. The paper's central differentiator, presented in the abstract and Sections 1 and 4, is that local dependence makes the sampling variance consistently estimable from a single realised experiment, in contrast to fixed-potential-outcome variance bounds. That claim fails to be operational for the aggregate estimand τ_n because the variance estimators (4.4)–(4.9) require the centred quantities ζ_i=θ̂_i−θ_i(ỹ_i), and θ_i(ỹ_i) is the true unit-level treatment effect, not a known design parameter. The manuscript itself confines ζ_i to \"sharp null hypotheses under which θ_i(˜yi) is specified\"; outside sharp nulls no data-based construction is offered, and the simulations use oracle centring. Consequently Corollaries 1–3 establish asymptotic normality of a statistic that an analyst cannot compute for a confidence interval for τ_n. This is a limitation of scope rather than a mathematical error in the oracle-level theorems: the consistency and asymptotic-normality results for point estimation are standard dependency-graph arguments and appear sound. But the headline claim as written overstates what is proved. The reader's conditions—revise the variance-estimation claim to state the sharp-null requirement, fix the simulation centring, and clarify the non-identifiability claim—are exactly the right remedial steps. I therefore do not move the verdict; I would keep it CONDITIONAL pending those revisions.","tokens_in":29167,"tokens_out":10304,"duration_ms":106567,"concrete_test":"Modify the public replication code for Figure 1 so that the variance routine is forbidden from using the true β_i. Construct ζ_i with a feasible centring, e.g. ζ_i=Y_iψ_i−bar{Y}_{z_i} (the treatment-arm mean in the Horvitz–Thompson representation (5.6)), and compare the resulting variance estimate with the Monte Carlo variance of τ̂_n over the same 2000 replications for n∈{100,200,500,1000} and d∈{0,0.1,0.2,0.25,0.3}. If the ratio of estimated variance to Monte Carlo variance does not approach 1 as n grows, then the consistency asserted in Theorem 5 depends on oracle knowledge of θ_i(ỹ_i), contradicting the abstract's claim of consistent variance estimation from a single experiment. As a complement, check whether Section 4 or Appendix C defines any estimator of θ_i(ỹ_i); if none is defined, ζ_i is not a statistic.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weakness is in the variance-estimation half of the central claim. Equation (4.1) defines ζ_i(z,ω)=θ̂_i(z,ω)−θ_i(ỹ_i), and the variance estimators (4.4)–(4.9) sum observed products ζ_iζ_j. These products are computable only when the true unit-level treatment effect θ_i(ỹ_i) is known. Section 4 states this restriction explicitly: \"We allow for sharp null hypotheses under which θ_i(˜yi) is specified, and under such nulls, the centred quantity ζ_i ... is available for variance estimation.\" No construction is given for non-sharp-null settings, and none can be: a single experiment observes each unit under one treatment at one ω, while θ_i(ỹ_i)=E[ỹ_i(z_1)−ỹ_i(z_0)] averages over counterfactual treatments and latent environments. The theorems therefore establish consistency of an oracle variance estimator, not feasible variance estimation for τ_n. The abstract and Section 1 claim the latter as the central advance over FPO. The simulations confirm the gap: Section 5 reports that in the size experiments \"the centring term is taken to be the true unit-level contrast θ_i(˜yi)\" and in the power experiments θ_i(˜yi)≡0 is imposed; neither is available when estimating an unknown average effect. An analyst who wants a confidence interval for τ_n cannot compute ζ_i. This invalidates Corollaries 1–3 as feasible inference statements for the aggregate estimand. A secondary related concern is that E_n in (4.3) is assumed known or conservatively specified (Assumptions 6 and 11), which is also not constructed from data. The point-estimation results may stand, but the variance-consistency claim, the paper's stated differentiator, is not operational.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an extension of design-based causal inference in which potential outcomes are random functions of a treatment assignment vector and a latent stochastic environment. The target estimand is a weighted average of unit-level treatment effect functionals θ_i(ỹ_i) represented through Riesz representers. Under dependency-graph sparsity conditions, the aggregate Riesz estimator is shown to be mean-square consistent (Theorem 3) and asymptotically normal (Theorem 4) using an external dependency-graph CLT. Section 4 introduces local-dependence variance estimators (4.4)–(4.9) and claims consistency (Theorems 5–7) and asymptotic normality of the studentized statistic (Corollaries 1–3). The simulation section studies size and power for baseline and network-interference designs.","tokens_in":29446,"tokens_out":6116,"duration_ms":64316,"significance":"The point-estimation component is largely sound: Theorem 3 follows from a valid variance bound (Proposition 4) and Theorem 4 is a careful adaptation of Ross (2011) to the RPO setting. The Riesz formulation provides a clean unification with Harshaw et al. (2022), and the conceptual distinction between estimates for a realized potential-outcome schedule and estimates for an ensemble-level target is valuable. The paper also states its assumptions clearly and ships reproducible simulation code. However, the central claimed advance—consistent variance estimation from a single realized experiment—is not operational as stated, because the proposed variance estimators require the true unit-level treatment effect θ_i(ỹ_i). The paper itself restricts availability of this term to sharp null hypotheses and the simulations implement exactly that oracle or null centring. If the variance claim is respecified as an oracle or sharp-null variance result, or if a feasible centring device is supplied, the theoretical core is publishable; in its current form the variance half of the central claim is not supported.","major_comments":[{"comment":"The variance estimators (4.4)–(4.9) are built from ζ_i = θ̂_i(z,ω) − θ_i(ỹ_i). A single experiment observes each unit under one treatment assignment and one latent environment, so the true unit-level contrast θ_i(ỹ_i) = E[˜y_i(z_1,ω) − ˜y_i(z_0,ω)] is not identified from the observed data. The manuscript explicitly states that ζ_i is available only under sharp null hypotheses ('We allow for sharp null hypotheses under which θ_i(˜yi) is specified'), and Section 5 confirms this: the size experiments use the true unit-level contrast and the power experiments impose θ_i(˜yi) ≡ 0. Consequently, Theorems 5–7 and Corollaries 1–3 establish consistency of an oracle or sharp-null variance estimator, not feasible variance estimation for the aggregate estimand τ_n promised in the abstract and Section 1.","section":"Section 4, Eq. (4.1); Corollaries 1–3"},{"comment":"The claim that the proposed variance estimators 'apply directly to the FPO framework as a special case' and are 'easier to implement' is not correct: Eq. (4.1) still requires θ_i(ỹ_i), which under FPO with binary treatment is the unobserved unit-level effect Y_i(1) − Y_i(0). Without sharp-null or oracle centring, the residual ζ_i is not computable, so the proposed estimators do not provide a practical alternative to Neyman-type conservative bounds in the classical fixed-outcome setting.","section":"Section 4, paragraph after Theorem 6"},{"comment":"The entire variance-estimation procedure assumes that the dependency graph E_n and the maximum neighbourhood size D_n are known or conservatively specified. Under RPO, the latent environment ω determines both the outcomes and possibly the interference structure, and the paper gives no data-based construction of E_n from a single realized experiment. Assumption 11 requires E_n^c ⊂ ˜E_n ⊂ E_n, which is a substantive structural condition that cannot be verified from the observed sample. The 'single experiment' feasibility claim therefore rests on unobserved structural knowledge, which should be stated as a limitation rather than as a demonstrated advance over FPO.","section":"Assumptions 6 and 11, Eq. (4.3)"},{"comment":"The paper repeatedly claims to delineate an 'identification boundary' for expectation-based estimands from a single experiment (e.g., 'we characterise local dependence structures', 'delineating an identification boundary'). However, no impossibility theorem or lower-bound result is proved; the paper provides only sufficient conditions for consistency. The statement that such estimands are 'generically impossible' absent dependence is asserted rather than established. Either a formal converse or a more modest claim of sufficient conditions would make the contribution precise.","section":"Sections 1 and 6"}],"minor_comments":[{"comment":"The notation 'Assumptions 8(r=3), 8(r=4)' is unusual and should be introduced clearly in the assumption statement; as written, the reader must infer that Assumption 8 is instantiated at two different values of r.","section":"Theorem 4"},{"comment":"In Eq. (B.6), the exponents p and q appear before they are defined; the statement should open with 'for any p,q ∈ [1,∞] satisfying 1/p + 1/q = 1/2' before displaying the inequality.","section":"Proposition 4"},{"comment":"The power experiments impose θ_i(˜yi) ≡ 0, which is a special sharp null rather than a general alternative for testing τ_n = 0; the text should explain what alternative is actually being detected, since under this centring the test statistics depend only on the difference between the null and true contrasts.","section":"Section 5"},{"comment":"The captions for Figures 1 and 2 are present but the plots themselves did not render in the manuscript text reviewed; please confirm the figures are included in the compiled version.","section":"Figures 1–2"}],"recommendation":"major_revision","confidential_remarks":"The paper's abstract and introduction overstate the variance-estimation contribution. If the author cannot provide a feasible centring device for θ_i(ỹ_i) or explicitly resubmit the work as a sharp-null/oracle variance result, the manuscript should not be accepted in its current form. I see no evidence of misconduct; the issue is scope and operationality of the main claim. The point-consistency and asymptotic-normality components are worth preserving."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know about arXiv:2505.01324: it has a genuinely new core idea — replacing fixed potential outcome schedules with a latent stochastic environment and showing that under sparse local dependence, a single experiment's cross-sectional average consistently estimates the expectation-based estimand. That is a real conceptual step beyond Harshaw et al. (2022). The Riesz-representation machinery is used cleanly, and the consistency and asymptotic normality results look like valid adaptations of Ross (2011) dependency-graph bounds. The unbiasedness is by construction, not by fitting, so the point-estimation half of the paper holds up.\n\nThe soft spot is exactly where the stress-test puts it. The abstract claims 'the variance becomes consistently estimable from a single experiment.' That is only true for an oracle. The variance estimators in (4.4)–(4.9) are built from ζ_i = θ̂_i − θ_i(ỹ_i), and θ_i(ỹ_i) is the true unit-level treatment effect, which is not known when the target is an average effect. Section 4 says so itself: those quantities are available under sharp null hypotheses. So the feasible variance claim reduces to testing a sharp null, not to confidence intervals for τ_n. The simulations quietly confirm this: the centring term is either the true contrast or zero. An analyst with a confidence interval target cannot compute ζ_i. That makes Corollaries 1–3 claims about feasible inference for an average effect. This is a load-bearing gap, because the paper's stated differentiator is the variance result.\n\nA secondary gap: the dependency graph E_n and its sparsity D_n are assumed known or conservatively specified (Assumptions 6 and 11), with no data-based construction. That is standard in this literature but worth flagging. The non-identifiability claim in Section 1 is stated, not proved; it deserves a precise statement or a reference. The simulation section also has a suspicious formula for the FPO estimand in (5.4) that looks dimensionally off if the treatment contrast is not from all-ones versus all-zeros.\n\nAll that said, this is a serious paper. The RPO formalism and the ergodic bridge are worth engaging with, and the point-estimation theory is sound. It deserves peer review, but the revision needs to reframe the variance claim as sharp-null-only (or oracle), state the non-identifiability result carefully, and clean up the simulation formulas.","headline":"Sound point-estimation theory for random potential outcomes, but the variance-consistency headline only works under sharp nulls; needs revision before publication.","tokens_in":30058,"tokens_out":3180,"would_cite":true,"duration_ms":29079,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62K10","62E20","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that under suitably sparse local dependence, cross-sectional averages in a single randomised experiment consistently estimate expectation-based causal estimands defined over a latent stochastic environment, and that the…","keywords":["random potential outcomes","design-based inference","local dependence","dependency graph","Riesz representation","variance estimation","asymptotic normality","causal estimands"],"falsifier":"Use the baseline design in Section 5 (Bernoulli assignment, outcome (5.3)) with block sizes $m_i = \\lfloor n^{0.2}\\rfloor$ and $m_i = \\lfloor n^{0.3}\\rfloor$, and compute coverage of 95% confidence intervals under the sharp null where $\\theta_i(\\tilde y_i)=\\beta_i$ is known. The paper predicts near-nominal coverage for $d=0.2$ and mild conservativeness for $d=0.3$; if coverage drops well below 0.95 for $d=0.2$ as $n$ grows, the consistency claim fails, whereas if coverage stays near nominal for $d=0.3$ uniformly, the stated $D_n=o(n^{1/4})$ threshold may be too strict.","tokens_in":28841,"feed_emoji":"🧪","tokens_out":12294,"duration_ms":108511,"temperature":0.7,"pith_summary":"The paper extends design-based causal inference from fixed potential outcomes (FPO) to random potential outcomes (RPO), where each unit's outcome is a function of treatment assignment and a latent stochastic environment. Its central claim is that under local dependence—sparse dependence among units encoded in a dependency graph—averaging across units within a single realised experiment behaves like averaging over repeated draws of the environment. Consequently, aggregate design-based estimators are consistent for the mechanism-level estimand and asymptotically normal, and the sampling variance can be consistently estimated from that one experiment. This matters because it offers a design-based route to ensemble-level causal targets (expected effects at scale, stochastic spillovers, measurement-noise-contaminated outcomes) without imposing outcome models, and it removes the structural conservativeness of Neyman-type variance bounds.","feed_headline":"One experiment can identify mechanism-level causal effects","feed_subtitle":"Under sparse local dependence, one experiment's data yields consistent variance estimates, not conservative bounds.","key_machinery":"The central object is the Riesz representer $\\psi_i$ of the unit-level treatment-effect functional $\\theta_i$: a stochastic element of the model space that converts an expectation-based causal estimand into an inner product $\\theta_i(\\tilde y_i)=\\mathbb{E}[\\tilde y_i(z,\\omega)\\psi_i(z,\\omega)]$. The machinery pairs this representer with a dependency graph capturing which outcome–representer pairs $(\\tilde y_i,\\psi_i)$ are independent of which others. Sparse growth of the dependency neighbourhoods ($D_n = o(n^d)$, $d>0$) gives cross-sectional averaging an ergodic property, so averaging over units replaces averaging over repeated experiments; the variance estimator then targets only those cross-terms indexed by the known dependency set $E_n$ or a conservative superset, which is what makes consistent uncertainty quantification possible from a single realisation.","core_discovery":"Working in a Hilbert space of outcome functions, the paper represents each unit's treatment effect $\\theta_i(\\tilde y_i)$ as an inner product $\\langle \\tilde y_i, \\psi_i\\rangle$ with a Riesz representer $\\psi_i$, making the Horvitz–Thompson-type estimator $\\hat\\tau_n = \\sum_i \\nu_{ni}\\tilde y_i(z,\\omega)\\psi_i(z,\\omega)$ unbiased for the weighted average $\\tau_n = \\sum_i \\nu_{ni}\\theta_i(\\tilde y_i)$. The substantive discovery is that local dependence supplies an ergodic bridge: if each outcome–representer pair $(\\tilde y_i,\\psi_i)$ is independent of units outside a small neighbourhood, then one realised experiment contains enough information for consistent estimation of, and inference on, the ensemble-level estimand. Theorems 3 and 4 give mean-square consistency with $O_p(n^{-1/2})$ rates and asymptotic normality under $D_n = o(n^{1/4})$, where $D_n$ is the maximum dependency neighbourhood size. Theorems 5–7 show that summing second-moment terms only over known dependent (or correlated) pairs yields a variance estimator consistent for the true sampling variance, in contrast to FPO where only conservative upper bounds are structurally available.","pith_inferences":["A practical consequence the paper does not develop: outside sharp-null specifications, the variance estimators in Section 4 are not directly implementable because $\\zeta_i = \\hat\\theta_i - \\theta_i(\\tilde y_i)$ requires knowing each unit-level effect; a workable extension would be a sensitivity analysis over plausible $\\theta_i$ values, but no such procedure is supplied here.","The dependency graph and its sparsity $D_n$ are taken as given; a data-driven procedure that learns $E_n$ from observed residual products (e.g., sparse covariance estimation) would make the framework operational when neighbourhood structure is unknown, but that extension is not in the paper.","The ergodic principle is not tied to the Riesz form, so it plausibly transfers to other design-based statistics—staggered-adoption difference-in-differences, cluster-randomised designs with sparse cross-cluster dependence—whenever the same local-dependence and moment conditions hold.","The simulations show only mild over-coverage at $d=0.3$, outside the proven $n^{1/4}$ threshold; this suggests the rate condition may be sufficient but not necessary, and pinning down the sharp threshold is a natural follow-up."],"forward_implications":["Horvitz–Thompson and other familiar design-based statistics now carry a mechanism-level interpretation: the same number computed in a single experiment estimates $\\tau_n$ over the latent environment, not just a fixed-schedule finite-population effect.","Under sparse local dependence, one can report confidence intervals with valid asymptotic coverage using a single experiment, something the classical FPO framework cannot deliver through Neyman-type bounds.","The $D_n=o(n^{1/4})$ condition and bounded fourth moments trace an identification boundary: without local dependence, expectation-based causal estimands are not recoverable from one realised experiment by design-based logic.","The variance estimators reduce to a diagonal form when units are independent, recovering standard i.i.d.-style inference as a special case.","The proposed variance estimators apply to the classical FPO setting as well, giving consistent variance estimates whenever the dependency structure is known."],"supporting_citations":[{"why":"Supplies the FPO Riesz-representation framework and the dependency-neighbourhood/variance-bound structure that the paper lifts to random potential outcomes.","marker":"Harshaw et al. (2022)"},{"why":"Provides the dependency-graph normal-approximation bound used in the proof of Theorem 4.","marker":"Ross (2011)"},{"why":"The aggregate estimator takes the Horvitz–Thompson form; in the binary case the Riesz representer is the HT weight.","marker":"Horvitz and Thompson (1952)"},{"why":"Defines the fixed-potential-outcome finite-population inference whose variance bounds the paper contrasts with its consistent estimators.","marker":"Neyman (1990)"},{"why":"Introduces the potential-outcome framework that the paper extends by making outcomes random.","marker":"Rubin (1974)"},{"why":"Representative interference-based variance estimation with known exposure mappings and positive joint inclusion probabilities, which the proposed dependency-graph estimators generalise.","marker":"Aronow and Samii (2017)"}],"fun_headline_variants":["One experiment, consistent causal variance","Sparse dependence empowers single-experiment inference","Ergodic bridge: one trial recovers mechanism averages","Local structure turns one experiment into many","Beyond conservative bounds: consistent variance from one trial"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The uncertainty-quantification claims require that the analyst knows the true unit-level treatment effect $\\theta_i(\\tilde y_i)$ (or specifies a sharp null that fixes it), because the variance estimators are centred on those values; without such knowledge, the variance estimates are not feasible.","fun_headline_variants_meta":{"raw":{"variants":["One experiment, consistent causal variance","Sparse dependence empowers single-experiment inference","Ergodic bridge: one trial recovers mechanism averages","Local structure turns one experiment into many","Beyond conservative bounds: consistent variance from one trial"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000247,"raw_usage":{"total_tokens":1533,"prompt_tokens":922,"completion_tokens":611,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":538,"completion_tokens_details":{"reasoning_tokens":544}},"tokens_in":538,"tokens_out":611,"duration_ms":6826,"temperature":1.0,"reasoning_tokens":544,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:22:33.520837+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use the baseline design in Section 5 (Bernoulli assignment, outcome (5.3)) with block sizes $m_i = \\lfloor n^{0.2}\\rfloor$ and $m_i = \\lfloor n^{0.3}\\rfloor$, and compute coverage of 95% confidence intervals under the sharp null where $\\theta_i(\\tilde y_i)=\\beta_i$ is known. The paper predicts near-nominal coverage for $d=0.2$ and mild conservativeness for $d=0.3$; if coverage drops well below 0.95 for $d=0.2$ as $n$ grows, the consistency claim fails, whereas if coverage stays near nominal for $d=0.3$ uniformly, the stated $D_n=o(n^{1/4})$ threshold may be too strict.","supporting_citations":[{"cited_title":"Wang, and F","cited_arxiv_id":null,"evidence_quote":"Supplies the FPO Riesz-representation framework and the dependency-neighbourhood/variance-bound structure that the paper lifts to random potential outcomes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the dependency-graph normal-approximation bound used in the proof of Theorem 4."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the fixed-potential-outcome finite-population inference whose variance bounds the paper contrasts with its consistent estimators."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Representative interference-based variance estimation with known exposure mappings and positive joint inclusion probabilities, which the proposed dependency-graph estimators generalise."}],"review_version":1}