{"id":"10fa68f1-79b2-4d78-a0b9-fb27aeb4cd4a","arxiv_id":"2505.20780","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":1,"one_line_summary":"The paper proposes unbiased dyadic-data estimators of the global average treatment effect under dyadic interference, with convergence rates, a central limit theorem, and conservative variance estimators.","lead":"This paper develops a causal inference framework for randomized experiments in which outcomes are pairwise interactions between users, such as messages or calls. It shows that using dyadic data can correct bias in standard unit-level estimators under network interference, and provides estimators, variance estimators, and asymptotic theory.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4's complete-randomization CLT is not established: the Hajek-coupling step in B.4 gives a variance-ratio bound that Assumption 4 does not force to vanish; for regular graphs with d=n^alpha, 0<alpha<1/3, the ratio is O(1) rather than o(1).","rationale":"The paper's central claim is that the dyadic Horvitz-Thompson and Hajek estimators are unbiased, consistent, and asymptotically normal under Assumptions 1-4. The reader identified Assumption 1 as the weakest assumption; that is a fair limitation and it is explicitly acknowledged in Section 7. My stress-test found a more specific internal issue: the proof of the complete-randomization CLT (Theorem 4) in Supplementary B.4 does not close. After constructing the Hajek coupling, the proof requires var(hat_tau_HT - tilde_tau_HT)=o(var(hat_tau_HT)), but the bound (S.16) is O(max_z d1(z)^2 n^{-1}). Under Assumption 4, var(hat_tau_HT) is only forced to be much larger than max_z d_infty(z)^2 n^{-4/3}; since d1 <= d_infty, the ratio of these two quantities is not o(1) in general. In a regular graph with degree d = n^alpha and 0<alpha<1/3, Assumption 4 holds at the natural variance scale d^2/n, and the coupling bound is also O(d^2/n), so the asserted o(1) does not follow. This does not prove the theorem false, but it means the central CLT claim for complete randomization rests on an unproved step. The paper otherwise gives careful design-based derivations and helpful simulations, and the Bernoulli-randomization CLT appears to have a valid Stein-method argument; hence the reader's conditional acceptance remains appropriate, with the additional condition that the complete-randomization argument in Theorem 4 be replaced or supplemented.","tokens_in":35380,"tokens_out":17498,"duration_ms":183087,"concrete_test":"Analytical check: for a regular directed graph with degree d = n^alpha, 0<alpha<1/3, and all nonzero dyadic outcomes normalized to 1, compute exact or sharp upper bounds for var(hat_tau_HT) under complete randomization and for the variance of the Hajek-coupled difference defined in B.4. If the ratio var(hat_tau_HT - tilde_tau_HT)/var(hat_tau_HT) does not tend to 0 while Assumption 4 holds, the proof of Theorem 4 for complete randomization is invalid as written; then the theorem needs a different proof, or Assumption 4 must be strengthened to something like max_z d_infty(z) n^{-1/2} var(hat_tau_HT)^{-1/2} = o(1).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most concrete weak point is the complete-randomization half of Theorem 4. In Supplementary B.4, the proof replaces the complete-randomization estimator by a Hajek-coupled Bernoulli estimator and asserts it suffices to show var(hat_tau_HT - tilde_tau_HT)/var(hat_tau_HT)=o(1). The bound actually derived in (S.16) is E var(...) = O(max_z d1(z)^2 n^{-1}) up to lower-order terms. Assumption 4 only implies var(hat_tau_HT) >> max_z d_infty(z)^2 n^{-4/3}. For a regular graph with degree d = n^alpha and 0<alpha<1/3, d1 = d_infty = d, the natural variance scale is d^2/n, Assumption 4 holds at that scale, and the coupling bound is also O(d^2/n); the displayed inequalities therefore do not imply the required o(1) ratio. This is not a claim that the theorem is false, but as written the central CLT result for complete randomization depends on an unproved step. The Bernoulli-randomization CLT and the consistency results appear to have valid arguments; the gap is specific to the complete-randomization coupling proof.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a design-based causal inference framework for randomized experiments with dyadic (pairwise) outcomes. Under a dyadic-interference assumption (Assumption 1), it defines the global average treatment effect, constructs Horvitz-Thompson and Hajek estimators based on dyadic outcomes, and derives convergence rates under Bernoulli, complete, and cluster randomization. It further develops conservative variance estimators and Wald-type confidence intervals, and reports simulation studies and a WeChat field experiment. The main theorems are stated for both Bernoulli and complete randomization, with complete-randomization results proved via Hajek coupling in the supplementary material.","tokens_in":35675,"tokens_out":5615,"duration_ms":52825,"significance":"If the results hold, the paper is a useful contribution to causal inference under network interference: it provides a clean formulation of dyadic interference, shows the bias of naive unit-level estimators, gives explicit convergence rates in terms of network degree statistics, and provides a conservative variance estimator with a transparent tuning parameter. The proofs in the supplement are detailed and the main consistency and unbiasedness results are directly tied to the stated assumptions. The empirical application and simulations are appropriate for illustrating the method. The main weakness is a specific gap in the proof of the complete-randomization central limit theorem (Theorem 4), which the stress-test concern identifies correctly; this does not undermine the Bernoulli-randomization results or the consistency claims, but it does affect a load-bearing theoretical claim.","major_comments":[{"comment":"The Hajek-coupling step for complete randomization is not established. The proof asserts that combining (S.14) and (S.16) gives var(hat_tau_HT - tilde_tau_HT)/var(hat_tau_HT) = o(1), but (S.16) only shows the numerator is O(max_z d1(z)^2 n^{-1}) (up to lower-order terms). Under Assumption 4, the denominator only satisfies var(hat_tau_HT) >> max_z d_infty(z)^2 n^{-4/3}. For a regular graph with degree d = n^alpha and 0 < alpha < 1/3, one has d1 = d_infty = d, the natural variance scale is d^2/n, Assumption 4 holds because d n^{-2/3} / (d n^{-1/2}) = n^{-1/6} = o(1), and the coupling bound is also of order d^2/n, so the displayed inequalities do not force the required o(1) ratio. The complete-randomization half of Theorem 4 therefore rests on an unproved assertion. The Bernoulli-randomization CLT and the consistency results appear valid; the gap is specific to this coupling argument.","section":"Supplementary B.4, Theorem 4"},{"comment":"The statement of Theorem 6 includes Assumption 5, but the proof of asymptotic conservativeness requires the additional condition Delta = o(max_z n^{1/2} d2(z)^{1/2}) in addition to Delta >= max_z d_infty(z). These two conditions can be incompatible: in a star network, max_z d_infty(z) is of order n while max_z n^{1/2} d2(z)^{1/2} is also of order n, so no Delta can satisfy both strict inequalities, and the theorem as stated does not cover such cases. The paper should either state the additional growth condition explicitly in the theorem or clarify the scope of the guarantee.","section":"Section 4, Theorem 6"}],"minor_comments":[{"comment":"The text contains a typographical error: 'SUTV A' should read 'SUTVA'.","section":"Section 1"},{"comment":"The section heading 'Dicussion' should be 'Discussion'.","section":"Section 7"},{"comment":"In the complete-randomization variance calculation, the denominator in one covariance expression is written as p_ij p_jk; this appears to be a typo for p_ij p_kl. The surrounding argument is otherwise clear.","section":"Supplementary B.2"},{"comment":"The equality tau = n^{-1} sum_i {U_i(1) - U_i(0)} = n^{-1} sum_i {D_i(1) - D_i(0)} is used without proof. It is true by exchanging the order of summation, but a one-line justification would help readers because the neighbor sets are treatment-dependent.","section":"Section 2, equation (2)"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern in Supplementary B.4 is well-founded and is the key reason for the major-revision recommendation. The gap is local and likely fixable -- for example, by proving the complete-randomization CLT under an additional variance lower bound or by a different coupling argument -- so I do not recommend rejection. The Bernoulli-randomization results and the overall framework are valuable and appear sound."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read on Li et al. (arXiv:2505.20780). The core contribution is real: a design-based framework for dyadic outcomes when treatment is assigned at the unit level, with HT and Hajek estimators, convergence rates, a CLT, and conservative variance estimators. That combination is new relative to D'Amour and Airoldi (dyad-level treatment) and Yu et al. (unit-level outcomes under additive network effects). The paper also shows cleanly that standard unit-level weighted estimators are biased unless treatment probability is 0.5, which is a useful negative result.\n\nThe strong parts: the setup is careful, Assumptions 1-4 are stated in a falsifiable way, and the proofs in the supplement are detailed. Theorem 2's rate with explicit dependence on the degree sequence and treatment fraction is a nice touch. The WeChat application is a good reality check, though the details are redacted and no code/data are provided.\n\nNow the soft spots. The most serious is the proof of Theorem 4 for complete randomization. The Hajek coupling step in Supplementary B.4 needs var(tau_hat - tilde_tau)/var(tau_hat) = o(1). The bound actually derived in (S.16) is O(max_z d1(z)^2 / n). Assumption 4 only yields var(tau_hat) >> max_z d_infty(z)^2 n^{-4/3}. For a regular graph with degree d = n^alpha and alpha < 1/3, Assumption 4 holds, the variance scale is d^2/n, and the coupling bound is also O(d^2/n) – so the ratio does not vanish. That means the printed proof does not establish the complete-randomization CLT. I don't think the theorem is false, but the argument as written is incomplete. Since the abstract only claims a CLT for Bernoulli randomization, one fix is to state Theorem 4 for Bernoulli only; another is to find a sharper coupling bound. Either way, a referee needs to check this.\n\nThe dyadic interference assumption (Assumption 1) is strong. The paper is upfront about it and discusses sensitivity analysis only as future work; formal bounds would strengthen the paper. The conservative variance estimator relies on a user-chosen Delta and Assumption 5; the guidance on choosing Delta is pragmatic but a bit ad hoc.\n\nWho's this for? Methodologists in causal inference and practitioners running online experiments with pairwise interaction metrics. The Bernoulli parts and the general framework are solid, and the paper is worth a serious referee. I'd recommend conditional acceptance: major revision to fix or scope the complete-randomization CLT, and ideally release the simulation code.\n\nBest.","headline":"Solid new framework for dyadic outcomes under unit-level randomization, but the complete-randomization CLT proof has a gap that needs fixing.","tokens_in":36164,"tokens_out":6106,"would_cite":true,"duration_ms":58737,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D05","62E20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Pairwise outcomes, weighted by their randomization probabilities, give unbiased estimates of the global average treatment effect even when unit-level estimators fail.","keywords":["dyadic data","network interference","randomized experiments","Horvitz-Thompson estimator","Hajek estimator","global average treatment effect","online controlled experiments","social networks"],"falsifier":"Take a small population, say $n=3$, choose bounded potential outcomes $Y_{ij}(z_i,z_j)$ satisfying Assumption 1, and enumerate all $2^3=8$ treatment assignments under Bernoulli randomization with $p_i=0.3$. Theorem 2 requires the exact expectation of $\\widehat{\\tau}_{\\mathrm{HT}}$ over these assignments to equal $\\tau=n^{-1}\\sum_i\\{U_i(1)-U_i(0)\\}$; if the enumeration shows any nonzero discrepancy, the unbiasedness claim is false. To test the robustness boundary, repeat the enumeration with $Y_{ij}$ shifted by $\\zeta Z_k$ for a third unit $k$, and compare the estimator's expectation with $\\tau$: the gap quantifies the bias that higher-order interference would introduce.","tokens_in":35213,"feed_emoji":"🕸️","tokens_out":15743,"duration_ms":141479,"temperature":0.7,"pith_summary":"Randomized experiments on social platforms have a network problem: treating one user can change the interactions of everyone that user talks to, so the usual no-interference assumption fails. This paper works with the pairwise interaction $Y_{ij}$ itself and assumes each such interaction depends only on the treatments of the two users in that pair. Under this dyadic-interference assumption, the paper builds Horvitz-Thompson and Hajek estimators of the global average treatment effect by weighting treated-treated dyads by $1/p_{ij}$ and control-control dyads by $1/q_{ij}$, where $p_{ij}$ and $q_{ij}$ are the known inclusion probabilities under the randomization design. It proves these estimators are unbiased and consistent under Bernoulli, complete, and (in the supplement) cluster randomization, and it shows that the usual unit-level weighted estimators fail in all but the perfectly balanced case $p_i=1/2$. The practical payoff is that dyad-level analysis can correct interference bias that unit-level A/B analysis misses, demonstrated in simulations and in a WeChat recommendation-algorithm experiment.","feed_headline":"Pairwise data fix interference bias in network experiments","feed_subtitle":"Dyadic Horvitz-Thompson and Hajek estimators stay unbiased and consistent where unit-level estimators fail.","key_machinery":"The load-bearing object is the dyadic interference assumption $Y_{ij}(\\mathbf{z})=Y_{ij}(z_i,z_j)$, which converts network interference into a pair-local structure while still allowing arbitrary interference at the unit level. Estimation rests on dyad inclusion probabilities $p_{ij}=\\Pr(Z_i=Z_j=1)$ and $q_{ij}=\\Pr(Z_i=Z_j=0)$, which are known from the randomization design, making the Horvitz-Thompson estimator a difference of two inverse-probability-weighted dyad sums. The rates are governed by the counterfactual network degree summaries $d_1(z)$, $d_2(z)$, and $d_\\infty(z)$—average, mean-squared, and maximum degree of the network that would form if all units received $z$—because dyads sharing a unit are dependent. The central limit theorem is proved with Stein's method on dependency neighbourhoods for Bernoulli randomization and transferred to complete randomization by Hajek coupling; the conservative variance estimator $\\widehat{V}(\\Delta)$ uses Young's inequality to dominate cross-world product terms, selecting $\\Delta$ at least as large as the maximum counterfactual degree.","core_discovery":"The central claim, stated on the paper's own terms, is that under Assumption 1, $Y_{ij}(\\mathbf{z})=Y_{ij}(z_i,z_j)$ for $i\\neq j$, the global average treatment effect $\\tau=n^{-1}\\sum_i\\{U_i(1)-U_i(0)\\}$ is estimable from dyadic outcomes. The estimator $\\widehat{\\tau}_{\\mathrm{HT}}=n^{-1}\\sum_i\\sum_{j\\in D_i}(Z_{ij}/p_{ij}-\\bar{Z}_{ij}/q_{ij})Y_{ij}$ contrasts treated-treated dyads with control-control dyads and is unbiased under Bernoulli and complete randomization, with convergence rate $r_n=\\max_{z\\in\\{0,1\\}}\\{P_z^{-1}d_1(z)^{1/2}, P_z^{-1/2}d_2(z)^{1/2}\\}n^{-1/2}$; the Hajek version has the same rate. A central limit theorem holds under a maximum-degree condition, and the paper constructs asymptotically conservative variance estimators using a sensitivity parameter $\\Delta$. The same framework shows that unit-level estimators $\\tilde{\\tau}(Y)=n^{-1}\\sum_i(Z_i/a_i-(1-Z_i)/b_i)Y_i$ are biased and inconsistent in general, becoming unbiased only in the special case $a_i=b_i=p_i=1/2$.","pith_inferences":["Because the dyadic estimators tolerate arbitrary known assignment probabilities, dyad-level data weaken the experimental-design pressure toward balanced 50/50 splits; this is directly relevant when treatment is costly and skewed allocation is standard.","The same inverse-probability weighting logic should carry over to observational studies with estimated assignment probabilities—the paper lists this as future work—and would give dyadic analogues of propensity-score weighting.","A formal sensitivity analysis for violations of dyadic interference, which the paper only sketches, could be built by bounding the higher-order spillover coefficient $\\zeta$ and propagating the resulting worst-case bias through the same Young's-inequality machinery used for the variance estimator.","Before applying the method, checking the observed degree distribution is wise: large hubs and near-star networks signal both inflated variance and potential failure of the central limit theorem, and the paper's Assumptions 4 and 5 make this diagnostic transparent."],"forward_implications":["Unit-level Horvitz-Thompson estimators are generally biased and inconsistent for the global average treatment effect under dyadic interference, while the dyadic estimators are unbiased and consistent in the same settings.","Under Bernoulli and complete randomization, both estimators converge at rate $r_n$; the supplement shows cluster randomization attains the same rate when inter-cluster leakage is negligible.","The convergence rate worsens as the network becomes denser: when the mean squared degree $d_2(z)$ grows with $n$, consistency slows, and a star-like degree distribution can destroy consistency entirely.","With $\\Delta\\ge\\max_z d_\\infty(z)$, the Wald confidence intervals based on $\\widehat{V}(\\Delta)$ are asymptotically conservative, so they protect nominal coverage even when the unadjusted variance estimator undercovers.","In the WeChat Channels experiment, dyad-based estimates of relative treatment effects (1.012%, 0.624%, 1.124%) exceed unit-level estimates (0.076%, -0.120%, 0.285%), changing the statistical conclusion for two of the three metrics when the unadjusted variance estimator is used."],"supporting_citations":[{"why":"Shows identification of the global effect is impossible under arbitrary interference, which motivates the dyadic interference assumption.","marker":"Basse and Airoldi (2018)"},{"why":"Provides the prior bias result for unit-level total-treatment-effect estimators under network interference that Theorem 1 extends to heterogeneous effects.","marker":"Yu et al. (2022)"},{"why":"Introduced the Hájek estimator form whose normalization by estimated treated and control counts gives stability in finite samples.","marker":"Hájek (1971)"},{"why":"Supplies Stein's method for dependent neighbours used to prove the central limit theorem under Bernoulli randomization.","marker":"Ross (2011)"},{"why":"Provides the coupling between Bernoulli and simple random sampling that transfers the central limit theorem to complete randomization.","marker":"Hájek (1960)"},{"why":"Gives the inconsistency argument for unit-level estimators with unknown interference and frames the variance estimation challenges under interference.","marker":"Sävje et al. (2021)"},{"why":"Uses a maximum-degree condition for dyadic central limit theorems that motivates Assumption 4.","marker":"Tabord-Meehan (2019)"}],"fun_headline_variants":["Dyadic estimators correct network interference in experiments","Network trials gain from dyadic causal inference methods","Pairwise data yield unbiased estimates in networked RCTs","Dyadic treatment effects cut bias in network experiment estimates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that each pairwise outcome is affected only by the treatments of the two people in the pair; if influences from friends-of-friends or wider parts of the network reach that outcome, the proposed estimators are generally biased, and the paper leaves formal sensitivity bounds to future work.","fun_headline_variants_meta":{"raw":{"variants":["Dyadic estimators correct network interference in experiments","Network trials gain from dyadic causal inference methods","Pairwise data yield unbiased estimates in networked RCTs","Dyadic treatment effects cut bias in network experiment estimates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000709,"raw_usage":{"total_tokens":3232,"prompt_tokens":1026,"completion_tokens":2206,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":642,"completion_tokens_details":{"reasoning_tokens":2146}},"tokens_in":642,"tokens_out":2206,"duration_ms":17548,"temperature":1.0,"reasoning_tokens":2146,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:46:39.831420+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a small population, say $n=3$, choose bounded potential outcomes $Y_{ij}(z_i,z_j)$ satisfying Assumption 1, and enumerate all $2^3=8$ treatment assignments under Bernoulli randomization with $p_i=0.3$. Theorem 2 requires the exact expectation of $\\widehat{\\tau}_{\\mathrm{HT}}$ over these assignments to equal $\\tau=n^{-1}\\sum_i\\{U_i(1)-U_i(0)\\}$; if the enumeration shows any nonzero discrepancy, the unbiasedness claim is false. To test the robustness boundary, repeat the enumeration with $Y_{ij}$ shifted by $\\zeta Z_k$ for a third unit $k$, and compare the estimator's expectation with $\\tau$: the gap quantifies the bias that higher-order interference would introduce.","supporting_citations":[],"review_version":1}