REVIEW 4 major objections 5 minor 1 cited by
Adjusting auxiliary variables under approximate neighborhood interference
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A normalizing map makes network regression adjustment safe and sharper.
desk verdict A genuinely useful new regression adjustment for network experiments under ANI, with a real proof gap in the central optimality theorem that needs patching before I'd trust it fully. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key machinery is the normalizing function $\phi_0$ combined with network-dependent regression that minimizes the HAC variance estimator. For each unit, $\phi_0$ regresses the auxiliary vector $G_i$ on the Horvitz-Thompson weight $w_{\mathrm{HT},i}$ and takes the residual, so that $E[w_{\mathrm{HT},i} \phi_0(G_i)] = 0$ at the potential-outcome level. This forces the unit-level treatment-effect term $\tau_{\phi_0(G),i}$ to be zero, which removes the dependence of the variance-estimator bias $R(\beta)$ on the regression coefficient $\beta$. Once the bias is $\beta$-free, minimizing $\hat{\sigma}^2_\star(\beta)$ asymptotically targets the minimizer of the true variance, yielding the optimal linear adjustment and a consistent variance estimate.
What would settle it
Fix a network where influence decays slowly enough that $\theta_{n,s}$ is not summable, and set $G_i$ to a global feature such as eigenvector centrality or an indicator of treatment at a distant hub; run the proposed $\mathrm{ND}$-$\phi_0(G)$ estimator and check whether its 95% empirical coverage stays at 0.95 and its asymptotic variance is no larger than the unadjusted Hajek estimator, since the theorem predicts that violating Assumption 10 should degrade coverage or worsen precision.
Extended reading notes
Core claim
The paper establishes that regression adjustment under network interference can be made simultaneously valid and efficiency-improving through two new devices: a class of auxiliary variables $G_i$ that may depend on the treatment vector and the network, and a normalizing map $\phi_0(G_i) = G_i - \gamma_i w_{\mathrm{HT},i}$, where $w_{\mathrm{HT},i}$ is the Horvitz-Thompson weight and $\gamma_i = E[w_{\mathrm{HT},i} G_i]/E[w^2_{\mathrm{HT},i}]$. With this normalization, the unit-level association between the auxiliary variable and the weighted estimator vanishes, which makes the bias of the HAC variance estimator independent of the regression coefficient $\beta$. The network-dependent estimator $\hat{\tau}_{\star,\mathrm{ND}}$ is then defined by minimizing the HAC variance estimator over $\beta$. Theorem 2 shows $n^{1/2}(\hat{\tau}_{\star,\mathrm{ND}}-\tau)/\sigma_\star(\tilde{\beta}^{\mathrm{opt}}_\star)$ converges in distribution to $N(0,1)$ and the estimated variance converges to $\sigma^2_\star(\tilde{\beta}^{\mathrm{opt}}_\star)+R$, so its asymptotic variance is the minimum over all linear adjustments and hence no larger than the unadjusted estimator.
Load-bearing premise
The load-bearing premise is Assumption 10: the chosen auxiliary variables must be bounded, and changing the treatment of a unit at graph distance $s$ must change each unit's auxiliary vector in expectation at the same rate $\theta_{n,s}$ that the ANI assumption imposes on outcomes; if a user builds $G$ from global network features that do not decay spatially, the normalization and optimality results can fail.
Editorial extensions
If this is right
- Practitioners may include network-based auxiliary variables such as neighbor covariate averages, treated proportions, or exposure indicators, and the resulting estimator is asymptotically normal with variance no larger than the unadjusted Horvitz-Thompson or Hajek estimator.
- Confidence intervals based on the estimated HAC variance are asymptotically valid and never asymptotically wider than intervals from unadjusted estimators, because $\sigma^2_\star(\tilde{\beta}^{\mathrm{opt}}_\star) \le \sigma^2_\star(0)$.
- Without the normalizing step, network-dependent regression can be badly inconsistent: in the paper's simulations, estimators lacking the normalization show large bias and severe under-coverage, so the normalization is essential rather than cosmetic.
- Adding network-informed auxiliary variables through the new method improves efficiency beyond covariate-only adjustment in the paper's simulations, and shortens estimated standard errors in the reanalysis of a field experiment.
Reading between the lines
- Extension: the normalizing map can be viewed as a unit-level projection of auxiliary variables onto the orthogonal complement of the Horvitz-Thompson weight, so the method is a covariate-balancing rather than an outcome-modeling device; this suggests combining it with rerandomized designs, which the paper itself conjectures may further improve precision.
- Extension: because the theory fixes the dimension of the auxiliary space, a practitioner with many candidate network features needs either dimension reduction or features constructed with geometrically decaying influence, such as powers of the row-normalized adjacency matrix, before the optimality result applies.
- Extension: one could devise a pre-analysis diagnostic that estimates the decay rate of candidate auxiliary variables under random reassignments of distant treatments; if the decay is not comparable to the outcome's $\theta_{n,s}$, Assumption 10 would be flagged as violated.
- Extension: under the no-interference benchmark with $T_i = D_i$ and $\pi_i(t) + \pi_i(t') = 1$, the normalizing function becomes a covariate-centering-type adjustment, so the framework contains classical Fisher and Lin regression adjustments as special cases.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops regression adjustment for randomized network experiments under approximate neighborhood interference (ANI). It defines auxiliary variables G_i that may depend on treatment assignments and network structure, introduces a normalization φ0(G_i) = G_i − γ_i w_HT,i with γ_i = E[w_HT,i G_i]/E[w_HT,i^2], and estimates the treatment effect by minimizing a HAC variance estimator over the regression coefficient β. The main theorem (Theorem 2) claims that the resulting estimator is asymptotically normal at the oracle-optimal variance and that its variance estimator is consistent up to a β-independent bias R, so adjusted confidence intervals are asymptotically no wider and often shorter than unadjusted ones. The paper also proposes a network-dependent version of Fisher's and Lin's regression, illustrates the methods in simulations, and reanalyzes the Cai et al. (2015) insurance experiment.
Significance. The contribution is potentially valuable: it unifies several existing regression-adjusted estimators, provides a principled way to use network features as auxiliary variables, and targets the desirable 'no efficiency loss' property in a design-based, model-free framework. The paper is transparent about its assumptions and includes a useful constructive counterexample showing that naive Fisher/Lin adjustment can hurt precision. However, the central theorem currently rests on an unproven ANI-type condition for the normalized variable, on the oracle availability of γ_i, and on an implicit nonnegativity of the variance bias R. These gaps are technical and likely repairable, but they must be addressed before the results can be fully relied upon.
major comments (4)
- [Supplement B, final verification paragraph; Theorem 2] The proof of Theorem 2 applies Theorem 3, which is stated under Assumption 10 with Gi replaced by φ(Gi). The closing sentence of Supplement B asserts that φ0(Gi) satisfies Assumption 10, but this is not demonstrated and is generally false for small s. For 1 ≤ s ≤ K, the term γ_i(w_HT,i(D) − w_HT,i(D^(i,s))) in φ0(Gi(D)) − φ0(Gi(D^(i,s))) can be O(1) even when the outcome ANI coefficient θ_{n,s} is small, because the exposure mapping T_i may depend on treatments to which Yi is insensitive. Thus max_i E||φ0(Gi(D)) − φ0(Gi(D^(i,s)))||∞ ≤ c_G θ_{n,s} need not hold. I do not think this destroys the CLT: in Assumptions 5 and 7 the dependence coefficients are truncated to 1 for s ≤ 2K, so a decay bound for s > K together with boundedness for small s is likely sufficient. The authors need to supply that truncated argument, or add an explicit assumption covering φ0(Gi), instead of the current verification.
- [Supplement A, proof of Proposition 1] The verification that (X_i^T 1(T_i=t), X_i^T 1(T_i=t'))^T satisfies Assumption 10 only treats s > 2 max{1,K}; for 1 ≤ s ≤ 2 max{1,K} the indicator difference 1(T_i(D)=t) − 1(T_i(D^(i,s))=t) is generally not O(θ_{n,s}). Since Proposition 1 and Theorem 1 rely on Proposition 2, which is stated under Assumption 10 for the transformed variable, the small-s range is again left unproved. The same truncated-ANI repair as in the previous comment is needed here.
- [Section 4.2, Eq. (1); Sections 5.1–5.2] Theorem 2 treats γ_i as known, but the method is implemented by Monte Carlo approximation of γ_i (10^5 draws in the simulation and 10^4 in the real-data analysis). The estimation error in γ_i is not covered by any theorem. If the number of Monte Carlo draws is fixed, the error is O_P(M^{-1/2}) and can affect β̂_ND and the variance estimator; in particular τ_{φ̂0(G),i} is no longer exactly zero, so R(β) is no longer β-independent and the optimality argument in Theorem 2 does not directly apply. The authors should either prove asymptotic equivalence under explicit conditions on M_n (for example M_n/n → ∞) or state clearly that the theoretical results concern the oracle normalization and that the simulations are a numerical check.
- [Section 3.2 and Theorem 2(ii)] Theorem 2(ii) establishes σ̂²_⋆(β̂_ND) = σ²_⋆(β̃_opt⋆) + R + o_P(1), but the paper does not assume R ≥ 0. If R is negative with non-negligible magnitude, Wald confidence intervals based on σ̂²_⋆(β̂_ND) will undercover, contradicting the paper's advertised claim of shorter but valid confidence intervals. Proposition 1 notes that R can be negative in general and refers to Leung (2019) for conditions under which the bias is nonnegative, but no such condition appears in Assumption 12 or Theorem 2. The authors should add an explicit nonnegativity or conservative-bias assumption for the settings where confidence intervals are recommended, or quantify the coverage distortion when R < 0.
minor comments (5)
- [Throughout] There are several typos: 'variacne' and 'asympototic' in Section 1, 'an a new normalization' in Section 1, 'Hovits-Thompson' in Section 3.1, 'regulatity' in Assumption 6, and 'coordinate' in Assumption 4. The manuscript should be carefully proofread.
- [Figure 1] The figure labels the network-dependent auxiliary-variable estimators as ND-G_1 and ND-G_2, while the text and tables use ND-φ0(G1) and ND-φ0(G2). Since ND-G1 and ND-G2 denote the inconsistent unnormalized estimators in Table 1, the figure labels are misleading and should be corrected.
- [Table 1] The table has formatting problems: several entries are run together (for example '0.8950.000' and '0.2300.259' in the linear-in-means panel), making the table hard to read.
- [Section 4.2] The sentence 'τ_{G,i} = E[w_HT,i G_i]. This is the correlation between w_HT,i and Gi' is imprecise: the quantity is an uncentered cross-expectation, not a correlation or covariance. Please rephrase.
- [Assumption 8] The definition of H_n(s,m) uses ℓ_A({i,k},{j,l}); the distance between two-element sets should be defined explicitly, for example as the minimum path distance between any element of {i,k} and any element of {j,l}.
Circularity Check
No significant circularity: the theoretical claims are derived from explicit assumptions and external ANI limit theorems; the normalization is design-based and not fitted to outcomes.
full rationale
The paper's derivation chain is self-contained in the relevant sense. The variance estimator and network-dependent regression estimator are defined from the known treatment assignment distribution and user-specified auxiliary variables. The normalizing function phi0(G_i) = G_i - gamma_i w_HT,i with gamma_i = E[w_HT,i G_i]/E[w_HT,i^2] is a projection coefficient computed from the design and the auxiliary variables, not from outcome data. The key identity tau_{phi0(G),i} = 0 is an algebraic consequence of the definition, and it is used to make the bias term R(beta) independent of beta; this is a construction, not a fitted parameter disguised as a prediction. The optimality statement sigma^2(beta_tilde_opt) <= sigma^2(0) is definitional, but the substantive content of Theorem 2 is the asymptotic normality of the network-dependent estimator and the consistency of its variance estimator, which are proved by reducing the problem to the external ANI results of Leung (2022) and Kojevnikov et al. (2021). No load-bearing self-citation chain is present: the citations to Lu et al. (2023) and other own prior work are not used to justify the central theorem, and Gao and Ding (2023) and Leung (2022) are independent external benchmarks. The main caveat is that Supplement B asserts without proof that phi0(G_i) satisfies Assumption 10; if this verification fails, Theorem 2 is unsupported. However, that is a technical correctness gap, not circular reasoning, because it does not reduce the theorem to its own conclusion or to a fitted input. Overall, no pattern of circular derivation is present.
Assumptions & free parameters
free parameters (2)
- HAC bandwidth b_n =
3 (simulations and real-data analysis); theory requires b_n → ∞
- Monte Carlo draws for gamma_i =
10^5 (simulation), 10^4 (real data)
assumptions (7)
- domain assumption Assumption 4 (ANI): theta_{n,s} := max_i E|Y_i(D)−Y_i(D^{(i,s)})| → 0 as s→∞.
- domain assumption Assumptions 1, 2, 6 (overlap and boundedness): pi_i(t) in [pi, pi_bar] subset (0,1), |Y_i(d)| <= c_Y, ||X_i||_inf <= c_X.
- domain assumption Assumption 3 (local exposure mapping): T(i,d,A) depends only on the K-neighborhood of i.
- domain assumption Assumptions 5 and 7 (weak dependency and HAC consistency): certain sums of neighborhood sizes times theta_{n,s} satisfy decay conditions; bandwidth b_n → ∞.
- domain assumption Assumption 10 (auxiliary ANI): ||G_i(d)||_inf <= c_G and max_i E||G_i(D)−G_i(D^{(i,s)})||_inf <= c_G theta_{n,s}.
- domain assumption Assumptions 8, 9, 11, 12 (nondegeneracy and bias magnitude): Hessians of sigma^2+R are positive definite; sigma^2 > 0; R = n^{-1} sum sum B_ij (tau_i−tau)(tau_j−tau) = O(1).
- standard math External asymptotic results: Leung (2022) Theorems 2 and 4; Kojevnikov et al. (2021) Proposition 4.1; Gao and Ding (2023) variance formulas.
Cite this review
Pith. "Pith review of Adjusting auxiliary variables under approximate neighborhood interference." pith.science (2026). https://pith.science/paper/BOE7WOFI
@misc{pith2026241119789,
author = {Pith},
title = {Pith review of: Adjusting auxiliary variables under approximate neighborhood interference},
year = {2026},
howpublished = {\url{https://pith.science/paper/BOE7WOFI}},
note = {Machine review of arXiv:2411.19789}
}
read the original abstract
Randomized experiments are the gold standard for causal inference. However, traditional assumptions, such as the Stable Unit Treatment Value Assumption (SUTVA), often fail in real-world settings where interference between units is present. Network interference, in particular, has garnered significant attention. Structural models, like the linear-in-means model, are commonly used to describe interference; but they rely on the correct specification of the model, which can be restrictive. Recent advancements in the literature, such as the Approximate Neighborhood Interference (ANI) framework, offer more flexible approaches by assuming negligible interference from distant units. In this paper, we introduce a general framework for regression adjustment for the network experiments under the ANI assumption. This framework expands traditional regression adjustment by accounting for imbalances in network-based covariates, ensuring precision improvement, and providing shorter confidence intervals. We establish the validity of our approach using a design-based inference framework, which relies solely on randomization of treatment assignments for inference without requiring correctly specified outcome models.
Figures
Forward citations
Cited by 1 Pith paper
-
Causal Inference under Interference: Regression Adjustment and Optimality
Under network interference, linear and kernel regression adjustments achieve the smallest asymptotic variance in their classes, and the paper provides consistent variance estimators.
Reference graph
Works this paper leans on
-
[1]
Aronow, P. M. and C. Samii (2017). Estimating average causal effects under general interference, with application to a social network experiment. The Annals of Applied Statistics 11 , 1912 –
work page 2017
-
[2]
For ease of notation, we write ˜Gi ≡ ϕ(Gi)
To do this, we first prove that Yi − ϕ(Gi)⊤β satisfies Assumptions 2, 4, 5 for Yi and then apply Theorem 4 from Leung (2022) with Yi ≡ Yi − ϕ(Gi)⊤β. For ease of notation, we write ˜Gi ≡ ϕ(Gi). Since Assumption 10 holds for ϕ(Gi), we have, for i ∈ Nn and d ∈ {0, 1}n, |Yi(d) − ˜Gi(d)⊤β| ≤ |Yi(d)| + ∥ ˜Gi(d)∥∞∥β∥1 ≤ cY + ∥β∥1cG and max i∈Nn E[|Yi(D) − Yi(D(i...
work page 2022
-
[3]
Define 1 HT(t) = n−1 Pn i=1 1(T i = t)/πi(t)
Let Gi = (Giq)Q q=1, µG(t) = ( µGq (t))Q q=1 and ˆµ⋆,G(t) = (ˆµ⋆,Gq (t))Q q=1, ⋆ ∈ {HT, Haj}. Define 1 HT(t) = n−1 Pn i=1 1(T i = t)/πi(t). By the proof of Leung (2022, Theorem
work page 2022
-
[4]
with Yi ≡ Yi − ϕ(Gi)⊤β, we have ˆσ2 HT(β) = σ2 HT(β) + R(β) + oP(1). For ˆσ2 Haj(β), define ˆV ⋆, ˜G,i and V ⋆, ˜G,i analogously as ˆV⋆,i and V⋆,i, with Yi replaced with ˜Gi. ˆVHaj,i(β) and VHaj,i(β) can be expressed as ˆVHaj,i(β) = ˆVHaj,i − ˆV ⊤ Haj, ˜G,iβ and VHaj,i(β) = VHaj,i − V ⊤ Haj, ˜G,iβ, respectively. We have EVHaj,i(β) = τi − τ − β⊤(τ ˜G,i − τ...
work page 2022
-
[6]
We define ψ-dependence in line with Definition 2.2 of Kojevnikov et al
and (ii) ˆσ2 ⋆( ˆβ⋆,ND) − σ2 ⋆( ˜β⋆,ND) − R( ˜β⋆,ND) = oP(1). We define ψ-dependence in line with Definition 2.2 of Kojevnikov et al. (2021) For d ∈ N, let Ld be the set of real-valued bounded Lipschitz functions on Rd: Ld := {f : Rd → R : ∥f ∥∞ < ∞, Lip(f ) < ∞}, where ∥f ∥∞ := supx∈Rd |f (x)| and Lip(f ) indicates the Lipschitz constant of f , that is |...
work page 2021
-
[7]
The proof is very similar to that of (Leung, 2022, Theorem 1). There- fore, we omit it. Proof of Theorem
work page 2022
-
[8]
We first prove that ˆβ⋆,ND = ˜β⋆,ND + oP(1). We see that ˆβ⋆,ND = n−1 nX i=1 nX j=1 Bij ˆV ⋆,ϕ(G),i ˆV ⊤ ⋆,ϕ(G),j −1 n−1 nX i=1 nX j=1 Bij ˆV ⋆,ϕ(G),i ˆV⋆,j . Let ˜Gi ≡ ϕ(Gi) = ( ˜Giq)Q q=1. Let ˆV ⋆,ϕ(G),i = ( ˆV⋆, ˜Gq,i)Q q=1 and V ⋆,ϕ(G),i = (V⋆, ˜Gq,i)Q q=1 and τ ϕ(G),i = (τ ˜Gq,i)Q q=1. For simplicity, we write ˆW i = ( ˆV⋆,i, ˆV ⊤ ⋆,ϕ(G),i)⊤ ∈ R1+Q ...
work page 2021
-
[9]
= oP(1). Similar as the proof of (Leung, 2022, Theorem 4), we have n−1 nX i=1 nX i=1 (Wiq1 − EWiq1)EWjq2Bij = oP(1). Putting together, we have n−1 nX i=1 nX j=1 ˆWiq1 ˆWiq2Bij = Cov(n−1/2 nX i=1 Wiq1, n−1/2 nX i=1 Wjq2)+ n−1 nX i=1 nX j=1 EWiq1EWjq2Bij + oP(1). 40 As a consequence, we have n−1 nX i=1 nX j=1 Bij ˆV ⋆,ϕ(G),i ˆV ⊤ ⋆,ϕ(G),j = Cov n−1/2 nX i=1...
work page 2022
Show all 13 references
-
[10]
41 In light of the above, applying (Leung, 2022, Theorem
=ˆτ⋆( ˜β⋆,ND) − τ + oP(n−1/2). 41 In light of the above, applying (Leung, 2022, Theorem
2022
-
[11]
Applying (Leung, 2022, Theorem
= ˆµHT(t) − µ(t)ˆ1HT(t) − ˜β ⊤ Haj,ND{ ˆµHT,ϕ(G)(t) − µϕ(G)(t)ˆ1HT(t)} ˆ1HT(t) − ˆµHT(t′) − µ(t′)ˆ1HT(t′) − ˜β ⊤ Haj,ND{ ˆµHT,ϕ(G)(t′) − µϕ(G)(t′)ˆ1HT(t′)} 1HT(t′) = h ˆµHT(t) − µ(t)ˆ1HT(t) − ˜β ⊤ Haj,ND{ ˆµHT,ϕ(G)(t) − µϕ(G)(t)ˆ1HT(t)} i | {z } =:T1 (1 + OP(n−1/2))− h ˆµHT(t′...
2022
-
[12]
As a consequence, we have n1/2(ˆτHaj( ˆβHaj,ND) − τ )/σHaj( ˜βHaj,ND) d − → N(0, 1)
with Yi ≡ Yi − µ(t)1(T i = t) − µ(t′)1(T i = t′) − (ϕ(Gi)−µϕ(G)(t)1(T i = t)−µϕ(G)(t′)1(T i = t′))⊤ ˜βHaj,ND, we haven1/2(T1−T2) d − → N(0, 1). As a consequence, we have n1/2(ˆτHaj( ˆβHaj,ND) − τ )/σHaj( ˜βHaj,ND) d − → N(0, 1). Now we prove Theorem 2 (ii). |ˆσ2 ⋆( ˆβ⋆,ND) − ˆ...
2023
-
[13]
ND-F , ND-L
(5) Now that ˜β(˜t) = 0, ˜t = 0, 1, we further have σ2 Haj − σ2 L = − 1 n nX i=1 πi(0)πi(1) X ⊤ i βL(1) πi(1) + X ⊤ i βL(0) πi(0) 2 = − βL(1) βL(0) ⊤ Pn i=1 πi(0) πi(1) X iX ⊤ i Pn i=1 X iX ⊤ i Pn i=1 X iX ⊤ i Pn i=1 πi(1) πi(0) X iX ⊤ i βL(1) βL(0). Thi...
2023
-
[1947]
Beaman, L. and A. Dillon (2018). Diffusion of agricultural information within social networks: Evidence on gender inequalities from mali. Journal of Development Eco- nomics 133 , 147–161. Bloniarz, A., H. Liu, C.-H. Zhang, J. S. Sekhon, and B. Yu (2016). Lasso adjustments of t...
2018 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.