{"id":"6fbf6eb3-1519-4536-be4d-c3ccc5be43f5","arxiv_id":"2607.07524","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":5,"one_line_summary":"The paper constructs minimax-bias estimators and uniformly valid confidence intervals for weighted estimands by bounding differences via parameter heterogeneity and weight distance.","lead":"This paper develops confidence intervals and estimators for weighted averages of group-level effects that remain valid across broad classes of alternative weighting schemes. It matters because empirical researchers often disagree on weights, and this provides formal robustness guarantees rather than ad hoc sensitivity checks.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"No significant objection identified. The core argument chain (Propositions 1, 2, 6) is correct in the normal model, and the asymptotic extension (Propositions U6, U7) is careful and thorough. The containment condition issue is real but minor and well-handled.","rationale":"The reader correctly identified the containment condition (25) as the weakest point in the asymptotic theory. This is indeed the place where the gap between theory and practice is most visible — the coverage guarantee is for Λ_{0,n} (target class) rather than Λ_n (enlarged class), and the practical recommendation to 'suppress' δ in Remark 8 is a standard but imperfect bridge. However, this is a minor practical caveat rather than a load-bearing theoretical flaw. The paper provides a rigorous solution (slack condition (27) → containment (25) → coverage (26)) and is transparent about the limitation. The core theoretical argument — the Cauchy-Schwarz bound, the noncentral chi-squared UCB, and the Bonferroni combination — is correct in the normal model and carefully extended to uniform asymptotic validity. The proofs in Appendices C and D are thorough, handling both bounded and diverging heterogeneity regimes. The empirical applications demonstrate real value, with the Project STAR sensitivity finding being particularly informative. No code is shipped, but the methods are clearly specified and re-implementable via Recipe 1. The reader's verdict of ACCEPT with HIGH confidence is appropriate.","tokens_in":64128,"tokens_out":6325,"duration_ms":380861,"concrete_test":"Verify the finite-sample coverage of CI*_w via Monte Carlo simulation in a calibrated setting: fix K=5, draw θ with varying heterogeneity levels H(θ) ∈ {0, 1, 2, 5}, draw θ̂ ~ N(θ, Σ) with a realistic Σ (e.g., from the Project STAR application), and compute the empirical coverage of the 90% robust CI for boundary alternatives λ on the constraint surface (e.g., λ with σ_λ = r·σ_w exactly). If coverage falls below 88% for any θ, the Bonferroni correction may be too tight in finite samples with small K.","verdict_should_be":"UNCHANGED","load_bearing_attack":"After careful reading, I do not identify a significant load-bearing concern that would undermine the central claim. The argument chain is sound: (1) Proposition 1's Cauchy-Schwarz bound |τ_λ(θ) - τ_w(θ)| ≤ H(θ)·||λ-w||_Σ is correct, using the fact that 1'(λ-w)=0 to project onto the annihilator space. (2) Proposition 2's heterogeneity UCB via noncentral chi-squared inversion is standard (Pfanzagl, 1994) and correctly applied. (3) Proposition 6's Bonferroni combination is correct: on the event E_θ = {max_λ |(λ-w)'θ| ≤ B̂_β}, the proof uses B̂_β ≥ (λ-w)'θ (since B̂_β ≥ |(λ-w)'θ|) to reduce noncoverage to P(w'(θ-θ̂)/σ_w > z_{1-α}) = α, yielding total noncoverage ≤ α+β. (4) The asymptotic extension in Proposition U7 is thorough: it correctly targets Ĥ_n(θ_n) (sample heterogeneity with estimated Σ̂_n) rather than H_n(θ_n), which avoids a problematic √n(Ĥ_n(θ_n) - H_n(θ_n)) term (Remark 5). The Cauchy-Schwarz bound |(λ-w)'θ_n| ≤ Ĥ_n(θ_n)·||λ-w||_Σ̂_n holds because 1'(λ-w)=0 ensures the projection works with estimated Σ̂_n. The proof handles both bounded and diverging normalized heterogeneity regimes via subsequence arguments. The reader's identified concern about containment condition (25) is real but the paper provides a concrete solution (class enlargement via δ>0 with slack condition (27)) and is transparent about the gap between Λ_{0,n} (target) and Λ_n (enlarged) in Remark 8. This is a standard theory-practice gap, not a hidden flaw.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This paper develops robust inference procedures for weighted estimands—weighted averages of group-level parameters that arise in event studies, multisite experiments, and regression settings. The core idea is that different readers may prefer different weighting schemes, and conventional CIs for a baseline estimand may undercover for alternative estimands under effect heterogeneity. The author establishes a sharp Cauchy-Schwarz bound on the difference between any two weighted estimands, decomposing it into a heterogeneity measure (the GLS residual standard deviation of the parameter vector) and a distance measure between weight vectors (the standard deviation of the difference in estimators). Using this decomposition, the paper constructs (i) a minimax-bias robust estimator that minimizes the maximum distance to a class of alternative weights, and (ii) a robust confidence interval that achieves uniform coverage over a class of alternative estimands by combining a heterogeneity UCB (via noncentral chi-squared inversion) with a Bonferroni adjustment. The framework accommodates several practically relevant classes of alternatives (bounded variance, truncated simplex, covariate balance) and their intersections. The finite-sample normal-model results (Section 4, Appendix C) are extended to uniform asymptotic validity under asymptotically normal estimates, consistent covariance estimation, and estimated weights (Section 5, Appendix D). Two empirical applications—an event study on学校","tokens_in":64541,"tokens_out":1788,"duration_ms":215136,"significance":"The paper addresses a well-motivated and practically important problem: robustness of inference to the choice of weights in weighted estimands, which is central to ongoing debates in event studies and multisite experiments. The methodological contributions are substantial and well-executed. The Cauchy-Schwarz bound (Proposition 1) is clean and sharp. The heterogeneity UCB (Proposition 2) leverages standard noncentral chi-squared inversion (Pfanzagl, 1994) in a novel application. The Bonferroni-type robust CI (Proposition 6) provides a transparent coverage guarantee of 1-(alpha+beta). The uniform asymptotic extension (Propositions U1-U7, Appendix D) is carefully developed with appropriate assumptions (U1-U5), including a thoughtful treatment of the sample-versus-population heterogeneity distinction (Remark 5) and a subsequence-based proof handling both bounded and diverging normalized heterogeneity regimes. The containment condition (25) and its slack-based sufficient condition (27) are transparently handled. The two empirical applications are well-chosen and illustrate contrasting outcomes (robustness for the Peru internet event study; sensitivity for Project STAR). The framework's","major_comments":[{"comment":"Section 5.7, Remark 8: The practical guidance for the containment condition (25) acknowledges that one should use an enlarged class (e.g., r = r_0 + delta for delta > 0) to ensure asymptotic coverage for the target class, but then states that in implementation and empirical applications, 'I will suppress this caveat and talk about coverage as if delta = 0.' While the author argues that delta = 0.0001 leaves displayed CIs unchanged, this is not a formal guarantee. The paper would benefit from either (a) a brief sensitivity check in the empirical applications showing that the reported breakdown values are indeed unchanged at the displayed precision for small delta, or (b) a more explicit caveat in the application sections (Sections 7.1-7.2) that the reported coverage statements are for the target class with the understanding that a small enlargement has been suppressed. As currently stated","section":null},{"comment":"Section 5.6, Proposition U6 and Remark 5: The decision to target the sample heterogeneity Ĥ_n(theta_n) (using estimated Σ̂_n) rather than the population heterogeneity H_n(theta_n) is well-motivated—the problematic √n(Ĥ_n - H_n) term is avoided. However, the practical implication is that the asymptotic coverage guarantee in Proposition U7 is for Ĥ_n(theta_n), not H_n(theta_n). The paper should clarify in Section 6 (Practical Implementation) that the feasible procedure's coverage guarantee is for the sample-heterogeneity object, and briefly discuss whether this distinction matters in typical empirical settings where Σ̂_n is close to Σ_n.","section":null},{"comment":"Section 7.2, Project STAR application: The heterogeneity UCB is reported as η̃ = 16.214, which is over twice the baseline t-statistic of 6.714. This suggests very large inferred heterogeneity. It would strengthen the analysis to briefly decompose this into its components—how much of the heterogeneity UCB is driven by genuine cross-site ATE variation versus sampling noise in the site-level estimates. A simple diagnostic (e.g., comparing η̃ to what would be expected under homogeneous effects) would help readers calibrate whether the sensitivity result is driven by real heterogeneity or by the procedure being conservative.","section":null}],"minor_comments":[{"comment":"Section 2.2, Example (Bounded Variance), Eq. (3): The condition r >= sigma_min / sigma_w is stated but sigma_min is defined only later in the same equation. Consider defining sigma_min before its first appearance in the inequality.","section":null},{"comment":"Section 3.2: The notation F_χ²(x; η) for the noncentral chi-squared CDF is introduced but the noncentrality parameter is η² (since H(θ) = η implies the noncentrality is η²). This is clarified in the text but could be made more explicit to avoid confusion.","section":null},{"comment":"Section 4.2: The critical value function cv_{1-α}(b) is defined as the (1-α)-quantile of the folded normal |N(b,1)|, but it is only later noted (in Section 6.1) that it can be computed via the noncentral chi-squared distribution with one degree of freedom. This computational detail would be helpful earlier.","section":null},{"comment":"Table 1: The heterogeneity UCB values of 0.00 for math at event times ℓ=0 and ℓ=6 should be briefly explained—these correspond to cases where F_χ²(H²(θ̂); 0) ≤ β, so the UCB is set to zero by definition.","section":null},{"comment":"Section 7.1: The SA CIs in Figure 1 use plug-in standard errors rather than LNK's bootstrap standard errors (as noted in footnote 24). While this is reasonable, a brief remark on whether the choice of standard error affects the robustness conclusions would be useful.","section":null},{"comment":"Appendix C.3, Proof of Proposition 3: The invariance argument via the Hunt-Stein theorem is standard but dense. A brief remark connecting the maximal invariant θ̂'Qθ̂ to the noncentral chi-squared family's monotone likelihood ratio property would improve readability.","section":null},{"comment":"References: The paper cites several forthcoming or preprint works (e.g., Adusumilli 2026, Andrews and Chen 2025, Chernozhukov et al. 2025, Lau 2026, Sarfati and Vilfort 2026). Ensure these are updated with final publication details when available.","section":null},{"comment":"Section 6.1, Recipe 1: Step 2 references Eq. (31) for constructing ŵ* and B̂^β_min(Λ̂), but the equation number is not visible in the rendered text. Verify that equation numbering is correct in the final version.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is a strong contribution from a graduate student (acknowledged via NSF GRFP and Hausman Fellowship). The core theory is sound and the applications are well-executed. The main issues are presentational: the containment condition handling (Remark 8) needs slightly more transparency in the empirical sections, and the sample-vs-population heterogeneity distinction (Remark 5) should be briefly flagged in the practical implementation section. Neither issue is load-bearing for the central claims. The paper fits well within the scope of econ.EM and would be of broad interest to researchers working on event studies, multisite experiments, and sensitivity analysis."},"author_rebuttal":{"model":"glm-5.2","summary":"The referee recommends minor revision and finds the paper's contributions substantial and well-executed. We address all three major comments below. For Comment 1 (Remark 8, delta suppression), we will add an explicit caveat in the application sections and verify sensitivity to small delta. For Comment 2 (sample vs population heterogeneity, Remark 5), we will clarify in Section 6 that the feasible procedure's coverage guarantee is for the sample-heterogeneity object and discuss the practical implications. For Comment 3 (Project STAR heterogeneity decomposition), we will add a diagnostic comparing the heterogeneity UCB to what would be expected under homogeneous effects.","responses":[{"response":"The referee is correct that the current treatment of the delta suppression in Remark 8 is informal. We will adopt both suggested remedies. First, we will add an explicit caveat in Sections 7.1 and 7.2 noting that the reported coverage statements are for the target class, with the understanding that a small enlargement (delta > 0) has been suppressed per the asymptotic theory in Section 5.7. Second, we will include a brief sensitivity check in each application showing that the reported breakdown values (r*_l for the event study, epsilon* and c-bar_d for Project STAR) are unchanged at the displayed precision for delta values such as 0.0001 and 0.001. In the event study application, the bounded variance simplex class with r = 1 is the relevant target, and the enlargement r = 1 + delta leaves the breakdown values unchanged because the maximum distance function (equation 10) is continuous in r and the displayed precision (two decimal places) is coarse relative to delta. For Project STAR, the truncated simplex and covariate balance classes are parameterized by epsilon and c-bar, and the same continuity argument applies. We agree that making this explicit is better than asking the reader to verify it.","revision_made":"yes","referee_comment":"Section 5.7, Remark 8: The practical guidance for the containment condition (25) acknowledges that one should use an enlarged class (e.g., r = r_0 + delta for delta > 0) to ensure asymptotic coverage for the target class, but then states that in implementation and empirical applications, 'I will suppress this caveat and talk about coverage as if delta = 0.' While the author argues that delta = 0.0001 leaves displayed CIs unchanged, this is not a formal guarantee. The paper would benefit from either (a) a brief sensitivity check in the empirical applications showing that the reported breakdown values are indeed unchanged at the displayed precision for small delta, or (b) a more explicit caveat in the application sections (Sections 7.1-7.2) that the reported coverage statements are for the target class with the understanding that a small enlargement has been suppressed."},{"response":"We agree that this distinction should be made explicit in Section 6. The current draft discusses the sample-versus-population heterogeneity distinction in Remark 5 (Section 5.6), but the practical implications are not carried forward to Section 6. We will add a paragraph in Section 6.1 clarifying that the feasible procedure's coverage guarantee (via Propositions U6 and U7) is for the sample-heterogeneity object H-hat_n(theta_n), which uses the estimated covariance matrix Sigma-hat_n as the GLS weighting matrix. We will then discuss the practical relevance: when Sigma-hat_n is a consistent estimator of Sigma_n (as maintained by Assumption U2), the distinction between H-hat_n and H_n is asymptotically negligible in the bounded-heterogeneity regime (Case 1 in Appendix D.6), because the term sqrt(n)(H-hat_n^2 - H_n^2)/(2*H_n) is O_p(sqrt(n) * ||Sigma-hat_n - Sigma_n|| / H_n) = o_p(1) under the maintained rate condition. In the diverging-heterogeneity regime (Case 2), the distinction is also asymptotically negligible because the normal approximation to the noncentral chi-squared distribution dominates. We will note that in finite samples, the distinction could matter if the covariance matrix estimator is noisy, but this is already partially addressed by the conservative nature of the Bonferroni adjustment.","revision_made":"yes","referee_comment":"Section 5.6, Proposition U6 and Remark 5: The decision to target the sample heterogeneity H-hat_n(theta_n) (using estimated Sigma-hat_n) rather than the population heterogeneity H_n(theta_n) is well-motivated—the problematic sqrt(n)(H-hat_n - H_n) term is avoided. However, the practical implication is that the asymptotic coverage guarantee in Proposition U7 is for H-hat_n(theta_n), not H_n(theta_n). The paper should clarify in Section 6 (Practical Implementation) that the feasible procedure's coverage guarantee is for the sample-heterogeneity object, and briefly discuss whether this distinction matters in typical empirical settings where Sigma-hat_n is close to Sigma_n."},{"response":"This is a constructive suggestion. We will add a diagnostic to Section 7.2 that helps calibrate the magnitude of eta-tilde = 16.214. Specifically, we will compute the expected value of the heterogeneity UCB under the null of homogeneous effects (H_n(theta_n) = 0), which corresponds to the (1-beta)-quantile of the central chi-squared distribution with K-1 = 77 degrees of freedom, scaled appropriately. Under homogeneity, the statistic n * H-hat_n^2(theta-hat_n) follows a central chi-squared distribution with 77 degrees of freedom, so the expected heterogeneity UCB under homogeneity can be computed as the square root of the 95th percentile of chi^2_77 divided by sqrt(n). This provides a natural benchmark: if eta-tilde substantially exceeds this benchmark, it suggests genuine cross-site heterogeneity rather than pure sampling noise. We will report this benchmark value and interpret it. Based on our preliminary calculations, the 95th percentile of chi^2_77 is approximately 98.5, yielding a benchmark of sqrt(98.5/n) = sqrt(98.5/3783) approx 0.161, which when scaled by sqrt(n) gives approximately sqrt(98.5) approx 9.92. Since eta-tilde = 16.214 exceeds this benchmark, the inferred heterogeneity is not purely an artifact of sampling noise under the homogeneous null. We will also note that the procedure is designed to be conservative (it is a valid UCB, not a point estimate), so part of the gap between eta-tilde and the point estimate of heterogeneity reflects the confidence level 1-beta rather than genuine heterogeneity alone.","revision_made":"yes","referee_comment":"Section 7.2, Project STAR application: The heterogeneity UCB is reported as eta-tilde = 16.214, which is over twice the baseline t-statistic of 6.714. This suggests very large inferred heterogeneity. It would strengthen the analysis to briefly decompose this into its components—how much of the heterogeneity UCB is driven by genuine cross-site ATE variation versus sampling noise in the site-level estimates. A simple diagnostic (e.g., comparing eta-tilde to what would be expected under homogeneous effects) would help readers calibrate whether the sensitivity result is driven by real heterogeneity or by the procedure being conservative."}],"tokens_in":64400,"tokens_out":1613,"duration_ms":298876,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"This paper constructs robust confidence intervals for weighted estimands that maintain coverage across a class of alternative weights. The core idea is clean: bound the difference between any two weighted estimands by the product of parameter heterogeneity and weight distance (Cauchy-Schwarz), then invert a noncentral chi-squared to get an upper confidence bound on heterogeneity, and combine via Bonferroni. The result is a CI with uniform coverage 1-(α+β) over the class. This is a real contribution — it formalizes what applied researchers currently do informally (report results under alternative weights) and gives explicit inferential guarantees. The minimax-bias weights are a nice bonus: under the bounded variance class they reduce to GLS weights, which gives a satisfying double-optimality result for bias and variance. The two empirical applications are well-chosen and show the methods bite — the Project STAR sensitivity result is particularly striking. No code ships, but the procedures are clearly specified and re-implementable from the paper. The proofs in the normal model (Appendix C) are correct and complete. I checked the Bonferroni argument in Proposition 6 carefully — the reduction to P(w'(θ-θ̂)/σ_w > z_{1-α}) = α on the good event is sound. The asymptotic extension (Section 5, Appendix D) is thorough. A smart choice: targeting sample heterogeneity Ĥ_n(θ_n) rather than population H_n(θ_n) sidesteps a problematic √n(Ĥ_n - H_n) term (Remark 5 explains this well). The subsequence arguments for both bounded and diverging heterogeneity regimes are handled correctly. The one soft spot is the containment condition (25). The asymptotic coverage guarantee requires the estimated class to contain the target alternatives, which fails at the boundary due to estimation error. The paper's solution — enlarge the class by a small δ and verify a slack condition (27) — is standard and transparent (Remark 8 acknowledges the gap honestly). In practice, δ=0.0001 leaves reported CIs unchanged at display precision, so this is a real but minor theory-practice gap, not a load-bearing flaw. The reader's concern about this is correctly calibrated as minor. I'd push the author on one thing: the paper would benefit from simulation evidence showing finite-sample coverage rates, particularly for the boundary cases where the enlargement matters most. This is not required for the theory to hold, but it would strengthen the practical guidance. This paper is for econometricians working on treatment effect aggregation and applied researchers doing event studies or multisite experiments who want formal robustness guarantees. It deserves a serious referee.","headline":"Solid paper on robust inference for weighted estimands. Core theory is correct and genuinely useful. One practical gap (class enlargement) is real but minor and well-handled.","tokens_in":64995,"tokens_out":1024,"would_cite":true,"duration_ms":102974,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["02.50.-r","02.50.Cw","02.50.Tt"],"model":"glm-5.2","headline":"Robust confidence intervals for weighted estimands","keywords":[],"falsifier":"If the containment condition fails (the estimated class of alternatives does not asymptotically contain the target weights), the robust CI can undercover. More fundamentally, if the heterogeneity UCB is badly calibrated (e.g., due to failure of asymptotic normality or covariance matrix inconsistency), the Bonferroni coverage guarantee breaks down.","tokens_in":64258,"feed_emoji":"⚖️","tokens_out":1064,"duration_ms":199342,"temperature":0.7,"pith_summary":"The paper addresses a problem that arises whenever researchers average group-level effects (like cohort-specific treatment effects in event studies or site-specific effects in experiments) using weights: different readers may prefer different weights, and under heterogeneous effects, different weights yield different conclusions. The author proves that the difference between any two weighted estimands is bounded by the product of two quantities: (1) the heterogeneity in the underlying parameters (how much group-level effects differ from each other, measured via a GLS residual) and (2) the distance between the weight vectors (measured as the standard deviation of the difference in the corresponding estimators). This bound is sharp. Using this decomposition, the author constructs two tools. First, minimax-bias weights that minimize the worst-case distance to any alternative weighting scheme in a user-specified class, yielding a robust point estimator. Second, a robust confidence interval that widens the conventional interval by an amount determined by an upper confidence bound on the heterogeneity (derived via inversion of a noncentral chi-squared distribution) and the maximum distance between the baseline and alternative weights. The resulting interval covers every alternative estimand in the class at confidence level one minus the sum of the conventional significance level alpha and the heterogeneity inference level beta (a Bonferroni-type adjustment). The framework accommodates several practically relevant classes of alternatives: a bounded-variance class (alternative estimands that remain precisely estimable), a truncated simplex class (nonnegative weights that stay close to the baseline), and a covariate-balance class (alternative populations whose covariate profiles do not differ too much from the baseline). Under the bounded-variance class, the robust estimator coincides with the GLS estimator, making it simultaneously optimal for bias and variance. The author establishes uniform asymptotic validity under standard conditions: asymptotically normal estimates, consistent covariance matrix estimation, and consistent estimation of the weights and alternative classes. Two empirical applications illustrate the methods: an event study of school-based internet access in Peru finds results robust to broad classes of nonnegative weights, while Project STAR findss","feed_headline":"One interval to cover every reader's preferred weights","feed_subtitle":"Sharp bounds on how much weighted averages can differ let researchers report confidence intervals valid across entire classes of alternative","key_machinery":"Weighted estimands tau_w(theta) = w'theta; GLS-based heterogeneity measure H(theta) = sqrt(theta'Q theta); weight distance ||lambda - w||_Sigma; noncentral chi-squared inversion for heterogeneity UCB; folded normal critical values for robust CI; minimax-bias weights w* minimizing max distance; bounded variance class, truncated simplex class, covariate balance class","core_discovery":"The central object is the sharp bound on the difference between two weighted estimands: it equals the product of the heterogeneity in parameters (the GLS residual standard deviation of the parameter vector) and the distance between weights (the standard deviation of the difference in estimators). This Cauchy-Schwarz-based decomposition separates the unknown (heterogeneity) from the contested (weight choice), allowing each to be handled independently. The heterogeneity is bounded above using a quantile-unbiased upper confidence bound derived from the noncentral chi-squared distribution of the GLS residual sum of squares. The weight disagreement is controlled by taking the maximum distance to ","pith_inferences":[],"forward_implications":["Researchers can report a single robust confidence interval alongside their conventional one, and readers who disagree with the baseline weights can still draw valid inferences at a known confidence level, formalizing what is currently an informal robustness-check exercise.","The bounded-variance class result shows that GLS (precision-weighted) estimators are not just variance-efficient but also minimax-bias-optimal when the only consensus is that alternative weights should yield estimators of bounded precision, giving GLS a double-optimality that provides a principled default when researchers face ambiguity over weight choice.","The breakdown-value framework (the smallest perturbation at which the robust CI includes a threshold like zero) gives practitioners a scalar summary of robustness analogous to stability concepts in other areas, making weight-sensitivity directly comparable across studies.","The Project STAR application demonstrates that even a well-known randomized experiment can have conclusions sensitive to small weight perturbations, suggesting that external validity concerns are quantifiable and sometimes binding even when internal validity is secure."],"fun_headline_variants":["Weighted estimands: sharp bounds on how much your weights matter","How much does weight choice matter? Sharp bounds on the difference","Bounds on weighted averages that hold across all reasonable weights","One confidence interval valid for any weights a reader might prefer","Weighted estimand differences decompose into heterogeneity times weight distance"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The asymptotic validity requires that the estimated class of alternative weights asymptotically contains the target alternatives. In practice this is handled by using a slightly enlarged class for estimation and then suppressing the enlargement in reported results, but if the enlargement is too small, coverage fails; if too large, intervals are unnecessarily wide.","fun_headline_variants_meta":{"raw":{"variants":["Weighted estimands: sharp bounds on how much your weights matter","How much does weight choice matter? Sharp bounds on the difference","Bounds on weighted averages that hold across all reasonable weights","One confidence interval valid for any weights a reader might prefer","Weighted estimand differences decompose into heterogeneity times weight distance"]},"model":"glm-5.2","effort":"low","cost_usd":0.0,"raw_usage":{"total_tokens":561,"prompt_tokens":493,"completion_tokens":68,"prompt_tokens_details":null},"tokens_in":493,"tokens_out":68,"duration_ms":64544,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T07:59:45.105272+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If the containment condition fails (the estimated class of alternatives does not asymptotically contain the target weights), the robust CI can undercover. More fundamentally, if the heterogeneity UCB is badly calibrated (e.g., due to failure of asymptotic normality or covariance matrix inconsistency), the Bonferroni coverage guarantee breaks down.","supporting_citations":[],"review_version":1}