{"id":"55bfb523-4bd4-485d-a35a-0840db412fa9","arxiv_id":"2607.21481","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A sampling-based confidence interval maintains nominal coverage and n^{-1/2} length for multiple treatment effects in linear IV models with possibly invalid instruments, under generalized majority/plurality conditions.","lead":"This paper gives a method for estimating multiple treatment effects in instrumental-variable studies when some instruments are invalid. It proposes sampling-based confidence intervals that keep nominal coverage even when invalid instruments are hard to detect.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 3's parametric-rate length guarantee rests on Assumption 3(39), a high-level local-plurality condition not implied by the identification conditions and not verifiable from data; when it fails, the SCI can include intervals centered at false effects and its length need not shrink at the param","rationale":"The reader identified Assumption 3, specifically inequality (39), as the weakest assumption, and I agree that it is the most load-bearing concern. Theorems 2 and 3 are the paper's central claims. Theorem 2's coverage guarantee is relatively robust and appears sound given Assumptions 1–2 and the generalized majority rule; it does not require (39). Theorem 3's parametric-rate length guarantee, however, is exactly the efficiency property that makes the SCI practically useful, and it rests entirely on (39). My constructed DGP satisfies the identification condition (16) and Condition 2 while violating (39), so the length result can fail under assumptions that are otherwise plausible. This is not an internal inconsistency—the theorem is stated conditionally—but it substantially narrows the scope of the central claim. The paper does not provide primitive sufficient conditions for (39), nor a data-based check, so practitioners cannot know whether the reported length guarantee applies. The adaptive C0 rule in Remark 4 is a second, real gap: the proof fixes C0, while the implementation chooses it adaptively; but this is secondary because it concerns implementation fidelity rather than the theoretical claim itself. I therefore recommend keeping the reader's CONDITIONAL verdict. The paper is valuable and the coverage result is credible, but the length guarantee should be presented as conditional on a strong, unverifiable local-plurality condition, and the implementation should be aligned with the theory or the theory extended to the adaptive rule.","tokens_in":36114,"tokens_out":18640,"duration_ms":180426,"concrete_test":"Simulate the DGP described above: p_z=10, p_d=2, |V*|=6 valid IVs, five invalid IVs whose hyperplanes intersect exactly at a false β^false, and two additional invalid IVs whose hyperplanes pass within C/√n of β^false; ensure all just-identifying subsets have full column rank. Apply Algorithm 1 with M=1000 and n=10^4, 10^5, 10^6. If the SCI length Len(α) does not decay as O(1/√n) but instead stabilizes at a positive constant comparable to ||β^false − β*||, then Assumption 3(39) is binding and Theorem 3's length guarantee does not extend to this identification-valid setting. A secondary check: run the same DGP with a fixed C0 set to the theoretical value in the proof of Proposition 3 versus the adaptive C0 rule from Remark 4; if coverage differs, the implementation is not covered by the theorem.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim includes the length guarantee of Theorem 3 (Eq. 41). That guarantee is secured entirely by Assumption 3, specifically inequality (39). This condition is strictly stronger than the generalized plurality rule: it replaces exact zeros in the voting set with a local window of width C1 log(C2/α0)/√n. It is not derived from more primitive conditions on the data-generating process, and it cannot be verified from the sample because the centering vectors π^(ℓ) are population limits associated with each just-identifying subset. A concrete failure is easy to construct. Take |S*|=10, pd=2, |V*|=6 valid IVs, so the generalized majority rule (16) holds because 6 > (10+1)/2 = 5.5. Arrange five invalid IVs to pass exactly through a false effect β^false, and two additional invalid IVs to pass within C/√n of β^false. Then Condition 2 holds: the true β* has 6 exact votes, the false β^false has 5 exact votes. But the left side of (39) for the strongly invalid just-identifying subset at β^false is at least 7 (two IVs in the subset plus five near-zero π^(ℓ)_k), which exceeds 5.5. The condition in (39) fails. The corresponding strongly invalid subset passes the majority threshold (33), so its TSLS interval is included in the SCI, and that interval is centered near β^false, far from β*. Consequently Len(α) is O(1), not O(log/√n). This shows the length guarantee is not a byproduct of identification but an extra, high-level, untestable assumption. The paper acknowledges (40) is 'slightly stronger' than Condition 2, but it does not offer primitive conditions or a way to check (39). The adaptive C0 procedure in Remark 4 is a secondary gap: the theorems are proved for a fixed C0, and the data-dependent rule is not shown to select a C0 satisfying the proof's requirements. The primary load-bearing concern remains (39).","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the linear instrumental variable model with multiple endogenous treatments and possibly invalid instruments. It introduces generalized plurality and majority rules for identifying the vector of treatment effects, emphasizing the geometric distinction from the single-treatment case. For inference, it proposes a Sampling Confidence Interval (SCI) that aggregates TSLS intervals over perturbed estimates of instrument validity, with the aim of remaining valid even when hard-thresholding cannot separate locally invalid from valid instruments. The main theoretical results are Theorem 2 (asymptotic coverage of the SCI under a generalized majority condition) and Theorem 3 (parametric-rate length under Assumption 3). The paper also reports simulations and a Mendelian randomization application.","tokens_in":1664,"tokens_out":2077,"duration_ms":105613,"significance":"The identification results are a useful extension of single-treatment plurality logic to multiple treatments, with a clear geometric interpretation and a transparent sufficient condition. The sampling scheme is a nontrivial generalization of Guo (2023) that avoids multidimensional grid search, and the supplementary proofs are detailed. However, the advertised parametric-rate length guarantee rests entirely on Assumption 3, Eq. (39), a high-level local-plurality condition that is neither derived from primitive conditions nor verifiable from data. In addition, the adaptive tuning of C0 in Remark 4 does not match the fixed constant used in the proof of Proposition 3. These issues prevent the paper from fully establishing its central efficiency claim, although the coverage theorem is more robust.","major_comments":[{"comment":"The parametric-rate length guarantee (41) is obtained only under a high-level condition not implied by the identification conditions. A concrete failure: |S*|=10, p_d=2, |V*|=6, so the generalized majority rule (16) holds. Let five invalid IVs pass exactly through a false effect beta_f, and two further invalid IVs have |pi_k^(ell)| <= C1 log(C2/alpha0)/sqrt(n) for the subset ell at beta_f. Then Condition 2 holds, but the left side of (39) is at least 7 > 5.5, so Assumption 3 fails. The corresponding TSLS interval is then included in the SCI and is centered near beta_f, so Len(alpha)=O(1), not O(log/sqrt(n)). Thus the length claim is not a byproduct of identification; it is an extra, untestable local-plurality assumption. Please derive (39) from primitives or reframe the length guarantee as conditional.","section":"Section 4, Assumption 3, Eq. (39); Theorem 3"},{"comment":"The proof sets a specific constant C0 in rho_n(M) that depends on alpha0 and the dimension, but Algorithm 1 uses an adaptive rule: start at 0.05 and multiply by 1.25 until more than 5% of the generalized-majority selections are nonempty. No argument shows that this adaptive choice preserves the event in Proposition 3 or the coverage/length bounds. The simulations and application implement the adaptive rule, so the reported performance is not the procedure for which theorems are proved. Either prove guarantees for the adaptive C0 or fix C0 as in the proof and report sensitivity.","section":"Section 3.3, Remark 4 and proof of Proposition 3"},{"comment":"Assumption 2(b) requires every just-identifying subset H in H* to have eigenvalue lambda_min(Upsilon_H^T Upsilon_H) bounded away from zero. This is much stronger than Condition 1, which only requires the full valid set to have full column rank. In Mendelian randomization, valid IVs are often weak individually, so there may exist just-identifying valid subsets with very small eigenvalues under Condition 1. Under such subsets, the TSLS estimator is not sqrt(n)-consistent, and Proposition 2, Proposition 3, and Theorem 2 are unavailable. Remark 5 acknowledges weak identification is out of scope, but the abstract and title do not state this restriction. It should be stated prominently.","section":"Section 4, Assumption 2 and Remark 5"}],"minor_comments":[{"comment":"The symbol M is used both for the number of resamples and for the filtered set in (28). Use a distinct notation such as N to avoid confusion.","section":"Section 3.3, Eq. (28)"},{"comment":"The definition of locally invalid IVs uses the TSHT candidate set Lhat from (26), but the sampling procedure later uses different candidate sets bT_m. Clarify whether the concept is relative to TSHT or to the SCI procedure.","section":"Section 3.2, Eq. (27)"},{"comment":"The method enumerates all subsets of size p_d from S_hat. For small p_z this is fine, but the paper does not discuss computational feasibility for larger p_z. A brief complexity discussion would be helpful.","section":"Section 5-6"},{"comment":"The constants C1 and C2 in (38)-(39) are introduced as absolute constants but never specified or connected to primitive parameters. The theorem would be more informative if these constants were made explicit or shown to depend only on identifiable quantities.","section":"Section 4, Assumption 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is potentially publishable after substantial revision. The main issue is that the parametric-rate length claim is conditional on a high-level, untestable local-plurality condition, and the implementation's adaptive C0 choice is not covered by the proofs. The editor should ask for either primitive conditions for (39) or a reframed, weaker length guarantee, as well as alignment between the theoretical constant and the implemented algorithm. The identification section and coverage theorem are valuable and should be preserved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth a serious look. The genuine novelty is the sampling confidence interval (SCI) for multiple treatment effects with possibly invalid instruments: instead of a grid search, it samples perturbations of the IV-validity estimates generated by just-identifying subsets, then aggregates TSLS intervals. That is a real step beyond Guo (2023), and it is implemented carefully. The identification results are mostly a clean geometric reframing of Andrews (1999), which the authors disclose, and the h0-quantified majority rule (15) is a useful way to track the rank structure of the relevance matrix. The coverage proof (Theorem 2) is sound, and Section C is detailed enough to check.\n\nThe soft spots are in the efficiency claims. Theorem 3 says the SCI shrinks at the parametric rate, but that relies entirely on Assumption 3(39), a high-level local-plurality condition that is not derived from primitives and cannot be verified from data. The stress-test example is valid: with 10 relevant IVs, 2 treatments, and 6 valid IVs, you can have 7 invalid IVs align near a false effect, so the majority threshold (33) admits intervals centered away from beta*, and the length is O(1), not O(log n / sqrt n). The paper calls (40) \"slightly stronger\" than Condition 2, but it is actually a substantial, untestable additional constraint. The adaptive C0 rule in Remark 4 is also not what the proofs use: Proposition 3 specifies a particular C0, while the implementation searches over C0 values until nonempty sets appear. That gap should be closed or the theory restated to match the implementation. Minor: the simulations do not report Monte Carlo error bars or replication counts, which makes the coverage comparisons harder to judge.\n\nNone of this invalidates the core idea. Theorem 2 shows coverage holds without Assumption 3, so the SCI is still a conservative, selection-robust interval. The parametric-rate length claim is oversold, but the method is useful, especially for multivariable Mendelian randomization where weak local violations are common. I would send this to a serious referee, but the referee should insist that the authors either derive (39) from more primitive conditions, or explicitly present Theorem 3 as conditional on a plausibility assumption, and align the tuning procedure with the proof. It is a credible, careful paper, and with that revision it would be a solid contribution to the IV literature.","headline":"Solid extension of single-treatment IV robustness to multiple exposures, with a genuinely new sampling CI, but the parametric-rate length claim leans on an untestable high-level condition (Assumption 3) that can fail; coverage is fine.","tokens_in":37125,"tokens_out":1989,"would_cite":true,"duration_ms":22335,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A generalized plurality rule identifies multiple treatment effects with invalid instruments, and a sampling confidence interval keeps nominal coverage at parametric rate.","keywords":["causal inference","invalid instruments","multiple treatments","generalized plurality rule","generalized majority rule","sampling confidence interval","locally invalid instruments","Mendelian randomization"],"falsifier":"Rerun the paper's own simulation design (S1) with n=5000 and τ=0.05, where one instrument is locally invalid and hard thresholding undercovers badly: the sampling interval must cover the true effect at about the nominal 95% level while its length shrinks with n. If empirical coverage is materially below 95% or the interval length does not shrink at the 1/√n rate, Theorems 2 and 3 are falsified. Alternatively, construct a design satisfying (39) but with p_d=2, p_z=9, and 5 valid instruments, and check that the SCI length remains O(1/√n).","tokens_in":35914,"feed_emoji":"📊","tokens_out":10901,"duration_ms":85675,"temperature":0.7,"pith_summary":"Instrumental-variable (IV) methods estimate causal effects when unmeasured confounders exist, but they require instruments that do not affect the outcome except through the treatment. This paper deals with the common situation of several treatments and some invalid instruments, where the classical single-treatment majority rule no longer identifies the effects. The authors show that identification holds under a 'generalized plurality rule' — the true effect vector must be supported by strictly more instrument-defined hyperplanes than any competing vector — and that a simpler 'generalized majority rule' on the count of valid instruments guarantees it. For finite samples, they propose a sampling confidence interval that perturbs the estimated instrument-validity measures, keeps only selections satisfying the majority rule, and unions the resulting two-stage least-squares intervals. If the paper is right, applied researchers can obtain honest confidence intervals for multiple causal effects without knowing which instruments are valid, even when some invalid instruments are nearly indistinguishable from valid ones.","feed_headline":"Sampling intervals fix instrument-selection errors at parametric rate","feed_subtitle":"Perturbing validity scores keeps coverage honest when hard thresholding cannot spot locally invalid instruments.","key_machinery":"The key geometric object is the hyperplane voting interpretation of instruments: in a multiple-treatment linear IV model, each relevant instrument defines a (p_d−1)-dimensional hyperplane in the p_d-dimensional effect space, and an instrument votes for any candidate effect lying on its hyperplane. The generalized plurality rule (Condition 2) requires β* to lie on more hyperplanes than any other candidate; Theorem 1 shows that the count condition |V*| > (|S*| + h_0 − 1)/2 suffices, where h_0 encodes the rank structure of the relevance matrix and reduces to p_d in the benign case. For inference, the mechanism is 'perturb-and-aggregate': Gaussian perturbations of the validity estimates generate","core_discovery":"On the paper's own terms, the central claim is that for p_d ≥ 1 treatments and p_z instruments, the vector of causal effects is identified by the generalized plurality rule, which says the true effect β* receives votes from more instruments than any other candidate effect, where each instrument 'votes' for every candidate on its hyperplane L_k = {β : Υ_{k·}(β−β*) = π_k}. A sufficient, easy-to-check condition is the generalized majority rule: |V*| > (|S*| + p_d − 1)/2, where V* is the set of truly valid instruments and S* the set of relevant ones. For inference, the paper constructs a sampling confidence interval (SCI): for each just-identifying subset it estimates the validity of every instr","pith_inferences":["The paper leaves open whether replacing the Gaussian perturbation with a bootstrap or wild bootstrap would improve finite-sample coverage; that trade-off is not explored.","The same perturb-and-aggregate scheme could be adapted to other post-selection inference problems, such as confidence intervals after moment selection in GMM, wherever a plurality condition can be formulated.","If the local plurality condition (39) fails, the length guarantee breaks down; an adaptive widening of the SCI could restore honesty at the cost of conservatism.","Proposition 1 suggests that under random coefficients, the generalized plurality rule holds almost surely once |V*| > p_d, implying that the majority rule may be unnecessarily restrictive for identification."],"forward_implications":["Applied researchers can report honest confidence intervals for multiple endogenous treatments without knowing which instruments are valid, as long as the valid instruments exceed the majority bound.","The method works with summary-level Mendelian randomization data, where only SNP-exposure and SNP-outcome coefficients and standard errors are available.","The SCI avoids a full grid search over the effect space by searching over just-identifying subsets, keeping the computation feasible for moderate dimensions.","The parametric-rate length guarantee means the procedure does not achieve coverage by inflating the interval width.","For a single treatment the generalized majority rule reduces to the classical majority rule, so the results unify single- and multi-treatment IV inference."],"fun_headline_variants":["Plurality rule identifies causal effects with invalid instruments","Sampling CIs beat selection errors at parametric rate","Robust CIs for multiple treatments under invalid IVs","Generalized plurality and sampling fix invalid instruments"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The procedure's parametric-rate length guarantee rests on Assumption 3's inequality (39), which says that for every just-identifying subset with strongly invalid instruments, at most (|S*|+p_d−1)/2 instruments can have centering bias inside the local-to-zero window C_1 log(C_2/α0)/√n.","fun_headline_variants_meta":{"raw":{"variants":["Plurality rule identifies causal effects with invalid instruments","Sampling CIs beat selection errors at parametric rate","Robust CIs for multiple treatments under invalid IVs","Generalized plurality and sampling fix invalid instruments"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000639,"raw_usage":{"total_tokens":2764,"prompt_tokens":715,"completion_tokens":2049,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":459,"completion_tokens_details":{"reasoning_tokens":1987}},"tokens_in":459,"tokens_out":2049,"duration_ms":12526,"temperature":1.0,"reasoning_tokens":1987,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T07:18:34.134645+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the paper's own simulation design (S1) with n=5000 and τ=0.05, where one instrument is locally invalid and hard thresholding undercovers badly: the sampling interval must cover the true effect at about the nominal 95% level while its length shrinks with n. If empirical coverage is materially below 95% or the interval length does not shrink at the 1/√n rate, Theorems 2 and 3 are falsified. Alternatively, construct a design satisfying (39) but with p_d=2, p_z=9, and 5 valid instruments, and check that the SCI length remains O(1/√n).","supporting_citations":[],"review_version":1}