{"id":"3ef462f1-1af5-4ab2-818a-72aef3df4e50","arxiv_id":"2607.29507","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Winsorized step-down multiple testing procedures control the familywise error rate under adversarial contamination in high-dimensional one- and two-sample mean testing with only 2+ moments.","lead":"Quantile-winsorized statistics are used to build multiple-testing procedures for high-dimensional means that tolerate adversarially corrupted observations, with finite-sample familywise-error-rate bounds and new two-sample Gaussian approximations. The methods let the number of hypotheses grow nearly exponentially in sample size while assuming only slightly more than two moments.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's finite-sample FWER claims all rest on the imported one-sample Gaussian approximation inequalities (Theorems A.1/A.2); any hidden extra condition there—stronger moments or eigenvalue bounds—would invalidate the 'slightly more than two moments' headline.","rationale":"The reader's weakest assumption already identifies the imported one-sample Gaussian approximation theorems as the load-bearing dependency. My reading confirms that these results are not proved or sketched in the present paper, and that the paper's own text describes them as restatements from companion work. The strongest claim—finite-sample strong FWER control with dimension growing exponentially under only slightly more than two moments—is exactly as strong as those imported inequalities. A hidden fourth-moment condition or a missing density assumption in the source would invalidate the headline, and the paper provides no independent check. I also considered the two-sample independence assumption (Assumption 3.2) as a candidate concern, since adversarial contamination of two samples could naturally induce dependence; however, that assumption is explicitly stated and only affects the two-sample extension, whereas the imported results underpin even the one-sample core. Thus the imported results are the single most load-bearing point. The reader's CONDITIONAL verdict already reflects this uncertainty, so no verdict change is warranted; an empirical simulation or a careful verification of the source proofs would be the appropriate next step.","tokens_in":36853,"tokens_out":12808,"duration_ms":129653,"concrete_test":"Obtain Kock and Preinerstorfer (2025a) Theorems 2.1 and 4.1 and re-derive (A.2)–(A.3) under exactly Assumption 2.1, checking that the moment bounds used are only m_X>2 and that the constants in A_X/B_X do not hide a fourth-moment or bounded-density condition. If such a condition is found, the dimension-growth claim fails for 2<m≤4. Complementary: simulate the one-sample procedure with n=200, d=1000, t-distribution with df=2.1, η=0.05 and verify the empirical FWER of Algorithm 3 is consistent with (25).","verdict_should_be":"UNCHANGED","load_bearing_attack":"Every finite-sample FWER bound in the paper—Theorems 2.1, 2.2, 2.4 and, via Theorem 3.4, the two-sample Theorems 3.6–3.8—uses Theorems A.1/A.2, which are restated without proof from Kock and Preinerstorfer (2025a). These imported results are the sole source of the key rates A_X and B_X in Assumption 2.2. The paper is a preliminary version (July 2026) and contains no numerical experiments or code. If either imported theorem secretly requires more than m_X>2 moments (e.g., finite fourth moments) or an extra condition such as a bounded density or a stronger eigenvalue separation than b1, then the exponential-dimension claim collapses: the rates in (20) would not vanish under the stated moment assumption, and the advertised 'slightly more than two moments' guarantee would be false. This is not an internal inconsistency but a serious unverified dependency; the central claim is fully conditional on an external result that the paper does not make available to the reader.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes step-down multiple testing procedures for coordinate-wise null hypotheses about high-dimensional means under adversarial contamination, using quantile-winsorized means and covariance estimators. In the one-sample case, finite-sample FWER bounds are derived from Gaussian approximation inequalities for winsorized means that are restated (without proof) from companion papers; in the two-sample case, new Gaussian approximation inequalities and covariance concentration bounds are developed. Theorems 2.1, 2.2, 2.4, 3.6, 3.7, and 3.8 provide finite-sample upper bounds of the form α + C(A + B [+ C]) that converge to α under Assumptions 2.2 and 3.3. Algorithms 3 and 6 explicitly account for the error from replacing exact Gaussian critical values by bootstrap order statistics, adding a B^{-1} term. The paper is a preliminary version with no numerical experiments.","tokens_in":37225,"tokens_out":8036,"duration_ms":84067,"significance":"If the imported one-sample Gaussian approximation results are correct under the stated moment conditions, the paper makes a useful contribution: it extends robust mean testing to multiple testing with strong FWER control, gives explicit finite-sample bounds that incorporate the cost of bootstrap-based critical values, and provides two-sample Gaussian approximation results that may be of independent interest. The treatment of data-dependent critical values in Propositions B.1 and C.1 is a genuine technical step. However, the advertised dimension-growth claim appears overstated relative to the paper's own rates, and the central one-sample inequalities are not proved or made available in the manuscript, so the significance is conditional on external results.","major_comments":[{"comment":"These one-sample Gaussian approximation inequalities are restated from Kock and Preinerstorfer (2025a) without proof. Every FWER bound in the paper — Theorems 2.1, 2.2, 2.4 and, through Theorem 3.4, the two-sample Theorems 3.6–3.8 — relies on them. The rates A_X, B_X in Eq. (20) are the mechanism by which the bounds converge under Assumption 2.2. Since the companion papers are also preprints, the reader cannot verify that the stated rates require only m_X > 2 moments and the eigenvalue lower bound b_1. Please include proofs of Theorems A.1 and A.2 in an appendix, or at minimum state explicitly and verify that no additional conditions (e.g., finite fourth moments, bounded density, stronger eigenvalue separation) are needed.","section":"Appendix A, Theorems A.1–A.2"},{"comment":"The abstract and Section 2.3.2 claim the procedures allow the number of hypotheses to grow 'exponentially with sample size.' But Assumption 2.2 requires log(d)/n^{(m_X-2)/(5m_X-2)} -> 0, and the rates in Eq. (20) are consistent with this. Since (m_X-2)/(5m_X-2) < 1/5 for every m_X, this permits at most log d = o(n^{1/5}), i.e., d = exp(o(n^{1/5})), which is super-polynomial but not exponential in the usual sense. The headline claim should be corrected to 'super-polynomially' or the dimension condition should be strengthened to actually allow log d = o(n) if that is the intended claim.","section":"Assumption 2.2 / Eq. (20)"},{"comment":"The two-sample results require the contaminated samples (X~_1,...,X~_nX) and (Y~_1,...,Y~_nY) to be independent. Under the stated adversarial contamination model, a single adversary who inspects both original samples before contaminating them may produce contaminated samples that are dependent, even when the original samples are independent. Independence of the original samples does not imply independence of the contaminated samples. Please clarify the threat model: is contamination assumed to be applied independently to each sample? If so, state this explicitly and discuss whether it is standard for two-sample adversarial robustness; if not, the two-sample theorems do not cover the stated model.","section":"Assumption 3.2 / Theorem 3.3"}],"minor_comments":[{"comment":"As noted above, 'exponentially with sample size' should be replaced by a precise statement such as 'super-polynomially' or 'faster than any polynomial rate'.","section":"Abstract / Section 2.3.2"},{"comment":"The manuscript reports no simulations or code. Given that Algorithms 3 and 6 introduce a specific bootstrap order statistic m(B,α,d), a small simulation study or at least a reproducible numerical example would help calibrate B and demonstrate finite-sample FWER behavior. This is not required for the theory but would improve the paper's usefulness.","section":"General"},{"comment":"There is a stray 's' before the bound on |[II]|; the line should read '|[II]| ≤ ...' instead of 's |[II]| ≤ ...'. This is typographical but should be fixed.","section":"Proof of Theorem 3.3 / Eq. (C.1)"},{"comment":"The notation 'w_XY' for sqrt(n_X n_Y/(n_X+n_Y)) is introduced without comment. It is clear from context, but a brief definition would avoid confusion with a product w_X w_Y.","section":"Eq. (30)"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest in stating that the one-sample Gaussian approximation results come from companion preprints, but those results are load-bearing for every theorem in the manuscript. I would not feel comfortable accepting without seeing proofs or confirmation that the rates hold under exactly the stated conditions. The exponential-dimension wording is a significant overstatement relative to Assumption 2.2 and should be corrected before reconsideration. No concerns about authorship or novelty beyond the dependency on the companion papers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: the two-sample Gaussian approximation theory is the real contribution; the one-sample results are repackaged versions of the authors' own companion papers, and every finite-sample FWER bound depends on those unproved-here theorems. Worth refereeing, but the referee must check the imports.\n\nWhat's new and good: Theorem 3.3, a two-sample Gaussian approximation for winsorized means, obtained by a conditioning argument; the covariance estimator concentration in Proposition 3.1; and the explicit handling of bootstrap-based critical value errors (Lemma 2.3, Theorems 2.4 and 3.8) which adds a clean B^-1 term. That last piece is genuinely useful—most bootstrap multiple testing theory ignores the discretization step. Propositions B.1 and C.1, which relate data-dependent quantiles to fixed quantiles, are careful and nontrivial. The step-down algorithms are standard Romano-Wolf, but the analysis of them in this contaminated setting is new.\n\nSoft spots: the load-bearing Theorems A.1/A.2 are restated from Kock and Preinerstorfer (2025a) without proof. The stress-test note is right that a hidden stronger moment condition there would affect the 'slightly more than two moments' claim. That is not evidence of error, but it is a serious unverified dependency. Also, there are no simulations, no real data, no code. For a methods paper this is a genuine gap; the winsorization tuning parameters are only supported by heuristics. Finally, it is a preliminary version; notation is heavy and constant tracking is hard, but those are minor.\n\nFor whom: someone working in robust high-dimensional inference or multiple testing will want to know these tools. I would not adopt the methods until the companion results are verified or simulations are provided. But the new two-sample theory and the bootstrap slack analysis deserve a serious referee.","headline":"Two-sample winsorized Gaussian approximation and bootstrap-critical-value control are the real contributions; the one-sample FWER bounds are repackaged from unproved companion results, so the headline moment/dimension claim is only as good as those imports.","tokens_in":37660,"tokens_out":3517,"would_cite":true,"duration_ms":37847,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H15","62G35","62G20","62E17"],"pacs":[],"model":"deepseek-v4-flash","headline":"Quantile-winsorized step-down multiple testing controls the familywise error rate in finite samples under adversarial contamination, with bounds approaching the nominal level even as the dimension grows super-polynomially under only slightl","keywords":["adversarial contamination","multiple testing","familywise error rate","winsorized means","high-dimensional Gaussian approximation","two-sample mean testing","bootstrap critical values","heavy-tailed distributions"],"falsifier":"Simulate one-sample data with exactly m=2.5 moments, choose d≈exp(n^{0.2}) and η≈n^{-0.4} so the stated regime holds, and compute the sup-over-hyperrectangles distance in the imported Gaussian approximation (Theorem A.2). If that distance does not vanish at the claimed rate, or if the empirical FWER of Algorithm 3 with a large B exceeds α plus a constant times the paper's error terms, the central claim collapses. A more targeted check is the covariance concentration bound: verify whether max_{j,k}|Σ̂_{j,k}-Σ_{j,k}| exceeds C·d_n with probability no greater than 24/n in the one-sample case.","tokens_in":36778,"feed_emoji":"🛡️","tokens_out":8396,"duration_ms":90121,"temperature":0.7,"pith_summary":"Adversarial contamination usually destroys tests built on coordinate averages; the paper claims that quantile-winsorized means and covariance estimators restore multiple testing guarantees. It builds step-down procedures whose probability of missing at least one true null hypothesis is at most the nominal level α plus explicit finite-sample terms A, B, C (and 1/B when bootstrap quantiles are used). These terms go to zero in regimes where the number of hypotheses d can grow faster than any polynomial in the sample size and the number of contaminated observations can diverge, while the data only need slightly more than two moments. The same construction is extended to two-sample equality-of-means testing, including a new Gaussian approximation for winsorized two-sample statistics and covariance concentration inequalities. A practitioner who trusts the imported one-sample Gaussian approximation gets robust, feasible algorithms with strong familywise error control, not just a heuristic.","feed_headline":"Winsorized tests keep familywise error at alpha under contamination","feed_subtitle":"Step-down procedures built on clipped means work with exponentially many hypotheses and just over two moments.","key_machinery":"The engine is coordinate-wise quantile winsorization: every coordinate of each contaminated vector is clipped at data-dependent order-statistic thresholds chosen slightly beyond the assumed contamination fraction, and the same clipping is used to build both a location estimator and a covariance estimator. A Gaussian approximation inequality over hyperrectangles transfers the distribution of the centered, self-normalized winsorized mean vector to a Gaussian with the true covariance (or correlation) matrix up to an explicit rate; a covariance concentration bound extends this to the self-normalized statistic. The step-down algorithm then compares test statistics against critical values that are","core_discovery":"The paper's central claim is that step-down multiple testing procedures based on quantile-winsorized statistics achieve strong familywise error rate control under adversarial contamination. For a target level α, the probability that the procedure's reported set excludes at least one genuinely null coordinate is bounded by α plus explicit finite-sample error terms; for the bootstrap implementation the additional term is B^{-1}. The bounds are uniform over distributions whose coordinates have only m>2 moments and whose variances are bounded away from zero and above, and they converge to α as n grows in regimes where √(n log d) η^{1−1/m}→0 and log(d)/n^{(m−2)/(5m−2)}→0. In the two-sample case t","pith_inferences":["The step-down construction could be repurposed to build simultaneous confidence intervals for the coordinates of a high-dimensional mean, since the FWER bound is uniform over the unknown set of true nulls; the paper does not pursue this.","The two-sample approximation is assembled by conditioning on one sample and reusing the one-sample result, which suggests a recursive template for k-sample or stratified robust tests, but the paper stops at two samples.","The finite-sample bounds are upper bounds; whether practical FWER is much smaller than α and how much power the winsorization costs near m=2, or with contamination at the boundary η≈1/2, are quantitative questions the paper does not answer and that simulation could settle.","The tuning parameters λ are user-chosen; the paper inherits practical defaults from numerical experience, so a systematic sensitivity analysis of the λ choices would be a natural follow-up."],"forward_implications":["Under Assumption 2.2 the one-sample algorithms' FWER upper bounds converge to α, so the step-down procedures become asymptotically exact while d grows faster than any polynomial and the number of contaminated observations can diverge.","Because only m>2 moments are required, the same guarantees apply to heavy-tailed data without sub-Gaussian or sub-exponential tail conditions.","The two-sample Gaussian approximation and covariance concentration inequalities imply finite-sample size control for robust global tests of equality of high-dimensional means, including a feasible bootstrap version.","The bootstrap algorithms (3 and 6) explicitly charge B^{-1} for estimating critical values, so a user can make the cost as small as desired by increasing B; the tabulated values suggest B = 5,000 or 10,000 makes the extra term negligible.","All six procedures provide strong control—the error is bounded no matter which set of coordinates is truly null—not merely weak control under the global null."],"fun_headline_variants":["Winsorized step-down tests beat contamination at scale","Clipped means give strong error control for many tests","Two moments enough for robust high-dimensional testing","Exponential hypotheses, no more than two moments: robust tests"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole edifice rests on the imported one-sample Gaussian approximation inequalities for quantile-winsorized means and their stated rates; if those inequalities are wrong, or secretly need stronger moment or eigenvalue conditions, the finite-sample FWER bounds and the super-polynomial-dimension convergence claims do not follow.","fun_headline_variants_meta":{"raw":{"variants":["Winsorized step-down tests beat contamination at scale","Clipped means give strong error control for many tests","Two moments enough for robust high-dimensional testing","Exponential hypotheses, no more than two moments: robust tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000506,"raw_usage":{"total_tokens":2246,"prompt_tokens":627,"completion_tokens":1619,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":371,"completion_tokens_details":{"reasoning_tokens":1555}},"tokens_in":371,"tokens_out":1619,"duration_ms":14372,"temperature":1.0,"reasoning_tokens":1555,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T05:36:42.815183+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate one-sample data with exactly m=2.5 moments, choose d≈exp(n^{0.2}) and η≈n^{-0.4} so the stated regime holds, and compute the sup-over-hyperrectangles distance in the imported Gaussian approximation (Theorem A.2). If that distance does not vanish at the claimed rate, or if the empirical FWER of Algorithm 3 with a large B exceeds α plus a constant times the paper's error terms, the central claim collapses. A more targeted check is the covariance concentration bound: verify whether max_{j,k}|Σ̂_{j,k}-Σ_{j,k}| exceeds C·d_n with probability no greater than 24/n in the one-sample case.","supporting_citations":[],"review_version":1}