{"id":"82e745b0-c027-4bde-89e4-71cb987de83a","arxiv_id":"2607.14812","paper_version":3,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"No universal multiplicative FDR bound holds for Benjamini-Hochberg under Gaussian dependence: worst-case FDR/q diverges like sqrt(log(1/q)), with exact constants in natural cases.","lead":"This paper shows that, for Gaussian test statistics with arbitrary correlations, the Benjamini-Hochberg procedure's false discovery rate can exceed its nominal level by an amount that grows without bound relative to that level as the level shrinks. It constructs explicit finite Gaussian models and proves matching upper bounds for a one-factor dependence class.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified; the main load-bearing step—regular first crossing of the limiting BH cutoff—is verified in every regime and appears sound.","rationale":"The reader's weakest-assumption analysis identified the regularity of the limiting BH crossing as the key risk; my independent read agrees. The paper addresses this risk with extensive local-monotonicity and derivative calculations, and I found no concrete failure. The two-step limit structure is legitimate: for fixed r, Theorem 2.8 gives convergence of the conditional FDP to D_r(Z) at regular z, and dominated convergence gives E[FDP_N]→E[D_r(Z)]; then Lemma 2.21 gives D_r(Z)→D*(Z) pointwise off a measure-zero set. The compact-uniform regularity proofs in Lemmas 2.15, 2.19, and 2.20 are sufficient because any non-exceptional z lies in some compact subset of its regime. The upper bounds in Section 4 are consistent with the lower-bound rates and do not affect the validity of the lower-bound construction. One editorial caveat: the manuscript's notation for \\Phi is used inconsistently in places (the two-sided p-value and the tail-ratio lemmas require \\Phi to be the upper-tail survival function, while the text calls it the standard normal CDF). This should be clarified, but it does not change the mathematical argument, which goes through with the survival-function convention. No ad hominem or manufactured concern is intended; the verdict remains unchanged.","tokens_in":49254,"tokens_out":42841,"duration_ms":350356,"concrete_test":"Run an independent interval-arithmetic or symbolic verification of the derivative positivity assertions in Lemma 2.19 (Eq. 114) and Lemma 2.20 (Eq. 132), starting from the displayed formulas (57), (64), (110). For q in, say, {0.001, 0.01, 0.1} and r=10^{-4}, compute \\bar R_{r,z}(c) and its derivative on the stated local intervals and confirm: (i) sup_{c<C_r} \\bar R_{r,z}<1, (ii) inf_{c>C_r} \\bar R_{r,z}>1, and (iii) the derivative is strictly positive on the interval required by Proposition 2.5. A sign violation at any tested point would indicate that the regularity step underpinning Theorem 2.8 needs revision; a clean pass would further support the current verdict.","verdict_should_be":"UNCHANGED","load_bearing_attack":"No significant objection identified. The central claim rests on a two-step limit: first N→∞ for fixed r (Theorem 2.8), then r→0 (Lemma 2.21), with dominated convergence to E[D*(Z)]. The load-bearing technical premise is that C_r(z) is a regular first crossing in the sense of Definition 2.4 for almost every z, since the finite-N FDP convergence proof uses the bracketing (30)–(31) and the identity D_r(z)=\\bar R_{0,r,z}(C_r(z)) only at regular crossings. I checked the regularity verifications: Lemma 2.15 (z>−1), Lemma 2.19 (−y_q<z<−1), Lemma 2.20 (z<−y_q), and the one-sided analogues in Lemmas 3.11–3.13. Each regime establishes derivative positivity on a local interval and uses Proposition 2.5, with the required uniformity on compacta. Since the two exceptional points {−1,−y_q} have measure zero, pointwise convergence for all other z suffices for dominated convergence. I did not find a hidden gap or an unverified assumption in these checks. The construction contains no fitted constants, and the upper bounds independently confirm the qualitative rate. The proof is long and not machine-checked, so the residual risk is confined to a possible subtle error in one of the long derivative estimates, but I did not locate one.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the worst-case FDR of the Benjamini–Hochberg procedure for Gaussian z-tests under arbitrary correlation. For two-sided tests it constructs a q-indexed family of finite Gaussian models with FDR at least an explicit ℓ=(q)>q, where ℓ=(q)=q√(log(1/q))/(2√π)+cℓ q+o(q); for one-sided tests, a sign-reversed common-factor construction gives ℓ≤(q)=q√(log(1/q))/√π+q/2+o(q). Together these disprove any universal multiplicative bound Cq. The paper also proves common-factor upper bounds: O(q√log(1/q)) for two-sided tests and q√log(1/q)/√π+O(q) for one-sided tests, the latter matching the one-sided lower constant. The proofs use a random two-level design, a regular-first-crossing analysis of the limiting BH cutoff via DKW-type empirical-CDF convergence, and Gaussian tail expansions; the lower constants are explicit functions of q and the standard normal distribution, with no fitted parameters.","tokens_in":848,"tokens_out":4348,"duration_ms":197535,"significance":"If correct, this settles a long-standing question: BH does not control FDR under arbitrary Gaussian dependence, and the inflation is worse than multiplicative. The explicit lower bounds are concrete and falsifiable, and the one-sided common-factor constant is proven sharp. The 'regular first crossing' machinery is a useful technical contribution. I checked the load-bearing regularity verifications (Lemmas 2.15, 2.19, 2.20, 3.11–3.13) and the dominated-convergence argument; I did not find a gap. The two-sided upper and lower constants are not matched, but the paper states this. The main fixes needed are presentational, especially notation.","major_comments":[{"comment":"The paper's notation for Φ is internally inconsistent. The text defines Φ and φ=Φ′ as the standard normal CDF and density, but then uses Φ(x) as the upper-tail survival function in p(x)=2Φ(x)=2{1−Φ(x)}, p+(x)=Φ(x)=1−Φ(x), λ(x)=φ(x)/Φ(x), and throughout the tail expansions. In particular, Eq. (5) defines ℓ=(q) with qΦ(1)+... and Φ(yq); using the stated CDF convention, ℓ=(q) can exceed 1 for some q∈(0,1), which is impossible for an FDR lower bound. This notational conflict affects Eqs. (141), (154), (161), (185), Lemma A.2, and many other places. The central proofs appear to be correct when read with Φ as the upper-tail function, and Eq. (9) requires the CDF elsewhere. The manuscript must be revised to introduce, say, Φ̄=1−Φ for the tail and use it consistently. This is a load-bearing notational issue because the main theorem's exact formula and the numerical tables are otherwise uninterpr","section":"§1, Eq. (1), (5), (9), (141); Appendix A"}],"minor_comments":[{"comment":"The constant cℓ=0.6492828... is reported with no indication of the quadrature method or error bound. Since cℓ is not needed for the divergence claim, this is a presentation issue, but a brief note on numerical accuracy would help.","section":"§2.7, Proposition 2.2"},{"comment":"The condition q≤2Φ(1)≈0.3173 should be restated after the Φ notation is fixed. It currently reads oddly because, under the paper's stated CDF convention, 2Φ(1) would be about 1.6826, not 0.3173.","section":"§4.2, Theorem 4.2"},{"comment":"The quoted phrase 'convincing simutheoretical evidence' should be marked as a quotation error or 'sic'.","section":"§1"}],"recommendation":"minor_revision","confidential_remarks":"The paper is substantial and the central argument appears sound. The only substantive concern is the pervasive inconsistency between CDF and survival notation for Φ; this is easily repaired but must be fixed before the exact formulas can be trusted as written. The regularity checks, which are the main technical risk, appear to withstand scrutiny. I would be comfortable with acceptance after a careful notation revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The punchline: this paper settles the long-running question of whether BH controls FDR under arbitrary Gaussian dependence. It does not just give a single counterexample—it constructs, for every q, finite Gaussian designs where the FDR exceeds q, and shows the ratio diverges like sqrt(log(1/q)) as q goes to 0. That kills any hope of a universal multiplicative bound Cq. For one-sided tests it even pins down the leading constant in the common-factor class.\n\nWhat is new and good: Dobriban (2026) had one counterexample at q=0.01. Lei gives a full q-indexed family, proves strict anticonservativeness for every q, and derives explicit small-q asymptotics. The two-sided lower bound is q sqrt(log(1/q))/(2 sqrt(pi)) plus a computable constant; the one-sided version is q sqrt(log(1/q))/sqrt(pi) + q/2. On the upper side, Theorem 4.2 and Corollary 4.5 give matching rates for the one-common-factor class, with constants that are sharp for one-sided tests. The construction is genuinely parameter-free: secondary masses solve an integral equation, not fit the target quantity. The one-sided mechanism is also illuminating—positive null–null correlations, negative null–non-null correlations, so the p-value vector is not PRDS, and the paper says so directly.\n\nSoft spots: the two-sided upper and lower leading constants do not match (1/sqrt(pi) versus 1/(2 sqrt(pi))), so the exact worst-case constant for two-sided tests remains open. The paper is explicit about this. The main risk is technical: the proof is long and relies on verifying that the limiting BH cutoff is a regular first crossing in several regimes. The regularity checks are careful and I did not find a gap, but the length means a subtle derivative estimate could be wrong. It is not machine-checked. That is the residual concern, not a conceptual one.\n\nWho it is for: anyone working on multiple testing under dependence, and practitioners who need to know whether BH's nominal level is trustworthy with correlated Gaussian tests. It is a cautionary result with real practical implications.\n\nRecommendation: send it to peer review. The claims are concrete, the constructions are explicit, and a referee can focus effort on the few long derivative lemmas. This deserves a serious referee, even though the proof is not machine-checked.","headline":"Disproves the long-open Gaussian FDR conjecture with sharp asymptotics; the proof is heavy but the central claims check out.","tokens_in":50037,"tokens_out":1972,"would_cite":true,"duration_ms":19384,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62F05","62H15"],"pacs":[],"model":"deepseek-v4-flash","headline":"For correlated Gaussian tests, the Benjamini-Hochberg FDR can exceed any fixed multiple of q, with the ratio growing like sqrt(log(1/q)).","keywords":["Benjamini-Hochberg procedure","false discovery rate","Gaussian dependence","two-sided tests","one-sided tests","one-common-factor model","worst-case FDR","positive regression dependence"],"falsifier":"Evaluate the normalized mixture CDF Rbar_z(c) on a fine grid in z around the boundary points -1 and -y_q (and around 0 and y_q^+ in the one-sided construction). If for some z in a set of positive measure the curve touches 1 before its first crossing or fails to exceed 1 immediately after, the regular-first-crossing condition fails and the dominated-convergence step cannot be applied; equally, a direct simulation of the paper's finite model at these parameters that clearly violates the stated lower bound would falsify the construction.","tokens_in":49199,"feed_emoji":"📊","tokens_out":7384,"duration_ms":64772,"temperature":0.7,"pith_summary":"The paper asks how far the Benjamini-Hochberg (BH) procedure's false discovery rate can rise above its nominal level q when test statistics are Gaussian with arbitrary correlations. It proves that for both two-sided and one-sided Gaussian tests, the worst-case FDR is at least an explicit function of q that is strictly larger than q for every q and whose ratio to q diverges as q tends to 0. This overturns a long-standing folklore conjecture that BH controls FDR for correlated two-sided Gaussian z-tests, and it rules out any universal multiplicative bound of the form Cq. The lower bounds are achieved by finite Gaussian models with covariance matrices of one-common-factor form, and matching upper bounds in that class show the q sqrt(log(1/q)) order is sharp; for one-sided tests the leading constant is determined exactly.","feed_headline":"Correlated Gaussian tests can push BH FDR past any fixed multiple of q","feed_subtitle":"Two-sided and one-sided Gaussian tests give worst-case FDR/q → ∞ as q → 0, disproving the Cq conjecture.","key_machinery":"The engine is a family of finite Gaussian designs (a common factor plus independent noise) with three population components: nulls with tiny loading, a point-mass primary non-null, and a continuum of 'secondary' non-nulls that shapes the conditional p-value CDF. The analysis centers on the limiting BH cutoff C(z) = inf{c >= u_q : Rbar_z(c) >= 1} for the normalized mixture CDF; the proof works only when the cutoff is a 'regular first crossing'—the CDF stays strictly below 1 before the cutoff, strictly above 1 just after it, and the null contribution is continuous there. A likelihood-ratio envelope M(y) = sup_{a>=0} e^{-a^2/2} cosh(ay), crossed via q M(y_q) = 1, fixes the two-sided secondary-m","core_discovery":"Two-sided: the supremum over N, mean vector, and correlation matrix of the FDR of BH at level q is at least l=(q) = q sqrt(log(1/q))/(2 sqrt(pi)) + c_l q + o(q) with c_l = 0.6492828..., and l=(q) > q for every q in (0,1). One-sided: at least l<=(q) = q sqrt(log(1/q))/sqrt(pi) + q/2 + o(q), also > q. Since these lower bounds divided by q diverge, no finite C can satisfy FDR <= Cq for all Gaussian dependence. Within the one-common-factor class, the paper proves FDR = O(q sqrt(log(1/q))) for two-sided tests and FDR = q sqrt(log(1/q))/sqrt(pi) + O(q) for one-sided tests, matching the one-sided lower-bound constant; the exact two-sided constant is left open.","pith_inferences":["Editorial inference: in high-throughput applications that report BH at very small q, the paper's numbers (roughly 1.36x inflation at q=0.01 two-sided and 1.83x one-sided for the constructed designs) imply that dependence-agnostic reporting could be materially anticonservative at even smaller q.","Editorial inference: the two-sided leading constant is open; the natural testable target is to decide whether the worst-case two-sided class has the larger 1/sqrt(pi) constant or a genuinely smaller one.","Editorial inference: the construction's dependence on the Gaussian envelope M(y) suggests the same divergence rate should be checked for heavier-tailed or non-normal symmetric likelihoods; the result may be qualitatively different there.","Editorial inference: a practical validation is to simulate the paper's one-common-factor designs with q as low as 10^-4 and compare empirical FDR against q sqrt(log(1/q))/sqrt(pi) to confirm the predicted rate."],"forward_implications":["At every q in (0,1), BH can be anticonservative under Gaussian dependence: worst-case FDR is at least l=(q) > q for two-sided tests and l<=(q) > q for one-sided tests.","The ratio of worst-case FDR to q diverges as q goes to 0, so no universal multiplicative constant C can bound the FDR uniformly over all Gaussian correlations.","Restricting to one-common-factor Gaussian models does not restore FDR control; the order q sqrt(log(1/q)) persists in both two-sided and one-sided settings.","For one-sided common-factor tests, the leading constant is exactly 1/sqrt(pi): upper and lower bounds match, so the inflation rate is determined within that class.","Dependence direction matters: the two-sided construction works with positive correlations because the absolute-value transformation breaks PRDS, while the one-sided construction requires negative null-to-non-null correlations to violate PRDS."],"fun_headline_variants":["Gaussian dependence inflates BH FDR beyond any Cq bound","No universal Cq bound: Gaussian dependence can multiply BH FDR","Gaussian correlation pushes BH FDR/q to infinity","BH FDR worst-case under Gaussian dependence: unbounded ratio","Cq conjecture disproved: Gaussian dependence inflates BH FDR"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing technical premise is that the limiting BH cutoff is a 'regular first crossing' for almost every common-factor value—the normalized p-value CDF must stay strictly below 1 up to the cutoff and strictly above 1 just after it, with the null contribution continuous there; the finite-sample FDR limit is obtained by dominated convergence only at such regular points.","fun_headline_variants_meta":{"raw":{"variants":["Gaussian dependence inflates BH FDR beyond any Cq bound","No universal Cq bound: Gaussian dependence can multiply BH FDR","Gaussian correlation pushes BH FDR/q to infinity","BH FDR worst-case under Gaussian dependence: unbounded ratio","Cq conjecture disproved: Gaussian dependence inflates BH FDR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000685,"raw_usage":{"total_tokens":3012,"prompt_tokens":877,"completion_tokens":2135,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":621,"completion_tokens_details":{"reasoning_tokens":2058}},"tokens_in":621,"tokens_out":2135,"duration_ms":13698,"temperature":1.0,"reasoning_tokens":2058,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T00:58:40.758039+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Evaluate the normalized mixture CDF Rbar_z(c) on a fine grid in z around the boundary points -1 and -y_q (and around 0 and y_q^+ in the one-sided construction). If for some z in a set of positive measure the curve touches 1 before its first crossing or fails to exceed 1 immediately after, the regular-first-crossing condition fails and the dominated-convergence step cannot be applied; equally, a direct simulation of the paper's finite model at these parameters that clearly violates the stated lower bound would falsify the construction.","supporting_citations":[],"review_version":1}