{"id":"bece6c9f-41da-4b59-b30b-bd4ebff62746","arxiv_id":"2607.12208","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"BH fails to control FDR at α=0.01 for a factor-model family of correlated two-sided Gaussian tests, with a rigorous certificate that FDR exceeds 0.0104 for large n.","lead":"The Benjamini–Hochberg procedure can exceed its nominal false-discovery-rate level for correlated two-sided Gaussian p-values. A factor-model counterexample with an interval-arithmetic certificate shows FDR > 0.0104 at α = 0.01 for large numbers of tests, disproving a twenty-year conjecture.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the abstract-only limitation already flagged by the reader; the claimed counterexample is well-scoped if the certificate holds.","rationale":"The abstract states a clean, high-stakes counterexample: a concrete factor model of correlated two-sided Gaussians for which an interval-arithmetic proof shows BH at α=0.01 yields FDR>0.0104 for all large m. That is exactly the setting of the long-standing conjecture, so membership in the conjectured class is the sole load-bearing assumption visible from the abstract. The reader already identified this assumption and correctly rendered a CONDITIONAL verdict with LOW confidence pending inspection of the certificate and model. No additional soft spot—circularity, regime violation, or internal contradiction—can be extracted without the full text. Monte-Carlo consistency and author verification of the AI-generated proof are supporting, not load-bearing, claims. Consequently the stress-test finds no reason to alter the reader’s verdict or to elevate correctness risk beyond the medium level already assigned.","tokens_in":1893,"tokens_out":543,"duration_ms":5388,"concrete_test":"Obtain the full paper (or arXiv source) and re-run the interval-arithmetic certificate on the stated factor model at α=0.01; confirm both that the lower bound on FDR remains strictly above 0.0104 for large m and that the correlation structure is a valid factor model of two-sided Gaussian p-values. If either fails, the counterexample collapses; if both hold, the conjecture is disproved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader’s weakest_assumption correctly isolates the only load-bearing condition that can be assessed from the abstract: that the constructed factor model belongs to the dependence class (correlated two-sided Gaussian p-values) for which the twenty-year conjecture asserted FDR control. Nothing in the abstract suggests the model is outside that class; the authors explicitly frame it as a factor model of correlated two-sided Gaussians and supply an interval-arithmetic certificate that FDR exceeds 0.0104 at α=0.01 for all large m. Because the full text, the precise factor loadings, and the certificate itself are unavailable, no further internal inconsistency or hidden assumption can be diagnosed. The AI-assisted derivation is a provenance detail, not a soundness defect once the author has checked the certificate. Thus the central claim stands or falls solely on verification of the (unseen) certificate and model membership—precisely the condition the reader already conditioned upon.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript asserts that the Benjamini–Hochberg (BH) procedure need not control the false discovery rate at its nominal level for correlated two-sided Gaussian p-values. The authors construct a factor model under which, at α = 0.01, a rigorous interval-arithmetic certificate establishes FDR > 0.0104 for all sufficiently large numbers of hypotheses, thereby disproving a conjecture widely believed for roughly twenty years. Monte Carlo experiments are reported as consistent with the asymptotic lower bound. The proof is stated to have been obtained by GPT-5.6 Pro and carefully checked by the author. Only the abstract is available for this review; the full construction, certificate, and Monte Carlo design are not present.","tokens_in":2111,"tokens_out":932,"duration_ms":16752,"significance":"If the claimed certificate is correct and the factor model lies inside the dependence class for which the conjecture asserted FDR control, the result is a high-impact negative finding in multiple-testing theory. BH is foundational, and a rigorous counterexample for two-sided Gaussian tests under factor dependence would sharpen the known sufficient conditions (e.g., PRDS) and guide practice. The use of interval arithmetic to produce a machine-checkable numerical lower bound, together with Monte Carlo consistency checks, is a methodological strength worth crediting once the certificate can be inspected. The AI-assisted provenance is secondary provided the author-checked certificate is independently verifiable.","major_comments":[{"comment":"The central claim rests on a rigorous interval-arithmetic certificate that FDR exceeds 0.0104 at α = 0.01 for all large m. With only the abstract available, the factor loadings, the asymptotic FDR expression, the interval bounds, and the software/implementation of the certificate cannot be examined. Without those details the load-bearing lower bound cannot be verified, so the disproof of the conjecture cannot yet be accepted.","section":"Abstract (certificate claim)"},{"comment":"For the construction to refute the twenty-year conjecture it must lie inside the dependence class the conjecture was understood to cover (correlated two-sided Gaussian p-values, typically under positive or factor dependence). The abstract asserts a “factor model” of such p-values but does not state the precise correlation structure or its relation to known sufficient conditions (PRDS, etc.). Explicit membership must be established so the example is not outside the conjecture’s scope.","section":"Abstract (factor-model claim)"},{"comment":"Because the proof was produced by GPT-5.6 Pro, the manuscript must document the author’s verification steps (interval-arithmetic code, bound propagation, and any hand checks) in enough detail for independent reproduction. An author statement that the proof was “carefully checked” is not itself a certificate; the verification trail is load-bearing for a machine-assisted numerical proof.","section":"Abstract (proof provenance)"}],"minor_comments":[{"comment":"The abstract should state the precise asymptotic regime (e.g., fixed factor loadings as m → ∞, or any sparsity/null proportion assumptions) so readers can immediately see the scope of the claimed FDR lower bound.","section":"Abstract"},{"comment":"When the full text is supplied, the Monte Carlo design (number of replications, m values, random-number generation, and how empirical FDR is estimated) should be reported with enough precision to allow independent reproduction of the consistency claim.","section":"Abstract (Monte Carlo sentence)"},{"comment":"The abstract would benefit from a one-sentence pointer to the known positive results (e.g., BH under PRDS or independence) that the counterexample is intended to sit outside, clarifying the logical gap being closed.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"Only the abstract was supplied; a full-text review is required before any accept/reject decision. The AI-generated provenance is not itself a soundness defect, but it raises the bar for documentation of the interval-arithmetic certificate. If the full paper’s certificate and model membership check out, the result is potentially suitable for a strong statistics journal; if either fails, the claim collapses. I recommend obtaining the full manuscript and, if possible, the certificate code before a second round."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple: Dobriban claims an explicit factor model of correlated two-sided Gaussian p-values where BH at α=0.01 has FDR strictly above 0.0104 for all large m, certified by interval arithmetic. That would falsify a conjecture the field has treated as true for two decades.\n\nWhat is new is the concrete construction plus the numerical certificate. A factor model that stays inside ordinary equicorrelated or low-rank Gaussian dependence, paired with a machine-checkable lower bound rather than pure simulation, is the right way to kill a positive conjecture. Monte Carlo consistency is a useful sanity check. Credit for shipping a rigorous bound instead of hand-waving asymptotics; the GPT provenance is just provenance once the author has verified the arithmetic.\n\nSoft spots are exactly the ones forced by an abstract-only read. We cannot inspect the loadings, the interval-arithmetic code, or the precise membership of the model in the class the conjecture covered. Those are the only load-bearing items, and they are currently invisible. Nothing in the abstract suggests the model is exotic or outside the usual two-sided Gaussian setting, so the claim is well-scoped if the certificate is real. Free parameters in the factor model are not a flaw for a counterexample; they are how you build one. No circularity, no invented entities.\n\nThis is for anyone who uses or teaches FDR methods under dependence—genomics, imaging, online A/B testing. A serious referee should see the full certificate and the model. If both check out, the paper belongs in the literature; if not, it dies cleanly. I would send it to review rather than desk-reject.","headline":"Abstract-only claim of a certified counterexample that BH can exceed nominal FDR for correlated two-sided Gaussians, killing a 20-year conjecture if the interval-arithmetic bound holds.","tokens_in":2705,"tokens_out":445,"would_cite":false,"duration_ms":14318,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62J15","62H15","62F03"],"pacs":[],"model":"grok-4.5","headline":"The Benjamini–Hochberg procedure can fail to control the false discovery rate for correlated two-sided Gaussian p-values.","keywords":["Benjamini-Hochberg","false discovery rate","FDR control","correlated p-values","two-sided Gaussian tests","factor model","interval arithmetic","multiple testing"],"falsifier":"Either a tighter analytic or numerical bound showing that the FDR of BH under the given factor model is at most 0.01 for all large n, or an independent re-verification of the interval-arithmetic certificate that fails to recover the strict lower bound 0.0104.","tokens_in":2764,"feed_emoji":"📊","tokens_out":893,"duration_ms":14802,"temperature":0.7,"pith_summary":"This paper establishes that the Benjamini–Hochberg (BH) procedure does not always control the false discovery rate at its nominal level when the underlying p-values come from correlated two-sided Gaussian tests. The authors exhibit an explicit factor model of dependence for which, at the conventional level α = 0.01, a machine-checked interval-arithmetic argument proves that the FDR stays strictly above 0.0104 once the number of hypotheses is large enough. The construction therefore supplies a concrete counter-example to a conjecture that had been widely regarded as true for two decades. Monte Carlo simulations of the same model are consistent with the certified lower bound. A sympathetic reader cares because BH is the default multiple-testing method in genomics, imaging, and many other high-dimensional settings that routinely produce two-sided Gaussian statistics; if the conjecture fails, those analyses may be reporting more false discoveries than their nominal level suggests.","feed_headline":"BH fails FDR control for correlated two-sided Gaussians","feed_subtitle":"Interval arithmetic certifies FDR > 0.0104 at level 0.01, ending a twenty-year conjecture","key_machinery":"An explicit Gaussian factor model that generates the joint distribution of the two-sided p-values, together with an interval-arithmetic certificate that rigorously lower-bounds the limiting FDR of BH under that model.","core_discovery":"There exists a factor model of correlated two-sided Gaussian p-values such that the Benjamini–Hochberg procedure run at nominal level α = 0.01 has FDR greater than 0.0104 for every sufficiently large number of hypotheses; the inequality is certified by rigorous interval arithmetic and thereby disproves the long-standing conjecture that BH controls FDR for all such tests.","pith_inferences":["The counter-example is specific to two-sided tests; the one-sided Gaussian case may still enjoy FDR control under the same dependence, suggesting a sharp distinction that future theory should map.","Because the excess is modest (only a few percent above nominal), the practical inflation may be small for many data sets, yet the logical failure already forces a revision of textbooks and software documentation that state the conjecture as fact.","Interval-arithmetic certification of asymptotic FDR bounds could become a standard tool for settling other open dependence questions in multiple testing."],"forward_implications":["BH cannot be invoked as a black-box FDR-controlling procedure for arbitrary positive dependence among two-sided Gaussian tests.","Any theoretical guarantee that previously relied on the disproved conjecture must be re-examined or restricted to narrower dependence classes.","Practitioners analyzing two-sided Gaussian statistics under factor-type correlation should either verify FDR control by other means or adopt more conservative multiple-testing procedures.","The same factor-model construction supplies a concrete test case against which any proposed extension of BH can be checked."],"fun_headline_variants":["BH can exceed FDR for correlated two-sided Gaussians","Factor model shows BH misses FDR at α=0.01 for large tests","Interval math certifies BH FDR >0.0104, killing 20-year belief","Correlated two-sided Gaussians let BH breach nominal FDR","BH fails FDR control under Gaussian factor correlation"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The constructed factor model must lie inside the precise dependence class (correlated two-sided Gaussians) for which the twenty-year conjecture claimed FDR control; otherwise the excess FDR would not constitute a genuine counter-example.","fun_headline_variants_meta":{"raw":{"variants":["BH can exceed FDR for correlated two-sided Gaussians","Factor model shows BH misses FDR at α=0.01 for large tests","Interval math certifies BH FDR >0.0104, killing 20-year belief","Correlated two-sided Gaussians let BH breach nominal FDR","BH fails FDR control under Gaussian factor correlation"]},"model":"grok-4.5","effort":"low","cost_usd":0.005056,"raw_usage":{"total_tokens":1284,"prompt_tokens":662,"num_sources_used":0,"completion_tokens":94,"cost_in_usd_ticks":50560000,"prompt_tokens_details":{"text_tokens":662,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":528,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":662,"tokens_out":94,"duration_ms":3967,"temperature":1.0,"reasoning_tokens":528,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-15T01:01:27.352110+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Either a tighter analytic or numerical bound showing that the FDR of BH under the given factor model is at most 0.01 for all large n, or an independent re-verification of the interval-arithmetic certificate that fails to recover the strict lower bound 0.0104.","supporting_citations":[],"review_version":1}