{"id":"fad6a080-cb77-493c-8966-2ec293b27f65","arxiv_id":"2607.17084","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":8.0,"correctness_risk":"low","formal_verification":"none","parameter_count":6,"one_line_summary":"Mirror and knockoff+ thresholds can have FDR far above nominal under dependence, even for uniform PRDS p-values and equicorrelated Gaussian scores.","lead":"This paper constructs explicit examples where mirror and knockoff+ thresholds, which compare the control side of a null distribution to the discovery side, fail to control the false discovery rate under dependence. A generalist should read it because it delimits when a widely used counting rule is valid: marginal symmetry, Gaussianity, PRDS, exchangeability, or zero pairwise correlation are not enough.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified","rationale":"The reader's acceptance is justified. The paper's central claim is a negative/existence result: certain familiar dependence properties do not suffice for the bare mirror/knockoff+ threshold. The counterexamples are exact, self-contained, and internally consistent. I checked the PRDS proof (mixture construction, monotonicity of Q(u)), the exact FDR computation for the eleven-hypothesis example, the Gaussian equicorrelation proof (tail-ratio Mills bound and the fixed-threshold-to-T step), the exchangeable pairwise-uncorrelated construction (covariance calculation, positive density, two-sided z argument, Fatou step), and Proposition 4's dichotomy. The only caveats are scope limitations the paper itself states: valid knockoff constructions restore the conditional sign-flip property, and Proposition 4 does not cover randomized or history-dependent rules. These are not defects in the argument. The numerical study lacks shipped code but is not load-bearing for the theorems. Thus no adjustment to the reader's verdict is needed.","tokens_in":12835,"tokens_out":22655,"duration_ms":191606,"concrete_test":"Independently verify the exact FDR formula (5) by Monte Carlo: simulate the m=11 PRDS mixture with epsilon=0.1 and apply the mirror threshold (1) at q=0.1 for 10^7 repetitions; the empirical rejection rate should equal 0.5*(0.9^10+0.1^10) approximately 0.1743 within 0.001. A mismatch would indicate an error in the example or the formula.","verdict_should_be":"UNCHANGED","load_bearing_attack":"After a careful check of the proofs, I find no load-bearing flaw in the central claim. Proposition 3's PRDS construction is valid: the mixture of product densities gives exactly uniform margins, full support, and the monotonicity argument for PRDS is correct. The exact FDR formula (5) for the m=11 example follows from the conditional independence of signs given the latent variable and the permutation argument. Theorem 1's Mills-ratio argument is sound once the survival-function notation (bar-Phi) is restored; the monotonicity in z is in the correct direction for upper/lower tail probabilities, and the fixed-threshold argument does yield a passing element of T. Theorem 2's exchangeable, pairwise-uncorrelated construction is internally consistent: the positive-definite covariance for sigma>0, the zero covariance, and the two-sided z argument all check out. Proposition 4's dichotomy is exact: if psi(m,0)<=q, the PRDS construction makes FDR arbitrarily close to 1/2; if psi(m,0)>q, monotonicity prevents any rejection. The only limitations are scope limitations the paper explicitly acknowledges: the results target the bare counting rule without valid-knockoff sign flips, and Proposition 4 is confined to deterministic monotone functions of the two current counts. These do not undermine the stated claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the mirror and knockoff+ thresholds — procedures that compare discovery-side counts with control-side counts — when applied to dependent scores or p-values that lack the conditional sign-flip property of valid knockoff statistics. It constructs four negative results: (i) an exactly uniform, full-support PRDS p-value model at nominal q=0.1 with FDR 17.4% for m=11, and FDR arbitrarily close to 1/2 within the same family; (ii) standard Gaussian null scores with any fixed positive equicorrelation ρ have liminf FDR at least 1/2 as m→∞; (iii) exchangeable, pairwise-uncorrelated, marginally symmetric scores can have FDR arbitrarily close to 1; and (iv) for q<1/2, no deterministic monotone threshold based only on the two current tail counts can give a nontrivial distribution-free repair over the full-support PRDS class. The paper carefully distinguishes these failures from valid knockoff theory, which relies on conditional coordinatewise sign flips rather than marginal symmetry, PRDS, exchangeability, or pairwise uncorrelatedness.","tokens_in":13177,"tokens_out":6295,"duration_ms":57682,"significance":"If the results hold, this is a valuable clarification of the limitations of count-comparison FDR methods. The paper provides exact, self-contained counterexamples with explicit formulas, and its scope limitations are stated clearly. The construction of a full-support PRDS example with uniform margins is technically clean and directly challenges the intuition that PRDS suffices for adaptive two-tail comparisons as it does for Benjamini–Hochberg. The Gaussian equicorrelation result is striking: no matter how small ρ>0 is, the FDR lower bound is 1/2 in the limit. The impossibility result for monotone count-only corrections is a useful negative result for attempts to fix the threshold by simple modifications. The paper also correctly credits existing methods (valid knockoffs, data splitting, conditional calibration) that add the additional structure needed for validity. Overall, the manuscript is a substantive theoretical contribution with machine-checkable-style proofs: all derivations are explicit and no parameters are fitted to data.","major_comments":[],"minor_comments":[{"comment":"The notation 'let \\Phi = 1 - \\Phi' is self-referential and should read 'let \\bar\\Phi = 1 - \\Phi'. This is a typographical issue and does not affect the argument.","section":"Section 6, proof of Theorem 1"},{"comment":"The displayed FDR calculation writes '0.910' and '0.110'; these are intended as powers 0.9^{10} and 0.1^{10}. Please fix the formatting and similarly in the exact formula (5).","section":"Section 1.1"},{"comment":"The sentence 'The corresponding scores have symmetric uniform margins and satisfy W d=-W' should read 'W \\stackrel{d}{=} -W' for clarity.","section":"Section 5"},{"comment":"There are frequent missing spaces in 'p-values' and 'pvalue' throughout the abstract and introduction; a copyedit pass would improve readability.","section":"General"}],"recommendation":"accept","confidential_remarks":"This is a clean theoretical paper. The proofs are self-contained and the central claims are sound. The stress-test concern about possible circularity does not land: the constructions are existential and no conclusion is assumed. The only caveat is the explicitly acknowledged scope limitation — the counterexamples apply to the bare threshold without valid-knockoff sign flips — but this is properly framed in the paper. I recommend acceptance after minor typographical corrections."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper's central claim is right. Apply the mirror/knockoff+ threshold as a bare counting rule to dependent null scores or p-values, and you do not get FDR control under marginal symmetry, PRDS, Gaussianity, exchangeability, or pairwise uncorrelatedness. The paper proves this with exact constructions, and I found no load-bearing flaw.\n\nWhat is actually new: the full-support uniform PRDS counterexample with exact FDR 17.4% at q=0.1; the fixed-rho Gaussian equicorrelation theorem (liminf FDR >= 1/2 for every rho>0 as m grows); the exchangeable, pairwise-uncorrelated construction with FDR arbitrarily close to 1; and the impossibility result for deterministic monotone two-count corrections over the PRDS class. These are not simple consequences of the earlier Storey/BH failures cited in the paper. The proofs are self-contained and internally consistent — the PRDS verification, the Mills-ratio tail bound, the block-majority argument, and the dichotomy in Proposition 4 all check out. The numerical study is a sanity check, not the argument.\n\nSoft spots are mostly scope limitations, and the paper states them clearly. Theorem 1 gives only a liminf, and the convergence is slow for small rho (order (log m)^{-1/2}), so finite-sample behavior at rho=0.05 can be mild. Theorem 2 requires taking m large and then sigma to 0 in a non-Gaussian mixture, so the near-one failure lives in an asymptotic regime. Proposition 4 confines the impossibility to deterministic monotone rules using the two current counts; randomized, non-monotone, or history-dependent rules are not covered. None of this undermines the stated claims. One minor point: the numerical study doesn't ship code, but the simulations are straightforward and the headline results are theorems.\n\nCitations: self-citations are contextual, not load-bearing. The acknowledgment of AI assistance is transparent and doesn't affect the mathematics.\n\nWho this is for: anyone working on FDR under dependence, knockoff theory, or mirror statistics. It deserves a serious referee. I'd send it out.","headline":"The central claim holds up: the bare mirror/knockoff+ count threshold is not an FDR guarantee under PRDS, Gaussian equicorrelation, exchangeability, or pairwise uncorrelatedness, and the paper proves it with exact, self-contained counterexamples.","tokens_in":13650,"tokens_out":2770,"would_cite":true,"duration_ms":25662,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62H15"],"pacs":[],"model":"deepseek-v4-flash","headline":"The mirror and knockoff+ thresholds do not control the false discovery rate under dependence unless null signs can be flipped independently.","keywords":["false discovery rate","mirror statistics","knockoff+","PRDS","dependence","sign flips","multiple testing","exchangeability"],"falsifier":"Run the knockoff+ threshold at q=0.1 on m=5000 all-null standard Gaussian equicorrelated scores with ρ=0.05 and repeat 10,000 times; Theorem 1 predicts an empirical FDR near 0.11 and rising with m, so observing the FDR stably below 0.1 would be a direct contradiction.","tokens_in":12727,"feed_emoji":"📊","tokens_out":4770,"duration_ms":43178,"temperature":0.7,"pith_summary":"The paper takes the mirror and knockoff+ thresholds—rules that reject when enough scores fall on the discovery side compared with the control side—and asks whether they still control the false discovery rate for dependent test statistics. It establishes that they do not, by constructing families of null distributions that satisfy the usual symmetry and dependence conditions (uniform margins, PRDS, Gaussianity, exchangeability, pairwise uncorrelatedness) yet drive the FDR above its nominal level. The central reason is that these thresholds adapt to the data by comparing only two current tail counts, which can all be elevated together by a shared latent factor. A sympathetic reader should care because these thresholds are used in practice outside the exact knockoff construction; the paper shows that the plus-one adjustment and symmetry alone are not protection. The paper is careful to note that valid knockoff statistics, which have conditional sign flips, are not contradicted.","feed_headline":"Dependent scores break mirror and knockoff+ FDR control","feed_subtitle":"Even tiny Gaussian correlation inflates false discoveries; no simple count-based fix can repair the threshold.","key_machinery":"The central objects are the two counting rules: the mirror threshold and the knockoff+ threshold, both of which compare control-side counts (L(u) or N_-(t)) against discovery-side counts (R(u) or N_+(t)). The proof mechanism that destroys validity is a lemma stating that if a block of all-null scores is positive together, the threshold is forced to pass, so the FDR equals the probability of such a block. Each counterexample builds a joint law that makes this block event likely while keeping the desired marginal or dependence properties: a latent Bernoulli mixture for the PRDS example, a common Gaussian factor for the equicorrelation result, and an exchangeable mixture LiZ + σε_i for the near","core_discovery":"The paper's central claim is that the two-count mirror/knockoff+ rule is not an FDR guarantee when applied to generic dependent scores, even if each null marginal is symmetric or uniform. It proves the claim with exact counterexamples: a full-support PRDS family of uniform p-values whose FDR at q=0.1 is 17.4% and can approach 1/2; standard equicorrelated Gaussian null scores for which liminf_m FDR_m ≥ 1/2 for every fixed ρ>0; and an exchangeable, pairwise-uncorrelated symmetric construction with FDR arbitrarily close to 1. The paper also proves an impossibility result: for q<1/2, no deterministic monotone function of the two current tail counts can repair the threshold over the full-support","pith_inferences":["Inference: the failure mechanism suggests that any two-count adaptive rule will be fragile whenever test statistics share a latent common factor, even one of small variance; applied users should screen for such factors before applying mirror-type thresholds.","Inference: the results point to a possible repair direction the paper leaves implicit: using the entire mirror process (all thresholds) or randomized thresholds could bypass the deterministic current-count limitation, at the cost of more complex theory.","Inference: the Gaussian equicorrelation theorem plausibly extends to other one-factor models with heavy-tailed factor loadings, where the liminf lower bound may be even closer to one; this is a testable extension.","Inference: for practitioners, the paper implies that the 'plus one' in knockoff+ is doing less protective work than is sometimes assumed when exchangeability is only approximate; correction factors from robust knockoff theory may need to be enforced rather than treated as negligible."],"forward_implications":["For any target level q<1/2, standard Gaussian null scores with a fixed positive equicorrelation ρ will eventually produce FDR at least 1/2 as m grows, regardless of how small ρ is.","The mirror threshold can fail within the PRDS family, a positive-dependence condition that is sufficient for some standard step-up procedures but not for this adaptive two-tail rule.","Exchangeability and pairwise uncorrelatedness do not imply that the threshold's signs align weakly; FDR can be arbitrarily close to one while every null marginal is the same continuous symmetric distribution.","No deterministic monotone rule based only on the two current tail counts can give a distribution-free FDR repair over the full-support PRDS class for q<1/2; power and validity cannot both be achieved without extra calibrated information.","The results leave valid knockoff+ theory intact: procedures that produce conditionally independent fair coin flip signs keep the FDR bound."],"fun_headline_variants":["FDR control fails for dependent scores","Mirror and knockoff+ break under dependence","Dependence kills FDR guarantees","No simple fix for dependent FDR","Dependent p-values shatter FDR control"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The conclusions rest on applying the bare two-count threshold to dependent scores that do not have the conditional sign-flip property; if a valid fixed-X or model-X knockoff construction is actually used, the negative results do not apply.","fun_headline_variants_meta":{"raw":{"variants":["FDR control fails for dependent scores","Mirror and knockoff+ break under dependence","Dependence kills FDR guarantees","No simple fix for dependent FDR","Dependent p-values shatter FDR control"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00019,"raw_usage":{"total_tokens":1232,"prompt_tokens":855,"completion_tokens":377,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":313}},"tokens_in":599,"tokens_out":377,"duration_ms":3383,"temperature":1.0,"reasoning_tokens":313,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T19:05:15.967004+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the knockoff+ threshold at q=0.1 on m=5000 all-null standard Gaussian equicorrelated scores with ρ=0.05 and repeat 10,000 times; Theorem 1 predicts an empirical FDR near 0.11 and rising with m, so observing the FDR stably below 0.1 would be a direct contradiction.","supporting_citations":[],"review_version":1}