REVIEW 4 minor
Mirror and knockoff+ thresholds under dependence
T0 review · 0 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read The mirror and knockoff+ thresholds do not control the false discovery rate under dependence unless null signs can be flipped independently.
desk verdict The central claim holds up: the bare mirror/knockoff+ count threshold is not an FDR guarantee under PRDS, Gaussian equicorrelation, exchangeability, or pairwise uncorrelatedness, and the paper proves it with exact, self-contained counterexamples. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central objects are the two counting rules: the mirror threshold and the knockoff+ threshold, both of which compare control-side counts (L(u) or N_-(t)) against discovery-side counts (R(u) or N_+(t)). The proof mechanism that destroys validity is a lemma stating that if a block of all-null scores is positive together, the threshold is forced to pass, so the FDR equals the probability of such a block. Each counterexample builds a joint law that makes this block event likely while keeping the desired marginal or dependence properties: a latent Bernoulli mixture for the PRDS example, a common Gaussian factor for the equicorrelation result, and an exchangeable mixture LiZ + σε_i for the near
What would settle it
Run the knockoff+ threshold at q=0.1 on m=5000 all-null standard Gaussian equicorrelated scores with ρ=0.05 and repeat 10,000 times; Theorem 1 predicts an empirical FDR near 0.11 and rising with m, so observing the FDR stably below 0.1 would be a direct contradiction.
Extended reading notes
Core claim
The paper's central claim is that the two-count mirror/knockoff+ rule is not an FDR guarantee when applied to generic dependent scores, even if each null marginal is symmetric or uniform. It proves the claim with exact counterexamples: a full-support PRDS family of uniform p-values whose FDR at q=0.1 is 17.4% and can approach 1/2; standard equicorrelated Gaussian null scores for which liminf_m FDR_m ≥ 1/2 for every fixed ρ>0; and an exchangeable, pairwise-uncorrelated symmetric construction with FDR arbitrarily close to 1. The paper also proves an impossibility result: for q<1/2, no deterministic monotone function of the two current tail counts can repair the threshold over the full-support
Load-bearing premise
The conclusions rest on applying the bare two-count threshold to dependent scores that do not have the conditional sign-flip property; if a valid fixed-X or model-X knockoff construction is actually used, the negative results do not apply.
Editorial extensions
If this is right
- For any target level q<1/2, standard Gaussian null scores with a fixed positive equicorrelation ρ will eventually produce FDR at least 1/2 as m grows, regardless of how small ρ is.
- The mirror threshold can fail within the PRDS family, a positive-dependence condition that is sufficient for some standard step-up procedures but not for this adaptive two-tail rule.
- Exchangeability and pairwise uncorrelatedness do not imply that the threshold's signs align weakly; FDR can be arbitrarily close to one while every null marginal is the same continuous symmetric distribution.
- No deterministic monotone rule based only on the two current tail counts can give a distribution-free FDR repair over the full-support PRDS class for q<1/2; power and validity cannot both be achieved without extra calibrated information.
- The results leave valid knockoff+ theory intact: procedures that produce conditionally independent fair coin flip signs keep the FDR bound.
Reading between the lines
- Inference: the failure mechanism suggests that any two-count adaptive rule will be fragile whenever test statistics share a latent common factor, even one of small variance; applied users should screen for such factors before applying mirror-type thresholds.
- Inference: the results point to a possible repair direction the paper leaves implicit: using the entire mirror process (all thresholds) or randomized thresholds could bypass the deterministic current-count limitation, at the cost of more complex theory.
- Inference: the Gaussian equicorrelation theorem plausibly extends to other one-factor models with heavy-tailed factor loadings, where the liminf lower bound may be even closer to one; this is a testable extension.
- Inference: for practitioners, the paper implies that the 'plus one' in knockoff+ is doing less protective work than is sometimes assumed when exchangeability is only approximate; correction factors from robust knockoff theory may need to be enforced rather than treated as negligible.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the mirror and knockoff+ thresholds — procedures that compare discovery-side counts with control-side counts — when applied to dependent scores or p-values that lack the conditional sign-flip property of valid knockoff statistics. It constructs four negative results: (i) an exactly uniform, full-support PRDS p-value model at nominal q=0.1 with FDR 17.4% for m=11, and FDR arbitrarily close to 1/2 within the same family; (ii) standard Gaussian null scores with any fixed positive equicorrelation ρ have liminf FDR at least 1/2 as m→∞; (iii) exchangeable, pairwise-uncorrelated, marginally symmetric scores can have FDR arbitrarily close to 1; and (iv) for q<1/2, no deterministic monotone threshold based only on the two current tail counts can give a nontrivial distribution-free repair over the full-support PRDS class. The paper carefully distinguishes these failures from valid knockoff theory, which relies on conditional coordinatewise sign flips rather than marginal symmetry, PRDS, exchangeability, or pairwise uncorrelatedness.
Significance. If the results hold, this is a valuable clarification of the limitations of count-comparison FDR methods. The paper provides exact, self-contained counterexamples with explicit formulas, and its scope limitations are stated clearly. The construction of a full-support PRDS example with uniform margins is technically clean and directly challenges the intuition that PRDS suffices for adaptive two-tail comparisons as it does for Benjamini–Hochberg. The Gaussian equicorrelation result is striking: no matter how small ρ>0 is, the FDR lower bound is 1/2 in the limit. The impossibility result for monotone count-only corrections is a useful negative result for attempts to fix the threshold by simple modifications. The paper also correctly credits existing methods (valid knockoffs, data splitting, conditional calibration) that add the additional structure needed for validity. Overall, the manuscript is a substantive theoretical contribution with machine-checkable-style proofs: all derivations are explicit and no parameters are fitted to data.
minor comments (4)
- [Section 6, proof of Theorem 1] The notation 'let \Phi = 1 - \Phi' is self-referential and should read 'let \bar\Phi = 1 - \Phi'. This is a typographical issue and does not affect the argument.
- [Section 1.1] The displayed FDR calculation writes '0.910' and '0.110'; these are intended as powers 0.9^{10} and 0.1^{10}. Please fix the formatting and similarly in the exact formula (5).
- [Section 5] The sentence 'The corresponding scores have symmetric uniform margins and satisfy W d=-W' should read 'W \stackrel{d}{=} -W' for clarity.
- [General] There are frequent missing spaces in 'p-values' and 'pvalue' throughout the abstract and introduction; a copyedit pass would improve readability.
Circularity Check
No significant circularity: the counterexamples are self-contained existence proofs; only self-citations are contextual.
full rationale
The paper's derivation chain is self-contained. Theorem 1 rests on Mills-ratio tail bounds and a conditional law-of-large-numbers argument; Theorem 2 constructs an explicit finite mixture and verifies exchangeability, zero pairwise covariance, and full-support density; Proposition 3 verifies PRDS by direct monotonicity calculations; Proposition 4 is a dichotomy following from monotonicity of the count-only rule. No parameter is fitted to data and then called a prediction, and no equation equates the target FDR failure with an input by definition. The only self-citations ([9], joint mirror; [24], covariate-adaptive FDR) occur in the related-work survey and are not used as load-bearing support for the paper's negative results. Proposition 2's proof is omitted and cited to Barber and Candès and Candès et al., but that is independent external support and is not what the paper claims to establish. The paper also explicitly limits its scope to the bare mirror/knockoff+ threshold without the conditional sign-flip property, so the apparent scope restriction is stated rather than hidden.
Assumptions & free parameters
free parameters (6)
- ε (mixture split in Proposition 3) =
0.1 in the m=11 example; ↓0 for the 1/2 limit
- m (number of hypotheses) =
11 in Proposition 3; large in Theorem 2
- π (mixing probability in Theorem 2) =
0.05 in the numerical study
- a and b=πa/(1−π) (scale constants in Theorem 2) =
a=1, b=1/19 in the numerical study
- σ (noise scale in Theorem 2) =
0.0005–0.2 in the numerical study; σ↓0 in the proof
- ρ (equicorrelation in Theorem 1)
assumptions (5)
- standard math Mills inequalities x/(1+x^2) φ(x) ≤ \barΦ(x) ≤ φ(x)/x for x>0.
- standard math Law of large numbers and Fatou's lemma.
- standard math PRDS sufficiency for one-sided Gaussian p-values with nonnegative correlations (Benjamini–Yekutieli, Sarkar).
- standard math Conditional sign-flip sufficiency for knockoff+ FDR control (Barber–Candès, Candès et al.).
- standard math Global-null identity FDR = P(any rejection).
Cite this review
Pith. "Pith review of Mirror and knockoff+ thresholds under dependence." pith.science (2026). https://pith.science/paper/2EGR5GLG
@misc{pith2026260717084,
author = {Pith},
title = {Pith review of: Mirror and knockoff+ thresholds under dependence},
year = {2026},
howpublished = {\url{https://pith.science/paper/2EGR5GLG}},
note = {Machine review of arXiv:2607.17084}
}
abstract
Many multiple-testing procedures control the false discovery rate (FDR) by comparing the two tails of a null distribution. At a fixed cutoff, marginal symmetry makes this natural. Mirror and knockoff+ thresholds select the cutoff from the same data, so the standard finite-sample guarantee uses a stronger property: conditional on magnitudes and nonnull scores, null signs are independent fair coins. Failure can be severe without this property. Models satisfying positive regression dependence on a subset (PRDS) can have exactly uniform null $p$-values and large FDR. Under every fixed positive Gaussian equicorrelation, the all-null FDR converges to one half. Opposing loadings in Gaussian factor models can make FDR and power arbitrarily close to one; near-total failure also occurs for exchangeable, pairwise-uncorrelated scores. At a nominal input level $q<1/2$, no deterministic rule based only on the two tail counts can both reject and control FDR uniformly over our class if more discoveries or fewer controls cannot make rejection harder. We give finite-sample repairs based on joint sign information. Conditional sign odds may be bounded outside an exceptional event or averaged over negative controls; neither route uniformly dominates, and the integrated bounds are sharp. Independent calibration data or a specified Gaussian joint model yield valid adjusted levels. Simulations show that integration retains more power under diffuse Gaussian dependence, whereas exceptional-event calibration is more powerful when very large odds occur only for rare aligned signs; unadjusted FDR exceeds the target in both settings. Covariance alone is insufficient outside a specified joint model. Thus the relevant boundary is not marginal symmetry but joint information that remains valid after adaptive cutoff selection.
Figures
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.