Pith. sign in

REVIEW 4 minor

Mirror and knockoff+ thresholds under dependence

T0 review · 0 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read The mirror and knockoff+ thresholds do not control the false discovery rate under dependence unless null signs can be flipped independently.

desk verdict The central claim holds up: the bare mirror/knockoff+ count threshold is not an FDR guarantee under PRDS, Gaussian equicorrelation, exchangeability, or pairwise uncorrelatedness, and the paper proves it with exact, self-contained counterexamples. read the letter →

arxiv 2607.17084 v3 pith:2EGR5GLG submitted 2026-07-19 math.ST stat.MEstat.TH

classification math.STstat.MEstat.TH MSC 62F0362H15
keywords falsediscoveryratemirrorstatisticsknockoff+PRDSdependencesignflipsmultipletestingexchangeability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper takes the mirror and knockoff+ thresholds—rules that reject when enough scores fall on the discovery side compared with the control side—and asks whether they still control the false discovery rate for dependent test statistics. It establishes that they do not, by constructing families of null distributions that satisfy the usual symmetry and dependence conditions (uniform margins, PRDS, Gaussianity, exchangeability, pairwise uncorrelatedness) yet drive the FDR above its nominal level. The central reason is that these thresholds adapt to the data by comparing only two current tail counts, which can all be elevated together by a shared latent factor. A sympathetic reader should care because these thresholds are used in practice outside the exact knockoff construction; the paper shows that the plus-one adjustment and symmetry alone are not protection. The paper is careful to note that valid knockoff statistics, which have conditional sign flips, are not contradicted.

What carries the argument

The central objects are the two counting rules: the mirror threshold and the knockoff+ threshold, both of which compare control-side counts (L(u) or N_-(t)) against discovery-side counts (R(u) or N_+(t)). The proof mechanism that destroys validity is a lemma stating that if a block of all-null scores is positive together, the threshold is forced to pass, so the FDR equals the probability of such a block. Each counterexample builds a joint law that makes this block event likely while keeping the desired marginal or dependence properties: a latent Bernoulli mixture for the PRDS example, a common Gaussian factor for the equicorrelation result, and an exchangeable mixture LiZ + σε_i for the near

What would settle it

Run the knockoff+ threshold at q=0.1 on m=5000 all-null standard Gaussian equicorrelated scores with ρ=0.05 and repeat 10,000 times; Theorem 1 predicts an empirical FDR near 0.11 and rising with m, so observing the FDR stably below 0.1 would be a direct contradiction.

Watch

Extended reading notes

Core claim

The paper's central claim is that the two-count mirror/knockoff+ rule is not an FDR guarantee when applied to generic dependent scores, even if each null marginal is symmetric or uniform. It proves the claim with exact counterexamples: a full-support PRDS family of uniform p-values whose FDR at q=0.1 is 17.4% and can approach 1/2; standard equicorrelated Gaussian null scores for which liminf_m FDR_m ≥ 1/2 for every fixed ρ>0; and an exchangeable, pairwise-uncorrelated symmetric construction with FDR arbitrarily close to 1. The paper also proves an impossibility result: for q<1/2, no deterministic monotone function of the two current tail counts can repair the threshold over the full-support

Load-bearing premise

The conclusions rest on applying the bare two-count threshold to dependent scores that do not have the conditional sign-flip property; if a valid fixed-X or model-X knockoff construction is actually used, the negative results do not apply.

Editorial extensions

If this is right

  • For any target level q<1/2, standard Gaussian null scores with a fixed positive equicorrelation ρ will eventually produce FDR at least 1/2 as m grows, regardless of how small ρ is.
  • The mirror threshold can fail within the PRDS family, a positive-dependence condition that is sufficient for some standard step-up procedures but not for this adaptive two-tail rule.
  • Exchangeability and pairwise uncorrelatedness do not imply that the threshold's signs align weakly; FDR can be arbitrarily close to one while every null marginal is the same continuous symmetric distribution.
  • No deterministic monotone rule based only on the two current tail counts can give a distribution-free FDR repair over the full-support PRDS class for q<1/2; power and validity cannot both be achieved without extra calibrated information.
  • The results leave valid knockoff+ theory intact: procedures that produce conditionally independent fair coin flip signs keep the FDR bound.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: the failure mechanism suggests that any two-count adaptive rule will be fragile whenever test statistics share a latent common factor, even one of small variance; applied users should screen for such factors before applying mirror-type thresholds.
  • Inference: the results point to a possible repair direction the paper leaves implicit: using the entire mirror process (all thresholds) or randomized thresholds could bypass the deterministic current-count limitation, at the cost of more complex theory.
  • Inference: the Gaussian equicorrelation theorem plausibly extends to other one-factor models with heavy-tailed factor loadings, where the liminf lower bound may be even closer to one; this is a testable extension.
  • Inference: for practitioners, the paper implies that the 'plus one' in knockoff+ is doing less protective work than is sometimes assumed when exchangeability is only approximate; correction factors from robust knockoff theory may need to be enforced rather than treated as negligible.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 4 minor

Summary. The paper studies the mirror and knockoff+ thresholds — procedures that compare discovery-side counts with control-side counts — when applied to dependent scores or p-values that lack the conditional sign-flip property of valid knockoff statistics. It constructs four negative results: (i) an exactly uniform, full-support PRDS p-value model at nominal q=0.1 with FDR 17.4% for m=11, and FDR arbitrarily close to 1/2 within the same family; (ii) standard Gaussian null scores with any fixed positive equicorrelation ρ have liminf FDR at least 1/2 as m→∞; (iii) exchangeable, pairwise-uncorrelated, marginally symmetric scores can have FDR arbitrarily close to 1; and (iv) for q<1/2, no deterministic monotone threshold based only on the two current tail counts can give a nontrivial distribution-free repair over the full-support PRDS class. The paper carefully distinguishes these failures from valid knockoff theory, which relies on conditional coordinatewise sign flips rather than marginal symmetry, PRDS, exchangeability, or pairwise uncorrelatedness.

Significance. If the results hold, this is a valuable clarification of the limitations of count-comparison FDR methods. The paper provides exact, self-contained counterexamples with explicit formulas, and its scope limitations are stated clearly. The construction of a full-support PRDS example with uniform margins is technically clean and directly challenges the intuition that PRDS suffices for adaptive two-tail comparisons as it does for Benjamini–Hochberg. The Gaussian equicorrelation result is striking: no matter how small ρ>0 is, the FDR lower bound is 1/2 in the limit. The impossibility result for monotone count-only corrections is a useful negative result for attempts to fix the threshold by simple modifications. The paper also correctly credits existing methods (valid knockoffs, data splitting, conditional calibration) that add the additional structure needed for validity. Overall, the manuscript is a substantive theoretical contribution with machine-checkable-style proofs: all derivations are explicit and no parameters are fitted to data.

minor comments (4)
  1. [Section 6, proof of Theorem 1] The notation 'let \Phi = 1 - \Phi' is self-referential and should read 'let \bar\Phi = 1 - \Phi'. This is a typographical issue and does not affect the argument.
  2. [Section 1.1] The displayed FDR calculation writes '0.910' and '0.110'; these are intended as powers 0.9^{10} and 0.1^{10}. Please fix the formatting and similarly in the exact formula (5).
  3. [Section 5] The sentence 'The corresponding scores have symmetric uniform margins and satisfy W d=-W' should read 'W \stackrel{d}{=} -W' for clarity.
  4. [General] There are frequent missing spaces in 'p-values' and 'pvalue' throughout the abstract and introduction; a copyedit pass would improve readability.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the counterexamples are self-contained existence proofs; only self-citations are contextual.

full rationale

The paper's derivation chain is self-contained. Theorem 1 rests on Mills-ratio tail bounds and a conditional law-of-large-numbers argument; Theorem 2 constructs an explicit finite mixture and verifies exchangeability, zero pairwise covariance, and full-support density; Proposition 3 verifies PRDS by direct monotonicity calculations; Proposition 4 is a dichotomy following from monotonicity of the count-only rule. No parameter is fitted to data and then called a prediction, and no equation equates the target FDR failure with an input by definition. The only self-citations ([9], joint mirror; [24], covariate-adaptive FDR) occur in the related-work survey and are not used as load-bearing support for the paper's negative results. Proposition 2's proof is omitted and cited to Barber and Candès and Candès et al., but that is independent external support and is not what the paper claims to establish. The paper also explicitly limits its scope to the bare mirror/knockoff+ threshold without the conditional sign-flip property, so the apparent scope restriction is stated rather than hidden.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The construction parameters (ε, π, a, b, σ, m) are legitimate degrees of freedom for existence proofs, not hidden fits; the paper's contribution is to show that within these families the threshold fails. The axioms are standard probabilistic tools and known external results. No new entities are introduced.

free parameters (6)
  • ε (mixture split in Proposition 3) = 0.1 in the m=11 example; ↓0 for the 1/2 limit
    Controls how much mass the two conditional densities put on each half; chosen small so the all-positive event has probability >q or near 1/2. This is a counterexample tuning parameter, not a data fit.
  • m (number of hypotheses) = 11 in Proposition 3; large in Theorem 2
    Required to satisfy m>1/q in Proposition 3 or to ensure K/m≈π in Theorem 2; chosen to make the block inequality (4) pass.
  • π (mixing probability in Theorem 2) = 0.05 in the numerical study
    Chosen so π/(1−π)<q, ensuring G_m has probability tending to 1; also makes E(L_i)=0, giving pairwise uncorrelated scores.
  • a and b=πa/(1−π) (scale constants in Theorem 2) = a=1, b=1/19 in the numerical study
    a>b gives separation: for z>0 the a-block dominates in magnitude, for z<0 the b-block does; b is determined by π and a.
  • σ (noise scale in Theorem 2) = 0.0005–0.2 in the numerical study; σ↓0 in the proof
    Small noise lets the deterministic block pattern determine signs and magnitudes; the theorem sends σ to 0.
  • ρ (equicorrelation in Theorem 1)
    Universal quantifier in the theorem (any fixed ρ>0); not fitted, but included because the lower bound depends on it.
assumptions (5)
  • standard math Mills inequalities x/(1+x^2) φ(x) ≤ \barΦ(x) ≤ φ(x)/x for x>0.
    Used in Theorem 1 to bound the lower-to-upper tail ratio at a fixed threshold and to choose tδ.
  • standard math Law of large numbers and Fatou's lemma.
    Used in Theorem 1 to pass from conditional tail probabilities to rejection probability and in Theorem 2 for the FDR limit as σ↓0.
  • standard math PRDS sufficiency for one-sided Gaussian p-values with nonnegative correlations (Benjamini–Yekutieli, Sarkar).
    Used in Corollary 1 to assert that the Gaussian example is PRDS.
  • standard math Conditional sign-flip sufficiency for knockoff+ FDR control (Barber–Candès, Candès et al.).
    Used as the external benchmark that the counterexamples do not contradict; not used to derive any counterexample.
  • standard math Global-null identity FDR = P(any rejection).
    Simplifies all FDR computations under the global null; standard for a fixed null set.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Mirror and knockoff+ thresholds under dependence." pith.science (2026). https://pith.science/paper/2EGR5GLG

@misc{pith2026260717084,
  author       = {Pith},
  title        = {Pith review of: Mirror and knockoff+ thresholds under dependence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2EGR5GLG}},
  note         = {Machine review of arXiv:2607.17084}
}
abstract

Many multiple-testing procedures control the false discovery rate (FDR) by comparing the two tails of a null distribution. At a fixed cutoff, marginal symmetry makes this natural. Mirror and knockoff+ thresholds select the cutoff from the same data, so the standard finite-sample guarantee uses a stronger property: conditional on magnitudes and nonnull scores, null signs are independent fair coins. Failure can be severe without this property. Models satisfying positive regression dependence on a subset (PRDS) can have exactly uniform null $p$-values and large FDR. Under every fixed positive Gaussian equicorrelation, the all-null FDR converges to one half. Opposing loadings in Gaussian factor models can make FDR and power arbitrarily close to one; near-total failure also occurs for exchangeable, pairwise-uncorrelated scores. At a nominal input level $q<1/2$, no deterministic rule based only on the two tail counts can both reject and control FDR uniformly over our class if more discoveries or fewer controls cannot make rejection harder. We give finite-sample repairs based on joint sign information. Conditional sign odds may be bounded outside an exceptional event or averaged over negative controls; neither route uniformly dominates, and the integrated bounds are sharp. Independent calibration data or a specified Gaussian joint model yield valid adjusted levels. Simulations show that integration retains more power under diffuse Gaussian dependence, whereas exceptional-event calibration is more powerful when very large odds occur only for rare aligned signs; unadjusted FDR exceeds the target in both settings. Covariance alone is insufficient outside a specified joint model. Thus the relevant boundary is not marginal symmetry but joint information that remains valid after adaptive cutoff selection.

Figures

Figures reproduced from arXiv: 2607.17084 by the authors.

Figure 1
Figure 1. Finite-sample behavior of the mirror/knockoff+ threshold under the global null, at [PITH_FULL_IMAGE:figures/full_fig_p015_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.