Pith. sign in

REVIEW 3 major objections 5 minor 2 cited by

Asymptotic FDR Control with Model-X Knockoffs: Is Moments Matching Sufficient?

T0 review · 3 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read The paper proves that the Gaussian knockoffs generator, matching only the first two moments of the covariate distribution, still controls FDR asymptotically for two-moment-based statistics, a first formal justification of its success.

desk verdict A real first formal justification for moments-matching Gaussian knockoffs, but the abstract overclaims: the proof excludes the all-null/weak-signal regime and the estimated-moments case is only sketched. read the letter →

arxiv 2502.05969 v1 pith:DG3UWJRB submitted 2025-02-09 stat.ML cs.LGmath.STstat.TH

classification stat.MLcs.LGmath.STstat.TH MSC 62E2062J1562J07
keywords model-XknockoffsGaussiangeneratormomentsmatchingasymptoticFDRcontrolapproximatemoderatedeviationshigh-dimensionalvariableselectiondebiasedLasso
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Model-X knockoffs only come with a finite-sample FDR guarantee when knockoff variables are drawn from the true covariate distribution, which is almost never known. This paper asks when substituting a user-specified working distribution still keeps FDR under control, and answers with a unified framework: three sufficient conditions on the approximate knockoff statistics that any working distribution can be checked against. The main application is the most popular practical construction — the Gaussian knockoffs generator, which matches only the first two moments of the covariates. The paper proves that for two-moment-based knockoff statistics (marginal correlation differences and regression coefficient differences) these Gaussian knockoffs achieve asymptotic FDR control, $\limsup_{n\to\infty}\text{FDR}\le q$, the first formal justification of the method's practical success despite an evidently misspecified covariate model. The cost is explicit: enough true signals must clear the noise floor, and the heavy-tailed case carries a dimensionality restriction $p\log n = o(\sqrt{n})$.

What carries the argument

The central object is the Gaussian knockoffs generator, which builds knockoffs from the first two moments of the covariates alone. Its load-bearing property is covariance exchangeability: for every swap permutation the joint covariance of $(X^\top, \tilde X^\top)^\top$ is invariant, so on null features the original/knockoff pair has swap-invariant covariance and the approximating Gaussian law satisfies $P_j(t) \equiv P(|G_1| - |G_2| \ge t) = P(|G_1| - |G_2| \le -t)$. The argument is carried by ratio-based moderate deviation theorems (Theorems 5, 7, and 8 of the supplement): they show $P(\pm \hat W_j \ge t)/P_j(t) \to 1$ uniformly in $t$ up to the localization threshold, which is exactly the accuracy the FDR proof needs because the denominator probabilities can vanish. The three sufficient conditions of Theorem 1 — approximate symmetry, indicator concentration, and threshold localization — then yield $\limsup_{n\to\infty}\text{FDR}\le q$, and a coupling proposition (Proposition 3) extends the result from population to estimated moments.

What would settle it

Run the approximate knockoffs procedure with Gaussian knockoffs and the marginal correlation statistic on $t$-distributed covariates with three degrees of freedom and $p\log n$ at or above $\sqrt{n}$, so the dimensionality condition of the heavy-tailed case in Theorem 2 fails while signal conditions hold; empirical FDR staying at or below $q$ would show that boundary is not sharp, while systematic inflation above $q$ confirms it is real. A complementary check targets the signal-coverage assumption: in the all-null model, where $a_n = 0$ and condition (C3)/(O4) fails, systematic FDR inflation would show the assumption is genuinely load-bearing, whereas clean control would show the stated conditions are sufficient but not necessary.

Watch

Extended reading notes

Core claim

The paper's central claim is that the distributional exchangeability of perfect model-X knockoffs can be relaxed, with no loss of asymptotic FDR control, to three conditions on the approximate knockoff statistics: asymptotic approximate symmetry of null statistics, ratio-consistent concentration of their empirical indicator sums, and localization of the knockoff threshold $T_q$. Verifying these conditions for the Gaussian knockoffs generator based on first-two-moment matching proves $\limsup_{n\to\infty}\text{FDR}\le q$ for the marginal correlation difference statistic (Theorem 2), the OLS regression coefficient difference statistic (Theorem 3), and the debiased Lasso coefficient difference statistic (Theorem 4), under the explicit conditions (C1)-(C4), (O1)-(O5), and (L1). The theorems cover covariates as non-Gaussian as Rademacher signs and $t$-distributions with three degrees of freedom — cases where the earlier coupling-based robustness proof provably cannot work — and the paper states this is the first theoretical justification of the Gaussian knockoffs generator's effectiveness and robustness.

Load-bearing premise

The guarantee requires the number of true signals strong enough to clear the noise floor to grow without bound (condition (C3) and its analogue (O4)); in the all-null or weak-signal regime where that fails, the threshold-localization lemmas prove nothing and the theorem is silent.

Editorial extensions

If this is right

  • Practitioners using the Gaussian knockoffs generator with marginal-correlation or regression-coefficient-difference statistics receive a formal asymptotic guarantee: $\limsup_{n\to\infty}\text{FDR}\le q$ under the stated conditions.
  • Any working distribution is certified the same way: it suffices to verify the three conditions of Theorem 1 on its knockoff statistics, giving a modular robustness check independent of the coupling construction.
  • The guarantee covers cases the coupling-based theory cannot, including Rademacher covariates and $t$-distributed covariates with as few as three degrees of freedom, under the stated dimensionality restrictions.
  • When the precision matrix is estimated from data, the sample-moment knockoff matrix inherits the asymptotic guarantee by coupling to the population-moment version (Proposition 3).
  • The conditions make the signal requirement explicit: enough features ($a_n\to\infty$) must have strengths clearing the noise floor for the threshold-localization step to work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's stated recommendation — match as many moments as the statistic uses — suggests a testable hierarchy: a statistic built from third or higher moments, such as the distance correlation difference used in the real-data example, should need higher-order moment matching for the same guarantee; the theorems here cover only first-two-moment statistics, and the distance-correlation results are
  • Because the theorem is silent when the number of strong signals stays bounded, and yet FDR control is the expected behavior of a valid procedure in the all-null regime, the signal-coverage condition (C3)/(O4) is plausibly an artifact of the threshold-localization proof rather than a real phenomenon; a sharper localization argument could remove it.
  • The proof requires of the noise vector $Z$ in the generator only moment conditions, so the construction should transfer to non-Gaussian generators with the same moment profile, demoting Gaussianity to a computational convenience.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies the asymptotic FDR of the model-X knockoffs procedure when knockoff variables are generated from a misspecified working distribution, focusing on the Gaussian knockoffs generator that matches only the first two moments. It first states a general sufficient-conditions framework (Theorem 1): if the null knockoff statistics satisfy approximate symmetry, indicator-function concentration, and threshold localization on a common interval, then limsup FDR is bounded by q. It then verifies these conditions for two-moment-based statistics: marginal correlation differences (Theorem 2), OLS coefficient differences (Theorem 3), and debiased Lasso coefficient differences (Theorem 4). The verification is carried out under explicit moment, sparsity, signal-strength, and dimensionality conditions. The paper also presents simulations and an HIV drug-resistance application illustrating the finite-sample behavior of the method.

Significance. If the formal results hold, this is a useful contribution: it provides the first modular sufficient conditions under which the practically ubiquitous Gaussian knockoffs construction can achieve asymptotic FDR control despite covariate-distribution misspecification. The proof of Theorem 1 is transparent and short, and the supplementary material contains detailed moderate-deviation and concentration arguments rather than leaving the verification as a black box. The empirical sections support but do not replace the theory. The main caveat is that the headline claim is conditional on a strong-signal regime that is not stated in the abstract, and one of the illustrative corollaries appears to conflict with its own moment assumption.

major comments (3)
  1. [§2.2–§3.1 (Theorem 1, Theorem 2, Lemma 2)] The abstract and introduction claim asymptotic FDR control for the Gaussian knockoffs generator without qualification, but the actual proof requires the strong-signal condition (C3): a_n = #{j ∈ H1 : \hat w_j ≥ C δ_n} → ∞. This condition enters the localization step in Lemma 2 through α_n = P^{-1}(q a_n/(2p)) and p α_n → ∞, and its analogues (O4) and (D3) play the same role in Lemmas 8 and 11. In the all-null or weak-signal regime a_n does not diverge, so the proof provides no localization of T_q and conditions (1)–(3) of Theorem 1 are not verified; yet a valid procedure is expected to control FDR in that regime as well. The formal theorems are stated conditionally, so this is fixable, but the abstract's unqualified "asymptotic FDR control" is misleading and should be revised to state the strong-signal restriction or supplemented by an argument covering the null/weak-signal case.
  2. [§3.1, condition (C1') and Corollary 1] Condition (C1') requires the entries of Q to have finite q-th moments for some q ≥ 3, while Corollary 1, part 2, claims coverage for X_i i.i.d. t-distributed with q degrees of freedom and q ≥ 3. A t-distribution with q degrees of freedom does not have a finite q-th moment, since E|X|^r < ∞ only for r < q. Thus the stated corollary is inconsistent with its own assumption. This can be repaired by requiring t-distributed covariates with more than q degrees of freedom, or by rephrasing the moment condition with a separate exponent, but the mismatch should be corrected.
  3. [§2.2 (Theorem 1, condition 2)] Condition (2) of Theorem 1 is stated as a uniform approximation over all t ∈ (0, α_n). In the subsequent verification, the width of this interval is ultimately tied to the signal count a_n through α_n = P^{-1}(q a_n/(2p)). This means that the theorem, as applied, cannot certify FDR control when the number of detectable signals is bounded, even if the null statistics themselves are well behaved. This is the same issue as the strong-signal restriction above, but it is worth emphasizing that the restriction is structural to the framework, not merely a technical tightening in one lemma: without a_n → ∞, the interval (0, α_n) on which conditions (1)–(3) are verified is not constructed at all.
minor comments (5)
  1. [§1.1 (Notation)] The notation section states "For any positive integer m ∈ Z+, let [m] ≡ {1, ..., n}"; the right-hand side should be {1, ..., m}.
  2. [§3.4 (Proposition 3)] In Proposition 3, the vector r is described as r ∈ R^n, but r in the Gaussian knockoff construction is a p-dimensional diagonal vector; it should be r ∈ R^p.
  3. [§2.2 (Theorem 1 proof)] In the proof of Theorem 1, the notation "1{#\hat W_j ≥ t}" appears where # is used as a placeholder for the sign; this is not defined in the main text and should be replaced by explicit positive and negative indicators.
  4. [§1 (Introduction)] The abstract and introduction repeatedly describe the Gaussian generator as "arguably the most popularly used" method; the qualifier is fine, but the phrasing is informal for a journal paper and could be made more precise by citing specific prevalence in the software ecosystem.
  5. [§4 (Numerical studies)] Table 1 appears in the Introduction before the data description in Section 4.2; this is unconventional and may confuse readers, though the results are eventually explained.

Circularity Check

1 steps flagged · score 4.0 of 10

Core verification is self-contained, but the estimated-moments extension is carried by an overlapping-authors citation.

  1. self citation load bearing [Section 3.4, 'Gaussian knockoffs generator using estimated moments', paragraph after Proposition 3]
    "Thus, using X̂ as a bridge, the coupling framework in Fan et al. (2025) can then be applied to show that knockoff statistics constructed using X̂_in also satisfy the conditions in Theorem 1 under mild additional assumptions, thereby ensuring asymptotic FDR control. Since the detailed proofs are very similar to those in Fan et al. (2025), for the sake of brevity and to focus on the main contributions of this paper, we omit further technical details."

    The sample-based 'estimated moments' generalization is a stated contribution of the abstract ('a user-specified distribution that can be learned using in-sample observations'), yet the derivation that X̂_in satisfies conditions (1)-(3) of Theorem 1 is not carried out in this paper. It is asserted to follow from the coupling framework of Fan et al. (2025), whose author list overlaps with the present paper. Proposition 3 supplies only a coupling-accuracy bound; the load-bearing step from that bound to the three sufficient conditions is delegated to a same-author citation rather than given as an independent derivation. This is not definitional circularity, but it makes this particular extension self-referential.

full rationale

The central results Theorems 2-4 are not circular: the three sufficient conditions of Theorem 1 are verified by dedicated moderate-deviation theorems (Theorems 5-8), variance and localization lemmas, and explicit concentration inequalities in the supplement. No fitted parameter is later renamed as a prediction, and the asymptotic approximate-symmetry condition is not defined in terms of the FDR target. The acknowledgment in Section 2.2 that conditions (2) and (3) were previously established in Fan et al. (2025) is not load-bearing because the approximate-statistics versions are reproved here. However, Section 3.4's extension to estimated moments is load-bearing and justified only by an overlapping-authors citation, for which the paper explicitly says it omits further technical details. In addition, the abstract's unqualified 'asymptotic FDR control' is broader than the strong-signal regime defined by conditions (C3), (O4), and (D3), under which localization of Tq is proved; this is a scope or correctness concern rather than a circularity concern. Overall, the main derivation chain has independent content, but one advertised contribution is carried by self-citation, so the paper is partially self-referential but not definitionally circular.

Assumptions & free parameters 2 free parameters · 6 assumptions · 0 invented entities

The theorems are conditional: they hold for factor-model or sub-Gaussian covariates, under strong-signal assumptions and regularity of precision matrices. No invented physical entities are introduced. The main hidden burden is the growing-signal condition a_n -> infinity and the unstated zero-mean condition in the correlated-features marginal-correlation proof.

free parameters (2)
  • r (Gaussian knockoff diagonal vector) = user-specified; in independent case r=(sigma_1^2,...,sigma_p^2)
    In equation (2), r controls the covariance between X and Xhat and is chosen to keep 2diag(r)-diag(r)Sigma^{-1}diag(r) positive definite. The theorems assume such a choice exists but do not analyze a data-driven rule for r.
  • Regularization constants in debiased Lasso (lambda_0 and lambda_j) = lambda_0 = C sqrt(n^{-1} log(2p)); lambda_j not fully specified
    These are hand-chosen tuning parameters assumed to satisfy the theory. They are not fitted to the response, but they are arbitrary constants in the construction of the knockoff statistics.
assumptions (6)
  • domain assumption Nonparametric model Y=F(X_{H1})+xi with xi independent of X and unique H1
    Section 2.1, equation (3). This defines the null set and is inherited from Candes et al. (2018).
  • domain assumption Factor-model representation X=Q Sigma_X^{1/2} with independent sub-exponential entries of Q
    Condition (C1)/(C1'), used in Theorem 2. It is needed for the CLT and moderate deviation results for marginal correlation statistics.
  • domain assumption Signal coverage: a_n = #{j in H1: signal >= C delta_n} -> infinity and negative-tail condition (C4)
    Conditions (C3)-(C4), Theorem 2. These ensure T_q <= alpha_n and that negative signals are rare; they exclude all-null and weak-signal regimes.
  • domain assumption For OLS and debiased Lasso: n/2p >= 1+tau, sub-Gaussianity, restricted eigenvalue, precision sparsity and signal strength (O1)-(O5)/(L1)
    Theorems 3 and 4. These are needed for Gaussian approximation of regression coefficient statistics and threshold localization.
  • ad hoc to paper Zero conditional mean for null marginal-correlation statistics in correlated-features analysis
    The Section D proof applies Theorem 10, which requires zero-mean summands. For correlated features and nonlinear Y, E[X_j | Y] need not vanish for null j. This assumption is not stated in the paper.
  • standard math Standard moderate deviation, concentration, and Berry-Esseen inequalities
    Used throughout Section 3 and the supplement, mostly citing Fang and Koike (2023), Gotze et al. (2021), and Rudelson and Vershynin (2013).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Asymptotic FDR Control with Model-X Knockoffs: Is Moments Matching Sufficient?." pith.science (2026). https://pith.science/paper/DG3UWJRB

@misc{pith2026250205969,
  author       = {Pith},
  title        = {Pith review of: Asymptotic FDR Control with Model-X Knockoffs: Is Moments Matching Sufficient?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DG3UWJRB}},
  note         = {Machine review of arXiv:2502.05969}
}
read the original abstract

We propose a unified theoretical framework for studying the robustness of the model-X knockoffs framework by investigating the asymptotic false discovery rate (FDR) control of the practically implemented approximate knockoffs procedure. This procedure deviates from the model-X knockoffs framework by substituting the true covariate distribution with a user-specified distribution that can be learned using in-sample observations. By replacing the distributional exchangeability condition of the model-X knockoff variables with three conditions on the approximate knockoff statistics, we establish that the approximate knockoffs procedure achieves the asymptotic FDR control. Using our unified framework, we further prove that an arguably most popularly used knockoff variable generation method--the Gaussian knockoffs generator based on the first two moments matching--achieves the asymptotic FDR control when the two-moment-based knockoff statistics are employed in the knockoffs inference procedure. For the first time in the literature, our theoretical results justify formally the effectiveness and robustness of the Gaussian knockoffs generator. Simulation and real data examples are conducted to validate the theoretical findings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semi-knockoffs: a model-agnostic conditional independence testing method with finite-sample guarantees

    math.ST 2026-01 conditional novelty 7.0 of 10

    Semi-knockoffs tests conditional independence with arbitrary pre-trained models and no train-test split by comparing losses on two resampled copies of each feature; oracle versions have finite-sample guarantees, estim...

  2. Knockoffs Inference under Privacy Constraints

    stat.ME 2025-06 reject novelty 7.0 of 10

    A differentially private mirror-peeling knockoff algorithm is introduced, with claimed exact FDR control and asymptotic power preservation.

Reference graph

Works this paper leans on

29 extracted references · 28 canonical work pages · cited by 2 Pith papers

  1. [1]

    Adamczak, R. (2015). A note on the Hanson-Wright inequality for random vectors with dependencies . Electronic Communications in Probability , 20:1--13

  2. [2]

    Bai, X., Ren, J., Fan, Y., and Sun, F. (2020). KIMI : Knockoff inference for motif identification from molecular sequences with controlled false discovery rate. Bioinformatics , 37(6):759--766

  3. [3]

    Barber, R. F. and Cand \`e s, E. J. (2015). Controlling the false discovery rate via knockoffs. The Annals of statistics , 43:2055--2085

  4. [4]

    F., Cand\`es, E

    Barber, R. F., Cand\`es, E. J., and Samworth, R. J. (2020). Robust inference with knockoffs. Ann. Statist. , 48(3):1409--1431

  5. [5]

    Boucheron, S., Lugosi, G., and Massart, P. (2013). Concentration Inequalities: A Nonasymptotic Theory of Independence . Oxford University Press, Oxford

  6. [6]

    Cand\`es, E., Fan, Y., Janson, L., and Lv, J. (2018). Panning for gold: ``model- X ' knockoffs for high dimensional controlled variable selection. Journal of the Royal Statistical Society Series B , 80(3):551--577

  7. [7]

    Chi, C.-M., Fan, Y., Ing, C.-K., and Lv, J. (2025). High-dimensional knockoffs inference for time series data. Journal of the American Statistical Association, to appear

  8. [8]

    Fan, J., Liao, Y., and Liu, H. (2016). An overview of the estimation of large covariance and precision matrices. The Econometrics Journal , 19(1):C1--C32

Show all 29 references
  1. [9]

    Fan, Y., Demirkaya, E., Li, G., and Lv, J. (2020a). R ANK : large-scale inference with graphical nonlinear knockoffs. J. Amer. Statist. Assoc. , 115(529):362--379

  2. [10]

    Fan, Y., Gao, L., and Lv, J. (2025). ARK : robust knockoffs inference with coupling. The Annals of Statistics, to appear

  3. [11]

    Fan, Y., Lv, J., Sharifvaghefi, M., and Uematsu, Y. (2020b). I PAD : stable interpretable forecasting with knockoffs inference. J. Amer. Statist. Assoc. , 115(532):1822--1834

  4. [12]

    and Koike, Y

    Fang, X. and Koike, Y. (2023). From p- W asserstein bounds to moderate deviations. Electronic Journal of Probability , 28:1--52

  5. [13]

    Fukumizu, K., Gretton, A., Sun, X., and Sch \"o lkopf, B. (2007). Kernel measures of conditional dependence. Advances in Neural Information Processing Systems , 20

  6. [14]

    Gao, L., Fan, Y., Lv, J., and Shao, Q.-M. (2021). Asymptotic distributions of high-dimensional distance correlation inference . The Annals of Statistics , 49(4):1999 -- 2020

  7. [15]

    G \"o tze, F., Sambale, H., and Sinulis, A. (2021). Concentration inequalities for polynomials in -sub-exponential random variables. Electronic Journal of Probability , 26:1--22

  8. [16]

    Ledoux, M. (2001). The Concentration of Measure Phenomenon . American Mathematical Society

  9. [17]

    Lu, Y., Fan, Y., Lv, J., and Stafford Noble, W. (2018). DeepPINK : reproducible feature selection in deep neural networks. Advances in Neural Information Processing Systems (NeurIPS 2018)

  10. [18]

    Niu, Z., Chakraborty*, A., Dukes, O., and Katsevich, E. (2024). Reconciling model-x and doubly robust approaches to conditional independence testing. Annals of Statistics , to appear

  11. [19]

    Pillai, N. S. and Yin, J. (2014). Universality of covariance matrices . The Annals of Applied Probability , 24(3):935--1001

  12. [20]

    L., and Shafer, R

    Rhee, S.-Y., Taylor, J., Wadhera, G., Ben-Hur, A., Brutlag, D. L., and Shafer, R. W. (2006). Genotypic predictors of human immunodeficiency virus type 1 drug resistance. Proceedings of the National Academy of Sciences , 103(46):17355--17360

  13. [21]

    Romano, Y., Sesia, M., and Cand \`e s, E. (2020). Deep knockoffs. Journal of the American Statistical Association , 115(532):1861--1872

  14. [22]

    Rosenthal, H. P. (1970). On the subspaces of L ^p ( p > 2 ) spanned by sequences of independent random variables. Israel J. Math. , 8:273--303

  15. [23]

    and Vershynin, R

    Rudelson, M. and Vershynin, R. (2009). Smallest singular value of a random rectangular matrix. Communications on Pure and Applied Mathematics , 62(12):1707--1739

  16. [24]

    and Vershynin, R

    Rudelson, M. and Vershynin, R. (2013). Hanson-Wright inequality and sub-gaussian concentration . Electronic Communications in Probability , 18:1--9

  17. [25]

    Sambale, H. (2023). Some notes on concentration for -subexponential random variables. In High Dimensional Probability IX: The Ethereal Volume , pages 167--192. Springer

  18. [26]

    J., Rizzo, M

    Sz\'ekely, G. J., Rizzo, M. L., and Bakirov, N. K. (2007). Measuring and testing dependence by correlation of distances. Ann. Statist. , 35(6):2769--2794

  19. [27]

    and Vu, V

    Tao, T. and Vu, V. (2015). Random matrices: universality of local spectral statistics of non-Hermitian matrices . The Annals of Probability , 43(2):782--874

  20. [28]

    and Zhang, S

    Zhang, C.-H. and Zhang, S. S. (2013). Confidence intervals for low dimensional parameters in high dimensional linear models. Journal of the Royal Statistical Society Series B: Statistical Methodology , 76(1):217--242

  21. [29]

    Zhu, Z., Fan, Y., Kong, Y., Lv, J., and Sun, F. (2021). DeepLINK : deep learning inference using knockoffs with applications to genomics. Proceedings of the National Academy of Sciences , 118(36):e2104683118

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.