{"id":"349887db-5a2a-439c-ac44-fb7cd7988490","arxiv_id":"2505.14675","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"TarGene applies semi-parametric efficient estimators, TMLE and OSE, to genetic effects and k-point interactions with population-dependence correction, yielding new software and biobank results.","lead":"TarGene is a statistical workflow for estimating genetic effects and gene-gene or gene-environment interactions in large biobanks without the linear-model assumptions of standard GWAS. It uses targeted machine learning to reduce bias and corrects for relatedness among participants, demonstrated on UK Biobank and All of Us cohorts.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Marginal MAF≥0.01 filter does not ensure joint positivity for k-point interactions, so the type I error guarantee for AIEs is not established.","rationale":"The reader's weakest assumption identified the positivity condition and the post hoc MAF filter. I agree that positivity is the load-bearing assumption, but the concern is more specific than the reader's formulation: the paper filters marginal genotype frequencies while the interaction estimand requires bounded joint propensities. The null simulation's low coverage without filtering confirms that the asymptotic regime is fragile, but the paper does not establish that the 0.01 marginal filter restores joint positivity for AIEs. This is directly relevant to a central claimed contribution (k-point interactions) and to the headline promise of controlled false discoveries. The concern does not invalidate the standard TMLE/OSE machinery; it identifies a gap in the interaction-specific guarantee that could likely be fixed by a joint positivity threshold or by reporting separate coverage results. The sieve plateau variance issue and the uncontrolled comparison with GeneATLAS are real but secondary: the former is a variance add-on whose negligible effect is acknowledged, and the latter is an application claim not needed for the core estimator. The mathematical derivations in Sections 2.2–2.4 appear standard, and the software is described in unusual detail; these strengthen the conditional-acceptance posture. A single targeted simulation on joint-sparse interaction estimands would settle whether the marginal filter is adequate.","tokens_in":38742,"tokens_out":11011,"duration_ms":111947,"concrete_test":"In the null simulation, construct AIE estimands whose marginal genotype changes all satisfy p(V=v)≥0.01 but whose minimum joint genotype frequency across the 2^k strata in Eq. 7 is below 0.01 (e.g., two biallelic variants with MAF 0.05, or one such variant times an environment quintile). Run n=500,000 with XGBoost nuisance estimation and both OSE and wTMLE, using the same 500-bootstrap coverage evaluation as Section 3.5. Report coverage separately for these estimands. If coverage falls below 90% (versus nominal 95%), the marginal MAF filter is insufficient for interaction estimands and the type I error guarantee must be restricted to estimands satisfying a joint positivity threshold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central promise is controlled type I error for semi-parametric estimates of genetic effects and interactions. The formal condition behind this is the positivity assumption δ < g0(w) < 1−δ stated in Section 2.1, where g0(w) is the joint propensity P(A=a(s)|W=w) for each combination in the interaction parameter (Eq. 10). The post hoc remedy in Section 3.5.1, however, is a marginal filter: genotype changes with p(V=v)<0.01 are excluded. For a k-point interaction, the relevant propensity is a joint probability over k variants (and, for G×E, the environment stratum). Even when every marginal genotype frequency exceeds 0.01, joint strata can be far smaller: two independent variants each with MAF 0.05 have a double-heterozygote frequency near 0.009, and conditioning on principal components can lower this further in stratified samples. The simulations report aggregate coverage after the marginal filter (Fig. 2B/3B), but do not report AIE coverage separately as a function of joint stratum frequency. The null simulation's sub-50% coverage for some estimands without filtering shows that the asymptotic regime is not reached under positivity violations; the weighted estimators mitigate but do not remove this. The paper's Discussion states the guarantee holds 'provided the minor allele frequency is bounded from below by 0.01', which switches the operative condition from joint propensity to marginal MAF. For interactions, that substitution is not justified and leaves the type I error claim for AIEs unsupported in sparse joint strata.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces TarGene, a targeted-learning workflow for semi-parametric efficient estimation of genetic main effects and k-point interactions at biobank scale. The estimand is an Average Interaction Effect, defined as a signed sum of covariate-adjusted genotype-contrast means; the authors derive its efficient influence function and exact second-order remainder, then construct one-step estimators, TMLE, weighted TMLE, and cross-validated versions. A sieve-plateau variance estimator based on the genetic relationship matrix is adapted to account for dependence among individuals. The method is evaluated in two UK Biobank-based simulation studies, one null and one with a neural-network generative model, and applied to FTO PheWAS, gene-environment, gene-gene, and three-way interaction analyses. The paper claims nominal coverage and type I error control provided a 1% minor allele frequency threshold is imposed, and argues that commonly reported LMM p-values may be inflated.","tokens_in":39025,"tokens_out":5672,"duration_ms":59035,"significance":"If the guarantees hold, this is a substantial contribution to statistical genetics: it brings semi-parametric efficient and double-robust estimation, with open-source scalable software, to genetic main effects and interactions in cohorts of hundreds of thousands of individuals. The mathematical derivations in Section 2 are standard but correctly assembled for the k-point interaction parameter, and the realistic simulation framework built on UK Biobank data is a genuine strength, as are the public implementations TMLE.jl and TarGene. The most valuable part of the paper is the combination of rigorous influence-function theory with a scalable pipeline. The main risk is that the headline type I error guarantee is conditional on a positivity assumption that the paper verifies only through a marginal minor-allele-frequency filter, which is not sufficient for interaction estimands.","major_comments":[{"comment":"The paper's type I error claim rests on a marginal filter p(V=v)≥0.01, but the formal positivity condition in Eq. (10) is joint: δ < P(A=a(s)|W=w) < 1−δ for every genotype combination in the interaction. For k-point interactions, a marginal MAF filter does not ensure joint genotype support. For example, two independent variants with MAF 0.05 have a double-heterozygote frequency of about 0.9%, and conditioning on principal components can lower this further; Table S1 lists joint (genotype,outcome) frequencies as low as 2.1×10^{-6}. The null simulation in Fig. 2A shows sub-50% coverage for some estimands without filtering, and Figs. 2B/3B report coverage aggregated after marginal filtering rather than AIE coverage as a function of joint stratum frequency. The Discussion's guarantee 'provided the minor allele frequency is bounded from below by 0.01' therefore substitutes a marginal condition for the joint condition in Eq. (10); that substitution is not justified for interactions. Please report coverage conditional on joint propensity or joint stratum frequency, or restrict the formal type I error guarantee to estimands satisfying joint positivity.","section":"§3.5.1 / §2.1, Eq. (10)"},{"comment":"The sieve-plateau variance estimator is a stated contribution, but it is not validated in either simulation. Sections 3.5.1 and 3.5.2 evaluate coverage only with the i.i.d. variance estimator; the application in Fig. 5D and Fig. S5B shows a variance increase of about 1.4% and 'little difference' in p-values, which does not establish that the SP correction restores nominal coverage under realistic relatedness. Since the abstract and contribution list emphasize variance correction for population dependence, the method should be evaluated in simulations with a known relatedness or stratification structure, or the claims about 'realistic p-values correctly accounting for population dependence' should be scaled back.","section":"§2.5 / §3.5 / Fig. 5D"},{"comment":"The statement that GeneATLAS results 'likely contain an inflated set of false discoveries' goes beyond the evidence shown. TarGene and GeneATLAS estimate different quantities: TarGene targets a genotype contrast with flexible nuisance learning and six principal components, while GeneATLAS uses a linear mixed model with an additive effect. No calibration simulation or ground-truth benchmark is supplied for the PheWAS comparison, so the observed p-value shift could reflect differences in estimand, sample filtering, or variance estimation rather than FDR inflation. This claim should be rephrased as a hypothesis or supported by a calibration analysis.","section":"§4.2.1, Fig. 5B and Introduction"},{"comment":"The Discussion states that 'TarGene estimators address any statistical gap due to model misspecification' despite the paper's own simulations showing sub-50% coverage for some estimands in the absence of positivity filtering (Figs. 2A and 3A) and despite the earlier caveat in the Introduction that the causal gap due to LD remains. The sentence should be qualified to the settings in which the required positivity and nuisance-rate conditions hold.","section":"§6 Discussion"}],"minor_comments":[{"comment":"The p-value '1.09×10^6' should read '1.09×10^{-6}'.","section":"§4.3"},{"comment":"The sieve-plateau estimator as written in Eq. (62) sums uncentered products D_i D_j; if the empirical mean of D is not exactly zero after targeting, this is not the covariance term in Eq. (60). Please add the appropriate centering or explain why it is negligible.","section":"§2.5, Eq. (62)"},{"comment":"The statement that 'we are conservative and simply do not test genotype changes for which the positivity threshold is not met' should be justified with respect to multiple testing: excluding unsupported contrasts changes the estimand set and the testing burden, which is not obviously conservative.","section":"Footnote 2, §2.4"},{"comment":"The text says that cross-validated XGBoost estimators have larger bias than their canonical counterparts and suggests this may be due to using only three folds; this is plausible but should be stated as a hypothesis, since the bootstrap bias estimates are not accompanied by standard errors.","section":"§3.5.2"},{"comment":"The paragraph reporting the replication of published epistatic pairs should state explicitly that two pairs were excluded because they failed the marginal positivity threshold, since this affects the interpretation of the replication rate.","section":"§4.4"}],"recommendation":"major_revision","confidential_remarks":"The mathematical core is sound and the software contribution is real. The main blocker is the mismatch between the joint positivity condition in Eq. (10) and the marginal MAF filter used to claim type I error control for interactions, plus the lack of simulation validation of the sieve-plateau variance estimator. If the authors can either demonstrate joint-positivity support or restrict the guarantees accordingly, I would be supportive."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a serious, well-executed application of targeted learning to genetic effect estimation, and it deserves a proper referee slot. The main methodological novelty is modest—the k-point interaction parameter and its influence function were already in Beentjes & Khamseh 2020, and the EIF here is just the linear combination of ATE EIFs. What is actually new is the engineering and the empirical work: TMLE/OSE variants with weighting and cross-validation for this estimand, the sieve plateau variance correction for genetic relatedness, and a large UK Biobank application with 39 traits showing non-linear allelic effects, 21 GxE signals, and 27 GxG signals. The simulations are extensive and more realistic than the usual parametric sims, using neural density estimators fit to UKB data. The paper also honestly discloses the positivity threshold required to achieve nominal coverage.\n\nThe math checks out as far as I can tell; the derivations are standard influence-function theory. The simulations are the strongest part: null and realistic, two sample sizes, four estimators, and they show coverage and type I error control only after filtering on minor allele frequency. Here is the real soft spot, and it is the one the stress-test flags. The formal positivity condition is on the joint propensity g0(w) for each genotype combination in the interaction. The post hoc filter is marginal MAF: exclude genotype changes with p(V=v) < 0.01. For k-point interactions, joint strata can be far sparser than any marginal frequency—two independent variants at MAF 0.05 have a double-heterozygote frequency near 0.009, and conditioning on PCs makes it worse. The paper reports aggregate coverage after the marginal filter, but does not report AIE coverage as a function of joint stratum frequency. So the type I error guarantee for interaction effects in sparse joint strata is not established. This is a limitation, not a fatal flaw, but the Discussion's phrase 'provided the minor allele frequency is bounded from below by 0.01' overstates what the filter delivers.\n\nAlso worth noting: the sieve plateau variance correction is not validated in simulations and shows little effect in the application; the paper acknowledges this, which is good, but the claim of accounting for population dependence is under-supported. The comparison with GeneATLAS p-values is suggestive but not a controlled baseline; the claim that LMM p-values are inflated would need a more careful analysis.\n\nWho is this for? Statistical geneticists who use TMLE/OSE and want interactions at biobank scale. The software is a real plus. I would send it to peer review, with a request to address the joint positivity issue and to tone down the 'p-values inflated' claim.","headline":"A solid, well-executed targeted-learning application to genetic effects; the joint-positivity gap for interactions is a real limitation but not a fatal one, and the paper deserves serious peer review.","tokens_in":39627,"tokens_out":2193,"would_cite":true,"duration_ms":21807,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G05","62G20","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"TarGene gives semi-parametric efficient, double-robust estimates of genetic effects and interactions.","keywords":["targeted minimum loss-based estimation","semi-parametric efficiency","average interaction effect","k-point interactions","epistasis","gene-environment interaction","sieve plateau variance","UK Biobank"],"falsifier":"Run the paper's null simulation for a variant with minor allele frequency 0.005 and 500,000 samples, using the canonical wTMLE with XGBoost nuisance fits; if the 95% confidence interval covers the true zero effect well below 95% and coverage returns to nominal only after applying the 0.01 filter, the positivity assumption is confirmed as the gatekeeper of the method's guarantees.","tokens_in":38533,"feed_emoji":"🧬","tokens_out":7371,"duration_ms":65152,"temperature":0.7,"pith_summary":"This paper introduces TarGene, a workflow for estimating genetic effects and interactions in biobank-scale cohorts without imposing a parametric form on the relationship between genotype, covariates, and trait. The target quantity is the k-point interaction, a direct generalisation of the average treatment effect, and it is estimated with one-step estimators and targeted minimum loss-based estimators that are doubly robust and asymptotically efficient. The authors argue that standard linear mixed model analyses can produce inflated p-values under model misspecification, and they support this with a phenome-wide FTO study in the UK Biobank where TarGene finds 63 significant traits at 5% FDR versus 159 for a linear mixed model. In simulations with a 1% minor allele frequency filter, the estimators achieve nominal coverage and controlled type I error, making the method a candidate replacement for LMM-based association and interaction testing in large cohorts.","feed_headline":"TarGene finds inflated p-values in linear-model GWAS","feed_subtitle":"A double-robust semi-parametric estimator for genetic effects and interactions in half-million-person cohorts.","key_machinery":"The machinery is the efficient influence function of the k-point interaction estimand $\\Psi^{(k)}_{a(0),a(1)}(P)$, built as a linear combination of average-treatment-effect influence functions, together with a clever covariate $H(g)(A,W)=\\sum_{s\\in\\{0,1\\}^k}(-1)^{k-(s_1+\\cdots+s_k)}\\frac{1\\{A=a(s)\\}}{g(A,W)}$ used in the TMLE targeting step. The paper also uses an exact second-order remainder whose double-robust product structure guarantees $o_P(n^{-1/2})$ bias when the outcome regression and propensity score are estimated at $n^{-1/4}$ rates, and sieve plateau variance estimators built from the genetic relationship matrix to account for relatedness among participants.","core_discovery":"The paper's central claim is that TarGene estimators close the statistical gap due to model misspecification and the causal gap due to population stratification, so that reported p-values from conventional GWAS may be inflated and contain more false discoveries than acknowledged. To show this, the paper defines the average interaction effect (AIE) as a k-point generalisation of the average treatment effect, derives its efficient influence function and double-robust second-order remainder, and constructs one-step, TMLE, weighted TMLE, and cross-validated versions of each. Simulation studies using the full UK Biobank structure show coverage at the nominal level once variants with minor allele frequency below 0.01 are excluded. In applications, TarGene replicates five of nine previously reported red-hair epistatic pairs, finds 39 traits with significant non-linear allelic effects at FTO, and demonstrates gene-by-environment interactions with deprivation indices, with effect-size estimates lower than those from linear mixed models.","pith_inferences":["A natural extension is to pair TarGene with rare-variant collapsing or burden tests that aggregate genotypes across a gene, raising stratum counts so the positivity assumption can hold without excluding low-frequency variants.","The five-of-nine replication of red-hair epistasis may reflect the additive rather than multiplicative interaction scale; a direct comparison of both scales on the same data would separate scale differences from true non-replication.","Because the causal gap from linkage disequilibrium remains open, TarGene's effect sizes are best read as locus-level associations; combining it with fine-mapping or knockoff-based selection would test whether the reported non-linear and interaction effects localise to the same causal variant.","The $2^k$ testing burden per locus rules out naive genome-wide interaction scans; a two-stage protocol that screens with light models and confirms with the full super learner, as the paper's runtime recommendations suggest, is the practical deployment path."],"forward_implications":["If the central claim holds, LMM-based GWAS p-values can be anti-conservative for variants with non-linear or non-additive effects, so some reported associations at a fixed FDR threshold would not survive a correctly specified model.","TarGene supplies estimates and confidence intervals for k-point gene-gene and gene-environment interactions on an additive scale, which is directly interpretable for public-health effect sizes rather than multiplicative logistic odds ratios.","The method requires a minor allele frequency filter near 1%, so its nominal coverage guarantee applies to common variants; rare-variant analyses need larger samples or genotype aggregation to satisfy positivity.","The sieve plateau variance correction is computationally feasible and changes UKB variance estimates by about 1.4%, suggesting it will matter more in family-based or more ancestrally diverse cohorts than in the UKB white subset.","The released pipeline makes a PheWAS with a comprehensive super learner run in about 30 hours and projects a 600,000-variant GWAS in about 375 hours on 200 cores, so semi-parametric estimation is practical at biobank scale."],"supporting_citations":[{"why":"Provides the TMLE framework and the fluctuation step that solves the efficient influence function.","marker":"Van der Laan and Rubin (2006)"},{"why":"Supplies the von Mises expansion and rate conditions used to establish asymptotic linearity of the one-step and TMLE estimators.","marker":"Benkeser et al. (2017)"},{"why":"Introduces the sieve plateau variance estimators adapted here to genetic relatedness.","marker":"Davies and van der Laan (2014)"},{"why":"Defines the k-point interaction parameter that the paper generalises as the average interaction effect.","marker":"Beentjes and Khamseh (2020)"},{"why":"Provides GeneATLAS linear mixed model results used as the main comparison baseline in the FTO PheWAS.","marker":"Canela-Xandri et al. (2018)"},{"why":"Supplies the UK Biobank cohort used for the realistic simulations and all applications.","marker":"Bycroft et al. (2018)"},{"why":"Reports the hair-colour epistatic pairs that TarGene attempts to reproduce.","marker":"Morgan et al. (2018b)"}],"fun_headline_variants":["TarGene shows standard GWAS p-values can be inflated","New estimator corrects inflated p-values in large-scale GWAS","TarGene double-robust method slashes false discoveries","TarGene corrects stratification bias in GWAS p-values","Efficient estimator for small genetic effects in big cohorts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is positivity: every genotype or genotype-environment combination being compared must occur with probability bounded away from zero in every covariate stratum, which fails for rare variants; the paper restores nominal coverage by excluding variants with minor allele frequency below 0.01.","fun_headline_variants_meta":{"raw":{"variants":["TarGene shows standard GWAS p-values can be inflated","New estimator corrects inflated p-values in large-scale GWAS","TarGene double-robust method slashes false discoveries","TarGene corrects stratification bias in GWAS p-values","Efficient estimator for small genetic effects in big cohorts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000784,"raw_usage":{"total_tokens":3501,"prompt_tokens":1023,"completion_tokens":2478,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":639,"completion_tokens_details":{"reasoning_tokens":2397}},"tokens_in":639,"tokens_out":2478,"duration_ms":18049,"temperature":1.0,"reasoning_tokens":2397,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:30:09.688100+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's null simulation for a variant with minor allele frequency 0.005 and 500,000 samples, using the canonical wTMLE with XGBoost nuisance fits; if the 95% confidence interval covers the true zero effect well below 95% and coverage returns to nominal only after applying the 0.01 filter, the positivity assumption is confirmed as the gatekeeper of the method's guarantees.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the TMLE framework and the fluctuation step that solves the efficient influence function."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the sieve plateau variance estimators adapted here to genetic relatedness."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides GeneATLAS linear mixed model results used as the main comparison baseline in the FTO PheWAS."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the UK Biobank cohort used for the realistic simulations and all applications."}],"review_version":1}