Pith. sign in

REVIEW 2 cited by

Floodgate: inference for model-free variable importance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.01283 v5 pith:BKSQLGWH submitted 2020-07-02 stat.ME

classification stat.ME
keywords floodgatevariableapplyemphconfidencedependenceerrorestimand
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Many modern applications seek to understand the relationship between an outcome variable $Y$ and a covariate $X$ in the presence of a (possibly high-dimensional) confounding variable $Z$. Although much attention has been paid to testing \emph{whether} $Y$ depends on $X$ given $Z$, in this paper we seek to go beyond testing by inferring the \emph{strength} of that dependence. We first define our estimand, the minimum mean squared error (mMSE) gap, which quantifies the conditional relationship between $Y$ and $X$ in a way that is deterministic, model-free, interpretable, and sensitive to nonlinearities and interactions. We then propose a new inferential approach called \emph{floodgate} that can leverage any working regression function chosen by the user (allowing, e.g., it to be fitted by a state-of-the-art machine learning algorithm or be derived from qualitative domain knowledge) to construct asymptotic confidence bounds, and we apply it to the mMSE gap. \acc{We additionally show that floodgate's accuracy (distance from confidence bound to estimand) is adaptive to the error of the working regression function.} We then show we can apply the same floodgate principle to a different measure of variable importance when $Y$ is binary. Finally, we demonstrate floodgate's performance in a series of simulations and apply it to data from the UK Biobank to infer the strengths of dependence of platelet count on various groups of genetic mutations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Conditional Mean Independence and Global Sensitivity Analysis using Nearest Neighbor Graphs

    stat.ME 2026-07 accept novelty 6.5 of 10

    A nearest-neighbor graph estimator of the normalized conditional mean discrepancy is consistent, rate-optimal in low dimension, asymptotically normal under the null, and yields a fast test and screening procedure.

  2. iLOCO: Distribution-Free Inference for Feature Interactions

    stat.ML 2025-02 conditional novelty 5.0 of 10

    iLOCO estimates the importance of pairwise and higher-order feature interactions in any predictive model and provides distribution-free confidence intervals via data splitting or minipatch ensembles.

Pith tools