Pith. sign in

REVIEW 3 cited by

Learning from a Biased Sample

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.01754 v5 pith:U7WAR3LY submitted 2022-09-05 stat.ME cs.LGstat.ML

classification stat.MEcs.LGstat.ML
keywords learningbiasedrisksamplesamplingtrainingdecisionmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The empirical risk minimization approach to data-driven decision making requires access to training data drawn under the same conditions as those that will be faced when the decision rule is deployed. However, in a number of settings, we may be concerned that our training sample is biased in the sense that some groups (characterized by either observable or unobservable attributes) may be under- or over-represented relative to the general population; and in this setting empirical risk minimization over the training set may fail to yield rules that perform well at deployment. We propose a model of sampling bias called conditional $\Gamma$-biased sampling, where observed covariates can affect the probability of sample selection arbitrarily much but the amount of unexplained variation in the probability of sample selection is bounded by a constant factor. Applying the distributionally robust optimization framework, we propose a method for learning a decision rule that minimizes the worst-case risk incurred under a family of test distributions that can generate the training distribution under $\Gamma$-biased sampling. We apply a result of Rockafellar and Uryasev to show that this problem is equivalent to an augmented convex risk minimization problem. We give statistical guarantees for learning a model that is robust to sampling bias via the method of sieves, and propose a deep learning algorithm whose loss function captures our robust learning target. We empirically validate our proposed method in a case study on prediction of mental health scores from health survey data and a case study on ICU length of stay prediction.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Practical Upper Bound on Selection Bias Effects in Medical Prediction Models

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    A new upper bound is derived for the worst-case effect of selection bias on medical prediction model performance under partial observation of the selection process and target data.

  2. Adversarially Robust Control of Conditional Value-at-Risk via Rockafellar-Uryasev Conformal Inference

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Online conformal framework for adversarial CVaR control with asymptotic guarantees and regret bounds, demonstrated on portfolio management and LLM toxicity mitigation.

  3. DRO: A Python Library for Distributionally Robust Optimization in Machine Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    The dro library provides a unified implementation of 14 DRO formulations across 9 model backbones, with claims of large speedups from vectorization and approximation.

Pith tools