Pith. sign in

REVIEW 1 cited by

Towards a statistical theory of data selection under weak supervision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.14563 v2 pith:WXCGVGMC submitted 2023-09-25 stat.ML cs.LG

classification stat.MLcs.LG
keywords dataselectiongivensizeboldsymbollabelslearningmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Given a sample of size $N$, it is often useful to select a subsample of smaller size $n<N$ to be used for statistical estimation or learning. Such a data selection step is useful to reduce the requirements of data labeling and the computational complexity of learning. We assume to be given $N$ unlabeled samples $\{{\boldsymbol x}_i\}_{i\le N}$, and to be given access to a `surrogate model' that can predict labels $y_i$ better than random guessing. Our goal is to select a subset of the samples, to be denoted by $\{{\boldsymbol x}_i\}_{i\in G}$, of size $|G|=n<N$. We then acquire labels for this set and we use them to train a model via regularized empirical risk minimization. By using a mixture of numerical experiments on real and synthetic data, and mathematical derivations under low- and high- dimensional asymptotics, we show that: $(i)$~Data selection can be very effective, in particular beating training on the full sample in some cases; $(ii)$~Certain popular choices in data selection methods (e.g. unbiased reweighted subsampling, or influence function-based subsampling) can be substantially suboptimal.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Influence Diagnostics in High-dimensional M-estimation: Precise Asymptotics

    stat.ML 2026-07 accept novelty 7.0 of 10

    Under Gaussian design with n ≍ d, the empirical distribution of leave-one-out influences for convex M-estimators converges to the pushforward of a four-dimensional Gaussian through an explicit nonlinear map built from...

Pith tools