Pith. sign in

REVIEW 2 cited by

Small Data, Big Decisions: Model Selection in the Small-Data Regime

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.12583 v1 pith:HPX4RAAQ submitted 2020-09-26 cs.LG stat.ML

classification cs.LGstat.ML
keywords modelperformanceselectiondatadecisionsexperimentsgeneralizationlead
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Highly overparametrized neural networks can display curiously strong generalization performance - a phenomenon that has recently garnered a wealth of theoretical and empirical research in order to better understand it. In contrast to most previous work, which typically considers the performance as a function of the model size, in this paper we empirically study the generalization performance as the size of the training set varies over multiple orders of magnitude. These systematic experiments lead to some interesting and potentially very useful observations; perhaps most notably that training on smaller subsets of the data can lead to more reliable model selection decisions whilst simultaneously enjoying smaller computational costs. Our experiments furthermore allow us to estimate Minimum Description Lengths for common datasets given modern neural network architectures, thereby paving the way for principled model selection taking into account Occams-razor.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semivalue-based data valuation is arbitrary and gameable

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Semivalue-based data valuations are shown to be highly sensitive to plausible utility-function choices and are gameable under the paper's weak definition of gameability.

  2. Small Data Explainer -- The impact of small data methods in everyday life

    cs.CY 2025-07 conditional novelty 3.0 of 10

    A review and explainer that frames small data methods through the recurring challenges of similarity, transfer, and uncertainty and maps them to application areas and technical approaches.

Pith tools