Pith. sign in

REVIEW 4 cited by

The Deep Bootstrap Framework: Good Online Learners are Good Offline Generalizers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.08127 v2 pith:I3EZRICW submitted 2020-10-16 cs.LG cs.CVcs.NEmath.STstat.MLstat.TH

classification cs.LGcs.CVcs.NEmath.STstat.MLstat.TH
keywords learningworlddeepframeworkgeneralizationideallossempirical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a new framework for reasoning about generalization in deep learning. The core idea is to couple the Real World, where optimizers take stochastic gradient steps on the empirical loss, to an Ideal World, where optimizers take steps on the population loss. This leads to an alternate decomposition of test error into: (1) the Ideal World test error plus (2) the gap between the two worlds. If the gap (2) is universally small, this reduces the problem of generalization in offline learning to the problem of optimization in online learning. We then give empirical evidence that this gap between worlds can be small in realistic deep learning settings, in particular supervised image classification. For example, CNNs generalize better than MLPs on image distributions in the Real World, but this is "because" they optimize faster on the population loss in the Ideal World. This suggests our framework is a useful tool for understanding generalization in deep learning, and lays a foundation for future research in the area.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Querying Kernel Methods Suffices for Reconstructing their Training Data

    cs.LG 2025-05 conditional novelty 8.0 of 10

    Query-only access to kernel regression, SVM and KDE models suffices to reconstruct their exact training points, via a measure-theoretic proof and image experiments.

  2. Model-based Bootstrap of Controlled Markov Chains

    stat.ML 2026-05 unverdicted novelty 7.0 of 10

    A model-based bootstrap for finite CMCs is distributionally consistent for transitions and, via the delta method, for OPE/OPR targets under nonstationary behavior policies.

  3. Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks

    cs.LG 2025-07 conditional novelty 7.0 of 10

    Compute-optimally trained networks of different sizes show loss curves that collapse onto one universal curve after normalization; with learning rate decay, the collapse is tighter than seed-to-seed noise, providing a...

  4. Not All Explanations for Deep Learning Phenomena Are Equally Valuable

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A position paper arguing that narrow, puzzle-solving explanations of deep learning edge case phenomena are low-value, and that these phenomena should instead be used to stress-test broad explanatory theories.

Pith tools