REVIEW 4 cited by
The Deep Bootstrap Framework: Good Online Learners are Good Offline Generalizers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose a new framework for reasoning about generalization in deep learning. The core idea is to couple the Real World, where optimizers take stochastic gradient steps on the empirical loss, to an Ideal World, where optimizers take steps on the population loss. This leads to an alternate decomposition of test error into: (1) the Ideal World test error plus (2) the gap between the two worlds. If the gap (2) is universally small, this reduces the problem of generalization in offline learning to the problem of optimization in online learning. We then give empirical evidence that this gap between worlds can be small in realistic deep learning settings, in particular supervised image classification. For example, CNNs generalize better than MLPs on image distributions in the Real World, but this is "because" they optimize faster on the population loss in the Ideal World. This suggests our framework is a useful tool for understanding generalization in deep learning, and lays a foundation for future research in the area.
Forward citations
Cited by 4 Pith papers
-
Querying Kernel Methods Suffices for Reconstructing their Training Data
Query-only access to kernel regression, SVM and KDE models suffices to reconstruct their exact training points, via a measure-theoretic proof and image experiments.
-
Model-based Bootstrap of Controlled Markov Chains
A model-based bootstrap for finite CMCs is distributionally consistent for transitions and, via the delta method, for OPE/OPR targets under nonstationary behavior policies.
-
Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
Compute-optimally trained networks of different sizes show loss curves that collapse onto one universal curve after normalization; with learning rate decay, the collapse is tighter than seed-to-seed noise, providing a...
-
Not All Explanations for Deep Learning Phenomena Are Equally Valuable
A position paper arguing that narrow, puzzle-solving explanations of deep learning edge case phenomena are low-value, and that these phenomena should instead be used to stress-test broad explanatory theories.
Discussion (0). Continue with ORCID to comment.