REVIEW 5 cited by
Deep Networks Always Grok and Here is Why
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Grokking, or delayed generalization, is a phenomenon where generalization in a deep neural network (DNN) occurs long after achieving near zero training error. Previous studies have reported the occurrence of grokking in specific controlled settings, such as DNNs initialized with large-norm parameters or transformers trained on algorithmic datasets. We demonstrate that grokking is actually much more widespread and materializes in a wide range of practical settings, such as training of a convolutional neural network (CNN) on CIFAR10 or a Resnet on Imagenette. We introduce the new concept of delayed robustness, whereby a DNN groks adversarial examples and becomes robust, long after interpolation and/or generalization. We develop an analytical explanation for the emergence of both delayed generalization and delayed robustness based on the local complexity of a DNN's input-output mapping. Our local complexity measures the density of so-called linear regions (aka, spline partition regions) that tile the DNN input space and serves as a utile progress measure for training. We provide the first evidence that, for classification problems, the linear regions undergo a phase transition during training whereafter they migrate away from the training samples (making the DNN mapping smoother there) and towards the decision boundary (making the DNN mapping less smooth there). Grokking occurs post phase transition as a robust partition of the input space thanks to the linearization of the DNN mapping around the training points. Website: https://bit.ly/grok-adversarial
Forward citations
Cited by 5 Pith papers
-
Learning words in groups: fusion algebras, tensor ranks and grokking
Group word operations can be learned by small two-layer networks because the associated word tensor has low rank, decomposable through the fusion algebra of the group's self-conjugate representations.
-
Emergent Generalization by Representation Learning in Artificial Neural Networks
An explicit low-dimensional bottleneck is necessary for OOD generalisation in reservoir networks, and the non-monotonic rise of causal emergence in the latent code predicts generalisation both in silico and in mouse CA1.
-
The Geometry of Grokking: Norm Minimization on the Zero-Loss Manifold
Post-memorization learning in grokking is equivalent to minimizing the weight norm on the zero-loss manifold, with a closed-form approximation for two-layer networks.
-
Grokking vs. Learning: Same Features, Different Encodings
Grokked and steadily trained models learn the same features, but steady training can produce much more compressible models in a parameter regime that grokking does not reach.
-
Mechanistic Insights into Grokking from the Embedding Layer
Trainable embeddings in a simple MLP cause delayed generalization (grokking) on modular arithmetic, and a higher embedding learning rate plus balanced sampling accelerates it.
Discussion (0). Continue with ORCID to comment.