REVIEW 5 cited by
Grokfast: Accelerated Grokking by Amplifying Slow Gradients
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
One puzzling artifact in machine learning dubbed grokking is where delayed generalization is achieved tenfolds of iterations after near perfect overfitting to the training data. Focusing on the long delay itself on behalf of machine learning practitioners, our goal is to accelerate generalization of a model under grokking phenomenon. By regarding a series of gradients of a parameter over training iterations as a random signal over time, we can spectrally decompose the parameter trajectories under gradient descent into two components: the fast-varying, overfitting-yielding component and the slow-varying, generalization-inducing component. This analysis allows us to accelerate the grokking phenomenon more than $\times 50$ with only a few lines of code that amplifies the slow-varying components of gradients. The experiments show that our algorithm applies to diverse tasks involving images, languages, and graphs, enabling practical availability of this peculiar artifact of sudden generalization. Our code is available at https://github.com/ironjr/grokfast.
Forward citations
Cited by 5 Pith papers
-
How to Tame Grokking: Representation Geometry as a Control Signal
Dimensionality collapse precedes grokking; GeomDR, a spectral regularizer on hidden covariances, accelerates it up to 52× on modular and permutation tasks for MLPs and transformers.
-
The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models
Capability formation in small transformers is claimed to obey a driven-nucleation rate law J = Nνσ(c)e^{−βK} − D, read forward as emergence, backward as plasticity loss, and completed as circuit control.
-
The Active Ingredient in Muon's Grokking
Orthogonalization—not spectral scaling—is the active ingredient in Muon's faster grokking, and reducing Newton-Schulz iterations trades first-crossing speed for solution stability.
-
GrokAlign: Geometric Characterisation and Acceleration of Grokking
GrokAlign, a Jacobian-norm regularizer, accelerates grokking by aligning Jacobians with training data, and centroid alignment tracks when generalization and robustness emerge.
-
Mechanistic Insights into Grokking from the Embedding Layer
Trainable embeddings in a simple MLP cause delayed generalization (grokking) on modular arithmetic, and a higher embedding learning rate plus balanced sampling accelerates it.
Discussion (0). Continue with ORCID to comment.