Pith. sign in

REVIEW 6 cited by

The Slingshot Mechanism: An Empirical Study of Adaptive Optimizers and the Grokking Phenomenon

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.04817 v2 pith:J3PUJK7A submitted 2022-06-10 cs.LG math.OC

classification cs.LGmath.OC
keywords grokkingmechanismslingshotadaptiveeasilyoptimizerstrainingwithout
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The grokking phenomenon as reported by Power et al. ( arXiv:2201.02177 ) refers to a regime where a long period of overfitting is followed by a seemingly sudden transition to perfect generalization. In this paper, we attempt to reveal the underpinnings of Grokking via a series of empirical studies. Specifically, we uncover an optimization anomaly plaguing adaptive optimizers at extremely late stages of training, referred to as the Slingshot Mechanism. A prominent artifact of the Slingshot Mechanism can be measured by the cyclic phase transitions between stable and unstable training regimes, and can be easily monitored by the cyclic behavior of the norm of the last layers weights. We empirically observe that without explicit regularization, Grokking as reported in ( arXiv:2201.02177 ) almost exclusively happens at the onset of Slingshots, and is absent without it. While common and easily reproduced in more general settings, the Slingshot Mechanism does not follow from any known optimization theories that we are aware of, and can be easily overlooked without an in depth examination. Our work points to a surprising and useful inductive bias of adaptive gradient optimizers at late stages of training, calling for a revised theoretical analysis of their origin.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning words in groups: fusion algebras, tensor ranks and grokking

    cs.LG 2025-09 conditional novelty 8.0 of 10

    Group word operations can be learned by small two-layer networks because the associated word tensor has low rank, decomposable through the fusion algebra of the group's self-conjugate representations.

  2. Thermodynamic Weight Decay: Exploring Grokking Acceleration via Attention Specific Heat

    cs.LG 2026-07 conditional novelty 6.0 of 10

    An optimizer that boosts weight decay when attention-logit variance spikes speeded grokking on a+b mod 97, but the effect is small, single-task, and not yet shown to come from the signal itself.

  3. The Active Ingredient in Muon's Grokking

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Orthogonalization—not spectral scaling—is the active ingredient in Muon's faster grokking, and reducing Newton-Schulz iterations trades first-crossing speed for solution stability.

  4. Grokking Is Conditional and Fragile: A Fully-Tractable, Multi-Seed Study at 12K Parameters

    cs.LG 2026-07 accept novelty 6.0 of 10

    In a fully tractable 12K Llama-style model, grokking is a conditional fragile phase transition gated by coverage (tracking modulus more than structure), weight decay, and floating-point reduction order, so evidence mu...

  5. Beamforming Feedback as a Novel Attack Surface for Wi-Fi Physical-Layer Security

    cs.CR 2026-04 unverdicted novelty 6.0 of 10

    BFIAttack reconstructs legitimate CSI from public beamforming feedback, reportedly succeeding over 93% (single-antenna) and about 73% (multi-antenna) against Wi-Fi physical-layer security.

  6. Grokking vs. Learning: Same Features, Different Encodings

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Grokked and steadily trained models learn the same features, but steady training can produce much more compressible models in a parameter regime that grokking does not reach.

Pith tools