Pith. sign in

REVIEW 5 cited by

Grokking as a First Order Phase Transition in Two Layer Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.03789 v3 pith:TZQHSYNF submitted 2023-10-05 stat.ML cond-mat.dis-nncs.LG

classification stat.MLcond-mat.dis-nncs.LG
keywords grokkingphaselearningfeaturetransitiondeepmixedmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A key property of deep neural networks (DNNs) is their ability to learn new features during training. This intriguing aspect of deep learning stands out most clearly in recently reported Grokking phenomena. While mainly reflected as a sudden increase in test accuracy, Grokking is also believed to be a beyond lazy-learning/Gaussian Process (GP) phenomenon involving feature learning. Here we apply a recent development in the theory of feature learning, the adaptive kernel approach, to two teacher-student models with cubic-polynomial and modular addition teachers. We provide analytical predictions on feature learning and Grokking properties of these models and demonstrate a mapping between Grokking and the theory of phase transitions. We show that after Grokking, the state of the DNN is analogous to the mixed phase following a first-order phase transition. In this mixed phase, the DNN generates useful internal representations of the teacher that are sharply distinct from those before the transition.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning words in groups: fusion algebras, tensor ranks and grokking

    cs.LG 2025-09 conditional novelty 8.0 of 10

    Group word operations can be learned by small two-layer networks because the associated word tensor has low rank, decomposable through the fusion algebra of the group's self-conjugate representations.

  2. Algebraic Representability as the Limiting Regime of Grokking: An Exactly Solvable Model with Holomorphic Activations

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A task ma+nb mod p is representable by a z^k holomorphic network iff m+n=k; non-representable tasks cannot be memorised at any width.

  3. Grokking vs. Learning: Same Features, Different Encodings

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Grokked and steadily trained models learn the same features, but steady training can produce much more compressible models in a parameter regime that grokking does not reach.

  4. Consciousness as a Jamming Phase

    cond-mat.dis-nn 2025-07 reject novelty 4.0 of 10

    Large language models are claimed to become conscious when their word embeddings jam into a critically correlated state, but no evidence is provided.

  5. Rethinking Over-Smoothing in Graph Neural Networks: A Perspective from Anderson Localization

    cs.LG 2025-06 reject novelty 4.0 of 10

    A single-author preprint re-frames GNN over-smoothing as Anderson localization, defining a participation-degree metric and proposing degree-dependent edge reweighting as mitigation, without proof or experiments.

Pith tools