Pith. sign in

REVIEW 4 cited by

Grokking phase transitions in learning local rules with gradient descent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.15435 v1 pith:HIMDRPSV submitted 2022-10-26 cond-mat.stat-mech cs.LG

classification cond-mat.stat-mechcs.LG
keywords grokkinglearninganalysecriticalmodelnumericallyphaseproposed
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We discuss two solvable grokking (generalisation beyond overfitting) models in a rule learning scenario. We show that grokking is a phase transition and find exact analytic expressions for the critical exponents, grokking probability, and grokking time distribution. Further, we introduce a tensor-network map that connects the proposed grokking setup with the standard (perceptron) statistical learning theory and show that grokking is a consequence of the locality of the teacher model. As an example, we analyse the cellular automata learning task, numerically determine the critical exponent and the grokking time distributions and compare them with the prediction of the proposed grokking model. Finally, we numerically analyse the connection between structure formation and grokking.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Algebraic Representability as the Limiting Regime of Grokking: An Exactly Solvable Model with Holomorphic Activations

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A task ma+nb mod p is representable by a z^k holomorphic network iff m+n=k; non-representable tasks cannot be memorised at any width.

  2. Let Me Grok for You: Accelerating Grokking via Embedding Transfer from a Weaker Model

    cs.LG 2025-04 conditional novelty 6.0 of 10

    GrokTransfer transfers an embedding learned by a small 'weaker' model to a larger model, eliminating the grokking delay so the target model generalizes almost immediately.

  3. BlueGlass: A Framework for Composite AI Safety

    cs.AI 2025-07 conditional novelty 5.0 of 10

    BlueGlass provides composite AI safety infrastructure; its case studies on object-detection VLMs reveal dataset trade-offs, a decoder-layer phase transition in probe accuracy, and SAE-discovered concepts including spu...

  4. Grokking Explained: A Statistical Phenomenon

    cs.LG 2025-02 reject novelty 5.0 of 10

    Grokking can be triggered systematically by shifting the training distribution through imbalanced subclass sampling, even with dense data and little hyperparameter tuning.

Pith tools