Pith. sign in

REVIEW 2 cited by

When Representations Align: Universality in Representation Learning Dynamics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.09142 v2 pith:LYENXNQS submitted 2024-02-14 cs.LG q-bio.NC

classification cs.LGq-bio.NC
keywords representationlearningarchitecturesrepresentationsdynamicstheoryarchitecturebehaviors
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep neural networks come in many sizes and architectures. The choice of architecture, in conjunction with the dataset and learning algorithm, is commonly understood to affect the learned neural representations. Yet, recent results have shown that different architectures learn representations with striking qualitative similarities. Here we derive an effective theory of representation learning under the assumption that the encoding map from input to hidden representation and the decoding map from representation to output are arbitrary smooth functions. This theory schematizes representation learning dynamics in the regime of complex, large architectures, where hidden representations are not strongly constrained by the parametrization. We show through experiments that the effective theory describes aspects of representation learning dynamics across a range of deep networks with different activation functions and architectures, and exhibits phenomena similar to the "rich" and "lazy" regime. While many network behaviors depend quantitatively on architecture, our findings point to certain behaviors that are widely conserved once models are sufficiently flexible.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Disentangling the Factors of Convergence between Brains and Computer Vision Models

    cs.AI 2025-08 unverdicted novelty 7.0 of 10

    By systematically varying model size, training amount, and image type in DINOv3 vision transformers, this paper shows that brain similarity increases with scale and human-centric data and emerges in a characteristic t...

  2. The Geometry of Grokking: Norm Minimization on the Zero-Loss Manifold

    cs.LG 2025-11 conditional novelty 6.0 of 10

    Post-memorization learning in grokking is equivalent to minimizing the weight norm on the zero-loss manifold, with a closed-form approximation for two-layer networks.

Pith tools