Pith. sign in

REVIEW 3 cited by

Dynamics of Transient Structure in In-Context Linear Regression Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.17745 v2 pith:UUWHL3U5 submitted 2025-01-29 cs.LG

classification cs.LG
keywords structuretransformersregressiontransientcomplexitydeepexplanationgeneral
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern deep neural networks display striking examples of rich internal computational structure. Uncovering principles governing the development of such structure is a priority for the science of deep learning. In this paper, we explore the transient ridge phenomenon: when transformers are trained on in-context linear regression tasks with intermediate task diversity, they initially behave like ridge regression before specializing to the tasks in their training distribution. This transition from a general solution to a specialized solution is revealed by joint trajectory principal component analysis. Further, we draw on the theory of Bayesian internal model selection to suggest a general explanation for the phenomena of transient structure in transformers, based on an evolving tradeoff between loss and complexity. We empirically validate this explanation by measuring the model complexity of our transformers as defined by the local learning coefficient.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Influence Dynamics and Stagewise Data Attribution

    cs.LG 2025-10 conditional novelty 7.0 of 10

    Using Bayesian influence functions and singular learning theory, the authors show that a sample's influence on a model varies non-monotonically over training, peaking and flipping sign at phase transitions.

  2. What does a Bayes-filtered transformer believe? A predictive Monte Carlo approach

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Predictive Monte Carlo recovers a Bayes-filtered transformer's implicit prior and posterior over latent tasks from next-token rollouts alone; the task-diversity threshold and transient generalization appear in this re...

  3. You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation

    cs.LG 2025-02 conditional novelty 3.0 of 10

    To align powerful AI, researchers must understand how statistical patterns in training data shape the internal structure of models, because that structure, not eval scores, determines generalization.

Pith tools