Pith. sign in

REVIEW 3 cited by

Transformers represent belief state geometry in their residual stream

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.15943 v3 pith:3P3TYENE submitted 2024-05-24 cs.LG cs.CL

classification cs.LGcs.CL
keywords beliefstructureresidualtransformersgeometrypredictionstatestates
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

What computational structure are we building into large language models when we train them on next-token prediction? Here, we present evidence that this structure is given by the meta-dynamics of belief updating over hidden states of the data-generating process. Leveraging the theory of optimal prediction, we anticipate and then find that belief states are linearly represented in the residual stream of transformers, even in cases where the predicted belief state geometry has highly nontrivial fractal structure. We investigate cases where the belief state geometry is represented in the final residual stream or distributed across the residual streams of multiple layers, providing a framework to explain these observations. Furthermore we demonstrate that the inferred belief states contain information about the entire future, beyond the local next-token prediction that the transformers are explicitly trained on. Our work provides a general framework connecting the structure of training data to the geometric structure of activations inside transformers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Neural networks leverage nominally quantum and post-quantum representations

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Pretrained transformers and RNNs linearly encode the Bayesian-updated belief geometry of the minimal classical, quantum, or post-quantum generator of their training data.

  2. Belief-Guided Decision Making with Uncertainty Gating in the Game of Go

    cs.AI 2026-07 reject novelty 4.0 of 10

    A disentangled Belief head with uncertainty gating is claimed to replace MCTS correction and enable professional-level search-free Go on consumer GPUs, but the reported experiments do not demonstrate that claim.

  3. A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks

    cs.CL 2025-02 conditional novelty 3.0 of 10

    A narrative review of behavioral and representational Theory of Mind in LLMs, with a taxonomy of safety risks and mitigation directions.

Pith tools