REVIEW 20 cited by
Emergent Linear Representations in World Models of Self-Supervised Sequence Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
How do sequence models represent their decision-making process? Prior work suggests that Othello-playing neural network learned nonlinear models of the board state (Li et al., 2023). In this work, we provide evidence of a closely related linear representation of the board. In particular, we show that probing for "my colour" vs. "opponent's colour" may be a simple yet powerful way to interpret the model's internal state. This precise understanding of the internal representations allows us to control the model's behaviour with simple vector arithmetic. Linear representations enable significant interpretability progress, which we demonstrate with further exploration of how the world model is computed.
Forward citations
Cited by 20 Pith papers
-
How are linear representations learned? Exact solutions to the dynamics of abstraction
Exact solutions show abstraction is set by input/target geometry, rises with depth, peaks under small init, and is attenuated by nonlinearities—improving LLM probes via GELU ablation.
-
Context Is King: How In-Context Specification Shapes the Geometry of Concepts
In capable Gemma and Qwen models, declarative in-context rules set the relational geometry and topology type that the model represents and causally uses, overriding strong pretrained priors.
-
When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary
High reward in sparse RL does not imply latent-state recovery; a hidden-DFA instrument separates perception from planning gaps and flags group-language structure as a pre-training warning.
-
A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
A single linear direction in an observer model's residual stream detects contextual hallucinations, transfers across models and datasets, and causally steers generation hallucination rates.
-
Convergent Linear Representations of Emergent Misalignment
A single activation direction extracted from one misaligned fine-tune can ablate emergent misalignment across models trained on different datasets with different LoRA setups.
-
When Do Neural Networks Learn World Models?
With Boolean variables, a low-degree bias, and a task distribution weighted toward simple functions of the latents, multi-task training provably recovers the latent world model up to permutations and negations.
-
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias
LLM-as-judge scoring biases concentrate in low-dimensional, type-specific activation subspaces that support bidirectional causal steering and cross-domain failure prediction.
-
Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.
-
Geometry-Guided Constraint Learning for LLM Safety Classification
Sparse-autoencoder features reduce the number of safety constraints needed to two for most BeaverTails categories, and a three-phase-trained cone constraint modestly beats a flat polytope on in-distribution Qwen3.5-9B...
-
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
When transformer models see easy component examples before a harder combined math problem in one prompt, they solve unseen versions of the combined problem and store intermediate steps internally, unlike models traine...
-
Model Organisms for Emergent Misalignment
Emergent misalignment can be induced in small models via a single rank-1 LoRA adapter, and its onset coincides with a phase transition in the adapter's weight direction.
-
The Origins of Representation Manifolds in Large Language Models
A proof and empirical check that cosine similarity in language model representations encodes the intrinsic geometry of features through shortest paths on manifolds.
-
Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task
In an RL-trained transformer solving adjacent-swap sorting, larger embedding dimensions improve the monotonicity of an internal order-encoding in attention weights and the match to a largest-adjacent-difference swap r...
-
The Geometry of Harmfulness in LLMs through Subconcept Probing
Fifty-five harmfulness subconcept directions in Llama-3.1-8B-Instruct form a nearly rank-1 subspace, and steering along the dominant direction cuts jailbreak success but costs accuracy and fails on Qwen.
-
Large Language Models and Emergence: A Complex Systems Perspective
A perspective paper arguing that LLM emergence claims are incomplete without evidence of internal coarse-grained representations, and that LLMs have not shown emergent intelligence.
-
Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey
A survey showing that common systematic generalization benchmarks measure behavioural systematicity, not the representational systematicity that Fodor and Pylyshyn's challenge requires, and mapping them onto Hadley's ...
-
A Statistical Physics of Language Model Reasoning
A switching linear dynamical system on a 40-dimensional projection of LLM hidden states captures about half the variance of reasoning trajectories and predicts belief shifts during adversarial prompts.
-
Linear Spatial World Models Emerge in Large Language Models
Spatial relation words in LLaMA and Qwen models form antipodal, roughly orthogonal directions in a low-dimensional subspace, and steering along these directions changes the model's output.
-
Understanding the learned look-ahead behavior of chess neural networks
The Leela Chess Zero policy network encodes information about destination squares of moves up to seven plies ahead, with attention heads that copy future-square information backward in time in a pattern-dependent way.
-
Towards Atoms of Large Language Models
The authors define 'atoms' as sparse, near-orthogonal directions in LLM representations under a data-adaptive inner product, and show threshold-activated sparse autoencoders can recover them with about 99.9% reconstru...
Discussion (0). Continue with ORCID to comment.