Pith. sign in

REVIEW 2 cited by

Structured World Representations in Maze-Solving Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.02566 v1 pith:OWKSA6D4 submitted 2023-12-05 cs.LG cs.AI

classification cs.LGcs.AI
keywords modelsfindheadsinternalmazerepresentationsstructuredtokens
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Transformer models underpin many recent advances in practical machine learning applications, yet understanding their internal behavior continues to elude researchers. Given the size and complexity of these models, forming a comprehensive picture of their inner workings remains a significant challenge. To this end, we set out to understand small transformer models in a more tractable setting: that of solving mazes. In this work, we focus on the abstractions formed by these models and find evidence for the consistent emergence of structured internal representations of maze topology and valid paths. We demonstrate this by showing that the residual stream of only a single token can be linearly decoded to faithfully reconstruct the entire maze. We also find that the learned embeddings of individual tokens have spatial structure. Furthermore, we take steps towards deciphering the circuity of path-following by identifying attention heads (dubbed $\textit{adjacency heads}$), which are implicated in finding valid subsequent tokens.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Linear Spatial World Models Emerge in Large Language Models

    cs.AI 2025-06 reject novelty 5.0 of 10

    Spatial relation words in LLaMA and Qwen models form antipodal, roughly orthogonal directions in a low-dimensional subspace, and steering along these directions changes the model's output.

  2. Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning

    cs.AI 2025-02 unverdicted novelty 4.0 of 10

    A perspective paper that advocates direct, post hoc interpretability for multi-agent deep reinforcement learning and offers a taxonomy of where those methods might apply.

Pith tools