Pith. sign in

REVIEW 5 cited by

Successor Heads: Recurring, Interpretable Attention Heads In The Wild

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.09230 v1 pith:CIILBVIT submitted 2023-12-14 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords headssuccessormodelsbehaviorlanguagellmsexplainhead
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work we present successor heads: attention heads that increment tokens with a natural ordering, such as numbers, months, and days. For example, successor heads increment 'Monday' into 'Tuesday'. We explain the successor head behavior with an approach rooted in mechanistic interpretability, the field that aims to explain how models complete tasks in human-understandable terms. Existing research in this area has found interpretable language model components in small toy models. However, results in toy models have not yet led to insights that explain the internals of frontier models and little is currently understood about the internal operations of large language models. In this paper, we analyze the behavior of successor heads in large language models (LLMs) and find that they implement abstract representations that are common to different architectures. They form in LLMs with as few as 31 million parameters, and at least as many as 12 billion parameters, such as GPT-2, Pythia, and Llama-2. We find a set of 'mod-10 features' that underlie how successor heads increment in LLMs across different architectures and sizes. We perform vector arithmetic with these features to edit head behavior and provide insights into numeric representations within LLMs. Additionally, we study the behavior of successor heads on natural language data, identifying interpretable polysemanticity in a Pythia successor head.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

    eess.AS 2025-08 conditional novelty 6.0 of 10

    Attention outputs in transformers occupy a subspace with about 60% effective rank, and starting sparse dictionaries inside that subspace reduces dead features from 87% to below 1%.

  2. Modular Arithmetic: Language Models Solve Math Digit by Digit

    cs.CL 2025-08 conditional novelty 6.0 of 10

    LLMs perform 3-digit addition and subtraction via digit-position-specific MLP circuits that can be intervened upon to change individual output digits.

  3. On Mechanistic Circuits for Extractive Question-Answering

    cs.CL 2025-02 conditional novelty 6.0 of 10

    One attention head from the extracted context-faithfulness circuit provides reliable extractive QA attribution and improves context faithfulness when its attributions are added to the prompt.

  4. GRAIT: Gradient-Driven Refusal-Aware Instruction Tuning for Effective Hallucination Mitigation

    cs.CL 2025-02 conditional novelty 6.0 of 10

    GRAIT selects and reweights refusal-training examples using gradient influence, reporting lower hallucination rates and better helpfulness scores than prior refusal-aware tuning baselines.

  5. Model Science: getting serious about verification, explanation and control of AI systems

    cs.AI 2025-08 conditional novelty 4.0 of 10

    Proposes 'Model Science' as a model-centric paradigm for AI with four pillars: verification, explanation, control, and interface.

Pith tools