Pith. sign in

REVIEW 15 cited by

Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.09458 v3 pith:AZZECWFC submitted 2023-07-18 cs.LG

classification cs.LG
keywords analysisanswerchinchillacircuitcorrectheadsattentionemph
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

\emph{Circuit analysis} is a promising technique for understanding the internal mechanisms of language models. However, existing analyses are done in small models far from the state of the art. To address this, we present a case study of circuit analysis in the 70B Chinchilla model, aiming to test the scalability of circuit analysis. In particular, we study multiple-choice question answering, and investigate Chinchilla's capability to identify the correct answer \emph{label} given knowledge of the correct answer \emph{text}. We find that the existing techniques of logit attribution, attention pattern visualization, and activation patching naturally scale to Chinchilla, allowing us to identify and categorize a small set of `output nodes' (attention heads and MLPs). We further study the `correct letter' category of attention heads aiming to understand the semantics of their features, with mixed results. For normal multiple-choice question answers, we significantly compress the query, key and value subspaces of the head without loss of performance when operating on the answer labels for multiple-choice questions, and we show that the query and key subspaces represent an `Nth item in an enumeration' feature to at least some extent. However, when we attempt to use this explanation to understand the heads' behaviour on a more general distribution including randomized answer labels, we find that it is only a partial explanation, suggesting there is more to learn about the operation of `correct letter' heads on multiple choice question answering.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Context Is King: How In-Context Specification Shapes the Geometry of Concepts

    cs.LG 2026-07 accept novelty 7.5 of 10

    In capable Gemma and Qwen models, declarative in-context rules set the relational geometry and topology type that the model represents and causally uses, overriding strong pretrained priors.

  2. Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Alignment on Llama 3 reduces explicit bias but amplifies implicit bias, because aligned models no longer represent 'black' and 'white' as racial concepts in ambiguous contexts.

  3. Pretraining Curricula Enable Selective Fine-tuning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Imbalanced pretraining curricula disentangle task circuits in transformers, improving in-context learning and the selectivity of refusal fine-tuning relative to balanced training.

  4. Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations

    cs.CL 2026-01 reject novelty 6.0 of 10

    A greedy gradient-component selection appears to beat full gradients and random projection for training-data retrieval, but the comparison is confounded because components are selected on the evaluation set itself.

  5. Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language Models

    cs.LG 2026-01 conditional novelty 6.0 of 10

    Masked diffusion post-training keeps autoregressive circuitry on locally causal tasks but reorganizes models into early-layer, distributed computation on global planning tasks.

  6. The Mechanistic Emergence of Symbol Grounding in Language Models

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Symbol grounding emerges in Transformers and state-space models through middle-layer 'aggregate' attention heads that connect environmental cues to words, but not in unidirectional LSTMs.

  7. BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A self-supervision method makes multimodal LLMs align their input image embeddings with the model's own refined internal representations, improving visual QA scores over LLaVA baselines.

  8. Spectral Principal Paths: A Spectral Perspective on Linear Representation Formation in LLMs

    cs.CV 2025-06 reject novelty 6.0 of 10

    The paper argues and partially tests that linear concept representations in LLMs originate in the input space and propagate through the leading singular directions of activation difference matrices.

  9. Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Over 40 checkpoints of OLMo-7B, attention heads and FFNs shift from general-purpose to specialized roles for factual recall, with location-based facts learned earlier and more stably than name-based facts.

  10. Circuit Stability Characterizes Language Model Generalization

    cs.CL 2025-05 reject novelty 6.0 of 10

    Circuit stability, measured as rank correlation between soft circuits across subtasks, is proposed as a predictor of language model generalization.

  11. On Mechanistic Circuits for Extractive Question-Answering

    cs.CL 2025-02 conditional novelty 6.0 of 10

    One attention head from the extracted context-faithfulness circuit provides reliable extractive QA attribution and improves context faithfulness when its attributions are added to the prompt.

  12. Mechanistic Interpretability of Emotion Inference in Large Language Models

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Emotion inference in LLMs is localized to mid-layer attention and feed-forward units, and steering learned appraisal directions shifts generated emotions in appraisal-theory-consistent ways.

  13. EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification

    cs.LG 2025-02 conditional novelty 5.0 of 10

    EAP-GP adapts the integration path in edge attribution patching to avoid gradient saturation, improving circuit faithfulness on GPT-2 models.

  14. Hierarchical Sparse Circuit Extraction from Billion-Parameter Language Models through Scalable Attribution Graph Decomposition

    cs.LG 2026-01 reject novelty 4.0 of 10

    HAGD claims to extract sparse circuits from billion-parameter LMs by hierarchical graph coarsening and GNN-guided search, but the O(n^2 log n) complexity guarantee rests on an unproven greedy-optimality assumption.

  15. You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation

    cs.LG 2025-02 conditional novelty 3.0 of 10

    To align powerful AI, researchers must understand how statistical patterns in training data shape the internal structure of models, because that structure, not eval scores, determines generalization.

Pith tools