REVIEW 15 cited by
Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
\emph{Circuit analysis} is a promising technique for understanding the internal mechanisms of language models. However, existing analyses are done in small models far from the state of the art. To address this, we present a case study of circuit analysis in the 70B Chinchilla model, aiming to test the scalability of circuit analysis. In particular, we study multiple-choice question answering, and investigate Chinchilla's capability to identify the correct answer \emph{label} given knowledge of the correct answer \emph{text}. We find that the existing techniques of logit attribution, attention pattern visualization, and activation patching naturally scale to Chinchilla, allowing us to identify and categorize a small set of `output nodes' (attention heads and MLPs). We further study the `correct letter' category of attention heads aiming to understand the semantics of their features, with mixed results. For normal multiple-choice question answers, we significantly compress the query, key and value subspaces of the head without loss of performance when operating on the answer labels for multiple-choice questions, and we show that the query and key subspaces represent an `Nth item in an enumeration' feature to at least some extent. However, when we attempt to use this explanation to understand the heads' behaviour on a more general distribution including randomized answer labels, we find that it is only a partial explanation, suggesting there is more to learn about the operation of `correct letter' heads on multiple choice question answering.
Forward citations
Cited by 15 Pith papers
-
Context Is King: How In-Context Specification Shapes the Geometry of Concepts
In capable Gemma and Qwen models, declarative in-context rules set the relational geometry and topology type that the model represents and causally uses, overriding strong pretrained priors.
-
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
Alignment on Llama 3 reduces explicit bias but amplifies implicit bias, because aligned models no longer represent 'black' and 'white' as racial concepts in ambiguous contexts.
-
Pretraining Curricula Enable Selective Fine-tuning
Imbalanced pretraining curricula disentangle task circuits in transformers, improving in-context learning and the selectivity of refusal fine-tuning relative to balanced training.
-
Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations
A greedy gradient-component selection appears to beat full gradients and random projection for training-data retrieval, but the comparison is confounded because components are selected on the evaluation set itself.
-
Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language Models
Masked diffusion post-training keeps autoregressive circuitry on locally causal tasks but reorganizes models into early-layer, distributed computation on global planning tasks.
-
The Mechanistic Emergence of Symbol Grounding in Language Models
Symbol grounding emerges in Transformers and state-space models through middle-layer 'aggregate' attention heads that connect environmental cues to words, but not in unidirectional LSTMs.
-
BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
A self-supervision method makes multimodal LLMs align their input image embeddings with the model's own refined internal representations, improving visual QA scores over LLaVA baselines.
-
Spectral Principal Paths: A Spectral Perspective on Linear Representation Formation in LLMs
The paper argues and partially tests that linear concept representations in LLMs originate in the input space and propagate through the leading singular directions of activation difference matrices.
-
Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models
Over 40 checkpoints of OLMo-7B, attention heads and FFNs shift from general-purpose to specialized roles for factual recall, with location-based facts learned earlier and more stably than name-based facts.
-
Circuit Stability Characterizes Language Model Generalization
Circuit stability, measured as rank correlation between soft circuits across subtasks, is proposed as a predictor of language model generalization.
-
On Mechanistic Circuits for Extractive Question-Answering
One attention head from the extracted context-faithfulness circuit provides reliable extractive QA attribution and improves context faithfulness when its attributions are added to the prompt.
-
Mechanistic Interpretability of Emotion Inference in Large Language Models
Emotion inference in LLMs is localized to mid-layer attention and feed-forward units, and steering learned appraisal directions shifts generated emotions in appraisal-theory-consistent ways.
-
EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification
EAP-GP adapts the integration path in edge attribution patching to avoid gradient saturation, improving circuit faithfulness on GPT-2 models.
-
Hierarchical Sparse Circuit Extraction from Billion-Parameter Language Models through Scalable Attribution Graph Decomposition
HAGD claims to extract sparse circuits from billion-parameter LMs by hierarchical graph coarsening and GNN-guided search, but the O(n^2 log n) complexity guarantee rests on an unproven greedy-optimality assumption.
-
You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation
To align powerful AI, researchers must understand how statistical patterns in training data shape the internal structure of models, because that structure, not eval scores, determines generalization.
Discussion (0). Continue with ORCID to comment.