Pith. sign in

REVIEW 9 cited by

Attribution Patching Outperforms Automated Circuit Discovery

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.10348 v2 pith:IWEM3X4D submitted 2023-10-16 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords patchingautomatedcircuitmethodactivationapproximationattributiondiscovery
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automated interpretability research has recently attracted attention as a potential research direction that could scale explanations of neural network behavior to large models. Existing automated circuit discovery work applies activation patching to identify subnetworks responsible for solving specific tasks (circuits). In this work, we show that a simple method based on attribution patching outperforms all existing methods while requiring just two forward passes and a backward pass. We apply a linear approximation to activation patching to estimate the importance of each edge in the computational subgraph. Using this approximation, we prune the least important edges of the network. We survey the performance and limitations of this method, finding that averaged over all tasks our method has greater AUC from circuit recovery than other methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LAWFUL: Law-Aligned Witness for Faithful Use of Latents

    cs.LG 2026-07 conditional novelty 7.0 of 10

    LAWFUL defines coverage-aware physical-consistency scores and circuit tests, reporting that a MoCap-to-Radar transformer's 9-component temporal circuit carries Doppler-law consistency via attention patterns.

  2. Verbalizable Representations Form a Global Workspace in Language Models

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Language models represent their current reasoning in a small, readable set of verbalizable vectors (the J-space) that functions like a global workspace.

  3. Faithfulness to Refusal: A Causal Audit of Neuron Selectors

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A causal audit via neuron-row zeroing shows attribution methods (LRP, IG) faithfully identify dispensable neurons and can install refusal behavior, while rank-stability proxies systematically miss selector failures.

  4. Mechanistic Interpretability as Statistical Estimation: A Variance Analysis

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Small changes in data or settings used to find a circuit in a language model often produce very different circuits: under bootstrap resampling, average pairwise overlap of EAP-IG circuits across tasks and models is on...

  5. Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Bias in GPT-2 and Llama-2 is localized to a small set of edges, and ablation of those edges reduces bias while impairing unrelated NLP tasks.

  6. Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis

    cs.CL 2025-05 conditional novelty 6.0 of 10

    PhantomCircuit traces knowledge overshadowing to attention circuits during training and prunes circuit edges to recover the overshadowed answer.

  7. BlueGlass: A Framework for Composite AI Safety

    cs.AI 2025-07 conditional novelty 5.0 of 10

    BlueGlass provides composite AI safety infrastructure; its case studies on object-detection VLMs reveal dataset trade-offs, a decoder-layer phase transition in probe accuracy, and SAE-discovered concepts including spu...

  8. Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers

    cs.LG 2025-07 reject novelty 4.0 of 10

    A framework that borrows activation patching to adversarially induce and measure deception, supported only by an underspecified toy network simulation.

  9. Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model Reasoning

    cs.AI 2025-02 reject novelty 4.0 of 10

    SICAF traces per-token self-influence inside extracted circuits to map GPT-2's reasoning on the IOI task.

Pith tools