Pith. sign in

REVIEW 6 major objections 7 minor 105 references

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii

T0 review · 6 major / 7 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper argues that a pluralist Explanatory Virtues Framework gives systematic criteria for choosing between competing mechanistic interpretability explanations, and that Compact Proofs, which embody many of these virtues, are a…

desk verdict A useful philosophical vocabulary for MI evaluation, but not yet the systematic framework it claims to be; the Compact Proofs verdict rests on a subjective rubric. read the letter →

arxiv 2505.01372 v1 pith:P7WATJPK submitted 2025-05-02 cs.LG cs.AIcs.CLcs.HC

classification cs.LGcs.AIcs.CLcs.HC
keywords mechanisticinterpretabilityexplanatoryvirtuestheorychoicephilosophyofscienceexplanationevaluationcompactproofssparseautoencoderscausalcircuits
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper seeks to establish that mechanistic interpretability can settle disputes between competing explanations of the same neural network by importing criteria for theory choice from the philosophy of science. It proposes a pluralist Explanatory Virtues Framework: a battery of formalisable properties—accuracy, precision, simplicity, unification, fruitfulness, hard-to-varyness, nomologicity, and others—that are claimed to be reliable indicators of whether an explanation is true. Applied to four common methods, the framework finds that clustering explanations, sparse autoencoder decompositions, and causal-circuit analyses each neglect one or more virtues, while Compact Proofs consider many and are therefore singled out as promising. If the framework is right, interpretability researchers gain an epistemic basis for preferring one explanation over another, and three research directions follow: clarifying what simplicity means, prioritising unifying and co-explanatory accounts, and seeking universal principles about neural networks.

What carries the argument

The central object is the Explanatory Virtues Framework itself, built from the Bayesian, Kuhnian, Deutschian, and Nomological accounts of explanation: a directed graph of virtues, each with a mathematical definition, that turns theory choice into a comparison problem. Its machinery includes Bayesian decompositions—$\mathrm{Acc}(E)=P(x_T\mid E)$ for accuracy, precision as expected log-likelihood, descriptiveness and co-explanation as additive and joint components of log-likelihood, power and unification as their theoretical counterparts—and a hard-to-varyness condition: an explanation is hard-to-vary if it sits at a local maximum of $\log(\mathrm{Acc}(E))-k(E)$, where $k$ is a complexity measure and a modification is a sequence of insertions, deletions, substitutions, or transpositions of symbols. The other load-bearing piece is the Compact Proofs evaluator, in which an explanation is converted into a verifier $V(\theta,E)$ that returns a worst-case performance bound; the tightness of the bound is accuracy, and the computational cost of checking the proof is simplicity, so a good explanation pushes out the (tightness, compactness) Pareto frontier. Validity conditions—model-level, ontic, causal-mechanistic, falsifiable—must be met before the virtues are compared.

What would settle it

Take one model, such as a small transformer trained to compose group operations, generate two valid MI explanations of the same behaviour that score very differently on the virtues, then run a battery of held-out intervention experiments that the two explanations predict differently; if the lower-scoring explanation predicts the held-out behaviour at least as well as the high-scoring one, the claim that the virtues are truth-conducive is falsified.

Watch

Extended reading notes

Core claim

The central claim is that 'what makes a good explanation?' can be answered by a pluralist set of explanatory virtues, and that these virtues—not subjective preference—should drive theory choice in mechanistic interpretability. To be valid, an MI explanation must be model-level, ontic, causal-mechanistic, and falsifiable. Among valid explanations, the paper maintains that the better explanation is the one embodying more of the virtues: empirical virtues such as accuracy, descriptiveness, co-explanation, and fruitfulness, and theoretical virtues such as precision, power, unification, consistency, simplicity, hard-to-varyness, and nomologicity. The framework attaches each virtue to one of four philosophy-of-science accounts and gives it a mathematical definition; for example, accuracy is a likelihood and hard-to-varyness is being a local maximum of log-accuracy minus complexity. When the rubric is applied to clustering, sparse autoencoders, causal-circuit analysis, and Compact Proofs, the paper finds that Compact Proofs are the method that considers many virtues and that current methods systematically neglect simplicity, unification, co-explanation, and nomological principles.

Load-bearing premise

The load-bearing premise is that the listed explanatory virtues are truth-conducive—that an explanation embodying more of them is genuinely more likely to be correct—because without that link the framework becomes a subjective preference list rather than a guide to which explanation of a neural network is actually true.

Editorial extensions

If this is right

  • MI researchers can replace subjective intuition with a shared rubric when two explanations of the same model conflict, giving explicit epistemic reasons for theory choice.
  • Methods can be improved on the axes the framework flags as neglected: simplicity, unification, co-explanation, and nomological principles.
  • Compact Proofs offer a concrete way to turn the accuracy–simplicity trade-off into a measurable Pareto frontier, with faithful explanations yielding tighter bounds at lower verification cost.
  • A clearer definition of explanatory simplicity would let different explanation methods be compared on a single accuracy–simplicity curve.
  • Seeking universal principles and reused building blocks would move interpretability from cataloguing individual cases toward nomological, predictive explanations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural test of the framework is to instantiate each virtue as a computable proxy—likelihood for accuracy, description length for simplicity, edit distance to a local maximum for hard-to-varyness—and score the same set of explanations with and without the proxies; convergence would suggest the rubric is measuring something real.
  • The truth-conduciveness claim is an empirical hypothesis: run a prediction tournament where rival explanations of the same model are scored on virtues and then probed by novel interventions, and see whether the virtue leader continues to predict behaviour outside its training distribution.
  • The framework's emphasis on unification suggests that shared, reused substructures should be privileged across tasks; if the same units keep predicting behaviour in new tasks, that would strengthen the link between unification and truth.
  • If truth-conduciveness fails, the framework still documents how interpretability researchers actually trade off values, but it would lose its normative force as a guide to which explanation is correct.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

6 major / 7 minor

Summary. This paper introduces an Explanatory Virtues Framework for evaluating explanations in mechanistic interpretability (MI), drawing on Bayesian, Kuhnian, Deutschian, and Nomological accounts from the philosophy of science. The authors define a set of virtues with mathematical formulas, apply the framework to four MI methods (Clustering, Sparse Autoencoders, Causal Circuits, and Compact Proofs), and conclude that Compact Proofs are a promising approach because they exhibit many virtues. The paper also proposes research directions emphasizing simplicity, unification, and nomological principles.

Significance. If the framework delivered what it promises—a canonical, computable way to compare competing MI explanations—it would be a valuable contribution to interpretability research, which currently lacks principled theory-choice criteria. The paper surveys relevant philosophy of science literature, and its taxonomy of virtues is genuinely useful as a vocabulary for discussing explanations. However, the paper does not provide machine-checked proofs, reproducible code, or quantitative data; the central evaluation in Table 1 is based on the authors' qualitative judgments. The paper's main value is therefore as a conceptual proposal, not as an established evaluation methodology.

major comments (6)
  1. [Section 3.1, glossary of Bayesian virtues] The definitions of Accuracy, Precision, Descriptiveness, Co-Explanation, Power, and Unification all presuppose that an explanation E is associated with a conditional probability P(x|E) over observational data. Yet none of the four methods analyzed in Section 4 is specified as a probabilistic model: a circuit is a structural causal model, an SAE is a dictionary plus feature descriptions, and a Compact Proof is a verifier program. The paper never states how to derive P(x|E) from these artifacts, so the promised 'consistent and canonical way to compute each virtue' (Section 3, p. 4) is not defined for the very explanations the framework is meant to compare.
  2. [Table 1 vs. Section 3.1] The framework's formal definitions imply that Co-Explanation and Unification are the same quantity—CoEx(E) = log(Acc(E)) − Desc(E) and Unif(E) = Prec(E) − Power(E) = E_{xT∼X}[log(P(xT|E)/∏P(xT,i|E))], which are identical up to the expectation. Yet Table 1 rates Compact Proofs as '✓' on Co-explanation and '✗' on Unification. This inconsistency is only possible because the Table 2 rubric is not derived from the formal definitions; it shows that the case-study ratings are subjective judgments rather than outcomes of the framework.
  3. [Section 4.2 and Abstract] The central conclusion that 'Compact Proofs consider many explanatory virtues and are hence a promising approach' (Abstract) is based entirely on the authors' own qualitative ratings in Table 1, not on any measured data or inter-rater agreement. The paper provides no evidence that another researcher applying the Table 2 rubric would reach the same ratings, so the claimed systematic comparison between Chughtai et al. (2023) and Stander et al. (2024) is not actually demonstrated.
  4. [Section 2.2] The definition of valid MI explanations (Model-level, Ontic, Causal-Mechanistic, Falsifiable) is imported from the authors' own forthcoming paper (Ayonrinde & Jaburi, 2025), which is not available for verification. Because this definition is the foundation upon which the Explanatory Virtues Framework is built, the framework's epistemic grounding is self-referential; the paper should either provide the argument for this definition in the present manuscript or cite a published, accessible source.
  5. [Section 3, 'Explanatory Virtues are properties that are reliable indicators of truth'] The truth-conduciveness of the listed virtues is asserted with only a reference to Schindler (2018). This premise is load-bearing: if virtues are not truth-conducive, the framework becomes a subjective preference list. The paper should give at least one concrete argument or empirical example for why these particular virtues indicate truth, or explicitly reframe the contribution as a descriptive account of explanatory values.
  6. [Section 3.3, Hard-to-Varyness definition] The formalization of hard-to-varyness as a local maximum of hv(E) = log(Acc(E)) − k(E) is under-specified because 'local' is only informally defined as 'a small number of edit operations apart' (footnote 9); without a metric on the space of explanations, the definition cannot be applied to decide whether a given explanation is hard-to-vary, which is exactly what Table 1 claims.
minor comments (7)
  1. [Section 3.1, notation] The subscript formatting is inconsistent: the text uses 'x T', 'x I', and 'xT,i' in different places, and the glossary uses 'x T∼X' without clearly defining the distribution over inference-time data; please standardize the notation.
  2. [Section 3.2, Pragmatic Utility paragraph] There is a typo: 'Sparse Auutoencoder' should be 'Sparse Autoencoder'.
  3. [Section 4.1.3, FCM criteria] The FCM criteria use expressions F(C\K) and F(M\K) without defining the function F; presumably F is the model output function on a data distribution, but this should be stated explicitly.
  4. [Section 3.5 and Figure 1] The caption for Figure 1 references colors, bold arrows, and dashed arrows, but the figure is not included in the manuscript text; please ensure the figure is present or describe the relationships in the text.
  5. [Section 4.2 and Table 1] Section 4.2 describes Compact Proofs as a method for evaluating other explanations, yet Table 1 treats Compact Proofs as an explanation method comparable to Clustering, SAEs, and Circuits; please clarify this distinction or reclassify the comparison.
  6. [Section 5, paragraph on universality] The claim that 'the MI community has sought to understand the universality ... with mixed results' lacks a citation supporting the 'mixed results' assertion; please add a reference or soften the claim.
  7. [References] Several references are to unpublished or forthcoming works (Ayonrinde & Jaburi 2025; Jaburi et al. 2025; Ayonrinde 2025), making it difficult for readers to verify the cited claims; consider providing preprints or including the relevant content in an appendix.

Circularity Check

2 steps flagged · score 4.0 of 10

Self-cited validity definition and validity-as-virtue double counting raise circularity concerns, but the core framework retains independent philosophical content.

  1. self citation load bearing [Section 2.2, Definition of Mechanistic Interpretability; used in Section 3.5]
    "Following Olah et al. (2020); Olsson et al. (2022), Ayonrinde &Jaburi (2025) define Mechanistic Interpretability as follows: Interpretability explanations are valid Mechanistic Interpretability explanations if they are Model-level, Ontic, Causal-Mechanistic, and Falsifiable."

    The paper's gate for applying the Explanatory Virtues Framework—what counts as a valid MI explanation—is taken from the authors' own forthcoming Part I.i, not from an independent or externally checkable source. Section 3.5 then makes this imported definition load-bearing: 'For an explanation to be a good explanation in Mechanistic Interpretability, it must first be a valid MI explanation. In Section 2.2 we identified valid MI explanations as those which are Model-Level, Ontic, Causal-Mechanistic, Falsifiable.' Thus the framework's scope and its later comparison of methods rest on a self-citation to work that is not yet available and is authored by the same people.

  2. self definitional [Figure 1 / Section 3.5 / Table 1 (Validity rows)]
    "The Explanatory Virtues which are essential for any scientific explanation (Falsifiability and Causal-Mechanisticity) to be valid are denoted with an exclamation mark; the most important virtues to decide between explanations (Simplicity, Hard-to-Varyness, and Fruitfulness) are marked with a star."

    The paper first defines a valid MI explanation as one that is Model-level, Ontic, Causal-Mechanistic, and Falsifiable (Section 2.2), then lists Falsifiability and Causal-Mechanisticity among the Explanatory Virtues in Figure 1 and Table 1. This means part of the framework's virtue score re-asserts the definition of the object being evaluated: a method scores well on these 'virtues' precisely because it is admitted as a valid MI explanation, not because the framework has discovered an independent criterion for choosing between competing valid explanations. The motivating comparison between Chughtai et al. and Stander et al. is between two explanations that both already satisfy these validity conditions, so these rows cannot decide between them.

full rationale

The central Bayesian, Kuhnian, Deutschian, and Nomological virtues are drawn from external philosophy-of-science sources, and their mathematical definitions (Accuracy, Precision, Descriptiveness, Co-Explanation, Power, Unification, Hard-to-Varyness, etc.) are not fitted to any dataset and do not themselves constitute a renamed prediction. The conclusion favoring Compact Proofs is based on the authors' rubric in Table 2 rather than on a computation using those equations; indeed, P(x|E) is not specified for circuits, SAE explanations, or Compact Proofs. That is a gap in derivational support, but it is not a case of a fitted parameter renamed as a prediction, so it does not count as circularity under the hard rules. The concrete circular elements are the self-cited definition of valid MI explanations and the double-counting of validity conditions (Falsifiability, Causal-Mechanisticity) as explanatory virtues. These are load-bearing for the framework's scope and for part of the Table 1 evaluation, but they do not consume the whole framework: the remaining virtues and the philosophical analysis retain independent content. Hence a score of 4 is appropriate.

Assumptions & free parameters 0 free parameters · 5 assumptions · 2 invented entities

The paper's central claims rest on philosophical assumptions (truth-conduciveness of virtues), on a definition of valid MI explanations imported from the authors' own forthcoming work, and on a qualitative rubric applied by the authors. There are no numerical free parameters, but the evaluation involves many unquantified judgment calls.

assumptions (5)
  • domain assumption Explanatory virtues are truth-conducive properties of explanations.
    Section 3 states that explanatory virtues are 'reliable indicators of truth' and cites Schindler (2018), but the claim is not derived or empirically demonstrated.
  • domain assumption The validity definition of Mecanistic Interpretability (Model-level, Ontic, Causal-Mechanistic, Falsifiable) is correct.
    Definition in Section 2.2 is taken from the authors' own forthcoming work (Ayonrinde and Jaburi 2025), so it is not independently established within this paper.
  • standard math The Bayesian decomposition of Accuracy into Descriptiveness and Co-Explanation, and Precision into Power and Unification, is meaningful.
    These are mathematical identities given the definitions in Section 3.1, but their epistemic significance is assumed. Under an i.i.d. data assumption, Co-Explanation and Unification are identically zero, which the paper does not discuss.
  • ad hoc to paper Hard-to-varyness is correctly formalized as being at a local maximum of log(Accuracy) minus complexity.
    This definition is introduced in Section 3.3 without evidence that it captures Deutsch's intended notion or that it is computable in practice.
  • domain assumption Nomologicity (appealing to general laws) is an explanatory virtue.
    Section 3.4 asserts this without proof, treating lawfulness as truth-conducive without argument.
invented entities (2)
  • Co-Explanation as a distinct explanatory virtue
    purpose: Measures the degree to which an explanation accounts for multiple data points jointly beyond individual predictions.
    Defined in Section 3.1 as log(Acc(E)) - Desc(E); no empirical demonstration that this construct is truth-conducive or practically measurable.
  • Formalization of Hard-to-varyness as a local maximum of hv(E)
    purpose: Captures Deutsch's notion of explanations that resist modification to fit new data.
    Introduced in Section 3.3; the connection to the philosophical concept is asserted, and no algorithm or empirical validation is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii." pith.science (2026). https://pith.science/paper/P7WATJPK

@misc{pith2026250501372,
  author       = {Pith},
  title        = {Pith review of: Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P7WATJPK}},
  note         = {Machine review of arXiv:2505.01372}
}
read the original abstract

Mechanistic Interpretability (MI) aims to understand neural networks through causal explanations. Though MI has many explanation-generating methods, progress has been limited by the lack of a universal approach to evaluating explanations. Here we analyse the fundamental question "What makes a good explanation?" We introduce a pluralist Explanatory Virtues Framework drawing on four perspectives from the Philosophy of Science - the Bayesian, Kuhnian, Deutschian, and Nomological - to systematically evaluate and improve explanations in MI. We find that Compact Proofs consider many explanatory virtues and are hence a promising approach. Fruitful research directions implied by our framework include (1) clearly defining explanatory simplicity, (2) focusing on unifying explanations and (3) deriving universal principles for neural networks. Improved MI methods enhance our ability to monitor, predict, and steer AI systems.

Figures

Figures reproduced from arXiv: 2505.01372 by the authors.

Figure 1
Figure 1. A Directed Acyclic Graph representation of the [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Given some (possibly intermediate) embeddings ( [PITH_FULL_IMAGE:figures/full_fig_p023_2.png] view at source ↗
Figure 3
Figure 3. (a) The SAE architecture. An encoder provides some set of latents (or feature [PITH_FULL_IMAGE:figures/full_fig_p024_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: A circuit explanation is a Causal-Mechanistic explanation such that the circuit C [PITH_FULL_IMAGE:figures/full_fig_p024_4.png]
Figure 5
Figure 5. Figure 5: (a) Compact Proofs evaluate explanations on two metrics, their compactness [PITH_FULL_IMAGE:figures/full_fig_p025_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

105 extracted references · 40 canonical work pages

  1. [1]

    The computational complexity of circuit discovery for inner interpretability

    Federico Adolfi, Martina G Vilas, and Todd Wareham. The computational complexity of circuit discovery for inner interpretability. arXiv preprint arXiv:2410.08025, 2024

  2. [2]

    Physics of language models: Part 3.1, knowledge storage and extraction

    Zeyuan Allen-Zhu and Yuanzhi Li. Physics of language models: Part 3.1, knowledge storage and extraction. arXiv preprint arXiv:2309.14316, 2023

  3. [3]

    Physics of language models: Part 3.1, knowledge storage and extraction

    Zeyuan Allen-Zhu and Yuanzhi Li. Physics of language models: Part 3.1, knowledge storage and extraction. In Forty-first International Conference on Machine Learning, 2024

  4. [4]

    The urgency of interpretability, 2025

    Dario Amodei. The urgency of interpretability, 2025. URL https://www.darioamodei.com/post/the-urgency-of-interpretability

  5. [5]

    Foundational challenges in assuring alignment and safety of large language models

    Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, et al. Foundational challenges in assuring alignment and safety of large language models. arXiv preprint arXiv:2404.09932, 2024

  6. [6]

    Ai as systems, not just models, 2024

    Andy Arditi. Ai as systems, not just models, 2024. URL https://www.lesswrong.com/posts/2po6bp2gCHzxaccNz/ai-as-systems-not-just-models

  7. [7]

    Standard saes might be incoherent: A choosing problem & a “concise” solution

    Kola Ayonrinde. Standard saes might be incoherent: A choosing problem & a “concise” solution. Blog post, 2024. URL https://www.lesswrong.com/posts/vNCAQLcJSzTgjPaWS/standard-saes-might-be-incoherent-a-choosing-problem-and-a

  8. [8]

    Position: Interpretability is a bidirectional communication problem

    Kola Ayonrinde. Position: Interpretability is a bidirectional communication problem. In ICLR 2025 Workshop on Bidirectional Human-AI Alignment, 2025. URL https://openreview.net/forum?id=O4LaRH4zSI

Show all 105 references
  1. [9]

    A mathematical philosophy of explanations in mechanistic interpretability: The strange science part i.i, 2025

    Kola Ayonrinde and Louis Jaburi. A mathematical philosophy of explanations in mechanistic interpretability: The strange science part i.i, 2025. forthcoming

  2. [10]

    Pearce, and Lee Sharkey

    Kola Ayonrinde, Michael T. Pearce, and Lee Sharkey. Interpretability as compression: Reconsidering sae explanations of neural activations with mdl-saes, 2024. URL https://arxiv.org/abs/2410.11179

  3. [11]

    Novum Organum

    Francis Bacon. Novum Organum. Clarendon Press, London, 1620. URL https://en.wikipedia.org/wiki/Novum_Organum. Part of the Instauratio Magna

  4. [12]

    Simplicity

    Alan Baker. Simplicity . In Edward N. Zalta (ed.), The Stanford Encyclopedia of Philosophy . Metaphysics Research Lab, Stanford University, S ummer 2022 edition, 2022

  5. [13]

    Design Rules: The Power of Modularity Volume 1

    Carliss Y Baldwin and Kim B Clark. Design Rules: The Power of Modularity Volume 1. MIT press, 1999

  6. [14]

    Icml 2024 mechanistic interpretability workshop, 2024

    Fazl Barez, Mor Geva, Lawrence Chan, Atticus Geiger, Kayo Yin, Neel Nanda, and Max Tegmark. Icml 2024 mechanistic interpretability workshop, 2024. URL https://icml2024mi.pages.dev/

  7. [15]

    Local vs

    Shahaf Bassan, Guy Amir, and Guy Katz. Local vs. global interpretability: A computational complexity perspective. arXiv preprint arXiv:2406.02981, 2024

  8. [16]

    Explanation: A mechanist alternative

    William Bechtel and Adele Abrahamsen. Explanation: A mechanist alternative. Studies in History and Philosophy of Science Part C: Studies in History and Philosophy of Biological and Biomedical Sciences, 36 0 (2): 0 421--441, 2005. doi:10.1016/j.shpsc.2005.03.010

  9. [17]

    Sander Beckers and Joseph Y. Halpern. Abstracting causal models. In Proceedings of the 33Rd Aaai Conference on Artificial Intelligence, pp.\ 2678--2685. 2019

  10. [18]

    International ai safety report

    Yoshua Bengio, S \"o ren Mindermann, Daniel Privitera, Tamay Besiroglu, Rishi Bommasani, Stephen Casper, Yejin Choi, Philip Fox, Ben Garfinkel, Danielle Goldfarb, et al. International ai safety report. arXiv preprint arXiv:2501.17805, 2025

  11. [19]

    Mechanistic Interpretability for AI Safety -- A Review , April 2024

    Leonard Bereska and Efstratios Gavves. Mechanistic Interpretability for AI Safety -- A Review , April 2024. URL http://arxiv.org/abs/2404.14082. arXiv:2404.14082 [cs]

  12. [20]

    J. Bernal. The social function of science. Philosophical Review, 49 0 (n/a): 0 377, 1940. doi:10.2307/2180883

  13. [21]

    Auditing local explanations is hard

    Robi Bhattacharjee and Ulrike von Luxburg. Auditing local explanations is hard. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/forum?id=ybMrn4tdn0

  14. [22]

    Language models can explain neurons in language models

    Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders. Language models can explain neurons in language models. https://openaipublic.blob.core.windows.net/neuron-explainer/paper/index.html, 2023

  15. [23]

    An Interpretability Illusion for BERT , April 2021

    Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Viégas, and Martin Wattenberg. An Interpretability Illusion for BERT , April 2021. URL http://arxiv.org/abs/2104.07143. arXiv:2104.07143 [cs]

  16. [24]

    Identifying Functionally Important Features with End -to- End Sparse Dictionary Learning , May 2024

    Dan Braun, Jordan Taylor, Nicholas Goldowsky-Dill, and Lee Sharkey. Identifying Functionally Important Features with End -to- End Sparse Dictionary Learning , May 2024. URL http://arxiv.org/abs/2405.12241. arXiv:2405.12241 [cs]

  17. [25]

    Interpretability in parameter space: Minimizing mechanistic description length with attribution-based parameter decomposition

    Dan Braun, Lucius Bushnaq, Stefan Heimersheim, Jake Mendel, and Lee Sharkey. Interpretability in parameter space: Minimizing mechanistic description length with attribution-based parameter decomposition. arXiv preprint arXiv:2501.14926, 2025

  18. [26]

    Towards Monosemanticity : Decomposing Language Models With Dictionary Learning

    Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen,...

  19. [27]

    Propositional interpretability in artificial intelligence

    David J Chalmers. Propositional interpretability in artificial intelligence. arXiv preprint arXiv:2501.15740, 2025

  20. [28]

    A toy model of universality: Reverse engineering how networks learn group operations

    Bilal Chughtai, Lawrence Chan, and Neel Nanda. A toy model of universality: Reverse engineering how networks learn group operations. In International Conference on Machine Learning, pp.\ 6243--6267. PMLR, 2023

  21. [29]

    The evolutionary origins of modularity

    Jeff Clune, Jean-Baptiste Mouret, and Hod Lipson. The evolutionary origins of modularity. Proceedings of the Royal Society b: Biological sciences, 280 0 (1755): 0 20122863, 2013

  22. [30]

    Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso

    Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso. Towards automated circuit discovery for mechanistic interpretability. 2023. URL https://arxiv.org/abs/2304.14997

  23. [31]

    Central dogma of molecular biology

    Francis Crick. Central dogma of molecular biology. Nature, 227 0 (5258): 0 561--563, 1970

  24. [32]

    The beginning of infinity: Explanations that transform the world

    David Deutsch. The beginning of infinity: Explanations that transform the world. penguin uK, 2011

  25. [33]

    Frank Watson Dyson, Arthur Stanley Eddington, and Charles Davidson. Ix. a determination of the deflection of light by the sun's gravitational field, from observations made at the total eclipse of may 29, 1919. Philosophical Transactions of the Royal Society of London. Series A...

  26. [34]

    Einstein

    A. Einstein. The foundation of the general theory of relativity. 1916

  27. [35]

    Richard P. Feynman. Cargo cult science. Engineering and Science, 37 0 (7): 0 10--13, 1974. ISSN 0013-7812. URL http://resolver.caltech.edu/CaltechES:37.7.CargoCult

  28. [36]

    Clusterability in neural networks

    Daniel Filan, Stephen Casper, Shlomi Hod, Cody Wild, Andrew Critch, and Stuart Russell. Clusterability in neural networks. arXiv preprint arXiv:2103.03386, 2021

  29. [37]

    Interpretability illusions in the generalization of simplified models

    Dan Friedman, Andrew Lampinen, Lucas Dixon, Danqi Chen, and Asma Ghandeharioun. Interpretability illusions in the generalization of simplified models. In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR.org, 2024

  30. [38]

    Scaling and evaluating sparse autoencoders, June 2024

    Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. Scaling and evaluating sparse autoencoders, June 2024. URL http://arxiv.org/abs/2406.04093. arXiv:2406.04093 [cs] version: 1

  31. [39]

    Causal abstractions of neural networks

    Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. Causal abstractions of neural networks. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (eds.), Advances in Neural Information Processing Systems, volume 34, pp.\ 9574--9586. Curran A...

  32. [40]

    Causal abstraction for faithful model interpretation

    Atticus Geiger, Chris Potts, and Thomas Icard. Causal abstraction for faithful model interpretation. arXiv preprint arXiv:2301.04709, 2023

  33. [41]

    Clustering algorithms

    Google Developers . Clustering algorithms. https://developers.google.com/machine-learning/clustering/clustering-algorithms, 2025. Accessed: 2025-02-23

  34. [42]

    Compact proofs of model performance via mechanistic interpretability

    Jason Gross, Rajashree Agrawal, Thomas Kwa, Euan Ong, Chun Hei Yip, Alex Gibson, Soufiane Noubir, and Lawrence Chan. Compact proofs of model performance via mechanistic interpretability. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

  35. [43]

    Interpbench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques

    Rohan Gupta, Iv\' a n Arcuschin, Thomas Kwa, and Adri\` a Garriga-Alonso. Interpbench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (eds.), Advances in N...

  36. [44]

    Hastie, R

    T. Hastie, R. Tibshirani, and J.H. Friedman. The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer series in statistics. Springer, 2009. ISBN 9780387848846. URL https://books.google.co.uk/books?id=eBSgoAEACAAJ

  37. [45]

    Hempel and Paul Oppenheim

    Carl G. Hempel and Paul Oppenheim. Studies in the logic of explanation. Philosophy of Science, 15 0 (2): 0 135--175, 1948. ISSN 00318248, 1539767X. URL http://www.jstor.org/stable/185169

  38. [46]

    Philosophy of Natural Science

    Carl Gustav Hempel. Philosophy of Natural Science. Prentice-Hall, Englewood Cliffs, N.J.,, 1966

  39. [47]

    Bayesianism and inference to the best explanation

    Leah Henderson. Bayesianism and inference to the best explanation. The British Journal for the Philosophy of Science, 2014

  40. [48]

    The developmental landscape of in-context learning

    Jesse Hoogland, George Wang, Matthew Farrugia-Roberts, Liam Carroll, Susan Wei, and Daniel Murfet. The developmental landscape of in-context learning. arXiv preprint arXiv:2402.02364, 2024

  41. [49]

    Sparse autoencoders find highly interpretable features in language models

    Robert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, and Lee Sharkey. Sparse autoencoders find highly interpretable features in language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=F76bwRSLeK

  42. [50]

    Hutter, E

    M. Hutter, E. Catt, and D. Quarel. An Introduction to Universal Artificial Intelligence. Chapman & Hall/CRC Artificial Intelligence and robotics series. Chapman & Hall/CRC Press, 2024. ISBN 9781003460299. URL https://books.google.co.uk/books?id=cfg60AEACAAJ

  43. [51]

    Fine-tuning neural networks to match their interpretation: Towards scaling compact proofs, 2025

    Louis Jaburi, Ronak Mehta, Soufiane Noubir, and Jason Gross. Fine-tuning neural networks to match their interpretation: Towards scaling compact proofs, 2025. forthcoming

  44. [52]

    Kandel, J.H

    E.R. Kandel, J.H. Schwartz, and T. Jessell. Principles of Neural Science, Fourth Edition. McGraw-Hill Companies,Incorporated, 2000. ISBN 9780838577011. URL https://books.google.co.uk/books?id=yzEFK7Xc87YC

  45. [53]

    Saebench: a comprehensive benchmark for sparse autoencoders, 2024

    Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Bloom, David Chanin, Yeu-Tong Lau, Eoin Farrell, Arthur Conmy, Callum McDougall, Kola Ayonrinde, Matthew Wearden, Samuel Marks, and Neel Nanda. Saebench: a comprehensive benchmark for sparse autoencoders, 2024. URL http...

  46. [54]

    Kennefick

    D. Kennefick. No Shadow of a Doubt: The 1919 Eclipse That Confirmed Einstein's Theory of Relativity. Princeton University Press, 2021. ISBN 9780691217154. URL https://books.google.co.uk/books?id=_Eb8DwAAQBAJ

  47. [55]

    Explanatory unification

    Philip Kitcher. Explanatory unification. Philosophy of science, 48 0 (4): 0 507--531, 1981

  48. [56]

    Three approaches to the quantitative definition ofinformation

    Andrei N Kolmogorov. Three approaches to the quantitative definition ofinformation. Problems of information transmission, 1 0 (1): 0 1--7, 1965

  49. [57]

    Thomas S. Kuhn. Objectivity, value judgment, and theory choice. In David Zaret (ed.), Review of Thomas S. Kuhn The Essential Tension: Selected Studies in Scientific Tradition and Change, pp.\ 320--39. Duke University Press, 1981

  50. [58]

    The Structure of Scientific Revolutions

    Thomas Samuel Kuhn. The Structure of Scientific Revolutions. University of Chicago Press, Chicago, 1962

  51. [59]

    Falsification and the methodology of scientific research programmes

    Imre Lakatos. Falsification and the methodology of scientific research programmes. In Imre Lakatos and Alan Musgrave (eds.), Criticism and the growth of knowledge, pp.\ 91--196. Cambridge University Press, 1970

  52. [60]

    The Methodology of Scientific Research Programmes

    Imre Lakatos. The Methodology of Scientific Research Programmes. Cambridge University Press, New York, 1978

  53. [61]

    Sparse autoencoders do not find canonical units of analysis

    Patrick Leask, Bart Bussmann, Michael Pearce, Joseph Bloom, Curt Tigges, Noura Al Moubayed, Lee Sharkey, and Neel Nanda. Sparse autoencoders do not find canonical units of analysis. arXiv preprint arXiv:2502.04878, 2025

  54. [62]

    Lindsay and David Bau

    Grace W. Lindsay and David Bau. Testing methods of neural systems understanding. Cogn. Syst. Res., 82: 0 101156, December 2023. URL https://doi.org/10.1016/j.cogsys.2023.101156

  55. [63]

    The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery

    Zachary C Lipton. The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery. Queue, 16 0 (3): 0 31--57, 2018

  56. [64]

    Mechanistic mode connectivity

    Ekdeep Singh Lubana, Eric J Bigelow, Robert P Dick, David Krueger, and Hidenori Tanaka. Mechanistic mode connectivity. In Proceedings of the 40th International Conference on Machine Learning, pp.\ 22965--23004, 2023

  57. [65]

    Information theory, inference and learning algorithms

    David JC MacKay. Information theory, inference and learning algorithms. Cambridge university press, 2003

  58. [66]

    Is this the subspace you are looking for? an interpretability illusion for subspace activation patching

    Aleksandar Makelov, Georg Lange, Atticus Geiger, and Neel Nanda. Is this the subspace you are looking for? an interpretability illusion for subspace activation patching. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum...

  59. [67]

    Downstream applications as validation of interpretability progress, March 2025

    Sam Marks. Downstream applications as validation of interpretability progress, March 2025. URL https://www.lesswrong.com/posts/wGRnzCFcowRCrpX4Y/downstream-applications-as-validation-of-interpretability

  60. [68]

    Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller

    Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller. Sparse Feature Circuits : Discovering and Editing Interpretable Causal Graphs in Language Models , March 2024. URL http://arxiv.org/abs/2403.19647. arXiv:2403.19647 [cs]

  61. [69]

    Cognitive styles in two cognitive sciences

    James Myers. Cognitive styles in two cognitive sciences. In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 34, 2012

  62. [70]

    Zoom in: An introduction to circuits

    Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. Zoom in: An introduction to circuits. Distill, 5 0 (3): 0 e00024--001, 2020

  63. [71]

    In-context learning and induction heads

    Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. In-context learning and induction heads. arXiv preprint arXiv:2209.11895, 2022

  64. [72]

    Automatically interpreting millions of features in large language models

    Gon c alo Paulo, Alex Mallen, Caden Juang, and Nora Belrose. Automatically interpreting millions of features in large language models. arXiv preprint arXiv:2410.13928, 2024

  65. [73]

    Causality

    Judea Pearl. Causality. Cambridge university press, 2009

  66. [74]

    Poincar \'e

    H. Poincar \'e . Science and Hypothesis. Library of philosophy, psychology and scientific methods. Science Press, 1905. URL https://books.google.co.uk/books?id=5nQSAAAAYAAJ

  67. [75]

    Karl R. Popper. The Logic of Scientific Discovery. Routledge, London, England, 1935

  68. [76]

    Hume on theoretical simplicity

    Hsueh Qu. Hume on theoretical simplicity. Philosophers' Imprint, 23 0 (1), 2023. doi:10.3998/phimp.1521

  69. [77]

    Escalation risks from language models in military and diplomatic decision-making

    Juan-Pablo Rivera, Gabriel Mukobi, Anka Reuel, Max Lamparth, Chandler Smith, and Jacquelyn Schneider. Escalation risks from language models in military and diplomatic decision-making. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, FAcc...

  70. [78]

    Four decades of scientific explanation

    Wesley Salmon. Four decades of scientific explanation. 1989. URL https://api.semanticscholar.org/CorpusID:46466034

  71. [79]

    Wesley C. Salmon. Scientific Explanation and the Causal Structure of the World. Princeton University Press, 1984. ISBN 9780691101705

  72. [80]

    Mechanistic?, 2024

    Naomi Saphra and Sarah Wiegreffe. Mechanistic?, 2024. URL https://arxiv.org/abs/2410.09087

  73. [81]

    Theoretical Virtues in Science: Uncovering Reality Through Theory

    Samuel Schindler. Theoretical Virtues in Science: Uncovering Reality Through Theory. Cambridge University Press, Cambridge, 2018

  74. [82]

    Riechers, Lucas Teixeira, Alexander Gietelink Oldenziel, and Sarah Marzen

    Adam Shai, Paul M. Riechers, Lucas Teixeira, Alexander Gietelink Oldenziel, and Sarah Marzen. Transformers represent belief state geometry in their residual stream. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.ne...

  75. [83]

    A mathematical theory of communication

    Claude Elwood Shannon. A mathematical theory of communication. The Bell system technical journal, 27 0 (3): 0 379--423, 1948

  76. [84]

    Michaud, Stephen Casper, Max Tegmark, William Saunders, David Bau, Eric Todd, Atticus Geiger, Mor Geva, Jesse Hoogland, Daniel Murfet, and Tom McGrath

    Lee Sharkey, Bilal Chughtai, Joshua Batson, Jack Lindsey, Jeff Wu, Lucius Bushnaq, Nicholas Goldowsky-Dill, Stefan Heimersheim, Alejandro Ortega, Joseph Bloom, Stella Biderman, Adria Garriga-Alonso, Arthur Conmy, Neel Nanda, Jessica Rumbelow, Martin Wattenberg, Nandi Schoots, ...

  77. [85]

    Hypothesis testing the circuit hypothesis in LLM s

    Claudia Shi, Nicolas Beltran-Velez, Achille Nazaret, Carolina Zheng, Adri \`a Garriga-Alonso, Andrew Jesson, Maggie Makar, and David Blei. Hypothesis testing the circuit hypothesis in LLM s. In ICML 2024 Workshop on Mechanistic Interpretability, 2024. URL https://openreview.ne...

  78. [86]

    The golden mean of scientific virtues, 2024

    Adam Shimi. The golden mean of scientific virtues, 2024. URL https://formethods.substack.com/p/the-golden-mean-of-scientific-virtues

  79. [87]

    Knowledge in Perspective: Selected Essays in Epistemology

    Ernest Sosa. Knowledge in Perspective: Selected Essays in Epistemology. Cambridge University Press, New York, 1991

  80. [88]

    Grokking group multiplication with cosets

    Dashiell Stander, Qinan Yu, Honglu Fan, and Stella Biderman. Grokking group multiplication with cosets. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp (eds.), Proceedings of the 41st International Co...

  81. [89]

    Simplicity as Evidence of Truth

    Richard Swinburne. Simplicity as Evidence of Truth. Marquette University Press, Milwaukee, 1997

  82. [90]

    Tracrbench: Generating interpretability testbeds with large language models, 2024

    Hannes Thurnherr and Jérémy Scheurer. Tracrbench: Generating interpretability testbeds with large language models, 2024. URL https://arxiv.org/abs/2409.13714

  83. [91]

    Interpretability in the wild: a circuit for indirect object identification in gpt-2 small

    Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. Interpretability in the wild: a circuit for indirect object identification in gpt-2 small. In The Eleventh International Conference on Learning Representations, 2023

  84. [92]

    Why favour simplicity? Analysis, 65 0 (3): 0 205--210, 2005

    Roger White. Why favour simplicity? Analysis, 65 0 (3): 0 205--210, 2005. doi:10.1093/analys/65.3.205

  85. [93]

    Understanding as compression

    Daniel A Wilkenfeld. Understanding as compression. Philosophical Studies, 176 0 (10): 0 2807--2831, 2019

  86. [94]

    Geschichte und Naturwissenschaft

    Wilhelm Windelband. Geschichte und Naturwissenschaft. Rede zum Antritt des Rectorats der Kaiser-Wilhelms-Universit \"a t Strassburg, geh. am 1. Mai 1894 . Heitz, 1894

  87. [95]

    From probability to consilience: How explanatory values implement bayesian reasoning

    Zachary Wojtowicz and Simon DeDeo. From probability to consilience: How explanatory values implement bayesian reasoning. Trends in Cognitive Sciences, 24 0 (12): 0 981--993, 2020

  88. [96]

    Woodward

    James F. Woodward. Making Things Happen: A Theory of Causal Explanation. Oxford University Press, New York, 2003

  89. [97]

    Unifying and verifying mechanistic interpretations: A case study with group operations

    Wilson Wu, Louis , Jacob Drori, and Jason Gross. Unifying and verifying mechanistic interpretations: A case study with group operations. arXiv preprint arXiv:2410.07476, 2024

  90. [98]

    Manning, and Christopher Potts

    Zhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang, Jing Huang, Dan Jurafsky, Christopher D. Manning, and Christopher Potts. Axbench: Steering llms? even simple baselines outperform sparse autoencoders, 2025. URL https://arxiv.org/abs/2501.17148

  91. [99]

    A theory of usable information under computational constraints

    Yilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart, and Stefano Ermon. A theory of usable information under computational constraints. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=r1eBeyHFDH

  92. [100]

    Locally decodable codes

    Sergey Yekhanin et al. Locally decodable codes. Foundations and Trends in Theoretical Computer Science , 6 0 (3): 0 139--255, 2012

  93. [101]

    The shift from models to compound ai systems

    Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, and Ali Ghodsi. The shift from models to compound ai systems. https://bair.berkeley.edu/blog/2024/02/18/compound-ai-systems/, 2024

  94. [102]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

  95. [103]

    @esa (Ref

    \@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...

  96. [104]

    \@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...

  97. [105]

    @open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.