REVIEW 6 major objections 7 minor 105 references
Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii
T0 review · 6 major / 7 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This paper argues that a pluralist Explanatory Virtues Framework gives systematic criteria for choosing between competing mechanistic interpretability explanations, and that Compact Proofs, which embody many of these virtues, are a…
desk verdict A useful philosophical vocabulary for MI evaluation, but not yet the systematic framework it claims to be; the Compact Proofs verdict rests on a subjective rubric. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Explanatory Virtues Framework itself, built from the Bayesian, Kuhnian, Deutschian, and Nomological accounts of explanation: a directed graph of virtues, each with a mathematical definition, that turns theory choice into a comparison problem. Its machinery includes Bayesian decompositions—$\mathrm{Acc}(E)=P(x_T\mid E)$ for accuracy, precision as expected log-likelihood, descriptiveness and co-explanation as additive and joint components of log-likelihood, power and unification as their theoretical counterparts—and a hard-to-varyness condition: an explanation is hard-to-vary if it sits at a local maximum of $\log(\mathrm{Acc}(E))-k(E)$, where $k$ is a complexity measure and a modification is a sequence of insertions, deletions, substitutions, or transpositions of symbols. The other load-bearing piece is the Compact Proofs evaluator, in which an explanation is converted into a verifier $V(\theta,E)$ that returns a worst-case performance bound; the tightness of the bound is accuracy, and the computational cost of checking the proof is simplicity, so a good explanation pushes out the (tightness, compactness) Pareto frontier. Validity conditions—model-level, ontic, causal-mechanistic, falsifiable—must be met before the virtues are compared.
What would settle it
Take one model, such as a small transformer trained to compose group operations, generate two valid MI explanations of the same behaviour that score very differently on the virtues, then run a battery of held-out intervention experiments that the two explanations predict differently; if the lower-scoring explanation predicts the held-out behaviour at least as well as the high-scoring one, the claim that the virtues are truth-conducive is falsified.
Extended reading notes
Core claim
The central claim is that 'what makes a good explanation?' can be answered by a pluralist set of explanatory virtues, and that these virtues—not subjective preference—should drive theory choice in mechanistic interpretability. To be valid, an MI explanation must be model-level, ontic, causal-mechanistic, and falsifiable. Among valid explanations, the paper maintains that the better explanation is the one embodying more of the virtues: empirical virtues such as accuracy, descriptiveness, co-explanation, and fruitfulness, and theoretical virtues such as precision, power, unification, consistency, simplicity, hard-to-varyness, and nomologicity. The framework attaches each virtue to one of four philosophy-of-science accounts and gives it a mathematical definition; for example, accuracy is a likelihood and hard-to-varyness is being a local maximum of log-accuracy minus complexity. When the rubric is applied to clustering, sparse autoencoders, causal-circuit analysis, and Compact Proofs, the paper finds that Compact Proofs are the method that considers many virtues and that current methods systematically neglect simplicity, unification, co-explanation, and nomological principles.
Load-bearing premise
The load-bearing premise is that the listed explanatory virtues are truth-conducive—that an explanation embodying more of them is genuinely more likely to be correct—because without that link the framework becomes a subjective preference list rather than a guide to which explanation of a neural network is actually true.
Editorial extensions
If this is right
- MI researchers can replace subjective intuition with a shared rubric when two explanations of the same model conflict, giving explicit epistemic reasons for theory choice.
- Methods can be improved on the axes the framework flags as neglected: simplicity, unification, co-explanation, and nomological principles.
- Compact Proofs offer a concrete way to turn the accuracy–simplicity trade-off into a measurable Pareto frontier, with faithful explanations yielding tighter bounds at lower verification cost.
- A clearer definition of explanatory simplicity would let different explanation methods be compared on a single accuracy–simplicity curve.
- Seeking universal principles and reused building blocks would move interpretability from cataloguing individual cases toward nomological, predictive explanations.
Reading between the lines
- A natural test of the framework is to instantiate each virtue as a computable proxy—likelihood for accuracy, description length for simplicity, edit distance to a local maximum for hard-to-varyness—and score the same set of explanations with and without the proxies; convergence would suggest the rubric is measuring something real.
- The truth-conduciveness claim is an empirical hypothesis: run a prediction tournament where rival explanations of the same model are scored on virtues and then probed by novel interventions, and see whether the virtue leader continues to predict behaviour outside its training distribution.
- The framework's emphasis on unification suggests that shared, reused substructures should be privileged across tasks; if the same units keep predicting behaviour in new tasks, that would strengthen the link between unification and truth.
- If truth-conduciveness fails, the framework still documents how interpretability researchers actually trade off values, but it would lose its normative force as a guide to which explanation is correct.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces an Explanatory Virtues Framework for evaluating explanations in mechanistic interpretability (MI), drawing on Bayesian, Kuhnian, Deutschian, and Nomological accounts from the philosophy of science. The authors define a set of virtues with mathematical formulas, apply the framework to four MI methods (Clustering, Sparse Autoencoders, Causal Circuits, and Compact Proofs), and conclude that Compact Proofs are a promising approach because they exhibit many virtues. The paper also proposes research directions emphasizing simplicity, unification, and nomological principles.
Significance. If the framework delivered what it promises—a canonical, computable way to compare competing MI explanations—it would be a valuable contribution to interpretability research, which currently lacks principled theory-choice criteria. The paper surveys relevant philosophy of science literature, and its taxonomy of virtues is genuinely useful as a vocabulary for discussing explanations. However, the paper does not provide machine-checked proofs, reproducible code, or quantitative data; the central evaluation in Table 1 is based on the authors' qualitative judgments. The paper's main value is therefore as a conceptual proposal, not as an established evaluation methodology.
major comments (6)
- [Section 3.1, glossary of Bayesian virtues] The definitions of Accuracy, Precision, Descriptiveness, Co-Explanation, Power, and Unification all presuppose that an explanation E is associated with a conditional probability P(x|E) over observational data. Yet none of the four methods analyzed in Section 4 is specified as a probabilistic model: a circuit is a structural causal model, an SAE is a dictionary plus feature descriptions, and a Compact Proof is a verifier program. The paper never states how to derive P(x|E) from these artifacts, so the promised 'consistent and canonical way to compute each virtue' (Section 3, p. 4) is not defined for the very explanations the framework is meant to compare.
- [Table 1 vs. Section 3.1] The framework's formal definitions imply that Co-Explanation and Unification are the same quantity—CoEx(E) = log(Acc(E)) − Desc(E) and Unif(E) = Prec(E) − Power(E) = E_{xT∼X}[log(P(xT|E)/∏P(xT,i|E))], which are identical up to the expectation. Yet Table 1 rates Compact Proofs as '✓' on Co-explanation and '✗' on Unification. This inconsistency is only possible because the Table 2 rubric is not derived from the formal definitions; it shows that the case-study ratings are subjective judgments rather than outcomes of the framework.
- [Section 4.2 and Abstract] The central conclusion that 'Compact Proofs consider many explanatory virtues and are hence a promising approach' (Abstract) is based entirely on the authors' own qualitative ratings in Table 1, not on any measured data or inter-rater agreement. The paper provides no evidence that another researcher applying the Table 2 rubric would reach the same ratings, so the claimed systematic comparison between Chughtai et al. (2023) and Stander et al. (2024) is not actually demonstrated.
- [Section 2.2] The definition of valid MI explanations (Model-level, Ontic, Causal-Mechanistic, Falsifiable) is imported from the authors' own forthcoming paper (Ayonrinde & Jaburi, 2025), which is not available for verification. Because this definition is the foundation upon which the Explanatory Virtues Framework is built, the framework's epistemic grounding is self-referential; the paper should either provide the argument for this definition in the present manuscript or cite a published, accessible source.
- [Section 3, 'Explanatory Virtues are properties that are reliable indicators of truth'] The truth-conduciveness of the listed virtues is asserted with only a reference to Schindler (2018). This premise is load-bearing: if virtues are not truth-conducive, the framework becomes a subjective preference list. The paper should give at least one concrete argument or empirical example for why these particular virtues indicate truth, or explicitly reframe the contribution as a descriptive account of explanatory values.
- [Section 3.3, Hard-to-Varyness definition] The formalization of hard-to-varyness as a local maximum of hv(E) = log(Acc(E)) − k(E) is under-specified because 'local' is only informally defined as 'a small number of edit operations apart' (footnote 9); without a metric on the space of explanations, the definition cannot be applied to decide whether a given explanation is hard-to-vary, which is exactly what Table 1 claims.
minor comments (7)
- [Section 3.1, notation] The subscript formatting is inconsistent: the text uses 'x T', 'x I', and 'xT,i' in different places, and the glossary uses 'x T∼X' without clearly defining the distribution over inference-time data; please standardize the notation.
- [Section 3.2, Pragmatic Utility paragraph] There is a typo: 'Sparse Auutoencoder' should be 'Sparse Autoencoder'.
- [Section 4.1.3, FCM criteria] The FCM criteria use expressions F(C\K) and F(M\K) without defining the function F; presumably F is the model output function on a data distribution, but this should be stated explicitly.
- [Section 3.5 and Figure 1] The caption for Figure 1 references colors, bold arrows, and dashed arrows, but the figure is not included in the manuscript text; please ensure the figure is present or describe the relationships in the text.
- [Section 4.2 and Table 1] Section 4.2 describes Compact Proofs as a method for evaluating other explanations, yet Table 1 treats Compact Proofs as an explanation method comparable to Clustering, SAEs, and Circuits; please clarify this distinction or reclassify the comparison.
- [Section 5, paragraph on universality] The claim that 'the MI community has sought to understand the universality ... with mixed results' lacks a citation supporting the 'mixed results' assertion; please add a reference or soften the claim.
- [References] Several references are to unpublished or forthcoming works (Ayonrinde & Jaburi 2025; Jaburi et al. 2025; Ayonrinde 2025), making it difficult for readers to verify the cited claims; consider providing preprints or including the relevant content in an appendix.
Circularity Check
Self-cited validity definition and validity-as-virtue double counting raise circularity concerns, but the core framework retains independent philosophical content.
-
self citation load bearing
[Section 2.2, Definition of Mechanistic Interpretability; used in Section 3.5]
"Following Olah et al. (2020); Olsson et al. (2022), Ayonrinde &Jaburi (2025) define Mechanistic Interpretability as follows: Interpretability explanations are valid Mechanistic Interpretability explanations if they are Model-level, Ontic, Causal-Mechanistic, and Falsifiable."
The paper's gate for applying the Explanatory Virtues Framework—what counts as a valid MI explanation—is taken from the authors' own forthcoming Part I.i, not from an independent or externally checkable source. Section 3.5 then makes this imported definition load-bearing: 'For an explanation to be a good explanation in Mechanistic Interpretability, it must first be a valid MI explanation. In Section 2.2 we identified valid MI explanations as those which are Model-Level, Ontic, Causal-Mechanistic, Falsifiable.' Thus the framework's scope and its later comparison of methods rest on a self-citation to work that is not yet available and is authored by the same people.
-
self definitional
[Figure 1 / Section 3.5 / Table 1 (Validity rows)]
"The Explanatory Virtues which are essential for any scientific explanation (Falsifiability and Causal-Mechanisticity) to be valid are denoted with an exclamation mark; the most important virtues to decide between explanations (Simplicity, Hard-to-Varyness, and Fruitfulness) are marked with a star."
The paper first defines a valid MI explanation as one that is Model-level, Ontic, Causal-Mechanistic, and Falsifiable (Section 2.2), then lists Falsifiability and Causal-Mechanisticity among the Explanatory Virtues in Figure 1 and Table 1. This means part of the framework's virtue score re-asserts the definition of the object being evaluated: a method scores well on these 'virtues' precisely because it is admitted as a valid MI explanation, not because the framework has discovered an independent criterion for choosing between competing valid explanations. The motivating comparison between Chughtai et al. and Stander et al. is between two explanations that both already satisfy these validity conditions, so these rows cannot decide between them.
full rationale
The central Bayesian, Kuhnian, Deutschian, and Nomological virtues are drawn from external philosophy-of-science sources, and their mathematical definitions (Accuracy, Precision, Descriptiveness, Co-Explanation, Power, Unification, Hard-to-Varyness, etc.) are not fitted to any dataset and do not themselves constitute a renamed prediction. The conclusion favoring Compact Proofs is based on the authors' rubric in Table 2 rather than on a computation using those equations; indeed, P(x|E) is not specified for circuits, SAE explanations, or Compact Proofs. That is a gap in derivational support, but it is not a case of a fitted parameter renamed as a prediction, so it does not count as circularity under the hard rules. The concrete circular elements are the self-cited definition of valid MI explanations and the double-counting of validity conditions (Falsifiability, Causal-Mechanisticity) as explanatory virtues. These are load-bearing for the framework's scope and for part of the Table 1 evaluation, but they do not consume the whole framework: the remaining virtues and the philosophical analysis retain independent content. Hence a score of 4 is appropriate.
Assumptions & free parameters
assumptions (5)
- domain assumption Explanatory virtues are truth-conducive properties of explanations.
- domain assumption The validity definition of Mecanistic Interpretability (Model-level, Ontic, Causal-Mechanistic, Falsifiable) is correct.
- standard math The Bayesian decomposition of Accuracy into Descriptiveness and Co-Explanation, and Precision into Power and Unification, is meaningful.
- ad hoc to paper Hard-to-varyness is correctly formalized as being at a local maximum of log(Accuracy) minus complexity.
- domain assumption Nomologicity (appealing to general laws) is an explanatory virtue.
invented entities (2)
-
Co-Explanation as a distinct explanatory virtue
-
Formalization of Hard-to-varyness as a local maximum of hv(E)
Cite this review
Pith. "Pith review of Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii." pith.science (2026). https://pith.science/paper/P7WATJPK
@misc{pith2026250501372,
author = {Pith},
title = {Pith review of: Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii},
year = {2026},
howpublished = {\url{https://pith.science/paper/P7WATJPK}},
note = {Machine review of arXiv:2505.01372}
}
read the original abstract
Mechanistic Interpretability (MI) aims to understand neural networks through causal explanations. Though MI has many explanation-generating methods, progress has been limited by the lack of a universal approach to evaluating explanations. Here we analyse the fundamental question "What makes a good explanation?" We introduce a pluralist Explanatory Virtues Framework drawing on four perspectives from the Philosophy of Science - the Bayesian, Kuhnian, Deutschian, and Nomological - to systematically evaluate and improve explanations in MI. We find that Compact Proofs consider many explanatory virtues and are hence a promising approach. Fruitful research directions implied by our framework include (1) clearly defining explanatory simplicity, (2) focusing on unifying explanations and (3) deriving universal principles for neural networks. Improved MI methods enhance our ability to monitor, predict, and steer AI systems.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
The computational complexity of circuit discovery for inner interpretability
Federico Adolfi, Martina G Vilas, and Todd Wareham. The computational complexity of circuit discovery for inner interpretability. arXiv preprint arXiv:2410.08025, 2024
arXiv 2024
-
[2]
Physics of language models: Part 3.1, knowledge storage and extraction
Zeyuan Allen-Zhu and Yuanzhi Li. Physics of language models: Part 3.1, knowledge storage and extraction. arXiv preprint arXiv:2309.14316, 2023
arXiv 2023
-
[3]
Physics of language models: Part 3.1, knowledge storage and extraction
Zeyuan Allen-Zhu and Yuanzhi Li. Physics of language models: Part 3.1, knowledge storage and extraction. In Forty-first International Conference on Machine Learning, 2024
2024
-
[4]
The urgency of interpretability, 2025
Dario Amodei. The urgency of interpretability, 2025. URL https://www.darioamodei.com/post/the-urgency-of-interpretability
2025
-
[5]
Foundational challenges in assuring alignment and safety of large language models
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, et al. Foundational challenges in assuring alignment and safety of large language models. arXiv preprint arXiv:2404.09932, 2024
arXiv 2024
-
[6]
Ai as systems, not just models, 2024
Andy Arditi. Ai as systems, not just models, 2024. URL https://www.lesswrong.com/posts/2po6bp2gCHzxaccNz/ai-as-systems-not-just-models
2024
-
[7]
Standard saes might be incoherent: A choosing problem & a “concise” solution
Kola Ayonrinde. Standard saes might be incoherent: A choosing problem & a “concise” solution. Blog post, 2024. URL https://www.lesswrong.com/posts/vNCAQLcJSzTgjPaWS/standard-saes-might-be-incoherent-a-choosing-problem-and-a
2024
-
[8]
Position: Interpretability is a bidirectional communication problem
Kola Ayonrinde. Position: Interpretability is a bidirectional communication problem. In ICLR 2025 Workshop on Bidirectional Human-AI Alignment, 2025. URL https://openreview.net/forum?id=O4LaRH4zSI
2025
Show all 105 references
-
[9]
A mathematical philosophy of explanations in mechanistic interpretability: The strange science part i.i, 2025
Kola Ayonrinde and Louis Jaburi. A mathematical philosophy of explanations in mechanistic interpretability: The strange science part i.i, 2025. forthcoming
2025
-
[10]
Pearce, and Lee Sharkey
Kola Ayonrinde, Michael T. Pearce, and Lee Sharkey. Interpretability as compression: Reconsidering sae explanations of neural activations with mdl-saes, 2024. URL https://arxiv.org/abs/2410.11179
2024 arXiv
-
[11]
Novum Organum
Francis Bacon. Novum Organum. Clarendon Press, London, 1620. URL https://en.wikipedia.org/wiki/Novum_Organum. Part of the Instauratio Magna
-
[12]
Simplicity
Alan Baker. Simplicity . In Edward N. Zalta (ed.), The Stanford Encyclopedia of Philosophy . Metaphysics Research Lab, Stanford University, S ummer 2022 edition, 2022
2022
-
[13]
Design Rules: The Power of Modularity Volume 1
Carliss Y Baldwin and Kim B Clark. Design Rules: The Power of Modularity Volume 1. MIT press, 1999
1999
-
[14]
Icml 2024 mechanistic interpretability workshop, 2024
Fazl Barez, Mor Geva, Lawrence Chan, Atticus Geiger, Kayo Yin, Neel Nanda, and Max Tegmark. Icml 2024 mechanistic interpretability workshop, 2024. URL https://icml2024mi.pages.dev/
2024
-
[15]
Local vs
Shahaf Bassan, Guy Amir, and Guy Katz. Local vs. global interpretability: A computational complexity perspective. arXiv preprint arXiv:2406.02981, 2024
2024 arXiv
-
[16]
Explanation: A mechanist alternative
William Bechtel and Adele Abrahamsen. Explanation: A mechanist alternative. Studies in History and Philosophy of Science Part C: Studies in History and Philosophy of Biological and Biomedical Sciences, 36 0 (2): 0 421--441, 2005. doi:10.1016/j.shpsc.2005.03.010
2005 doi
-
[17]
Sander Beckers and Joseph Y. Halpern. Abstracting causal models. In Proceedings of the 33Rd Aaai Conference on Artificial Intelligence, pp.\ 2678--2685. 2019
2019
-
[18]
International ai safety report
Yoshua Bengio, S \"o ren Mindermann, Daniel Privitera, Tamay Besiroglu, Rishi Bommasani, Stephen Casper, Yejin Choi, Philip Fox, Ben Garfinkel, Danielle Goldfarb, et al. International ai safety report. arXiv preprint arXiv:2501.17805, 2025
2025 arXiv
-
[19]
Mechanistic Interpretability for AI Safety -- A Review , April 2024
Leonard Bereska and Efstratios Gavves. Mechanistic Interpretability for AI Safety -- A Review , April 2024. URL http://arxiv.org/abs/2404.14082. arXiv:2404.14082 [cs]
2024 arXiv
-
[20]
J. Bernal. The social function of science. Philosophical Review, 49 0 (n/a): 0 377, 1940. doi:10.2307/2180883
1940 doi
-
[21]
Auditing local explanations is hard
Robi Bhattacharjee and Ulrike von Luxburg. Auditing local explanations is hard. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/forum?id=ybMrn4tdn0
2024
-
[22]
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders. Language models can explain neurons in language models. https://openaipublic.blob.core.windows.net/neuron-explainer/paper/index.html, 2023
2023
-
[23]
An Interpretability Illusion for BERT , April 2021
Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Viégas, and Martin Wattenberg. An Interpretability Illusion for BERT , April 2021. URL http://arxiv.org/abs/2104.07143. arXiv:2104.07143 [cs]
2021 arXiv
-
[24]
Identifying Functionally Important Features with End -to- End Sparse Dictionary Learning , May 2024
Dan Braun, Jordan Taylor, Nicholas Goldowsky-Dill, and Lee Sharkey. Identifying Functionally Important Features with End -to- End Sparse Dictionary Learning , May 2024. URL http://arxiv.org/abs/2405.12241. arXiv:2405.12241 [cs]
2024 arXiv
-
[25]
Interpretability in parameter space: Minimizing mechanistic description length with attribution-based parameter decomposition
Dan Braun, Lucius Bushnaq, Stefan Heimersheim, Jake Mendel, and Lee Sharkey. Interpretability in parameter space: Minimizing mechanistic description length with attribution-based parameter decomposition. arXiv preprint arXiv:2501.14926, 2025
2025 arXiv
-
[26]
Towards Monosemanticity : Decomposing Language Models With Dictionary Learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen,...
2023
-
[27]
Propositional interpretability in artificial intelligence
David J Chalmers. Propositional interpretability in artificial intelligence. arXiv preprint arXiv:2501.15740, 2025
2025 arXiv
-
[28]
A toy model of universality: Reverse engineering how networks learn group operations
Bilal Chughtai, Lawrence Chan, and Neel Nanda. A toy model of universality: Reverse engineering how networks learn group operations. In International Conference on Machine Learning, pp.\ 6243--6267. PMLR, 2023
2023
-
[29]
The evolutionary origins of modularity
Jeff Clune, Jean-Baptiste Mouret, and Hod Lipson. The evolutionary origins of modularity. Proceedings of the Royal Society b: Biological sciences, 280 0 (1755): 0 20122863, 2013
2013
-
[30]
Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso. Towards automated circuit discovery for mechanistic interpretability. 2023. URL https://arxiv.org/abs/2304.14997
2023 arXiv
-
[31]
Central dogma of molecular biology
Francis Crick. Central dogma of molecular biology. Nature, 227 0 (5258): 0 561--563, 1970
1970
-
[32]
The beginning of infinity: Explanations that transform the world
David Deutsch. The beginning of infinity: Explanations that transform the world. penguin uK, 2011
2011
-
[33]
Frank Watson Dyson, Arthur Stanley Eddington, and Charles Davidson. Ix. a determination of the deflection of light by the sun's gravitational field, from observations made at the total eclipse of may 29, 1919. Philosophical Transactions of the Royal Society of London. Series A...
1919
-
[34]
Einstein
A. Einstein. The foundation of the general theory of relativity. 1916
1916
-
[35]
Richard P. Feynman. Cargo cult science. Engineering and Science, 37 0 (7): 0 10--13, 1974. ISSN 0013-7812. URL http://resolver.caltech.edu/CaltechES:37.7.CargoCult
1974
-
[36]
Clusterability in neural networks
Daniel Filan, Stephen Casper, Shlomi Hod, Cody Wild, Andrew Critch, and Stuart Russell. Clusterability in neural networks. arXiv preprint arXiv:2103.03386, 2021
2021 arXiv
-
[37]
Interpretability illusions in the generalization of simplified models
Dan Friedman, Andrew Lampinen, Lucas Dixon, Danqi Chen, and Asma Ghandeharioun. Interpretability illusions in the generalization of simplified models. In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR.org, 2024
2024
-
[38]
Scaling and evaluating sparse autoencoders, June 2024
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. Scaling and evaluating sparse autoencoders, June 2024. URL http://arxiv.org/abs/2406.04093. arXiv:2406.04093 [cs] version: 1
2024 arXiv
-
[39]
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. Causal abstractions of neural networks. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (eds.), Advances in Neural Information Processing Systems, volume 34, pp.\ 9574--9586. Curran A...
2021
-
[40]
Causal abstraction for faithful model interpretation
Atticus Geiger, Chris Potts, and Thomas Icard. Causal abstraction for faithful model interpretation. arXiv preprint arXiv:2301.04709, 2023
2023 arXiv
-
[41]
Clustering algorithms
Google Developers . Clustering algorithms. https://developers.google.com/machine-learning/clustering/clustering-algorithms, 2025. Accessed: 2025-02-23
2025
-
[42]
Compact proofs of model performance via mechanistic interpretability
Jason Gross, Rajashree Agrawal, Thomas Kwa, Euan Ong, Chun Hei Yip, Alex Gibson, Soufiane Noubir, and Lawrence Chan. Compact proofs of model performance via mechanistic interpretability. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024
2024
-
[43]
Interpbench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques
Rohan Gupta, Iv\' a n Arcuschin, Thomas Kwa, and Adri\` a Garriga-Alonso. Interpbench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (eds.), Advances in N...
2024
-
[44]
Hastie, R
T. Hastie, R. Tibshirani, and J.H. Friedman. The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer series in statistics. Springer, 2009. ISBN 9780387848846. URL https://books.google.co.uk/books?id=eBSgoAEACAAJ
2009
-
[45]
Hempel and Paul Oppenheim
Carl G. Hempel and Paul Oppenheim. Studies in the logic of explanation. Philosophy of Science, 15 0 (2): 0 135--175, 1948. ISSN 00318248, 1539767X. URL http://www.jstor.org/stable/185169
1948
-
[46]
Philosophy of Natural Science
Carl Gustav Hempel. Philosophy of Natural Science. Prentice-Hall, Englewood Cliffs, N.J.,, 1966
1966
-
[47]
Bayesianism and inference to the best explanation
Leah Henderson. Bayesianism and inference to the best explanation. The British Journal for the Philosophy of Science, 2014
2014
-
[48]
The developmental landscape of in-context learning
Jesse Hoogland, George Wang, Matthew Farrugia-Roberts, Liam Carroll, Susan Wei, and Daniel Murfet. The developmental landscape of in-context learning. arXiv preprint arXiv:2402.02364, 2024
2024 arXiv
-
[49]
Sparse autoencoders find highly interpretable features in language models
Robert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, and Lee Sharkey. Sparse autoencoders find highly interpretable features in language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=F76bwRSLeK
2024
-
[50]
Hutter, E
M. Hutter, E. Catt, and D. Quarel. An Introduction to Universal Artificial Intelligence. Chapman & Hall/CRC Artificial Intelligence and robotics series. Chapman & Hall/CRC Press, 2024. ISBN 9781003460299. URL https://books.google.co.uk/books?id=cfg60AEACAAJ
2024
-
[51]
Fine-tuning neural networks to match their interpretation: Towards scaling compact proofs, 2025
Louis Jaburi, Ronak Mehta, Soufiane Noubir, and Jason Gross. Fine-tuning neural networks to match their interpretation: Towards scaling compact proofs, 2025. forthcoming
2025
-
[52]
Kandel, J.H
E.R. Kandel, J.H. Schwartz, and T. Jessell. Principles of Neural Science, Fourth Edition. McGraw-Hill Companies,Incorporated, 2000. ISBN 9780838577011. URL https://books.google.co.uk/books?id=yzEFK7Xc87YC
2000
-
[53]
Saebench: a comprehensive benchmark for sparse autoencoders, 2024
Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Bloom, David Chanin, Yeu-Tong Lau, Eoin Farrell, Arthur Conmy, Callum McDougall, Kola Ayonrinde, Matthew Wearden, Samuel Marks, and Neel Nanda. Saebench: a comprehensive benchmark for sparse autoencoders, 2024. URL http...
2024
-
[54]
Kennefick
D. Kennefick. No Shadow of a Doubt: The 1919 Eclipse That Confirmed Einstein's Theory of Relativity. Princeton University Press, 2021. ISBN 9780691217154. URL https://books.google.co.uk/books?id=_Eb8DwAAQBAJ
1919
-
[55]
Explanatory unification
Philip Kitcher. Explanatory unification. Philosophy of science, 48 0 (4): 0 507--531, 1981
1981
-
[56]
Three approaches to the quantitative definition ofinformation
Andrei N Kolmogorov. Three approaches to the quantitative definition ofinformation. Problems of information transmission, 1 0 (1): 0 1--7, 1965
1965
-
[57]
Thomas S. Kuhn. Objectivity, value judgment, and theory choice. In David Zaret (ed.), Review of Thomas S. Kuhn The Essential Tension: Selected Studies in Scientific Tradition and Change, pp.\ 320--39. Duke University Press, 1981
1981
-
[58]
The Structure of Scientific Revolutions
Thomas Samuel Kuhn. The Structure of Scientific Revolutions. University of Chicago Press, Chicago, 1962
1962
-
[59]
Falsification and the methodology of scientific research programmes
Imre Lakatos. Falsification and the methodology of scientific research programmes. In Imre Lakatos and Alan Musgrave (eds.), Criticism and the growth of knowledge, pp.\ 91--196. Cambridge University Press, 1970
1970
-
[60]
The Methodology of Scientific Research Programmes
Imre Lakatos. The Methodology of Scientific Research Programmes. Cambridge University Press, New York, 1978
1978
-
[61]
Sparse autoencoders do not find canonical units of analysis
Patrick Leask, Bart Bussmann, Michael Pearce, Joseph Bloom, Curt Tigges, Noura Al Moubayed, Lee Sharkey, and Neel Nanda. Sparse autoencoders do not find canonical units of analysis. arXiv preprint arXiv:2502.04878, 2025
2025 arXiv
-
[62]
Lindsay and David Bau
Grace W. Lindsay and David Bau. Testing methods of neural systems understanding. Cogn. Syst. Res., 82: 0 101156, December 2023. URL https://doi.org/10.1016/j.cogsys.2023.101156
2023
-
[63]
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C Lipton. The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery. Queue, 16 0 (3): 0 31--57, 2018
2018
-
[64]
Mechanistic mode connectivity
Ekdeep Singh Lubana, Eric J Bigelow, Robert P Dick, David Krueger, and Hidenori Tanaka. Mechanistic mode connectivity. In Proceedings of the 40th International Conference on Machine Learning, pp.\ 22965--23004, 2023
2023
-
[65]
Information theory, inference and learning algorithms
David JC MacKay. Information theory, inference and learning algorithms. Cambridge university press, 2003
2003
-
[66]
Is this the subspace you are looking for? an interpretability illusion for subspace activation patching
Aleksandar Makelov, Georg Lange, Atticus Geiger, and Neel Nanda. Is this the subspace you are looking for? an interpretability illusion for subspace activation patching. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum...
2024
-
[67]
Downstream applications as validation of interpretability progress, March 2025
Sam Marks. Downstream applications as validation of interpretability progress, March 2025. URL https://www.lesswrong.com/posts/wGRnzCFcowRCrpX4Y/downstream-applications-as-validation-of-interpretability
2025
-
[68]
Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller
Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller. Sparse Feature Circuits : Discovering and Editing Interpretable Causal Graphs in Language Models , March 2024. URL http://arxiv.org/abs/2403.19647. arXiv:2403.19647 [cs]
2024 arXiv
-
[69]
Cognitive styles in two cognitive sciences
James Myers. Cognitive styles in two cognitive sciences. In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 34, 2012
2012
-
[70]
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. Zoom in: An introduction to circuits. Distill, 5 0 (3): 0 e00024--001, 2020
2020
-
[71]
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. In-context learning and induction heads. arXiv preprint arXiv:2209.11895, 2022
2022 arXiv
-
[72]
Automatically interpreting millions of features in large language models
Gon c alo Paulo, Alex Mallen, Caden Juang, and Nora Belrose. Automatically interpreting millions of features in large language models. arXiv preprint arXiv:2410.13928, 2024
2024 arXiv
-
[73]
Causality
Judea Pearl. Causality. Cambridge university press, 2009
2009
-
[74]
Poincar \'e
H. Poincar \'e . Science and Hypothesis. Library of philosophy, psychology and scientific methods. Science Press, 1905. URL https://books.google.co.uk/books?id=5nQSAAAAYAAJ
1905
-
[75]
Karl R. Popper. The Logic of Scientific Discovery. Routledge, London, England, 1935
1935
-
[76]
Hume on theoretical simplicity
Hsueh Qu. Hume on theoretical simplicity. Philosophers' Imprint, 23 0 (1), 2023. doi:10.3998/phimp.1521
2023 doi
-
[77]
Escalation risks from language models in military and diplomatic decision-making
Juan-Pablo Rivera, Gabriel Mukobi, Anka Reuel, Max Lamparth, Chandler Smith, and Jacquelyn Schneider. Escalation risks from language models in military and diplomatic decision-making. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, FAcc...
2024
-
[78]
Four decades of scientific explanation
Wesley Salmon. Four decades of scientific explanation. 1989. URL https://api.semanticscholar.org/CorpusID:46466034
1989
-
[79]
Wesley C. Salmon. Scientific Explanation and the Causal Structure of the World. Princeton University Press, 1984. ISBN 9780691101705
1984
-
[80]
Mechanistic?, 2024
Naomi Saphra and Sarah Wiegreffe. Mechanistic?, 2024. URL https://arxiv.org/abs/2410.09087
2024 arXiv
-
[81]
Theoretical Virtues in Science: Uncovering Reality Through Theory
Samuel Schindler. Theoretical Virtues in Science: Uncovering Reality Through Theory. Cambridge University Press, Cambridge, 2018
2018
-
[82]
Riechers, Lucas Teixeira, Alexander Gietelink Oldenziel, and Sarah Marzen
Adam Shai, Paul M. Riechers, Lucas Teixeira, Alexander Gietelink Oldenziel, and Sarah Marzen. Transformers represent belief state geometry in their residual stream. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.ne...
2024
-
[83]
A mathematical theory of communication
Claude Elwood Shannon. A mathematical theory of communication. The Bell system technical journal, 27 0 (3): 0 379--423, 1948
1948
-
[84]
Michaud, Stephen Casper, Max Tegmark, William Saunders, David Bau, Eric Todd, Atticus Geiger, Mor Geva, Jesse Hoogland, Daniel Murfet, and Tom McGrath
Lee Sharkey, Bilal Chughtai, Joshua Batson, Jack Lindsey, Jeff Wu, Lucius Bushnaq, Nicholas Goldowsky-Dill, Stefan Heimersheim, Alejandro Ortega, Joseph Bloom, Stella Biderman, Adria Garriga-Alonso, Arthur Conmy, Neel Nanda, Jessica Rumbelow, Martin Wattenberg, Nandi Schoots, ...
2025 arXiv
-
[85]
Hypothesis testing the circuit hypothesis in LLM s
Claudia Shi, Nicolas Beltran-Velez, Achille Nazaret, Carolina Zheng, Adri \`a Garriga-Alonso, Andrew Jesson, Maggie Makar, and David Blei. Hypothesis testing the circuit hypothesis in LLM s. In ICML 2024 Workshop on Mechanistic Interpretability, 2024. URL https://openreview.ne...
2024
-
[86]
The golden mean of scientific virtues, 2024
Adam Shimi. The golden mean of scientific virtues, 2024. URL https://formethods.substack.com/p/the-golden-mean-of-scientific-virtues
2024
-
[87]
Knowledge in Perspective: Selected Essays in Epistemology
Ernest Sosa. Knowledge in Perspective: Selected Essays in Epistemology. Cambridge University Press, New York, 1991
1991
-
[88]
Grokking group multiplication with cosets
Dashiell Stander, Qinan Yu, Honglu Fan, and Stella Biderman. Grokking group multiplication with cosets. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp (eds.), Proceedings of the 41st International Co...
2024
-
[89]
Simplicity as Evidence of Truth
Richard Swinburne. Simplicity as Evidence of Truth. Marquette University Press, Milwaukee, 1997
1997
-
[90]
Tracrbench: Generating interpretability testbeds with large language models, 2024
Hannes Thurnherr and Jérémy Scheurer. Tracrbench: Generating interpretability testbeds with large language models, 2024. URL https://arxiv.org/abs/2409.13714
2024 arXiv
-
[91]
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. Interpretability in the wild: a circuit for indirect object identification in gpt-2 small. In The Eleventh International Conference on Learning Representations, 2023
2023
-
[92]
Why favour simplicity? Analysis, 65 0 (3): 0 205--210, 2005
Roger White. Why favour simplicity? Analysis, 65 0 (3): 0 205--210, 2005. doi:10.1093/analys/65.3.205
2005 doi
-
[93]
Understanding as compression
Daniel A Wilkenfeld. Understanding as compression. Philosophical Studies, 176 0 (10): 0 2807--2831, 2019
2019
-
[94]
Geschichte und Naturwissenschaft
Wilhelm Windelband. Geschichte und Naturwissenschaft. Rede zum Antritt des Rectorats der Kaiser-Wilhelms-Universit \"a t Strassburg, geh. am 1. Mai 1894 . Heitz, 1894
-
[95]
From probability to consilience: How explanatory values implement bayesian reasoning
Zachary Wojtowicz and Simon DeDeo. From probability to consilience: How explanatory values implement bayesian reasoning. Trends in Cognitive Sciences, 24 0 (12): 0 981--993, 2020
2020
-
[96]
Woodward
James F. Woodward. Making Things Happen: A Theory of Causal Explanation. Oxford University Press, New York, 2003
2003
-
[97]
Unifying and verifying mechanistic interpretations: A case study with group operations
Wilson Wu, Louis , Jacob Drori, and Jason Gross. Unifying and verifying mechanistic interpretations: A case study with group operations. arXiv preprint arXiv:2410.07476, 2024
2024 arXiv
-
[98]
Manning, and Christopher Potts
Zhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang, Jing Huang, Dan Jurafsky, Christopher D. Manning, and Christopher Potts. Axbench: Steering llms? even simple baselines outperform sparse autoencoders, 2025. URL https://arxiv.org/abs/2501.17148
2025 arXiv
-
[99]
A theory of usable information under computational constraints
Yilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart, and Stefano Ermon. A theory of usable information under computational constraints. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=r1eBeyHFDH
2020
-
[100]
Locally decodable codes
Sergey Yekhanin et al. Locally decodable codes. Foundations and Trends in Theoretical Computer Science , 6 0 (3): 0 139--255, 2012
2012
-
[101]
The shift from models to compound ai systems
Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, and Ali Ghodsi. The shift from models to compound ai systems. https://bair.berkeley.edu/blog/2024/02/18/compound-ai-systems/, 2024
2024
-
[102]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
-
[103]
@esa (Ref
\@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...
-
[104]
\@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...
-
[105]
@open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.