REVIEW 3 major objections 6 minor 1 cited by
How Causal Abstraction Underpins Computational Explanation
T0 review · 3 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read A system implements a computation only if the algorithm is a causal abstraction of it.
desk verdict A clear, honest programmatic paper connecting causal abstraction to implementation, but the sufficiency claim rests on an unformalized 'native vehicles' constraint and the necessary condition is nearly vacuous under arbitrary translations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the abstraction-under-translation relation built from three ingredients: exact transformation, constructive abstraction, and translation. An exact transformation is a pair of maps $(\tau,\omega)$ from low-level values and interventions to high-level values and interventions such that running the low-level model, intervening, and translating always matches the high-level model's run under the corresponding intervention. A constructive abstraction ignores distinctions by partitioning low-level variables into macro-variables. A translation is a bijective recarving $\tau$ of the variable space, which induces a canonical high-level model and an intervention algebra $I_\tau$ of pu
What would settle it
Take a trained network that solves the hierarchical equality task and compute the gerrymandered translation $\tau$ witnessing the XNOR circuit as an abstraction-under-translation. Then test the network on novel stimulus pairs, such as arrows instead of faces, and check whether its outputs track the circuit's algorithm; a divergence would falsify the sufficiency claim. Conversely, if researchers agree a system implements a computation but no abstraction-under-translation can be exhibited, the necessity claim is falsified.
Extended reading notes
Core claim
The central claim is the 'No Computation without Abstraction' Principle: given a computational model $H$, a system $L$ implements $H$ only if $H$ is an abstraction-under-translation of $L$. Abstraction-under-translation is a two-step relation: translate $L$ by a bijective recarving $\tau$ of its variable space (with induced interventionals $I_\tau$), then take a constructive abstraction of $\tau(L)$ by partitioning low-level variables into macro-variables and checking that the induced maps form an exact transformation, so that low-level interventions reproduce high-level interventions. The authors illustrate the principle with a neural network trained on the hierarchical equality task whose
Load-bearing premise
The account assumes that any bijective recarving of the low-level system counts as a legitimate translation, which lets a known triviality theorem show that almost every network can be made to abstract every algorithm; the proposed necessary condition is then nearly vacuous unless the extra constraints mentioned in the paper are actually formalized.
Editorial extensions
If this is right
- If the principle is right, claims that a neural network implements a cognitive algorithm should be backed by an explicit abstraction-under-translation, not just by matching input-output behavior.
- Distributed representations in trained networks may require a translation such as a rotation before the algorithm's causal variables become visible, giving a causal rationale for linear representation hypotheses.
- The triviality theorem becomes a positive internal-model result: whenever a network solves a task, there exist modular low-level interventions that realize the algorithm, but those interventions may be too complex to find or use.
- Representational vehicles in implemented computations need not be individual neurons or regions; they can be abstract operations on the system, so content attaches to interventionals rather than to states alone.
- Prediction and generalization beyond observed inputs will require restricting admissible translations, for instance to mappings the system itself can natively support.
Reading between the lines
- The paper leaves the key constraint on translations informal: if 'native' vehicles could be formalized, the triviality theorem's gerrymandered translations would be excluded and the necessity claim would regain bite.
- A testable extension is to impose linearity on $\tau$ and check whether the recent triviality construction still goes through; if it does, linearity alone is too weak to ground implementation.
- The account suggests a link between implementation and compression: algorithmic causal variables that support generalization may be exactly those that yield compact descriptions of behavior, which could turn 'native vehicles' into a measurable quantity.
- The treatment of representational content implies that content is not an intrinsic property of a neural population, but a product of the intervention toolbox a scientist is willing to use; different toolboxes will yield different representational claims.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that computational implementation should be understood through the lens of causal abstraction. It reviews and uses formal notions from previous work—exact transformation, constructive abstraction, and abstraction-under-translation—and states a 'No Computation without Abstraction' principle: a system L implements a computational model H only if H is an abstraction-under-translation of L. The paper illustrates the framework with a neural network for the hierarchical equality task, discusses how representational content can be assigned to the abstract variables via causal role, and then confronts a triviality result by Sutter et al. (2025), which shows that under minimal assumptions every sufficiently large network with the right input/output behavior is an abstraction-under-translation of every algorithm. The paper interprets this result partly as a positive modular-control finding, but then argues that explanatory goals—especially generalization and prediction—require further constraints on admissible translations and vehicles, such as 'native' vehicles and linearity. The final sections present the account as a promising framework rather than a completed theory, explicitly leaving the admissibility problem open.
Significance. If the framework succeeds, it offers a precise causal formalization of a central concept in the philosophy of computation and connects long-standing debates about computational implementation, triviality, and representation to contemporary mechanistic interpretability. The paper's strengths are its precise definitions, the concrete neural-network case study, and its unusually honest treatment of the triviality problem: rather than ignoring Sutter et al., it engages directly and explores what a positive reading would require. It is also valuable in separating the purely structural notion of abstraction-under-translation from broader explanatory demands. The main limitation is that the proposed remedy to triviality—restricting to 'native' vehicles and generalization-supporting mappings—is not formalized, so the central positive claim remains programmatic. The necessary condition is defensible, but it is weak in light of the triviality theorem, and the sufficiency inclination in Section 7 is not a claim about the defined formalism as it stands.
major comments (3)
- [§7, with reference to §3.2 and §6] The paper states: 'we are inclined to accept abstraction-under-implementation not only as a necessary condition for computational implementation, but also as a sufficient one.' But the formal notion from §3.2 allows arbitrary bijective translations τ and their associated interventionals Iτ. As the paper itself reports in §6, Sutter et al. (2025) show that under minimal assumptions every sufficiently large network with the right task behavior is an abstraction-under-translation of every algorithm. Taken literally, the sufficiency claim would therefore assert that all such networks implement all algorithms. The response that vehicles must be 'native to how the system works' is not formalized: no constraint on τ or Iτ is added to the definition in §3.2, and no criterion for 'native' is given. The sufficiency claim is thus not a consequence of the formal framework; it is a gesture toward an
- [§6] The triviality theorem is treated as a positive finding about modular manipulation and control, but it undercuts the discriminatory content of the 'No Computation without Abstraction' principle. If abstraction-under-translation is satisfied by every sufficiently large network for every algorithm, then the necessary condition cannot distinguish genuine implementations from gerrymandered ones. The observation that Sutter et al. found the interventionals difficult to extract is epistemic and do not constitute a constitutive constraint on implementation; it says nothing about which translations are admissible. The paper needs to make a choice: either (a) concede that abstraction-under-translation is merely a necessary condition and that the full implementation concept requires additional, not-yet-specified conditions, or (b) prove that a restricted class of translations avoids the triviality
- [§2 and §7] The account depends on the substantive claim that computational models can and must be construed causally. The paper explicitly says this is 'not defended here' and that it is a stance adopted. Since the paper is making a philosophical claim about computational explanation, this is a load-bearing assumption: if the causal construal of computation is rejected, the entire abstraction-under-translation framework ceases to be an account of computational implementation. At minimum, the paper should provide a brief defense or a more precise scope condition stating that the argument is conditional on the causal construal. As written, the independence of the main principle from this assumption is unclear, and the reader is asked to accept a foundational premise without support.
minor comments (6)
- [§3.2] Typo: 'another casual model' should be 'another causal model'.
- [§7] The phrase 'each of these lays bear' is ungrammatical; it should be 'these lay bare' or 'this lays bare'.
- [§7] The term 'abstraction-under-implementation' is introduced without a definition and appears to be used synonymously with 'abstraction-under-translation.' Define the term or use a single consistent label.
- [§3.1] The demonstration that the XNOR circuit M is a constructive abstraction of the network N is very compressed. A more explicit step-by-step verification, or a precise pointer to the corresponding example in Geiger et al. (2025), would help the reader check the central example.
- [§2] The definitions of causal model, intervention, intervention algebra, and exact transformation are informal and not numbered. Since they carry the weight of the paper, a formal appendix with numbered definitions and statements of the relevant theorems from Geiger et al. (2025) would substantially improve usability.
- [References] The reference for Sutter et al. (2025) lacks a venue, arXiv identifier, or other locator, making it difficult to verify the statement of the triviality theorem cited in §6.
Circularity Check
Necessary condition is explicitly stipulated as a definition; central discussion remains independent, so partial circularity.
-
self definitional
[§3.2, Abstraction-Under-Translation definition and 'No Computation without Abstraction' Principle]
"With this much we can formulate what we take to be a necessary condition on any claim of computational implementation: “No Computation without Abstraction” Principle: Given a computational model H, another system L implements H only if H is an abstraction-under-translation of L. In other words, we claim that to implement a computation, the (causal construal of that) computation must be a constructive abstraction of a translation of the system in question."
The principle is stated immediately after the paper itself defines 'abstraction-under-translation' as 'H is an abstraction-under-translation of L if there is a translation τpLq of L such that H is a constructive abstraction of τpLq.' The necessary condition is thus an analytic consequence of that stipulated definition plus the additional stipulation that implementation entails it; it is not derived from independent premises about computation. The surrounding text presents it as capturing 'the spirit of claims in the literature,' i.e., as a characterization rather than a derived result. This is definitional circularity, though the paper is transparent about the stipulation.
full rationale
The only load-bearing step with a circular flavor is the 'No Computation without Abstraction' principle in §3.2: it is an immediate reformulation of the paper's own definition of abstraction-under-translation, prefaced by 'we take to be a necessary condition.' No independent argument establishes why implementation must coincide with this particular relation; the definition was expressly designed to capture prior implementation talk. This is a mild self-definitional move. However, the paper does not hide it, and the rest of the paper contains substantial independent content: the constructive-abstraction and translation examples are worked out, the engagement with Sutter et al.'s external triviality theorem is genuine, and the §7 discussion of generalization and 'native vehicles' is a substantive open problem rather than a circular derivation. The self-citations to Geiger et al. (2025) supply formal definitions and theorems that are checkable, not a uniqueness proof invoked to forbid alternatives. No parameter is fitted and no prediction is generated from data, so patterns 2, 5, and 6 do not apply. The unformalized 'native vehicles' constraint is a gap in the sufficiency argument, not a circularity.
Assumptions & free parameters
free parameters (3)
- translation τ
- partition Π and component maps π for constructive abstraction
- admissible interventionals Iτ
assumptions (4)
- domain assumption Computational models can be construed as causal models.
- standard math Causal models are acyclic functional models with mechanism replacement interventions.
- ad hoc to paper The interventionals IL and IH form intervention algebras.
- domain assumption Representational content is captured by the three criteria: Information, Use, Misrepresentation.
Cite this review
Pith. "Pith review of How Causal Abstraction Underpins Computational Explanation." pith.science (2026). https://pith.science/paper/TJIXTSOR
@misc{pith2026250811214,
author = {Pith},
title = {Pith review of: How Causal Abstraction Underpins Computational Explanation},
year = {2026},
howpublished = {\url{https://pith.science/paper/TJIXTSOR}},
note = {Machine review of arXiv:2508.11214}
}
read the original abstract
Explanations of cognitive behavior often appeal to computations over representations. What does it take for a system to implement a given computation over suitable representational vehicles within that system? We argue that the language of causality -- and specifically the theory of causal abstraction -- provides a fruitful lens on this topic. Drawing on current discussions in deep learning with artificial neural networks, we illustrate how classical themes in the philosophy of computation and cognition resurface in contemporary machine learning. We offer an account of computational implementation grounded in causal abstraction, and examine the role for representation in the resulting picture. We argue that these issues are most profitably explored in connection with generalization and prediction.
Figures
Forward citations
Cited by 1 Pith paper
-
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
In-context entity retrieval in LMs is a mixture of positional, lexical, and reflexive mechanisms; the pure positional view fails in middle positions of long lists.
Reference graph
Works this paper leans on
-
[1]
Alhama, R. G. and Zuidema, W. (2019). A review of computational models of basic rule learning: The neural-symbolic debate and beyond. Psychonomic Bulletin & Review , 26(4):1174--1194
2019
-
[2]
Arora, A., Jurafsky, D., and Potts, C. (2024). C ausal G ym: Benchmarking causal interpretability methods on linguistic tasks. In Ku, L.-W., Martins, A., and Srikumar, V., editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages 14638--14663, Bangkok, Thailand. Association for Computa...
2024
-
[3]
Beckers, S., Eberhardt, F., and Halpern, J. Y. (2019). Approximate causal abstractions. In Proceedings of The 35th Uncertainty in Artificial Intelligence Conference
2019
-
[4]
and Halpern, J
Beckers, S. and Halpern, J. (2019). Abstracting causal models. In AAAI Conference on Artificial Intelligence
2019
-
[5]
Block, N. (1986). Advertisement for a semantics for psychology. Midwest Studies in Philosophy , 10:615–678
1986
-
[6]
Chalmers, D. (1996). Does a rock implement every finite-state automaton? Synthese , 108:310--333
1996
-
[7]
Chalmers, D. (2011a). A computational foundation for the study of cognition. Journal of Cognitive Science , 12(4):323--357
-
[8]
Chalmers, D. (2011b). The varieties of computation: A reply. Journal of Cognitive Science , 13(3):211--248
Show all 120 references
-
[9]
Chalupka, K., Eberhardt, F., and Perona, P. (2017). Causal feature learning: an overview. Behaviormetrika , 44:137–164
2017
-
[10]
Chan, L., Garriga-Alonso, A., Goldwosky-Dill, N., Greenblatt, R., Nitishinskaya, J., Radhakrishnan, A., Shlegeris, B., and Thomas, N. (2022). Causal scrubbing, a method for rigorously testing interpretability hypotheses. AI Alignment Forum . https://www.alignmentforum.org/post...
2022
-
[11]
and Gentner, D
Christie, S. and Gentner, D. (2014). Language helps children succeed on a classic analogy task. Cognitive Science , 38(2):383--397
2014
-
[12]
Conant, R. C. and Ashby, W. R. (1970). Every good regulator of a system must be a model of that system. International Journal of Systems Science , 1(2):89--97
1970
-
[13]
D., and Geiger, A
Csord \'a s, R., Potts, C., Manning, C. D., and Geiger, A. (2024). Recurrent neural networks learn to store and generate sequences using non-linear representations. In Belinkov, Y., Kim, N., Jumelet, J., Mohebbi, H., Mueller, A., and Chen, H., editors, Proceedings of the 7th B...
2024
-
[14]
Cummins, R. (1975). Functional analysis. Journal of Philosophy , 72:741--765
1975
-
[15]
Cummins, R. (1977). Programs in the explanation of behavior. Philosophy of Science , 44(2):269--287
1977
-
[16]
Dai, Q., Heinzerling, B., and Inui, K. (2024). Representational analysis of binding in language models. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N., editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages 17468--17493, Miami, ...
2024
-
[17]
R., and Bau, D
Davies, X., Nadeau, M., Prakash, N., Shaham, T. R., and Bau, D. (2023). Discovering variable binding circuitry with desiderata. arXiv preprint arXiv:2307.03637
2023 arXiv
-
[18]
K., Aitchison, M., Orseau, L., Hutter, M., and Veness, J
Deletang, G., Ruoss, A., Duquenne, P.-A., Catt, E., Genewein, T., Mattern, C., Grau-Moya, J., Wenliang, L. K., Aitchison, M., Orseau, L., Hutter, M., and Veness, J. (2024). Language modeling is compression. In The Twelfth International Conference on Learning Representations
2024
-
[19]
Dennett, D. C. (1978). Toward a cognitive theory of consciousness. In Savage, C. W., editor, Minnesota Studies in the Philosophy of Science, Volume 9: Perception and Cognition, Issues in the Foundation of Psychology , pages 201--228. University of Minnesota Press
1978
-
[20]
Douglas, H. E. (2009). Reintroducing prediction to explanation. Philosophy of Science , 76(4):444--463
2009
-
[21]
Dretske, F. (1988). Explaining Behavior: Reasons in a World of Causes . MIT Press
1988
-
[22]
Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., Grosse, R., McCandlish, S., Kaplan, J., Amodei, D., Wattenberg, M., and Olah, C. (2022). Toy models of superposition. Transformer Circuits Thread
2022
-
[23]
J., Liao, I., Gurnee, W., and Tegmark, M
Engels, J., Michaud, E. J., Liao, I., Gurnee, W., and Tegmark, M. (2025). Not all language model features are one-dimensionally linear. In The Thirteenth International Conference on Learning Representations
2025
-
[24]
Feldman, J. (2016). The simplicity principle in perception and cognition. WIREs Cognitive Science , 7(5):330--340
2016
-
[25]
and Steinhardt, J
Feng, J. and Steinhardt, J. (2024). How do language models bind entities in context? In The Twelfth International Conference on Learning Representations
2024
-
[26]
Finlayson, M., Mueller, A., Gehrmann, S., Shieber, S., Linzen, T., and Belinkov, Y. (2021). Causal analysis of syntactic agreement mechanisms in neural language models. In Zong, C., Xia, F., Li, W., and Navigli, R., editors, Proceedings of the 59th Annual Meeting of the Associ...
2021
-
[27]
Fodor, J. A. (1975). The Language of Thought . Harvard University Press
1975
-
[28]
Frank, M. C. and Goodman, N. D. (2025). Cognitive modeling using artificial intelligence. Annual Review of Psychology . Forthcoming
2025
-
[29]
C., and Potts, C
Geiger, A., Carstensen, A., Frank, M. C., and Potts, C. (2023). Relational reasoning and generalization using nonsymbolic neural networks. Psychological Review , 130(2):308--333
2023
-
[30]
Geiger, A., Ibeling, D., Zur, A., Chaudhary, M., Chauhan, S., Huang, J., Arora, A., Wu, Z., Goodman, N., Potts, C., and Icard, T. (2025). Causal abstraction: A theoretical foundation for mechanistic interpretability. Journal of Machine Learning Research , 26(83):1--64
2025
-
[31]
F., and Potts, C
Geiger, A., Lu, H., Icard, T. F., and Potts, C. (2021). Causal abstractions of neural networks. In Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W., editors, Advances in Neural Information Processing Systems
2021
-
[32]
Geiger, A., Richardson, K., and Potts, C. (2020). Neural natural language inference models partially embed theories of lexical entailment and negation. In Alishahi, A., Belinkov, Y., Chrupa a, G., Hupkes, D., Pinter, Y., and Sajjad, H., editors, Proceedings of the Third Blackb...
2020
-
[33]
Geiger, A., Wu, Z., Potts, C., Icard, T., and Goodman, N. D. (2024). Finding alignments between interpretable causal variables and distributed neural representations. In Proceedings of the 3rd Conference on Causal Learning and Reasoning (CLeaR) , pages 160--187
2024
-
[34]
Geva, M., Bastings, J., Filippova, K., and Globerson, A. (2023). Dissecting recall of factual associations in auto-regressive language models. In The 2023 Conference on Empirical Methods in Natural Language Processing
2023
-
[35]
Godfrey-Smith, P. (2009). Triviality arguments against functionalism. Philosophical Studies , 145(2):273–295
2009
-
[36]
Guerner, C., Svete, A., Liu, T., Warstadt, A., and Cotterell, R. (2023). A geometric notion of causal probing. CoRR , abs/2307.15054
2023 arXiv
-
[37]
Hanna, M., Liu, O., and Variengien, A. (2023). How does GPT -2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model. In Thirty-seventh Conference on Neural Information Processing Systems
2023
-
[38]
Harding, J. (2023). Operationalising representation in natural language processing. The British Journal for the Philosophy of Science
2023
-
[39]
Harding, J., Gerstenberg, T., and Icard, T. (2025). A communication-first account of explanation
2025
-
[40]
and Sharadin, N
Harding, J. and Sharadin, N. (2025). What is it for a machine learning model to have a capability? British Journal for the Philosophy of Science
2025
-
[41]
Hempel, C. G. and Oppenheim, P. (1948). Studies in the Logic of Explanation . Philosophy of Science , 15(2):135--175
1948
-
[42]
E., McClelland, J
Hinton, G. E., McClelland, J. L., and Rumelhart, D. E. (1986). Distributed representations. In Rumelhart, D. E., McClelland, J. L., and the PDP Research Group, editors, Parallel Distributed Processing: Explorations in the Microstructure of Cognition: Psychological and Biologic...
1986
-
[43]
Holland, P. W. (1986). Statistics and causal inference. Journal of the American Statistical Association , 81(396):945--960
1986
-
[44]
Huang, J., Tao, J., Icard, T., Yang, D., and Potts, C. (2025). Internal causal mechanisms robustly predict language model out-of-distribution behaviors. In Forty-second International Conference on Machine Learning
2025
-
[45]
Huang, J., Wu, Z., Potts, C., Geva, M., and Geiger, A. (2024). RAVEL : Evaluating interpretability methods on disentangling language model representations. In Ku, L.-W., Martins, A., and Srikumar, V., editors, Proceedings of the 62nd Annual Meeting of the Association for Compu...
2024
-
[46]
Hupkes, D., Giulianelli, M., Dankers, V., Artetxe, M., Elazar, Y., Pimentel, T., Christodoulopoulos, C., Lasri, K., Saphra, N., Sinclair, A., Ulmer, D., Schottmann, F., Batsuren, K., Sun, K., Sinha, K., Khalatbari, L., Ryskina, M., Frieske, R., Cotterell, R., and Jin, Z. (2023...
2023
-
[47]
and Icard, T
Ibeling, D. and Icard, T. F. (2019). On open-universe causal reasoning. In Proceedings of the 35th Conference on Uncertainty in Artificial Intelligence (UAI)
2019
-
[48]
Icard, T. F. (2017). From programs to causal models. In Proceedings of the 21st Amsterdam Colloquium
2017
-
[49]
A., Schrimpf, M., Anzellotti, S., Zaslavsky, N., Fedorenko, E., and Isik, L
Ivanova, A. A., Schrimpf, M., Anzellotti, S., Zaslavsky, N., Fedorenko, E., and Isik, L. (2022). Beyond linear regression: Mapping models in cognitive neuroscience should align with research goals. Neurons, Behavior, Data analysis, and Theory , 1
2022
-
[50]
James, W. (1890). Principles of Psychology , volume 1. New York: Holt
-
[51]
and Mejia, S
Janzing, D. and Mejia, S. H. G. (2022). Phenomenological causality
2022
-
[52]
and Liang, P
Jia, R. and Liang, P. (2017). Adversarial examples for evaluating reading comprehension systems. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , pages 2021--2031
2017
-
[53]
and Kording, K
Jonas, E. and Kording, K. P. (2017). Could a neuroscientist understand a microprocessor? PloS Computational Biology , 13(1)
2017
-
[54]
Kiela, D., Bartolo, M., Nie, Y., Kaushik, D., Geiger, A., Wu, Z., Vidgen, B., Prasad, G., Singh, A., Ringshia, P., Ma, Z., Thrush, T., Riedel, S., Waseem, Z., Stenetorp, P., Jia, R., Bansal, M., Potts, C., and Williams, A. (2021). Dynabench: Rethinking benchmarking in NLP . In...
2021
-
[55]
Kriegeskorte, N. (2011). Pattern-information analysis: From stimulus decoding to computational-model testing. NeuroImage , 56(2):411--421
2011
-
[56]
Lashley, K. S. (1929/1948). Brain mechanisms and intelligence. In Dennis, W., editor, Readings in the History of Psychology , page 557–570. Appleton-Century-Crofts
1929
-
[57]
K., Bau, D., Vi \'e gas, F., Pfister, H., and Wattenberg, M
Li, K., Hopkins, A. K., Bau, D., Vi \'e gas, F., Pfister, H., and Wattenberg, M. (2023). Emergent world representations: Exploring a sequence model trained on a synthetic task. In The Eleventh International Conference on Learning Representations
2023
-
[58]
Longino, H. (2002). The Fate of Knowledge . Princeton University Press
2002
-
[59]
Marcus, G. (2001). The Algebraic Mind: Integrating Connectionism and Cognitive Science . MIT Press
2001
-
[60]
and Tegmark, M
Marks, S. and Tegmark, M. (2024). The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. In First Conference on Language Modeling
2024
-
[61]
Marr, D. (1982). Vision . W.H. Freeman and Company
1982
-
[62]
Meng, K., Bau, D., Andonian, A., and Belinkov, Y. (2022). Locating and editing factual associations in gpt. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A., editors, Advances in Neural Information Processing Systems , volume 35, pages 17359--17372. C...
2022
-
[63]
S., and Dean, J
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. (2013). Distributed representations of words and phrases and their compositionality. In Burges, C., Bottou, L., Welling, M., Ghahramani, Z., and Weinberger, K., editors, Advances in Neural Information Processin...
2013
-
[64]
Mingard, C., Rees, H., Valle-P\' e rez, G., and Louis, A. A. (2025). Deep neural networks have an inbuilt O ccam’s razor. Nature Communications , 16(220):1--9
2025
-
[65]
Moschovakis, Y. N. (2001). What is an algorithm? In Engquist, B. and Schmid, W., editors, Mathematics Unlimited — 2001 and Beyond , page 919–936. Springer
2001
-
[66]
S., Sun, J., Todd, E., Bau, D., and Belinkov, Y
Mueller, A., Brinkmann, J., Li, M., Marks, S., Pal, K., Prakash, N., Rager, C., Sankaranarayanan, A., Sharma, A. S., Sun, J., Todd, E., Bau, D., and Belinkov, Y. (2024). The quest for the right mediator: A history, survey, and theoretical grounding of causal interpretability
2024
-
[67]
S., Fiotto-Kaufman, J
Mueller, A., Geiger, A., Wiegreffe, S., Arad, D., Arcuschin, I., Belfki, A., Chan, Y. S., Fiotto-Kaufman, J. F., Haklay, T., Hanna, M., Huang, J., Gupta, R., Nikankin, Y., Orgad, H., Prakash, N., Reusch, A., Sankaranarayanan, A., Shao, S., Stolfo, A., Tutek, M., Zur, A., Bau, ...
2025
-
[68]
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J. (2023a). Progress measures for grokking via mechanistic interpretability. In The Eleventh International Conference on Learning Representations
-
[69]
Nanda, N., Lee, A., and Wattenberg, M. (2023b). Emergent linear representations in world models of self-supervised sequence models. In Belinkov, Y., Hao, S., Jumelet, J., Kim, N., McCarthy, A., and Mohebbi, H., editors, Proceedings of the 6th BlackboxNLP Workshop: Analyzing an...
2023
-
[70]
Neander, K. (2017). A Mark of the Mental: A Defence of Informational Teleosemantics . MIT Press, Cambridge, MA
2017
-
[71]
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S. (2020). Zoom in: An introduction to circuits. Distill . https://distill.pub/2020/circuits/zoom-in
2020
-
[72]
J., and Veitch, V
Park, K., Choe, Y. J., and Veitch, V. (2023). The linear representation hypothesis and the geometry of large language models. CoRR , abs/2311.03658
2023 arXiv
-
[73]
Pearl, J. (2001). Direct and indirect effects. In Breese, J. S. and Koller, D., editors, UAI '01: Proceedings of the 17th Conference in Uncertainty in Artificial Intelligence, University of Washington, Seattle, Washington, USA, August 2-5, 2001 , pages 411--420. Morgan Kaufmann
2001
-
[74]
Pearl, J. (2009). Causality . Cambridge University Press
2009
-
[75]
Peters, J., Janzing, D., and Schölkopf, B. (2017). Elements of Causal Inference: Foundations and Learning Algorithms . The MIT Press, Cambridge, MA
2017
-
[76]
Piantadosi, S. T. and Gallistel, C. R. (2024). Formalising the role of behaviour in neuroscience. European Journal of Neuroscience , 60(5):4756–4770
2024
-
[77]
Piantadosi, S. T. and Hill, F. (2022). Meaning without reference in large language models
2022
-
[78]
Piccinini, G. (2015). Physical Computation: A Mechanistic Account . Oxford University Press
2015
-
[79]
Potochnik, A. (2017). Idealization and the Aims of Science . University of Chicago Press , Chicago
2017
-
[80]
S., Riedl, C., Belinkov, Y., Shaham, T
Prakash, N., Shapira, N., Sharma, A. S., Riedl, C., Belinkov, Y., Shaham, T. R., Bau, D., and Geiger, A. (2025). Language models use lookbacks to track beliefs
2025
-
[81]
Premack, D. (1983). The codes of man and beasts. Behavioral and Brain Sciences , 6(1):125--136
1983
-
[82]
Putnam, H. (1967). Psychological predicates. In Capitan, W. H. and Merrill, D. D., editors, Art, Mind, and Religion . Pittsburgh University Press
1967
-
[83]
Putnam, H. (1975). Philosophy and our mental life. In Mind, Language, and Reality . Cambridge University Press
1975
-
[84]
Putnam, H. (1988). Representation and Reality . MIT Press
1988
-
[85]
Pylyshyn, Z. W. (1984). Computation and Cognition . MIT Press
1984
-
[86]
Rescorla, M. (2013). Against structuralist theories of computational implementation. British Journal for the Philosophy of Science , 64(4):681--707
2013
-
[87]
Rescorla, M. (2015). The representational foundations of computation. Philosophia Mathematica , 23(3):338--366
2015
-
[88]
and Everitt, T
Richens, J. and Everitt, T. (2024). Robust agents learn causal world models. In Kim, B., Yue, Y., Chaudhuri, S., Fragkiadaki, K., Khan, M., and Sun, Y., editors, International Conference on Representation Learning , volume 2024, pages 15786--15817
2024
-
[89]
D., Mueller, A., and Misra, K
Rodriguez, J. D., Mueller, A., and Misra, K. (2025). Characterizing the role of similarity in the property inferences of language models. In Chiruzzo, L., Ritter, A., and Wang, L., editors, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Associ...
2025
-
[90]
K., Weichwald, S., Bongers, S., Mooij, J
Rubenstein, P. K., Weichwald, S., Bongers, S., Mooij, J. M., Janzing, D., Grosse-Wentrup, M., and Schölkopf, B. (2017). Causal consistency of structural equation models. In Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence (UAI)
2017
-
[91]
and Lappi, O
Rusanen, A.-M. and Lappi, O. (2016). On computational explanations. Synthese , 193:3931–3949
2016
-
[92]
and Wiegreffe, S
Saphra, N. and Wiegreffe, S. (2024). Mechanistic? In Belinkov, Y., Kim, N., Jumelet, J., Mohebbi, H., Mueller, A., and Chen, H., editors, Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP , pages 480--498, Miami, Florida, US. Assoc...
2024
-
[93]
S., Benton, J., and Shlegeris, B
Scherlis, A., Sachan, K., Jermyn, A. S., Benton, J., and Shlegeris, B. (2022). Polysemanticity and capacity in neural networks. CoRR , abs/2210.01892
2022 arXiv
-
[94]
Scheutz, M. (2001). Computational vs. causal complexity. Minds and Machines , 11(4):543--566
2001
-
[95]
Searle, J. R. (1992). The Rediscovery of the Mind . MIT Press
1992
-
[96]
Shagrir, O. (2001). Content, computation and externalism. Mind , 110(438):369--400
2001
-
[97]
Shea, N. (2007). Content and its vehicles in connectionist systems. Mind & Language , 22:246--269
2007
-
[98]
Shea, N. (2018). Representation in Cognitive Science . Oxford University Press
2018
-
[99]
Simon, H. A. and Ando, A. (1961). Aggregation of variables in dynamic systems. Econometrica , 29(2):111--138
1961
-
[100]
Smolensky, P. (1986). Neural and conceptual interpretation of PDP models. In McClelland, J. L., Rumelhart, D. E., and the PDP Research Group, editors, Parallel Distributed Processing: Explorations in the Microstructure of Cognition: Psychological and Biological Models , volume...
1986
-
[101]
Spirtes, P., Glymour, C., and Scheines, R. (2000). Causation, Prediction, and Search . MIT Press
2000
-
[102]
Sprevak, M. (2010). Computation, individuation, and the received view on representation. Studies in History and Philosophy of Science , 41:260–270
2010
-
[103]
Sprevak, M. (2018). Triviality arguments about computational implementation. In Sprevak, M. and Colombo, M., editors, The Routledge Handbook of the Computational Mind , pages 175--191. Routledge
2018
-
[104]
Sutter, D., Minder, J., Hofmann, T., and Pimentel, T. (2025). The non-linear representation dilemma: Is causal abstraction enough for mechanistic interpretability?
2025
-
[105]
S., Cao, R., and Yamins, D
Thobani, I., Sagastuy-Brena, J., Nayebi, A., Prince, J. S., Cao, R., and Yamins, D. L. (2025). Model-brain comparison using inter-animal transforms. In 8th Annual Conference on Cognitive Computational Neuroscience
2025
-
[106]
Thompson, D. (2023). Algorithms and Execution Traces . PhD thesis, Stanford University
2023
-
[107]
K., Rattermann, M
Thompson, R. K., Rattermann, M. J., and Oden, D. L. (2001). Perception and judgement of abstract same-different relations by monkeys, apes and children: Do symbols make explicit only that which is implicit? Croatian Review of Rehabilitation Research , 37(1):9--22
2001
-
[108]
J., Geiger, A., and Nanda, N
Tigges, C., Hollinsworth, O. J., Geiger, A., and Nanda, N. (2024). Language models linearly represent sentiment. In Belinkov, Y., Kim, N., Jumelet, J., Mohebbi, H., Mueller, A., and Chen, H., editors, Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neur...
2024
-
[109]
S., Mueller, A., Wallace, B
Todd, E., Li, M., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D. (2024). Function vectors in large language models. In The Twelfth International Conference on Learning Representations
2024
-
[110]
Turing, A. M. (1936). On computable numbers, with an application to the E ntscheidungsproblem. Proceedings of the London Mathematical Society , s2-42:230–265
1936
-
[111]
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S. (2020). Causal mediation analysis for interpreting neural NLP : The case of gender bias
2020
-
[112]
R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J
Wang, K. R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J. (2023). Interpretability in the wild: a circuit for indirect object identification in GPT -2 small. In The Eleventh International Conference on Learning Representations
2023
-
[113]
Wasserman, E. A. and Young, M. E. (2010). Same-different discrimination: the keel and backbone of thought and reasoning. Journal of Experimental Psychology: Animal Behavior Processes , 36(1):3--22
2010
-
[114]
Wendong, L., Buchholz, S., and Sch\" o lkopf, B. (2025). Algorithmic causal structure emerging through compression. In Huang, B. and Drton, M., editors, Proceedings of the Fourth Conference on Causal Learning and Reasoning , pages 201--242
2025
-
[115]
Woodward, J. (2003). Making Things Happen: A Theory of Causal Explanation . Oxford University Press
2003
-
[116]
Wu, Z., Geiger, A., Icard, T., Potts, C., and Goodman, N. (2023). Interpretability at scale: Identifying causal mechanisms in A lpaca. In Thirty-seventh Conference on Neural Information Processing Systems
2023
-
[117]
and Kuorikoski, J
Ylikoski, P. and Kuorikoski, J. (2010). Dissecting explanatory power. Philosophical Studies , 148(2):201--219
2010
-
[118]
Zeller, A. (2002). Isolating cause-effect chains from computer programs. In Proceedings of the 10th ACM SIGSOFT Symposium on Foundations of Software Engineering , SIGSOFT '02/FSE-10, page 1–10. Association for Computing Machinery
2002
-
[119]
Zennaro, F. M. (2022). Abstraction between structural causal models: A review of definitions and properties. In UAI 2022 Workshop on Causal Representation Learning
2022
-
[120]
and Nanda, N
Zhang, F. and Nanda, N. (2024). Towards best practices of activation patching in language models: Metrics and methods. In The Twelfth International Conference on Learning Representations
2024
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.