Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

How Causal Abstraction Underpins Computational Explanation

T0 review · 3 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read A system implements a computation only if the algorithm is a causal abstraction of it.

desk verdict A clear, honest programmatic paper connecting causal abstraction to implementation, but the sufficiency claim rests on an unformalized 'native vehicles' constraint and the necessary condition is nearly vacuous under arbitrary translations. read the letter →

arxiv 2508.11214 v2 pith:TJIXTSOR submitted 2025-08-15 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords causalabstractioncomputationalimplementationmechanisticinterpretabilityinterventionalgebratrivialityargumentrepresentationneuralnetworksgeneralization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to sharpen the old idea that a physical system implements an algorithm when its causal structure mirrors the algorithm's causal structure. It proposes a precise necessary condition: a low-level system implements a computational model only if the model is an abstraction-under-translation of the system, meaning the system can be recarved by a bijective translation and then partitioned into macro-variables so that interventions line up exactly. If this is right, the language of causal abstraction gives cognitive science and mechanistic interpretability a common criterion for when a neural network genuinely runs a hypothesized algorithm. The paper also confronts a triviality result: under its own permissive translations, almost any network can be made to implement almost any algorithm, so it argues that genuine explanation and prediction require additional restrictions on admissible translations, such as vehicles native to the system.

What carries the argument

The machinery is the abstraction-under-translation relation built from three ingredients: exact transformation, constructive abstraction, and translation. An exact transformation is a pair of maps $(\tau,\omega)$ from low-level values and interventions to high-level values and interventions such that running the low-level model, intervening, and translating always matches the high-level model's run under the corresponding intervention. A constructive abstraction ignores distinctions by partitioning low-level variables into macro-variables. A translation is a bijective recarving $\tau$ of the variable space, which induces a canonical high-level model and an intervention algebra $I_\tau$ of pu

What would settle it

Take a trained network that solves the hierarchical equality task and compute the gerrymandered translation $\tau$ witnessing the XNOR circuit as an abstraction-under-translation. Then test the network on novel stimulus pairs, such as arrows instead of faces, and check whether its outputs track the circuit's algorithm; a divergence would falsify the sufficiency claim. Conversely, if researchers agree a system implements a computation but no abstraction-under-translation can be exhibited, the necessity claim is falsified.

Watch

Extended reading notes

Core claim

The central claim is the 'No Computation without Abstraction' Principle: given a computational model $H$, a system $L$ implements $H$ only if $H$ is an abstraction-under-translation of $L$. Abstraction-under-translation is a two-step relation: translate $L$ by a bijective recarving $\tau$ of its variable space (with induced interventionals $I_\tau$), then take a constructive abstraction of $\tau(L)$ by partitioning low-level variables into macro-variables and checking that the induced maps form an exact transformation, so that low-level interventions reproduce high-level interventions. The authors illustrate the principle with a neural network trained on the hierarchical equality task whose

Load-bearing premise

The account assumes that any bijective recarving of the low-level system counts as a legitimate translation, which lets a known triviality theorem show that almost every network can be made to abstract every algorithm; the proposed necessary condition is then nearly vacuous unless the extra constraints mentioned in the paper are actually formalized.

Editorial extensions

If this is right

  • If the principle is right, claims that a neural network implements a cognitive algorithm should be backed by an explicit abstraction-under-translation, not just by matching input-output behavior.
  • Distributed representations in trained networks may require a translation such as a rotation before the algorithm's causal variables become visible, giving a causal rationale for linear representation hypotheses.
  • The triviality theorem becomes a positive internal-model result: whenever a network solves a task, there exist modular low-level interventions that realize the algorithm, but those interventions may be too complex to find or use.
  • Representational vehicles in implemented computations need not be individual neurons or regions; they can be abstract operations on the system, so content attaches to interventionals rather than to states alone.
  • Prediction and generalization beyond observed inputs will require restricting admissible translations, for instance to mappings the system itself can natively support.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves the key constraint on translations informal: if 'native' vehicles could be formalized, the triviality theorem's gerrymandered translations would be excluded and the necessity claim would regain bite.
  • A testable extension is to impose linearity on $\tau$ and check whether the recent triviality construction still goes through; if it does, linearity alone is too weak to ground implementation.
  • The account suggests a link between implementation and compression: algorithmic causal variables that support generalization may be exactly those that yield compact descriptions of behavior, which could turn 'native vehicles' into a measurable quantity.
  • The treatment of representational content implies that content is not an intrinsic property of a neural population, but a product of the intervention toolbox a scientist is willing to use; different toolboxes will yield different representational claims.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper argues that computational implementation should be understood through the lens of causal abstraction. It reviews and uses formal notions from previous work—exact transformation, constructive abstraction, and abstraction-under-translation—and states a 'No Computation without Abstraction' principle: a system L implements a computational model H only if H is an abstraction-under-translation of L. The paper illustrates the framework with a neural network for the hierarchical equality task, discusses how representational content can be assigned to the abstract variables via causal role, and then confronts a triviality result by Sutter et al. (2025), which shows that under minimal assumptions every sufficiently large network with the right input/output behavior is an abstraction-under-translation of every algorithm. The paper interprets this result partly as a positive modular-control finding, but then argues that explanatory goals—especially generalization and prediction—require further constraints on admissible translations and vehicles, such as 'native' vehicles and linearity. The final sections present the account as a promising framework rather than a completed theory, explicitly leaving the admissibility problem open.

Significance. If the framework succeeds, it offers a precise causal formalization of a central concept in the philosophy of computation and connects long-standing debates about computational implementation, triviality, and representation to contemporary mechanistic interpretability. The paper's strengths are its precise definitions, the concrete neural-network case study, and its unusually honest treatment of the triviality problem: rather than ignoring Sutter et al., it engages directly and explores what a positive reading would require. It is also valuable in separating the purely structural notion of abstraction-under-translation from broader explanatory demands. The main limitation is that the proposed remedy to triviality—restricting to 'native' vehicles and generalization-supporting mappings—is not formalized, so the central positive claim remains programmatic. The necessary condition is defensible, but it is weak in light of the triviality theorem, and the sufficiency inclination in Section 7 is not a claim about the defined formalism as it stands.

major comments (3)
  1. [§7, with reference to §3.2 and §6] The paper states: 'we are inclined to accept abstraction-under-implementation not only as a necessary condition for computational implementation, but also as a sufficient one.' But the formal notion from §3.2 allows arbitrary bijective translations τ and their associated interventionals Iτ. As the paper itself reports in §6, Sutter et al. (2025) show that under minimal assumptions every sufficiently large network with the right task behavior is an abstraction-under-translation of every algorithm. Taken literally, the sufficiency claim would therefore assert that all such networks implement all algorithms. The response that vehicles must be 'native to how the system works' is not formalized: no constraint on τ or Iτ is added to the definition in §3.2, and no criterion for 'native' is given. The sufficiency claim is thus not a consequence of the formal framework; it is a gesture toward an
  2. [§6] The triviality theorem is treated as a positive finding about modular manipulation and control, but it undercuts the discriminatory content of the 'No Computation without Abstraction' principle. If abstraction-under-translation is satisfied by every sufficiently large network for every algorithm, then the necessary condition cannot distinguish genuine implementations from gerrymandered ones. The observation that Sutter et al. found the interventionals difficult to extract is epistemic and do not constitute a constitutive constraint on implementation; it says nothing about which translations are admissible. The paper needs to make a choice: either (a) concede that abstraction-under-translation is merely a necessary condition and that the full implementation concept requires additional, not-yet-specified conditions, or (b) prove that a restricted class of translations avoids the triviality
  3. [§2 and §7] The account depends on the substantive claim that computational models can and must be construed causally. The paper explicitly says this is 'not defended here' and that it is a stance adopted. Since the paper is making a philosophical claim about computational explanation, this is a load-bearing assumption: if the causal construal of computation is rejected, the entire abstraction-under-translation framework ceases to be an account of computational implementation. At minimum, the paper should provide a brief defense or a more precise scope condition stating that the argument is conditional on the causal construal. As written, the independence of the main principle from this assumption is unclear, and the reader is asked to accept a foundational premise without support.
minor comments (6)
  1. [§3.2] Typo: 'another casual model' should be 'another causal model'.
  2. [§7] The phrase 'each of these lays bear' is ungrammatical; it should be 'these lay bare' or 'this lays bare'.
  3. [§7] The term 'abstraction-under-implementation' is introduced without a definition and appears to be used synonymously with 'abstraction-under-translation.' Define the term or use a single consistent label.
  4. [§3.1] The demonstration that the XNOR circuit M is a constructive abstraction of the network N is very compressed. A more explicit step-by-step verification, or a precise pointer to the corresponding example in Geiger et al. (2025), would help the reader check the central example.
  5. [§2] The definitions of causal model, intervention, intervention algebra, and exact transformation are informal and not numbered. Since they carry the weight of the paper, a formal appendix with numbered definitions and statements of the relevant theorems from Geiger et al. (2025) would substantially improve usability.
  6. [References] The reference for Sutter et al. (2025) lacks a venue, arXiv identifier, or other locator, making it difficult to verify the statement of the triviality theorem cited in §6.

Circularity Check

1 steps flagged · score 4.0 of 10

Necessary condition is explicitly stipulated as a definition; central discussion remains independent, so partial circularity.

  1. self definitional [§3.2, Abstraction-Under-Translation definition and 'No Computation without Abstraction' Principle]
    "With this much we can formulate what we take to be a necessary condition on any claim of computational implementation: “No Computation without Abstraction” Principle: Given a computational model H, another system L implements H only if H is an abstraction-under-translation of L. In other words, we claim that to implement a computation, the (causal construal of that) computation must be a constructive abstraction of a translation of the system in question."

    The principle is stated immediately after the paper itself defines 'abstraction-under-translation' as 'H is an abstraction-under-translation of L if there is a translation τpLq of L such that H is a constructive abstraction of τpLq.' The necessary condition is thus an analytic consequence of that stipulated definition plus the additional stipulation that implementation entails it; it is not derived from independent premises about computation. The surrounding text presents it as capturing 'the spirit of claims in the literature,' i.e., as a characterization rather than a derived result. This is definitional circularity, though the paper is transparent about the stipulation.

full rationale

The only load-bearing step with a circular flavor is the 'No Computation without Abstraction' principle in §3.2: it is an immediate reformulation of the paper's own definition of abstraction-under-translation, prefaced by 'we take to be a necessary condition.' No independent argument establishes why implementation must coincide with this particular relation; the definition was expressly designed to capture prior implementation talk. This is a mild self-definitional move. However, the paper does not hide it, and the rest of the paper contains substantial independent content: the constructive-abstraction and translation examples are worked out, the engagement with Sutter et al.'s external triviality theorem is genuine, and the §7 discussion of generalization and 'native vehicles' is a substantive open problem rather than a circular derivation. The self-citations to Geiger et al. (2025) supply formal definitions and theorems that are checkable, not a uniqueness proof invoked to forbid alternatives. No parameter is fitted and no prediction is generated from data, so patterns 2, 5, and 6 do not apply. The unformalized 'native vehicles' constraint is a gap in the sufficiency argument, not a circularity.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The framework rests on the assumption that computations are causal models and on the definition of abstraction-under-translation from the authors' prior work. There are no fitted parameters, but the choices of translation, partition, and admissible interventions are free parameters that determine whether an implementation claim holds.

free parameters (3)
  • translation τ
    The paper allows any bijective map from the low-level variable space; the choice of τ determines which interventionals are admissible and drives the triviality result (§3.2, §6).
  • partition Π and component maps π for constructive abstraction
    The grouping of low-level variables into high-level variables is chosen to make the abstraction exact (§3.1).
  • admissible interventionals Iτ
    The set of operations witnessing the abstraction is not constrained except by algebraic requirements, which is what permits gerrymandered implementations (§6).
assumptions (4)
  • domain assumption Computational models can be construed as causal models.
    Stated in §2: 'we do, however, insist that a computational model can be construed in causal terms.' This is the load-bearing premise for the whole framework.
  • standard math Causal models are acyclic functional models with mechanism replacement interventions.
    Assumed in §2, standard for structural causal models.
  • ad hoc to paper The interventionals IL and IH form intervention algebras.
    Assumed in §3. The algebra condition is needed for the exact transformation definition and for compatibility of sets of interventions.
  • domain assumption Representational content is captured by the three criteria: Information, Use, Misrepresentation.
    Appealed to in §5, following Harding (2023). This is a substantive philosophical assumption about representation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of How Causal Abstraction Underpins Computational Explanation." pith.science (2026). https://pith.science/paper/TJIXTSOR

@misc{pith2026250811214,
  author       = {Pith},
  title        = {Pith review of: How Causal Abstraction Underpins Computational Explanation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TJIXTSOR}},
  note         = {Machine review of arXiv:2508.11214}
}
read the original abstract

Explanations of cognitive behavior often appeal to computations over representations. What does it take for a system to implement a given computation over suitable representational vehicles within that system? We argue that the language of causality -- and specifically the theory of causal abstraction -- provides a fruitful lens on this topic. Drawing on current discussions in deep learning with artificial neural networks, we illustrate how classical themes in the philosophy of computation and cognition resurface in contemporary machine learning. We offer an account of computational implementation grounded in causal abstraction, and examine the role for representation in the resulting picture. We argue that these issues are most profitably explored in connection with generalization and prediction.

Figures

Figures reproduced from arXiv: 2508.11214 by the authors.

Figure 1
Figure 1. A simple circuit construed as a causal model M. The arrows denote functional dependence. For instance, there is an arrow from A4 to B2 because a change to A4 can (in some context) bring about a change to B2. 3The reader may consult Geiger et al. (2025) for many more details on the following. See also Pearl (2009); Peters et al. (2017) for more general treatments of causal models, especially as they feature in causal… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context

    cs.CL 2025-10 conditional novelty 7.0 of 10

    In-context entity retrieval in LMs is a mixture of positional, lexical, and reflexive mechanisms; the pure positional view fails in middle positions of long lists.

Reference graph

Works this paper leans on

120 extracted references · 75 canonical work pages · cited by 1 Pith paper

  1. [1]

    Alhama, R. G. and Zuidema, W. (2019). A review of computational models of basic rule learning: The neural-symbolic debate and beyond. Psychonomic Bulletin & Review , 26(4):1174--1194

  2. [2]

    Arora, A., Jurafsky, D., and Potts, C. (2024). C ausal G ym: Benchmarking causal interpretability methods on linguistic tasks. In Ku, L.-W., Martins, A., and Srikumar, V., editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages 14638--14663, Bangkok, Thailand. Association for Computa...

  3. [3]

    Beckers, S., Eberhardt, F., and Halpern, J. Y. (2019). Approximate causal abstractions. In Proceedings of The 35th Uncertainty in Artificial Intelligence Conference

  4. [4]

    and Halpern, J

    Beckers, S. and Halpern, J. (2019). Abstracting causal models. In AAAI Conference on Artificial Intelligence

  5. [5]

    Block, N. (1986). Advertisement for a semantics for psychology. Midwest Studies in Philosophy , 10:615–678

  6. [6]

    Chalmers, D. (1996). Does a rock implement every finite-state automaton? Synthese , 108:310--333

  7. [7]

    Chalmers, D. (2011a). A computational foundation for the study of cognition. Journal of Cognitive Science , 12(4):323--357

  8. [8]

    Chalmers, D. (2011b). The varieties of computation: A reply. Journal of Cognitive Science , 13(3):211--248

Show all 120 references
  1. [9]

    Chalupka, K., Eberhardt, F., and Perona, P. (2017). Causal feature learning: an overview. Behaviormetrika , 44:137–164

  2. [10]

    Chan, L., Garriga-Alonso, A., Goldwosky-Dill, N., Greenblatt, R., Nitishinskaya, J., Radhakrishnan, A., Shlegeris, B., and Thomas, N. (2022). Causal scrubbing, a method for rigorously testing interpretability hypotheses. AI Alignment Forum . https://www.alignmentforum.org/post...

  3. [11]

    and Gentner, D

    Christie, S. and Gentner, D. (2014). Language helps children succeed on a classic analogy task. Cognitive Science , 38(2):383--397

  4. [12]

    Conant, R. C. and Ashby, W. R. (1970). Every good regulator of a system must be a model of that system. International Journal of Systems Science , 1(2):89--97

  5. [13]

    D., and Geiger, A

    Csord \'a s, R., Potts, C., Manning, C. D., and Geiger, A. (2024). Recurrent neural networks learn to store and generate sequences using non-linear representations. In Belinkov, Y., Kim, N., Jumelet, J., Mohebbi, H., Mueller, A., and Chen, H., editors, Proceedings of the 7th B...

  6. [14]

    Cummins, R. (1975). Functional analysis. Journal of Philosophy , 72:741--765

  7. [15]

    Cummins, R. (1977). Programs in the explanation of behavior. Philosophy of Science , 44(2):269--287

  8. [16]

    Dai, Q., Heinzerling, B., and Inui, K. (2024). Representational analysis of binding in language models. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N., editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages 17468--17493, Miami, ...

  9. [17]

    R., and Bau, D

    Davies, X., Nadeau, M., Prakash, N., Shaham, T. R., and Bau, D. (2023). Discovering variable binding circuitry with desiderata. arXiv preprint arXiv:2307.03637

  10. [18]

    K., Aitchison, M., Orseau, L., Hutter, M., and Veness, J

    Deletang, G., Ruoss, A., Duquenne, P.-A., Catt, E., Genewein, T., Mattern, C., Grau-Moya, J., Wenliang, L. K., Aitchison, M., Orseau, L., Hutter, M., and Veness, J. (2024). Language modeling is compression. In The Twelfth International Conference on Learning Representations

  11. [19]

    Dennett, D. C. (1978). Toward a cognitive theory of consciousness. In Savage, C. W., editor, Minnesota Studies in the Philosophy of Science, Volume 9: Perception and Cognition, Issues in the Foundation of Psychology , pages 201--228. University of Minnesota Press

  12. [20]

    Douglas, H. E. (2009). Reintroducing prediction to explanation. Philosophy of Science , 76(4):444--463

  13. [21]

    Dretske, F. (1988). Explaining Behavior: Reasons in a World of Causes . MIT Press

  14. [22]

    Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., Grosse, R., McCandlish, S., Kaplan, J., Amodei, D., Wattenberg, M., and Olah, C. (2022). Toy models of superposition. Transformer Circuits Thread

  15. [23]

    J., Liao, I., Gurnee, W., and Tegmark, M

    Engels, J., Michaud, E. J., Liao, I., Gurnee, W., and Tegmark, M. (2025). Not all language model features are one-dimensionally linear. In The Thirteenth International Conference on Learning Representations

  16. [24]

    Feldman, J. (2016). The simplicity principle in perception and cognition. WIREs Cognitive Science , 7(5):330--340

  17. [25]

    and Steinhardt, J

    Feng, J. and Steinhardt, J. (2024). How do language models bind entities in context? In The Twelfth International Conference on Learning Representations

  18. [26]

    Finlayson, M., Mueller, A., Gehrmann, S., Shieber, S., Linzen, T., and Belinkov, Y. (2021). Causal analysis of syntactic agreement mechanisms in neural language models. In Zong, C., Xia, F., Li, W., and Navigli, R., editors, Proceedings of the 59th Annual Meeting of the Associ...

  19. [27]

    Fodor, J. A. (1975). The Language of Thought . Harvard University Press

  20. [28]

    Frank, M. C. and Goodman, N. D. (2025). Cognitive modeling using artificial intelligence. Annual Review of Psychology . Forthcoming

  21. [29]

    C., and Potts, C

    Geiger, A., Carstensen, A., Frank, M. C., and Potts, C. (2023). Relational reasoning and generalization using nonsymbolic neural networks. Psychological Review , 130(2):308--333

  22. [30]

    Geiger, A., Ibeling, D., Zur, A., Chaudhary, M., Chauhan, S., Huang, J., Arora, A., Wu, Z., Goodman, N., Potts, C., and Icard, T. (2025). Causal abstraction: A theoretical foundation for mechanistic interpretability. Journal of Machine Learning Research , 26(83):1--64

  23. [31]

    F., and Potts, C

    Geiger, A., Lu, H., Icard, T. F., and Potts, C. (2021). Causal abstractions of neural networks. In Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W., editors, Advances in Neural Information Processing Systems

  24. [32]

    Geiger, A., Richardson, K., and Potts, C. (2020). Neural natural language inference models partially embed theories of lexical entailment and negation. In Alishahi, A., Belinkov, Y., Chrupa a, G., Hupkes, D., Pinter, Y., and Sajjad, H., editors, Proceedings of the Third Blackb...

  25. [33]

    Geiger, A., Wu, Z., Potts, C., Icard, T., and Goodman, N. D. (2024). Finding alignments between interpretable causal variables and distributed neural representations. In Proceedings of the 3rd Conference on Causal Learning and Reasoning (CLeaR) , pages 160--187

  26. [34]

    Geva, M., Bastings, J., Filippova, K., and Globerson, A. (2023). Dissecting recall of factual associations in auto-regressive language models. In The 2023 Conference on Empirical Methods in Natural Language Processing

  27. [35]

    Godfrey-Smith, P. (2009). Triviality arguments against functionalism. Philosophical Studies , 145(2):273–295

  28. [36]

    Guerner, C., Svete, A., Liu, T., Warstadt, A., and Cotterell, R. (2023). A geometric notion of causal probing. CoRR , abs/2307.15054

  29. [37]

    Hanna, M., Liu, O., and Variengien, A. (2023). How does GPT -2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model. In Thirty-seventh Conference on Neural Information Processing Systems

  30. [38]

    Harding, J. (2023). Operationalising representation in natural language processing. The British Journal for the Philosophy of Science

  31. [39]

    Harding, J., Gerstenberg, T., and Icard, T. (2025). A communication-first account of explanation

  32. [40]

    and Sharadin, N

    Harding, J. and Sharadin, N. (2025). What is it for a machine learning model to have a capability? British Journal for the Philosophy of Science

  33. [41]

    Hempel, C. G. and Oppenheim, P. (1948). Studies in the Logic of Explanation . Philosophy of Science , 15(2):135--175

  34. [42]

    E., McClelland, J

    Hinton, G. E., McClelland, J. L., and Rumelhart, D. E. (1986). Distributed representations. In Rumelhart, D. E., McClelland, J. L., and the PDP Research Group, editors, Parallel Distributed Processing: Explorations in the Microstructure of Cognition: Psychological and Biologic...

  35. [43]

    Holland, P. W. (1986). Statistics and causal inference. Journal of the American Statistical Association , 81(396):945--960

  36. [44]

    Huang, J., Tao, J., Icard, T., Yang, D., and Potts, C. (2025). Internal causal mechanisms robustly predict language model out-of-distribution behaviors. In Forty-second International Conference on Machine Learning

  37. [45]

    Huang, J., Wu, Z., Potts, C., Geva, M., and Geiger, A. (2024). RAVEL : Evaluating interpretability methods on disentangling language model representations. In Ku, L.-W., Martins, A., and Srikumar, V., editors, Proceedings of the 62nd Annual Meeting of the Association for Compu...

  38. [46]

    Hupkes, D., Giulianelli, M., Dankers, V., Artetxe, M., Elazar, Y., Pimentel, T., Christodoulopoulos, C., Lasri, K., Saphra, N., Sinclair, A., Ulmer, D., Schottmann, F., Batsuren, K., Sun, K., Sinha, K., Khalatbari, L., Ryskina, M., Frieske, R., Cotterell, R., and Jin, Z. (2023...

  39. [47]

    and Icard, T

    Ibeling, D. and Icard, T. F. (2019). On open-universe causal reasoning. In Proceedings of the 35th Conference on Uncertainty in Artificial Intelligence (UAI)

  40. [48]

    Icard, T. F. (2017). From programs to causal models. In Proceedings of the 21st Amsterdam Colloquium

  41. [49]

    A., Schrimpf, M., Anzellotti, S., Zaslavsky, N., Fedorenko, E., and Isik, L

    Ivanova, A. A., Schrimpf, M., Anzellotti, S., Zaslavsky, N., Fedorenko, E., and Isik, L. (2022). Beyond linear regression: Mapping models in cognitive neuroscience should align with research goals. Neurons, Behavior, Data analysis, and Theory , 1

  42. [50]

    James, W. (1890). Principles of Psychology , volume 1. New York: Holt

  43. [51]

    and Mejia, S

    Janzing, D. and Mejia, S. H. G. (2022). Phenomenological causality

  44. [52]

    and Liang, P

    Jia, R. and Liang, P. (2017). Adversarial examples for evaluating reading comprehension systems. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , pages 2021--2031

  45. [53]

    and Kording, K

    Jonas, E. and Kording, K. P. (2017). Could a neuroscientist understand a microprocessor? PloS Computational Biology , 13(1)

  46. [54]

    Kiela, D., Bartolo, M., Nie, Y., Kaushik, D., Geiger, A., Wu, Z., Vidgen, B., Prasad, G., Singh, A., Ringshia, P., Ma, Z., Thrush, T., Riedel, S., Waseem, Z., Stenetorp, P., Jia, R., Bansal, M., Potts, C., and Williams, A. (2021). Dynabench: Rethinking benchmarking in NLP . In...

  47. [55]

    Kriegeskorte, N. (2011). Pattern-information analysis: From stimulus decoding to computational-model testing. NeuroImage , 56(2):411--421

  48. [56]

    Lashley, K. S. (1929/1948). Brain mechanisms and intelligence. In Dennis, W., editor, Readings in the History of Psychology , page 557–570. Appleton-Century-Crofts

  49. [57]

    K., Bau, D., Vi \'e gas, F., Pfister, H., and Wattenberg, M

    Li, K., Hopkins, A. K., Bau, D., Vi \'e gas, F., Pfister, H., and Wattenberg, M. (2023). Emergent world representations: Exploring a sequence model trained on a synthetic task. In The Eleventh International Conference on Learning Representations

  50. [58]

    Longino, H. (2002). The Fate of Knowledge . Princeton University Press

  51. [59]

    Marcus, G. (2001). The Algebraic Mind: Integrating Connectionism and Cognitive Science . MIT Press

  52. [60]

    and Tegmark, M

    Marks, S. and Tegmark, M. (2024). The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. In First Conference on Language Modeling

  53. [61]

    Marr, D. (1982). Vision . W.H. Freeman and Company

  54. [62]

    Meng, K., Bau, D., Andonian, A., and Belinkov, Y. (2022). Locating and editing factual associations in gpt. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A., editors, Advances in Neural Information Processing Systems , volume 35, pages 17359--17372. C...

  55. [63]

    S., and Dean, J

    Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. (2013). Distributed representations of words and phrases and their compositionality. In Burges, C., Bottou, L., Welling, M., Ghahramani, Z., and Weinberger, K., editors, Advances in Neural Information Processin...

  56. [64]

    Mingard, C., Rees, H., Valle-P\' e rez, G., and Louis, A. A. (2025). Deep neural networks have an inbuilt O ccam’s razor. Nature Communications , 16(220):1--9

  57. [65]

    Moschovakis, Y. N. (2001). What is an algorithm? In Engquist, B. and Schmid, W., editors, Mathematics Unlimited — 2001 and Beyond , page 919–936. Springer

  58. [66]

    S., Sun, J., Todd, E., Bau, D., and Belinkov, Y

    Mueller, A., Brinkmann, J., Li, M., Marks, S., Pal, K., Prakash, N., Rager, C., Sankaranarayanan, A., Sharma, A. S., Sun, J., Todd, E., Bau, D., and Belinkov, Y. (2024). The quest for the right mediator: A history, survey, and theoretical grounding of causal interpretability

  59. [67]

    S., Fiotto-Kaufman, J

    Mueller, A., Geiger, A., Wiegreffe, S., Arad, D., Arcuschin, I., Belfki, A., Chan, Y. S., Fiotto-Kaufman, J. F., Haklay, T., Hanna, M., Huang, J., Gupta, R., Nikankin, Y., Orgad, H., Prakash, N., Reusch, A., Sankaranarayanan, A., Shao, S., Stolfo, A., Tutek, M., Zur, A., Bau, ...

  60. [68]

    Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J. (2023a). Progress measures for grokking via mechanistic interpretability. In The Eleventh International Conference on Learning Representations

  61. [69]

    Nanda, N., Lee, A., and Wattenberg, M. (2023b). Emergent linear representations in world models of self-supervised sequence models. In Belinkov, Y., Hao, S., Jumelet, J., Kim, N., McCarthy, A., and Mohebbi, H., editors, Proceedings of the 6th BlackboxNLP Workshop: Analyzing an...

  62. [70]

    Neander, K. (2017). A Mark of the Mental: A Defence of Informational Teleosemantics . MIT Press, Cambridge, MA

  63. [71]

    Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S. (2020). Zoom in: An introduction to circuits. Distill . https://distill.pub/2020/circuits/zoom-in

  64. [72]

    J., and Veitch, V

    Park, K., Choe, Y. J., and Veitch, V. (2023). The linear representation hypothesis and the geometry of large language models. CoRR , abs/2311.03658

  65. [73]

    Pearl, J. (2001). Direct and indirect effects. In Breese, J. S. and Koller, D., editors, UAI '01: Proceedings of the 17th Conference in Uncertainty in Artificial Intelligence, University of Washington, Seattle, Washington, USA, August 2-5, 2001 , pages 411--420. Morgan Kaufmann

  66. [74]

    Pearl, J. (2009). Causality . Cambridge University Press

  67. [75]

    Peters, J., Janzing, D., and Schölkopf, B. (2017). Elements of Causal Inference: Foundations and Learning Algorithms . The MIT Press, Cambridge, MA

  68. [76]

    Piantadosi, S. T. and Gallistel, C. R. (2024). Formalising the role of behaviour in neuroscience. European Journal of Neuroscience , 60(5):4756–4770

  69. [77]

    Piantadosi, S. T. and Hill, F. (2022). Meaning without reference in large language models

  70. [78]

    Piccinini, G. (2015). Physical Computation: A Mechanistic Account . Oxford University Press

  71. [79]

    Potochnik, A. (2017). Idealization and the Aims of Science . University of Chicago Press , Chicago

  72. [80]

    S., Riedl, C., Belinkov, Y., Shaham, T

    Prakash, N., Shapira, N., Sharma, A. S., Riedl, C., Belinkov, Y., Shaham, T. R., Bau, D., and Geiger, A. (2025). Language models use lookbacks to track beliefs

  73. [81]

    Premack, D. (1983). The codes of man and beasts. Behavioral and Brain Sciences , 6(1):125--136

  74. [82]

    Putnam, H. (1967). Psychological predicates. In Capitan, W. H. and Merrill, D. D., editors, Art, Mind, and Religion . Pittsburgh University Press

  75. [83]

    Putnam, H. (1975). Philosophy and our mental life. In Mind, Language, and Reality . Cambridge University Press

  76. [84]

    Putnam, H. (1988). Representation and Reality . MIT Press

  77. [85]

    Pylyshyn, Z. W. (1984). Computation and Cognition . MIT Press

  78. [86]

    Rescorla, M. (2013). Against structuralist theories of computational implementation. British Journal for the Philosophy of Science , 64(4):681--707

  79. [87]

    Rescorla, M. (2015). The representational foundations of computation. Philosophia Mathematica , 23(3):338--366

  80. [88]

    and Everitt, T

    Richens, J. and Everitt, T. (2024). Robust agents learn causal world models. In Kim, B., Yue, Y., Chaudhuri, S., Fragkiadaki, K., Khan, M., and Sun, Y., editors, International Conference on Representation Learning , volume 2024, pages 15786--15817

  81. [89]

    D., Mueller, A., and Misra, K

    Rodriguez, J. D., Mueller, A., and Misra, K. (2025). Characterizing the role of similarity in the property inferences of language models. In Chiruzzo, L., Ritter, A., and Wang, L., editors, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Associ...

  82. [90]

    K., Weichwald, S., Bongers, S., Mooij, J

    Rubenstein, P. K., Weichwald, S., Bongers, S., Mooij, J. M., Janzing, D., Grosse-Wentrup, M., and Schölkopf, B. (2017). Causal consistency of structural equation models. In Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence (UAI)

  83. [91]

    and Lappi, O

    Rusanen, A.-M. and Lappi, O. (2016). On computational explanations. Synthese , 193:3931–3949

  84. [92]

    and Wiegreffe, S

    Saphra, N. and Wiegreffe, S. (2024). Mechanistic? In Belinkov, Y., Kim, N., Jumelet, J., Mohebbi, H., Mueller, A., and Chen, H., editors, Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP , pages 480--498, Miami, Florida, US. Assoc...

  85. [93]

    S., Benton, J., and Shlegeris, B

    Scherlis, A., Sachan, K., Jermyn, A. S., Benton, J., and Shlegeris, B. (2022). Polysemanticity and capacity in neural networks. CoRR , abs/2210.01892

  86. [94]

    Scheutz, M. (2001). Computational vs. causal complexity. Minds and Machines , 11(4):543--566

  87. [95]

    Searle, J. R. (1992). The Rediscovery of the Mind . MIT Press

  88. [96]

    Shagrir, O. (2001). Content, computation and externalism. Mind , 110(438):369--400

  89. [97]

    Shea, N. (2007). Content and its vehicles in connectionist systems. Mind & Language , 22:246--269

  90. [98]

    Shea, N. (2018). Representation in Cognitive Science . Oxford University Press

  91. [99]

    Simon, H. A. and Ando, A. (1961). Aggregation of variables in dynamic systems. Econometrica , 29(2):111--138

  92. [100]

    Smolensky, P. (1986). Neural and conceptual interpretation of PDP models. In McClelland, J. L., Rumelhart, D. E., and the PDP Research Group, editors, Parallel Distributed Processing: Explorations in the Microstructure of Cognition: Psychological and Biological Models , volume...

  93. [101]

    Spirtes, P., Glymour, C., and Scheines, R. (2000). Causation, Prediction, and Search . MIT Press

  94. [102]

    Sprevak, M. (2010). Computation, individuation, and the received view on representation. Studies in History and Philosophy of Science , 41:260–270

  95. [103]

    Sprevak, M. (2018). Triviality arguments about computational implementation. In Sprevak, M. and Colombo, M., editors, The Routledge Handbook of the Computational Mind , pages 175--191. Routledge

  96. [104]

    Sutter, D., Minder, J., Hofmann, T., and Pimentel, T. (2025). The non-linear representation dilemma: Is causal abstraction enough for mechanistic interpretability?

  97. [105]

    S., Cao, R., and Yamins, D

    Thobani, I., Sagastuy-Brena, J., Nayebi, A., Prince, J. S., Cao, R., and Yamins, D. L. (2025). Model-brain comparison using inter-animal transforms. In 8th Annual Conference on Cognitive Computational Neuroscience

  98. [106]

    Thompson, D. (2023). Algorithms and Execution Traces . PhD thesis, Stanford University

  99. [107]

    K., Rattermann, M

    Thompson, R. K., Rattermann, M. J., and Oden, D. L. (2001). Perception and judgement of abstract same-different relations by monkeys, apes and children: Do symbols make explicit only that which is implicit? Croatian Review of Rehabilitation Research , 37(1):9--22

  100. [108]

    J., Geiger, A., and Nanda, N

    Tigges, C., Hollinsworth, O. J., Geiger, A., and Nanda, N. (2024). Language models linearly represent sentiment. In Belinkov, Y., Kim, N., Jumelet, J., Mohebbi, H., Mueller, A., and Chen, H., editors, Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neur...

  101. [109]

    S., Mueller, A., Wallace, B

    Todd, E., Li, M., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D. (2024). Function vectors in large language models. In The Twelfth International Conference on Learning Representations

  102. [110]

    Turing, A. M. (1936). On computable numbers, with an application to the E ntscheidungsproblem. Proceedings of the London Mathematical Society , s2-42:230–265

  103. [111]

    Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S. (2020). Causal mediation analysis for interpreting neural NLP : The case of gender bias

  104. [112]

    R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J

    Wang, K. R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J. (2023). Interpretability in the wild: a circuit for indirect object identification in GPT -2 small. In The Eleventh International Conference on Learning Representations

  105. [113]

    Wasserman, E. A. and Young, M. E. (2010). Same-different discrimination: the keel and backbone of thought and reasoning. Journal of Experimental Psychology: Animal Behavior Processes , 36(1):3--22

  106. [114]

    Wendong, L., Buchholz, S., and Sch\" o lkopf, B. (2025). Algorithmic causal structure emerging through compression. In Huang, B. and Drton, M., editors, Proceedings of the Fourth Conference on Causal Learning and Reasoning , pages 201--242

  107. [115]

    Woodward, J. (2003). Making Things Happen: A Theory of Causal Explanation . Oxford University Press

  108. [116]

    Wu, Z., Geiger, A., Icard, T., Potts, C., and Goodman, N. (2023). Interpretability at scale: Identifying causal mechanisms in A lpaca. In Thirty-seventh Conference on Neural Information Processing Systems

  109. [117]

    and Kuorikoski, J

    Ylikoski, P. and Kuorikoski, J. (2010). Dissecting explanatory power. Philosophical Studies , 148(2):201--219

  110. [118]

    Zeller, A. (2002). Isolating cause-effect chains from computer programs. In Proceedings of the 10th ACM SIGSOFT Symposium on Foundations of Software Engineering , SIGSOFT '02/FSE-10, page 1–10. Association for Computing Machinery

  111. [119]

    Zennaro, F. M. (2022). Abstraction between structural causal models: A review of definitions and properties. In UAI 2022 Workshop on Causal Representation Learning

  112. [120]

    and Nanda, N

    Zhang, F. and Nanda, N. (2024). Towards best practices of activation patching in language models: Metrics and methods. In The Twelfth International Conference on Learning Representations

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.