REVIEW 3 major objections 6 minor 67 references
"Cause" is Mechanistic Narrative within Scientific Domains: An Ordinary Language Philosophical Critique of "Causal Machine Learning"
T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Causal machine learning redefines 'cause' and overstates what statistical graphs can certify, a philosophical critique argues.
desk verdict A well-read philosophical essay whose descriptive core about domain-specific causal language largely works, but whose strong conclusion that causal ML cannot certify causes rests on an under-defended ordinary-language premise. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the Ordinary Language method, drawn from Wittgenstein's language-game analysis, which treats the meaning of 'cause' as its actual use within each scientific community's practices. This method lets the paper show that the grammar of causal claims differs across physics, biology, and social science, and it uses Lakatos's 'hard core' idea to locate causal mechanisms within each domain's foundational assumptions. The argument then proceeds by contrasting this domain-relative meaning with the graphical independence model of causal learning, which the paper says amounts to redefining the word 'cause'.
What would settle it
A single well-documented case where a causal DAG discovered purely from conditional independence in an open biological or social system led to a verified intervention that the field's mechanistic theories had rejected, and where the mechanism was later confirmed, would directly contradict the paper's claim that independence-based discovery cannot certify real causes outside physics and engineering.
Extended reading notes
Core claim
The authors discover that causality is a functional concept whose form is fixed by the grammatical conventions of each scientific community, not a single relation that statistical independence can certify. They argue that causal learning takes the exact equation systems of physics as an analogy and applies that analogy to open, emergent, and interpretive systems where it does not carry the same authority. Consequently, a fitted DAG with passing conditional-independence tests is at best a plausibility suggestion, not a certified cause. Definitive causality requires what the authors call an agglomeration of consistent evidence across domains, scales, and narrative forms.
Load-bearing premise
The argument depends on the premise that the everyday meaning of 'cause' used inside each scientific community is the correct standard for what a true causal claim must mean, which is a normative interpretive choice rather than a mathematical fact.
Editorial extensions
If this is right
- Causal discovery results on real datasets should be reported as generating hypotheses, not as certified causes.
- Research evaluating causal machine learning should demand domain-specific mechanistic validation before accepting a graph as evidence of causation.
- The replication crisis in social science will not be fixed by more sophisticated statistics alone, but by lowering confidence and seeking convergence across multiple disciplines.
- Causal statements in physics and engineering remain legitimate where the equation models are known and validated.
- The burden of proof in medicine and policy shifts from a single DAG or p-value to multi-scale, multi-domain evidence.
Reading between the lines
- A concrete test of the paper's position would be to run causal discovery on well-understood physics benchmarks versus open biological or social datasets and compare whether the discovered graphs align with the domains' known mechanisms.
- The argument implies a communication standard for high-stakes 'causal AI': outputs should be phrased as 'consistent with' rather than 'causes' unless the domain mechanism is known and identified.
- The authors' sketch of analogical reasoning via CP-logic suggests a research program for formalizing cross-domain evidence integration, which could be tested by building systems that transfer causal rules between modeled domains.
- The critique can be read as a call for pluralistic causal models, where physics, biology, and social science each get a distinct causal calculus rather than one universal statistical framework.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This is an argumentative, interdisciplinary paper that critiques the foundations of causal machine learning (causal inference/discovery as implemented via Bayesian networks, DAGs, and conditional independence tests). The authors argue, using ordinary language philosophy, that the word "cause" has no single formal definition; its meaning is fixed by the language games of distinct scientific communities. They then demarcate physics/engineering (where mathematical models can fully capture causality), biology (where emergence requires multi-scale evidence), and the social sciences (where hermeneutic interpretation is valuable but precise causal claims require multi-domain convergence). Their central critical claim is that DAG-based causal learning should not be read as certifying that 'A causes B', and they argue that definitive causal claims about open systems require an agglomeration of evidence across multiple domains and levels of abstraction.
Significance. If its central argument were fully supported, the paper would be a significant contribution to the ongoing debate about the overinterpretation of causal machine learning results. Its main strengths are its synthetic use of philosophy of science (Wittgenstein, Kuhn, Lakatos, Quine) to frame the issue, its clear typology of causal semantics across physics, biology, and social science, and its explicit concession that causal formalism is correct for physics/engineering when the underlying equation model is known and exact. The paper also makes a useful practical recommendation: that causal discovery results in open, complex systems should be treated as hypothesis-generating rather than as definitive certification. The paper does not claim to provide mathematical proofs; its value is conceptual, and its credibility hinges on the normative premise about the authority of ordinary language in fixing the meaning of 'cause'.
major comments (3)
- [§3.4–§3.5, §6.2] The paper moves from the descriptive observation that 'cause' is used differently in different scientific domains to the normative conclusion that the ordinary-language usage within each scientific tribe fixes what a causal claim must mean. This premise is never defended. The strongest statement in §6.2 that DAG-based causal discovery 'nor can it' certify causes, and in §3.5 that it 'amounts to redefining the word "cause"', depends entirely on that premise. Furthermore, the paper's own functional definition in §1—'the mechanism underlying fundamental forces of influence'—is precisely what a structural causal model's assignment X_i := f_i(Pa_i, U_i) formalizes. The graphical model is therefore not a redefinition but an abstraction of the paper's own concept; the remaining dispute is about whether the level of abstraction preserves the domain-relevant mechanisms, which is a pragmatic and gradable epistemic question. The categorical 'cannot' is not supported. I recommend reframing the conclusion as a counsel of caution about overclaiming from DAG+CI in open systems, rather than an in-principle impossibility.
- [§6.3, §3.3] The positive thesis that definitive causal claims require an 'agglomeration of consistent evidence across multiple domains' is asserted but not operationalized. If, as the paper argues with Kuhn in §3.3, scientific paradigms are incommensurable language games, then it is nontrivial to say what counts as consistency of evidence across those games. The examples (smoking, Weber's Protestant Ethic, cognitive science) are suggestive, but no criterion is given for when cross-domain results harmonize rather than merely coexist. Without such a criterion, the proposed mixed-methods framework is underspecified as a research program.
- [§2.1 and §6.2] Several empirical generalizations are made without supporting evidence. The claim in §2.1 that the authors could not identify a single real-data causal-learning paper with all conditional independence tests passing is anecdotal, not a systematic survey. Similarly, §6.2 states that sparse causal models in social sciences are 'uncommon in the literature' and that the few that exist 'present type-1 error concerns', but no citation is given. These empirical claims are used to bolster the critique of causal learning's practical value and should either be substantiated with a proper literature review or removed.
minor comments (6)
- [Abstract] The final sentence of the abstract is a fragment: 'Given the role of epistemic hubris ... optimizing integration of different findings.' It needs a main clause to be a complete sentence.
- [§3.4] The name 'Halpern' is misspelled as 'Harpen' twice in the discussion of actual causality.
- [§5.3] 'paropagation' should be 'propagation'.
- [§5.4] 'feedforwark' should be 'feedforward'.
- [§6.2] The quotation 'bewitchment of intelligence by language' is a paraphrase of Wittgenstein's Philosophical Investigations §109; an exact citation would be helpful.
- [§2.1] The phrase 'temporal difference in differences at interventions' is unclear and should be clarified.
Circularity Check
No significant circularity; the central philosophical claims are argued from independent sources, and the only self-citations are to the authors' own extended version for auxiliary detail.
full rationale
This paper is a philosophical critique, not a formal derivation chain. It contains no fitted parameters, no quantitative predictions, and no theorem whose conclusion is assumed in its premises. The central thesis—that "cause" has distinct grammatical forms across scientific domains while serving a common functional role—is argued from ordinary-language observations and from external philosophical authorities (Wittgenstein, Kuhn, Quine, Gadamer, Lakatos), not from the conclusion itself. The only self-citations are references to the authors' extended version (Kungurtsev et al., 2025) for additional formulations and a taxonomy of cognitive science subdomains; these are ancillary and not load-bearing. No uniqueness theorem or exclusion result is imported from the authors' prior work. The strongest skeptical objection is that the impossibility claim—that causal machine learning "cannot" certify causes and "amounts to redefining the word 'cause'"—depends on an asserted normative premise that ordinary language inside each scientific tribe fixes the legitimate meaning of "cause." That is a substantive epistemic criticism, but it is not circularity: the paper does not define "cause" as whatever causal ML cannot certify; it attempts to establish its semantic thesis from usage and then draws consequences. One might also note an internal tension: the paper's own functional definition of cause as "the mechanism underlying fundamental forces of influence" is close to what structural causal models formalize, which could undercut the critique, but tension is not self-derivation. Overall, the argument is self-contained in the sense required here, with only a minor, non-load-bearing self-citation, so the circularity score is low.
Assumptions & free parameters
free parameters (1)
- domain demarcation thresholds =
physics/engineering, biology, social sciences
assumptions (4)
- domain assumption Ordinary language use is the correct standard for philosophical meaning.
- domain assumption Scientific disciplines function as separate language games with distinct grammars of causality.
- domain assumption A spurious statistical association can never be upgraded to a causal relation without mechanistic narrative.
- domain assumption Psychotherapy outcome equivalence across modalities is an established empirical fact.
invented entities (1)
-
hermeneutic truth as a distinct epistemic category
Cite this review
Pith. "Pith review of "Cause" is Mechanistic Narrative within Scientific Domains: An Ordinary Language Philosophical Critique of "Causal Machine Learning"." pith.science (2026). https://pith.science/paper/5RTXYI2L
@misc{pith2026250105844,
author = {Pith},
title = {Pith review of: "Cause" is Mechanistic Narrative within Scientific Domains: An Ordinary Language Philosophical Critique of "Causal Machine Learning"},
year = {2026},
howpublished = {\url{https://pith.science/paper/5RTXYI2L}},
note = {Machine review of arXiv:2501.05844}
}
read the original abstract
Causal Learning has emerged as a major theme of research in statistics and machine learning in recent years, promising specific computational techniques to apply to datasets that reveal the true nature of cause and effect in a number of important domains. In this paper we consider the epistemology of recognizing true cause and effect phenomena. We apply the Ordinary Language method of engaging on the customary use of the word 'cause' to investigate valid semantics of reasoning about cause and effect. We recognize that the grammars of cause and effect are fundamentally distinct in form across scientific domains, yet they maintain a consistent and central function. This function can best be described as the mechanism underlying fundamental forces of influence as considered prominent in the respective scientific domain. We demarcate 1) physics and engineering as domains wherein mathematical models are sufficient to comprehensively describe causality, 2) biology as introducing challenges of emergence while providing opportunities for showing consistent mechanisms across scale, and 3) the social sciences as introducing grander difficulties for establishing models of low prediction error but providing, through Hermeneutics, the potential for findings that are still instrumentally useful to individuals. We posit that definitive causal claims regarding a given phenomenon (writ large) can only come through an agglomeration of consistent evidence across multiple domains. This presents important methodological questions as far as harmonizing between language games and emergence across scales. Given the role of epistemic hubris in the contemporary crisis of credibility in the sciences, exercising greater caution as far as communicating precision as to the real degree of certainty certain evidence provides for rich collections of open problems in optimizing integration of different findings.
Figures
Reference graph
Works this paper leans on
-
[1]
Abdelfattah, A. M. and Krumnack, U. (2019). Semantics of analogies from a logical perspective. KI-K \"u nstliche Intelligenz , 33:243--251
work page 2019
-
[2]
Badiou, A. (2007). Being and event . A&C Black
work page 2007
-
[3]
Baudrillard, J. (1994). Simulacra and simulation. U of Michigan P
work page 1994
-
[4]
Bechtel, W., Abrahamsen, A., and Graham, G. (2017). The Life of Cognitive Science , pages 1--104. Blackwell Publishing Ltd. First published 1998, reprinted 2017
work page 2017
-
[5]
Bergel, D. H. (2012). Cardiovascular fluid dynamics . Elsevier
work page 2012
-
[6]
Besold, T. R., Garcez, A. d., Stenning, K., van der Torre, L., and van Lambalgen, M. (2017). Reasoning in non-probabilistic uncertainty: Logic programming and neural-symbolic computing as examples. Minds and Machines , 27:37--77
work page 2017
-
[7]
Blease, C., Colagiuri, B., and Locher, C. (2023). Replication crisis and placebo studies: rebooting the bioethical debate. Journal of Medical Ethics , 49(10):663--669
work page 2023
-
[8]
Bongers, S., Forr \'e , P., Peters, J., and Mooij, J. M. (2021). Foundations of structural causal models with cycles and latent variables. The Annals of Statistics , 49(5):2885--2915
work page 2021
Show all 67 references
-
[9]
Bottomore, T. B. (2002). The Frankfurt School and its critics . Psychology Press
2002
-
[10]
Bourdieu, P. (1984). Distinction: A Social Critique of the Judgement of Taste . Harvard University Press
1984
-
[11]
Bourdieu, P. (1991). Language and symbolic power . Polity
1991
-
[12]
Caverni, J.-P., Fabre, J.-M., and Gonzalez, M. (1990). Cognitive biases . Elsevier
1990
-
[13]
Clark, A. (2013). Whatever next? Predictive brains, situated agents, and the future of cognitive science. Behavioral and Brain Sciences , 36(3):181--204
2013
-
[14]
Coen, E. (2019). The storytelling arms race: Origin of human intelligence and the scientific mind. Heredity , 123(1):67--78
2019
-
[15]
and Kimmig, A
De Raedt, L. and Kimmig, A. (2015). Probabilistic (logic) programming concepts. Machine Learning , 100:5--47
2015
-
[16]
Doets, K. (1994). From logic to logic programming . Mit Press
1994
-
[17]
H., and Williams, R
Dyson, P., Ransing, R., Williams, P. H., and Williams, R. (2008). Fluid properties at nano/meso scale: a numerical treatment . John Wiley & Sons
2008
-
[18]
Eberhardt, F. (2009). Introduction to the epistemology of causation. Philosophy Compass , 4(6):913--925
2009
-
[19]
Echebur \'u a, E., B \'a ez, C., and Fern \'a ndez-Montalvo, J. (1996). Comparative effectiveness of three therapeutic modalities in the psychological treatment of pathological gambling: Long-term outcome. Behavioural and cognitive psychotherapy , 24(1):51--72
1996
-
[20]
Enfield, N. J. (2024). Language vs. reality: Why language is good for lawyers and bad for scientists . MIT Press
2024
-
[21]
Gadamer, H.-G. (2013). Truth and method . A&C Black
2013
-
[22]
Gintis, H. (2000). Game theory evolving: A problem-centered introduction to modeling strategic behavior . Princeton university press
2000
-
[23]
Graeber, D. (2018). Bullshit Jobs . E mploi
2018
-
[24]
Halpern, J. Y. (2016). Actual causality . MiT Press
2016
-
[25]
Hecht-Nielsen, R. (1989). Neurocomputing . Addison-Wesley Longman Publishing Co., Inc
1989
-
[26]
Hohwy, J. (2013). The Predictive Mind . Oxford University Press
2013
-
[27]
W., and Noeri, G
Horkheimer, M., Adorno, T. W., and Noeri, G. (2002). Dialectic of enlightenment . Stanford University Press
2002
-
[28]
M., Peters, J., and Sch \"o lkopf, B
Hoyer, P., Janzing, D., Mooij, J. M., Peters, J., and Sch \"o lkopf, B. (2008). Nonlinear causal discovery with additive noise models. Advances in neural information processing systems , 21
2008
-
[29]
F., Akyol, A., Zhaksylyk, A., Seiil, B., and Yessirkepov, M
Kocyigit, B. F., Akyol, A., Zhaksylyk, A., Seiil, B., and Yessirkepov, M. (2023). Analysis of retracted publications in medical literature due to ethical violations. Journal of Korean Medical Science , 38(40)
2023
-
[30]
Kueffner, K. R. (2021). A comprehensive Survey of the Actual Causality Literature . PhD thesis, Wien
2021
-
[31]
cause" is mechanistic narrative within scientific domains: An ordinary language philosophical critique of
Kungurtsev, V., Moore, L. C., Krutsky, M., et al. (2025). "cause" is mechanistic narrative within scientific domains: An ordinary language philosophical critique of" causal machine learning. arXiv preprint arXiv:2501.05844v2
2025 arXiv
-
[32]
and Feyerabend, P
Lakatos, I. and Feyerabend, P. (2019). For and against method: including Lakatos's lectures on scientific method and the Lakatos-Feyerabend correspondence . University of Chicago Press
2019
-
[33]
Lechner, M. (2023). Causal machine learning and its use for public policy. Swiss Journal of Economics and Statistics , 159(1):8
2023
-
[34]
Locatello, F., Bauer, S., Lucic, M., Raetsch, G., Gelly, S., Sch \"o lkopf, B., and Bachem, O. (2019). Challenging common assumptions in the unsupervised learning of disentangled representations. In international conference on machine learning , pages 4114--4124. PMLR
2019
-
[35]
Logothetis, N. K. (2003). The underpinnings of the BOLD functional magnetic resonance imaging signal. Journal of Neuroscience , 23(10):3963--3971
2003
-
[36]
Mitchell, M. (2001). Analogy making as a complex adaptive system. In SANTA FE INSTITUTE STUDIES IN THE SCIENCES OF COMPLEXITY-PROCEEDINGS VOLUME- , pages 335--360. Reading, Mass.; Addison-Wesley; 1998
2001
-
[37]
Montmarquet, J. A. (1987). Epistemic virtue. Mind , 96(384):482--497
1987
-
[38]
G., Tabacco, L., and Bilotta, F
Nato, C. G., Tabacco, L., and Bilotta, F. (2022). Fraud and retraction in perioperative medicine publications: what we learned and what can be implemented to prevent future recurrence. Journal of Medical Ethics , 48(7):479--484
2022
-
[39]
Orwell, G. (2021). Politics and the English language and other essays . epubli
2021
-
[40]
Pearl, J. (2009). Causality . Cambridge university press
2009
-
[41]
and Cooper, R
Peebles, D. and Cooper, R. P. (2015). Thirty years after marr's vision: levels of analysis in cognitive science
2015
-
[42]
Peirce, C. S. (1972). Charles s. peirce: the essential writings
1972
-
[43]
Mechanistic neural networks for scientific machine learning
Pervez, A., Locatello, F., and Gavves, S. Mechanistic neural networks for scientific machine learning. In ICLR 2024 Workshop on AI4DifferentialEquations In Science
2024
-
[44]
Peters, J., Janzing, D., and Sch \"o lkopf, B. (2017). Elements of causal inference: foundations and learning algorithms . The MIT Press
2017
-
[45]
and Cushing, J
Pickering, A. and Cushing, J. T. (1986). Constructing quarks: A sociological history of particle physics
1986
-
[46]
A., Fletcher, P
Poldrack, R. A., Fletcher, P. C., Henson, R. N., Worsley, K. J., Brett, M., and Nichols, T. E. (2008). Guidelines for reporting an fMRI study. NeuroImage , 40(2):409--414
2008
-
[47]
Popper, K. (2005). The logic of scientific discovery . Routledge
2005
-
[48]
D., Hoffman, D
Prakash, C., Stephens, K. D., Hoffman, D. D., Singh, M., and Fields, C. (2021). Fitness beats truth in the evolution of perception. Acta Biotheoretica , 69:319--341
2021
-
[49]
Risjord, M. (2022). Philosophy of social science: A contemporary introduction . Routledge
2022
-
[50]
Rosenberg, A. (1988). Philosophy of social science , volume 2. Westview Press Boulder, CO
1988
-
[51]
P., Xia, T., Watson, H
Sanchez, P., Voisey, J. P., Xia, T., Watson, H. I., O’Neil, A. Q., and Tsaftaris, S. A. (2022). Causal machine learning for healthcare and precision medicine. Royal Society Open Science , 9(8):220638
2022
-
[52]
R., Kalchbrenner, N., Goyal, A., and Bengio, Y
Sch \"o lkopf, B., Locatello, F., Bauer, S., Ke, N. R., Kalchbrenner, N., Goyal, A., and Bengio, Y. (2021). Toward causal representation learning. Proceedings of the IEEE , 109(5):612--634
2021
-
[53]
and Zayed, A
Shen, X. and Zayed, A. I. (2012). Multiscale signal analysis and modeling . Springer Science & Business Media
2012
-
[54]
Stupple, A., Singerman, D., and Celi, L. A. (2019). The reproducibility crisis in the age of digital medicine. NPJ digital medicine , 2(1):2
2019
-
[55]
Svedin, U. et al. (2005). Micro, meso, macro: Addressing complex systems couplings . World Scientific
2005
-
[56]
and Smith, L
Thelen, E. and Smith, L. B. (1994). A Dynamic Systems Approach to the Development of Cognition and Action . MIT Press
1994
-
[57]
Tsotsos, J. K. (1995). Behaviorist intelligence and the scaling problem. Artificial Intelligence , 75(2):135--160
1995
-
[58]
van Orman Quine, W. (1976). Two dogmas of empiricism. In Can Theories be Refuted? Essays on the Duhem-Quine Thesis , pages 41--64. Springer
1976
-
[59]
Van Steenkiste, S., Locatello, F., Schmidhuber, J., and Bachem, O. (2019). Are disentangled representations helpful for abstract visual reasoning? Advances in neural information processing systems , 32
2019
-
[60]
J., Thompson, E., and Rosch, E
Varela, F. J., Thompson, E., and Rosch, E. (1991). The Embodied Mind: Cognitive Science and Human Experience . MIT Press
1991
-
[61]
Veblen, T. (2017). The theory of the leisure class . Routledge
2017
-
[62]
Vennekens, J., Denecker, M., and Bruynooghe, M. (2009). Cp-logic: A language of causal probabilistic events and its relation to logic programming. Theory and practice of logic programming , 9(3):245--308
2009
-
[63]
Wang, S., Sankaran, S., and Perdikaris, P. (2022). Respecting causality is all you need for training physics-informed neural networks. arXiv preprint arXiv:2203.07404
2022 arXiv
-
[64]
and Kalberg, S
Weber, M. and Kalberg, S. (2013). The Protestant ethic and the spirit of capitalism . Routledge
2013
-
[65]
Wittgenstein, L. (2009). Philosophical investigations . John Wiley & Sons
2009
-
[66]
Wittgenstein, L. (2023). Tractatus logico-philosophicus
2023
-
[67]
Yao, D., Muller, C., and Locatello, F. (2024). Marrying causal representation learning with dynamical systems for science. arXiv preprint arXiv:2405.13888
2024 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.