Pith. sign in

REVIEW 4 major objections 5 minor 4 references

Putnam's Critical and Explanatory Tendencies Interpreted from a Machine Learning Perspective

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper argues that Putnam's critical and explanatory tendencies are necessarily interdependent, and that deep learning models illustrate the interdependence in practice.

desk verdict A readable, honest reconstruction of Putnam that overstates its case: the necessity claim doesn't follow from the premises, and the author's own concessions confirm it. read the letter →

arxiv 2501.03026 v1 pith:UA4KECGK submitted 2025-01-06 cs.CY cs.AIcs.LG

classification cs.CYcs.AIcs.LG
keywords HilaryPutnamcriticaltendencyexplanatorytheorychoicescientificexplanationmachinelearninginterpretabilitydeepphilosophyofscience
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that two ways of representing scientific problems that Hilary Putnam distinguished—the critical tendency, where theory plus auxiliary statements yields a prediction to be tested, and the explanatory tendency, where theory plus missing auxiliary statements yields an explanation of a fact—cannot function apart. Its central claim is a biconditional: any successful critical representation already depends on explanatory content in its auxiliary statements, and any successful explanatory representation already depends on predictive content. The author uses Putnam's own examples (predicting Earth's orbit; explaining Uranus's orbit) to show the two schemata interlock, then maps them onto deep learning: a model plus input produces a prediction, and explaining a model's trained features requires successful input–output predictions. If the argument is right, prediction and explanation are two aspects of the same underlying scientific act, and machine learning gives a concrete place to watch that interdependence.

What carries the argument

The central machinery is Putnam's paired schemata as modified by the author's biconditional reading. Schema I reads THEORY + AUXILIARY STATEMENTS → PREDICTION (true or false?), and schema II reads THEORY + ?????? → FACT TO BE EXPLAINED. The argument's engine is the identification of the missing clause in each schema with content of the other: the auxiliary statements in schema I are said to supply explanatory power, and the missing auxiliary in schema II is said to be a predictive statement of schema I form, as in the Uranus example. The author also introduces a machine-learning analogue: a trained model plus an input produces an output/prediction (schema I), and extracting meaning from learned parameters via successful input–output tests forms a schema II-like explanatory loop. This analogue lets the author claim that the interdependence is not just a philosophical reconstruction but observable in working AI systems.

What would settle it

Find one successful scientific prediction whose auxiliary statements contribute no explanatory content at all—for example, a purely conventional coordinate choice that makes a theory computable but that no one would claim explains anything—and the paper's claim that explanatory power is necessary for schema I collapses.

Watch

Extended reading notes

Core claim

The paper's central claim is that the critical tendency and the explanatory tendency are necessarily interdependent, so that a scientific problem represented by one schema is always already entangled with the other. For the critical tendency this means the auxiliary statements that make a prediction possible must also carry explanatory power—otherwise the prediction is not meaningful—and for the explanatory tendency it means the missing auxiliary statements used to explain a fact must themselves have predictive power, often taking the form of a lower-level schema I prediction. The author supports the first half by extending Putnam's observation that a theory never predicts alone: just as a theory needs auxiliaries to yield a prediction, it needs auxiliaries with explanatory content to make the prediction meaningful, illustrated by the Copernican case where two theories had equal predictive power but different explanatory reach. The author supports the second half through Putnam's Uranus example, where the auxiliary that completes the explanation is itself a schema I prediction about the existence of a planet. The paper then reads deep learning models through the same two schemata, arguing that model predictions depend on contextualizing inputs and that model explainability depends on predictive success, making machine learning a live instance of the interdependence.

Load-bearing premise

The argument assumes that every scientific virtue can be boiled down to either predictive power or explanatory power; if that reduction fails, the claimed necessity in each schema loses its support.

Editorial extensions

If this is right

  • If every successful prediction already rests on explanatory auxiliaries, debates about theory choice cannot reduce to predictive accuracy alone; explanatory value is already inside the predictive act.
  • If every successful explanation already relies on a predictive auxiliary, then explanations that float free of any possible prediction are not usable by a scientific community.
  • Machine learning use, schematised as model plus input to output, inherits the same structure: an LLM's output is not a prediction from the model alone but from model plus prompt, so the prompt functions as an explanatory auxiliary.
  • Model explainability, interpreted as a schema II problem, depends on predictive success: we explain model features by testing what input–output pairs they make true, so the two tendencies interlock inside AI systems as well.
  • If machine learning models continue to improve as the paper assumes, future extraordinary science may be driven by new parameter structures rather than by new observations, a possibility the author calls the 'parameter-ladenness of theories'.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the paper does not pursue: if the biconditional holds, removing the contextual prompt from an LLM should degrade prediction quality in a way that tracks the loss of explanatory auxiliary content, not just loss of information.
  • The author's reduction of epistemic values could be made quantitative by measuring explanatory and predictive power per parameter in trained models, turning simplicity and scope into efficiency measures rather than separate values.
  • The parameter-ladenness idea suggests a sharper research question than the paper states: which scientific explanations would change if a foundation model's internal representation, rather than an explicit law, became the standard auxiliary statement?
  • If the interdependence is real, purely rationalizing AI explanations that make no predictive difference should be judged as failed science, a criterion the paper gestures at but does not state.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper reconstructs Hilary Putnam's distinction between the critical tendency (schema I: theory plus auxiliary statements yielding a prediction to be judged true/false) and the explanatory tendency (schema II: theory plus auxiliary statements explaining a fact). It then argues for a biconditional necessity claim (C): the two tendencies are necessarily interdependent. The argument proceeds through two premises, P1 (successful schema I representations require auxiliary statements with explanatory power) and P2 (successful schema II representations require conjoining theory with prediction), and is then illustrated through an analogy with machine learning models, where model parameters, inputs, and outputs are mapped onto the schemata. The conclusion reflects on machine learning as a 'wrench' in theory-choice debates and suggests the possible discovery of 'model-independent objective parameters.' The paper is candid about its limitations, including an explicit admission that the term 'successfully' may be a weasel word and that representation without the other tendency may still be possible.

Significance. If the necessity claim C were established, the paper would offer a substantive philosophical thesis: every predictive act in science carries explanatory presuppositions, and every explanation depends on predictive auxiliaries. That would give a principled unity to two tendencies usually treated as complementary but separable. The paper has notable strengths: it accurately reconstructs Putnam's schemata and examples (Earth orbit, Uranus/Neptune), uses standard literature (Kuhn, McMullin, Putnam), and introduces an original and thought-provoking machine learning analogy. It also explicitly flags its own weakest points, which is a commendable scholarly practice. However, the central modal inference is not demonstrated; the paper shows at most that interdependence is valuable and frequently present, not that it is necessary. The epistemic-values reduction that would support P1 is asserted rather than argued. The machine learning analogy, while suggestive, does not close the modal gap because it establishes a computational dependence on input rather than an explanatory power carried by that input.

major comments (4)
  1. [Introduction, argument block (P1, P2, C) and Conclusion] The inference from the conditional premises to the necessity claim C is not valid. Both P1 and P2 are conditional on a scientific problem being 'successfully represented' by one of the schemata, but C asserts an unconditional necessity: that representation in either tendency is dependent on the other. The author's own concluding sentence concedes the gap: 'It can be put forth that it is possible to represent a scientific problem without the other tendency; however, I suggest that it is the worse for it.' That sentence directly contradicts C, since a representation that is merely worse is still possible. To repair this, the author must either define 'successfully' so that it already includes the presence of the other tendency (which would make the argument circular) or weaken C to a claim of valuable or typical interdependence.
  2. [Support for P1, third paragraph] The load-bearing premise that all epistemic values reduce to predictive or explanatory power is asserted rather than demonstrated. The author explicitly says that reducing scope or simplicity 'can't be approached in the same manner, but I believe it is possible.' Since P1's 'must' depends on this reduction (auxiliary statements must have explanatory power because only predictive and explanatory power matter), the premise is left unsupported at exactly the point where the necessity claim is supposed to enter. The paper needs at least a sketch of a reduction argument for simplicity and scope, or an alternative argument for why auxiliary statements are indispensable for successful schema I representation.
  3. [Support for P2, Uranus example] Putnam's Uranus/Neptune example, as reconstructed here, demonstrates that a schema II problem can be resolved by introducing a schema I-type auxiliary statement (S3), but it does not demonstrate that such a resolution is the only possible one or that every successful schema II representation must contain a predictive auxiliary. The later claim that 'the predictive power must be there if the representation is to be used by the scientific community' is a sociological or pragmatic assertion about use conditions, not an argument that a schema II representation without predictive power is impossible. The author also acknowledges that 'it is debatable whether a prediction must be based on observation,' which further weakens the necessity claim. P2 therefore remains a claim about a common and valuable pattern, not an established necessity.
  4. [Machine learning schematizations (pp. 3-5)] The machine learning analogy does not repair the modal gap in the argument. The claim that a model cannot produce an output without an input establishes a computational dependence, but it does not establish that the input has explanatory power in Putnam's sense; an input vector or a prompt is not thereby an 'auxiliary statement' that explains where the theory is applied. Similarly, the explainability schema (pp. 4-5) shows that predictions can be used to probe model features, but that is a methodological point about model interpretation, not a demonstration that the critical and explanatory tendencies are necessarily interdependent. The analogy is suggestive, but it trades on an equivocation between 'input as a necessary condition for computation' and 'auxiliary statements as providing explanatory meaning.'
minor comments (5)
  1. [Support for P1, p. 3] There is a typo in 'Khun's fruitfulness'; it should be 'Kuhn's fruitfulness.'
  2. [Machine learning schema, p. 5] The diagram heading reads 'Molel (Parameters)'; it should be 'Model (Parameters).'
  3. [References and citations] The in-text citation 'Curd & Clover, 1998' does not match the reference list entry 'Curd, Martin, J. A. Cover, and Chris Pincock'; the surname should be 'Cover', not 'Clover', and the year in the text should be reconciled with the reference entry.
  4. [References] The final reference for the Explainability survey ends with a duplicated fragment '/10.48550/arXiv.2309.01029.' that appears to be a leftover from the previous line and should be deleted.
  5. [References] McMullin is cited in the text (in 'Rationality and Paradigm Change') but does not appear in the reference list; a full bibliographic entry should be added.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is a reconstruction of Putnam with an external ML analogy; the modal gap in C is a support failure, not a definitional reduction.

full rationale

This paper contains no fitted parameters, no self-citation chain, and no imported uniqueness theorem. The central argument is a philosophical reconstruction of Putnam's two schemata, and the author explicitly concedes this: 'At the core, it looks like my philosophical claim becomes one of emphasis, and otherwise is a reconstruction of many of the same movements as Putnam's.' That admission is an acknowledgment of intellectual debt, not a circular reduction. The ML discussion is an external analogy and does not feed back into the premises that establish C. The main weakness is logical rather than circular: P1 and P2 are conditionals on 'successfully represented,' while C asserts unconditional necessary interdependence. The author even concedes the modal gap: 'It can be put forth that it is possible to represent a scientific problem without the other tendency; however, I suggest that it is the worse for it.' This contradicts the necessity claim, but it also shows that 'successfully' was not defined so as to make C true by construction. No equation or definition is shown to be equivalent to its own input, and no prediction is fitted and then relabeled as an independent result. Therefore no circular step is present.

Assumptions & free parameters 0 free parameters · 4 assumptions · 2 invented entities

The paper contains no quantitative fits, so the free-parameter list is empty; its load-bearing premises are philosophical assumptions and speculative entities.

assumptions (4)
  • domain assumption Neural networks are scalable: more compute directly yields higher prediction accuracy.
    Stated in the Introduction ('I assume that neural networks are scalable'); load-bearing for the speculation that ML will produce qualitatively better predictions and for the 'wrench' significance claim.
  • ad hoc to paper All other epistemic values (fertility, simplicity, scope) reduce to predictive or explanatory power.
    Asserted in Support for P1 without argument ('I believe it is possible'); used to elevate explanatory power to the essential cognitive value and thus to support P1's 'must.'
  • domain assumption Schema II representations may legitimately use predictive law-like auxiliary statements.
    Assumed in Support for P2; the author offers Carnap's Aufbau as a possible rather than demonstrated defense.
  • domain assumption McMullin's interpretation of the Copernican revolution (explanatory power drove the shift) is accepted as given.
    Used in Support for P1 to show explanatory power decides when predictive power is equal; no historical analysis is provided.
invented entities (2)
  • The 'wrench' (ML as disrupter of theory-choice debates)
    purpose: Metaphor for the paper's claimed contribution: ML models as a new force in philosophy of science debates on theory choice.
    The paper states the wrench is 'that small wrench that I hope to better understand'; no concrete mechanism is given.
  • Model-independent objective parameters of theories
    purpose: Speculated discovery target: objective parameters of scientific theories that analysis of ML models might reveal, leading to 'parameter-ladenness of theories.'
    Concluding speculation with no falsifiable handle; explicitly 'may conceivably be discovered.'

how reviews work

0 comments
Cite this review

Pith. "Pith review of Putnam's Critical and Explanatory Tendencies Interpreted from a Machine Learning Perspective." pith.science (2026). https://pith.science/paper/UA4KECGK

@misc{pith2026250103026,
  author       = {Pith},
  title        = {Pith review of: Putnam's Critical and Explanatory Tendencies Interpreted from a Machine Learning Perspective},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UA4KECGK}},
  note         = {Machine review of arXiv:2501.03026}
}
read the original abstract

Making sense of theory choice in normal and across extraordinary science is central to philosophy of science. The emergence of machine learning models has the potential to act as a wrench in the gears of current debates. In this paper, I will attempt to reconstruct the main movements that lead to and came out of Putnam's critical and explanatory tendency distinction, argue for the biconditional necessity of the tendencies, and conceptualize that wrench through a machine learning interpretation of my claim.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

4 extracted references · 1 canonical work pages

  1. [4]

    /10.48550/arXiv.2309.01029

    https://doi.org/10.48550/arXiv.2309.01029. /10.48550/arXiv.2309.01029

  2. [1979]

    Attention Is All You Need

    https://doi.org/10.1017/CBO9780511625268.018. Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. “Attention Is All You Need.” arXiv, August 2,

  3. [2023]

    Explainability for Large Language Models: A Survey

    https://doi.org/10.48550/arXiv.1706.03762. Zhao, Haiyan, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. “Explainability for Large Language Models: A Survey.” arXiv, November 28,

  4. [2024]

    AI-Generated Poetry Is Indistinguishable from Human-Written Poetry and Is Rated More Favorably

    https://doi.org/10.48550/arXiv.2410.14724. Porter, Brian, and Edouard Machery. “AI-Generated Poetry Is Indistinguishable from Human-Written Poetry and Is Rated More Favorably.” Scientific Reports 14, no. 1 (November 14, 2024): 26133. https://doi.org/10.1038/s41598-024-76900-1. Putnam, Hilary, ed. “The ‘Corroboration’ of Theories.” In Mathematics, Matter a...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.