{"id":"36d93621-b323-40c4-93ca-66dca7da1f53","arxiv_id":"2608.07812","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Behavioral fit alone never justifies treating a foundation model as an explanatory cognitive model; explicit theoretical commitments, diagnostic tasks, and contrastive evaluation are required.","lead":"This perspective lays out a four-stage framework for deciding when a foundation model's match to human behavior counts as a cognitive explanation rather than a clever imitation. It argues that behavioral alignment is only scientifically meaningful when paired with explicit links to theory, diagnostic tasks, and comparisons across models.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Underdetermination persists after all four stages: Stage 4's finite contrastive comparisons cannot establish the necessity of model features, so the framework's promise of explanatory alignment is unsupported.","rationale":"The paper is a thoughtful perspective that correctly argues behavioral fit alone is insufficient, and its four-stage framework usefully systematizes best practices. My stress-test targeted the positive inference claimed for Stage 4: that contrastive evaluation across models and manipulations licenses conclusions about which computational features are necessary for human-like behavior. This inference is the backbone of the claim that alignment becomes 'scientifically meaningful' under the framework. The paper acknowledges multiple realizability and unfaithful traces in Section 4.2, but it does not show how these are resolved by the framework; it merely asserts that interpretability and Stage 4 comparisons 'can' help. A finite set of comparisons is not an abstract argument against underdetermination. Thus, the central methodological prescription is not guaranteed to deliver explanatory insight. However, the paper's explicit claim is a necessary-condition claim ('only when'), not a sufficiency claim. The concern that the conditions may not be sufficient is a limitation, but not a falsification of the central claim. Moreover, the paper repeatedly hedges ('explanatory value increases', 'treated as a starting point'), so the concern is largely acknowledged. For these reasons, the verdict should remain ACCEPT, though the limitations could be stated more sharply. I agree with the reader's identification of the weakest assumption.","tokens_in":14125,"tokens_out":7793,"duration_ms":72739,"concrete_test":"Construct two minimal models of the same benchmark task (e.g., digit-matrix analogical reasoning): a transformer fine-tuned on the task and a transparent symbolic rule-based system with the same input-output accuracy. Run the full four-stage pipeline on both, including ablations within each model family (e.g., removing attention heads vs. removing rules). If both models pass Stage 4 relative to their own families yet implement fundamentally different algorithms, and if both fit the human data equally well, then the framework's contrastive inference cannot resolve multiple realizability. This simulation would show that 'necessity' claims are model-family-relative and hence insufficient to ground explanatory claims about human cognition.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that alignment becomes scientifically meaningful when embedded in explicit theory, diagnostic contrasts, and comparative evaluation hinges on Stage 4: the inference from contrastive model comparisons or ablations to the necessity of particular computational features for reproducing human behavior. This inference is not warranted by the arguments given. Because of multiple realizability (Section 4.2), many distinct algorithms can in principle produce the same behavioral profile on any finite set of tasks. A contrastive comparison over a few architectures, scales, or ablations can establish at most that a feature is necessary within that model family; it cannot establish that the feature is necessary for any computational account of the behavior, still less that it corresponds to a human mechanism. The paper acknowledges the problem ('two systems might produce similar outputs while relying on different internal computations') but treats it as a motivation for more interpretability work rather than as a structural limit on the framework. Without a well-defined model space or a formal criterion for what constitutes an alternative mechanistic account, 'necessity' conclusions from Stage 4 are underdetermined. Consequently, the framework may not deliver the promised elevation from descriptive fit to explanatory adequacy, even when all four stages are followed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a perspective article that asks under what conditions behavioral alignment between foundation models (FMs) and human performance can justify treating FMs as explanatory models of cognition. It proposes a four-stage inferential framework: (1) adapting human tasks to model-compatible formats, (2) specifying linking hypotheses that map model outputs to human measures, (3) evaluating behavioral correspondence with attention to diagnostic contrasts, and (4) comparing across candidate models or manipulations. The authors argue that behavioral fit alone is insufficient and that meaningful alignment requires explicit theoretical commitments, theory-diagnostic tasks, and contrastive evaluation. They discuss four linking hypotheses (similarity, surprisal, prompting, process-trace) and four challenges (theoretical underdetermination, mechanistic opacity, training/developmental mismatch, and population variability), then distill the framework into five research guidelines, using the digit-matrix analogical reasoning task as a running example.","tokens_in":14279,"tokens_out":6227,"duration_ms":55484,"significance":"If the framework is adopted, it would provide a common vocabulary and set of standards for a rapidly growing literature that evaluates FMs against human and developmental data. The paper's main contributions are conceptual: it separates task adaptation from linking hypotheses and evaluation, emphasizes diagnostic contrasts over aggregate fit, and insists that alignment is a relation between model, task, linking hypothesis, and theoretical claim rather than a property of the model alone. The treatment is careful and self-consciously hedged: the authors acknowledge multiple realizability, unfaithful chain-of-thought traces, and the correlational status of developmental correspondences. The paper also offers concrete, actionable guidelines and a running example that makes the abstract stages easy to follow. No new empirical validation is provided, but for a perspective article this is appropriate; the value lies in organizing and constraining future practice.","major_comments":[{"comment":"Stage 4 is described as identifying computational features that are 'necessary' to reproduce a behavioral signature, and Guideline 4 repeats this language. Finite contrastive comparisons over a few architectures, scales, ablations, or training regimes can establish at most that a feature is necessary within that particular model family and manipulation set; because of multiple realizability, which the paper itself acknowledges in Section 4.2, they cannot establish that the feature is necessary for any computational account of the behavior, let alone that it corresponds to a human mechanism. I recommend rewording these passages to say 'necessary within the class of models and manipulations under consideration' and adding a sentence that unconditional mechanistic necessity would require additional theoretical constraints beyond contrastive evaluation. This is a local but important precision issue, because the abstract and conclusion also use 'necessary' when describing what the framework can illuminate.","section":"Section 2, Stage 4; Guideline 4"},{"comment":"The paper reviews four linking hypotheses and explains their commitments, but it does not give the reader guidance for choosing among them beyond saying that appropriateness is 'empirical and task-dependent' in the surprisal section. Since the framework makes the linking hypothesis the central theoretical commitment, I would welcome a brief selection principle, for example, prefer the most proximal mapping that preserves the theoretical construct of interest, and triangulate across at least two linking hypotheses when possible. This would make the framework more actionable.","section":"Section 3, intro and Section 3.2"}],"minor_comments":[{"comment":"The reference list contains several formatting artifacts from LaTeX source, such as 'Y u', 'V arma', 'F orty-third', and 'ET AL .' in headings; these should be cleaned before publication.","section":"References"},{"comment":"The term 'proximal' linking hypothesis is used without a definition; a brief gloss (e.g., mapping inputs and outputs directly rather than through internal representations) would help readers who are not familiar with the distal/proximal distinction.","section":"Section 3.3"},{"comment":"The guideline notes that adaptation artifacts can lower observed alignment, but it is equally possible for surface cues to inflate alignment; adding this symmetric warning would make the point more complete.","section":"Section 5, Guideline 2"},{"comment":"The discussion of persona-based prompting cites conflicting findings, but does not specify which findings conflict or how a reader should interpret them; one or two concrete examples would strengthen the caution.","section":"Section 4.4"}],"recommendation":"minor_revision","confidential_remarks":"The paper is a well-scoped perspective piece and does not require empirical validation. Its self-citations are used as illustrative examples and do not create circularity. The main issue is the 'necessary' wording in Stage 4, which should be qualified; after that, I would be happy to see it published."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the paper. The short version: this is a genuinely useful perspective piece. It gives a four-stage framework (task adaptation, linking hypothesis, behavioral evaluation, contrastive comparison) and a four-way taxonomy of linking hypotheses (similarity, surprisal, prompting, process-tracing). Both are presented as distillations of prior commentary, not raw inventions, and the authors are honest about that debt. That is a credit, not a defect.\n\nWhat it does well: it is clear, well-organized, and disciplined about its own limits. The running example of digit-matrix Raven's problems keeps the abstraction concrete. The linking-hypothesis taxonomy is the strongest part; the discussion of how each mapping makes different theoretical commitments is genuinely useful, and the point that alignment is not a property of the model alone is well made. The paper also flags its own weaknesses—mechanistic opacity, multiple realizability, task adaptation risks, and the danger of reading intermediate traces as faithful process measures. That is real intellectual honesty.\n\nWhere the soft spots are, in proportion. The main one is the stress-test worry about Stage 4. Contrastive comparisons over a handful of architectures and ablations cannot establish necessity in any strong sense: any finite set of alternatives leaves open infinitely many other algorithms that reproduce the behavior. That is true. But the paper is careful not to claim more than it needs to. It says contrastive evaluation 'helps identify' features that are necessary, and it explicitly cabins the promise in the conclusion ('sufficient, and perhaps necessary'). The central thesis—behavioral fit alone is insufficient—does not depend on a strong necessity inference. So the stress-test does not land as a load-bearing flaw; it is a genuine limitation that the authors largely acknowledge and then bracket. Worth saying in a referee comment, but not grounds for rejection.\n\nTwo smaller concerns. First, the framework is a synthesis, and its components will feel familiar to anyone who reads the cited commentaries; the novelty is packaging, which is fine for a perspective but should be characterized as such. Second, there is no worked empirical validation of the four-stage framework end-to-end. For a perspective paper, that is acceptable; the guidelines are plausible and actionable. The authors do not overclaim.\n\nWho this is for: anyone running model-human alignment studies, especially in cognitive science and NLP. It deserves a serious referee, not a desk reject. I would send it out.","headline":"A clear, honest synthesis of the inferential steps between foundation-model outputs and human behavior; the framework is useful even though the strongest necessity claims stay underdetermined.","tokens_in":14809,"tokens_out":1728,"would_cite":true,"duration_ms":16549,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that behavioral fit alone cannot justify treating foundation models as explanatory cognitive models; alignment becomes meaningful only within explicit theory, diagnostic tasks, and contrastive evaluation.","keywords":["foundation models","cognitive modeling","linking hypotheses","behavioral alignment","model evaluation","cognitive science","large language models","developmental alignment"],"falsifier":"Find two foundation models with architecturally distinct mechanisms that both pass all four stages on the same human dataset, and show that no manipulation of task, linking hypothesis, or training regime separates their predictions; such a result would show that the contrastive stage cannot identify which computational features are necessary, undermining the framework's explanatory criterion.","tokens_in":13914,"feed_emoji":"🧠","tokens_out":8374,"duration_ms":64219,"temperature":0.7,"pith_summary":"Foundation models can reproduce many behavioral signatures of human cognition, but the paper argues that such behavioral fit is not itself evidence that a model explains how humans think. The authors propose a four-stage inferential framework—task adaptation, linking hypothesis specification, behavioral evaluation, and contrastive model comparison—as the minimal structure under which alignment becomes scientifically meaningful. The central message is that alignment is a property of a model evaluated under explicit theoretical commitments and diagnostic contrasts, not a property of the model alone. A sympathetic reader would take away that most current demonstrations of model–human similarity are proof-of-possibility results, not explanatory claims, until embedded in this structure.","feed_headline":"Matching human behavior is not enough to call an AI cognitive","feed_subtitle":"New framework: alignment counts only with explicit theory, diagnostic tasks, and contrastive model comparison.","key_machinery":"The central object is the four-stage inferential framework, split into an inner alignment loop (Stages 1–3: task adaptation, linking hypothesis, goodness-of-fit evaluation) and an outer contrastive loop (Stage 4: cross-model and manipulation comparison). The linking hypothesis is the load-bearing pivot of the framework: it defines what counts as evidence of alignment by mapping model outputs to human behavioral measures, and the paper argues that a strong fit under one linking hypothesis can vanish under another. The running example is the translation of progressive matrix analogies into symbolic digit matrices [107], which the paper uses to show how each stage changes the interpretation of an alignment result.","core_discovery":"The paper's central claim is that behavioral alignment justifies treating a foundation model as an explanatory cognitive model only when the evaluation embeds the model in a theory-diagnostic design: adapt the task so model and humans perform functionally equivalent problems, specify a linking hypothesis that maps model outputs to human measures, evaluate fit on theoretically diagnostic contrasts rather than aggregate scores, and compare across candidate models or manipulations to identify which computational features are necessary for the fit. Under this view, a model that simply gets high accuracy or matches average human judgments remains a behavioral proxy. The paper supports the claim by showing how each of four common linking hypotheses—similarity in representational space, surprisal, prompting, and process-trace analysis—carries different theoretical commitments, and by identifying four challenges that constrain alignment claims: theoretical underdetermination, mechanistic opacity, training and developmental mismatch, and population-level variability.","pith_inferences":["If the framework is right, a large share of published model–behavior benchmarks should be reinterpreted as capability or similarity studies rather than cognitive models, unless they already include diagnostic contrasts and model comparison.","A testable extension: reanalyzing landmark alignment results under alternative linking hypotheses should change the apparent fit, and the direction of change would reveal which theoretical commitment is doing the work.","The framework implies that the field's next bottleneck is not larger models but better theory: designing tasks whose contrasts discriminate between computational mechanisms, and reporting null or negative contrasts as informative.","The four-stage structure could generalize to other 'black box' scientific models, wherever mere behavioral fit risks being mistaken for mechanistic insight."],"forward_implications":["Single-model demonstrations, however strong, establish at most that a cognitive signature is reproducible, not that the model's mechanisms explain it.","Evaluation of alignment without diagnostic contrasts—comparing conditions that discriminate between theories—cannot adjudicate between cognitive accounts.","Developmental alignment claims require more than matching learning curves with children; they need causal manipulations of training regime or data ordering.","Because linking hypotheses are not neutral, results should be triangulated across multiple linking assumptions before drawing explanatory conclusions.","The framework gives concrete shape to reporting standards for model–cognition studies: state the theory, the adaptation, the linking hypothesis, the contrast, and the comparison set."],"supporting_citations":[{"why":"Provides the running example of progressive matrix analogies as digit matrices and the proof-of-possibility result the framework reinterprets.","marker":"[107]"},{"why":"Supplies the concept of linking hypotheses and the requirement that computational models test theoretically motivated hypotheses.","marker":"[29]"},{"why":"Frames the use of neural networks as models of human language acquisition and motivates constraints on training and evaluation.","marker":"[105]"},{"why":"Exemplifies a broad-coverage foundation model for predicting human cognition that the paper cites to argue broader evidential coverage strengthens alignment.","marker":"[8]"},{"why":"Documents unfaithful chain-of-thought rationales, the evidence that constrains process-trace linking hypotheses.","marker":"[96]"},{"why":"Provides the three-level analysis distinction used to argue a model can align computationally while diverging algorithmically.","marker":"[62]"},{"why":"States the multiple realizability principle, used to argue that behavioral alignment can overstate explanatory equivalence.","marker":"[82]"},{"why":"Provides the developmental checkpoints analysis used in the training and developmental mismatch challenge.","marker":"[88]"}],"fun_headline_variants":["Behavioral mimicry isn't cognition: a framework for AI models","AI models need theory, not just matching human answers","Four-stage test for when AI models count as cognitive","Why behavioral alignment isn't proof of AI cognition","For AI to be cognitive, theory must drive the test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework assumes that adding explicit theory, diagnostic contrasts, and contrastive model comparison can overcome underdetermination and mechanistic opacity—if the same behavior can arise from very different internal computations, even a model that passes all four stages may not reveal the mechanisms of human cognition.","fun_headline_variants_meta":{"raw":{"variants":["Behavioral mimicry isn't cognition: a framework for AI models","AI models need theory, not just matching human answers","Four-stage test for when AI models count as cognitive","Why behavioral alignment isn't proof of AI cognition","For AI to be cognitive, theory must drive the test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00063,"raw_usage":{"total_tokens":2885,"prompt_tokens":897,"completion_tokens":1988,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":1909}},"tokens_in":513,"tokens_out":1988,"duration_ms":12945,"temperature":1.0,"reasoning_tokens":1909,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:12:26.912096+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Find two foundation models with architecturally distinct mechanisms that both pass all four stages on the same human dataset, and show that no manipulation of task, linking hypothesis, or training regime separates their predictions; such a result would show that the contrastive stage cannot identify which computational features are necessary, undermining the framework's explanatory criterion.","supporting_citations":[{"cited_title":"What artiﬁcial neural networks can tell us about human language ac- quisition","cited_arxiv_id":null,"evidence_quote":"Frames the use of neural networks as models of human language acquisition and motivates constraints on training and evaluation."},{"cited_title":"Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting","cited_arxiv_id":null,"evidence_quote":"Documents unfaithful chain-of-thought rationales, the evidence that constrains process-trace linking hypotheses."},{"cited_title":"Psychological predicates","cited_arxiv_id":null,"evidence_quote":"States the multiple realizability principle, used to argue that behavioral alignment can overstate explanatory equivalence."},{"cited_title":"Development of cognitive intelligence in pre-trained language models","cited_arxiv_id":null,"evidence_quote":"Provides the developmental checkpoints analysis used in the training and developmental mismatch challenge."}],"review_version":1}