REVIEW 2 major objections 4 minor 112 references
On the use of foundation models in cognitive science
T0 review · 2 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper argues that behavioral fit alone cannot justify treating foundation models as explanatory cognitive models; alignment becomes meaningful only within explicit theory, diagnostic tasks, and contrastive evaluation.
desk verdict A clear, honest synthesis of the inferential steps between foundation-model outputs and human behavior; the framework is useful even though the strongest necessity claims stay underdetermined. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the four-stage inferential framework, split into an inner alignment loop (Stages 1–3: task adaptation, linking hypothesis, goodness-of-fit evaluation) and an outer contrastive loop (Stage 4: cross-model and manipulation comparison). The linking hypothesis is the load-bearing pivot of the framework: it defines what counts as evidence of alignment by mapping model outputs to human behavioral measures, and the paper argues that a strong fit under one linking hypothesis can vanish under another. The running example is the translation of progressive matrix analogies into symbolic digit matrices [107], which the paper uses to show how each stage changes the interpretation of an alignment result.
What would settle it
Find two foundation models with architecturally distinct mechanisms that both pass all four stages on the same human dataset, and show that no manipulation of task, linking hypothesis, or training regime separates their predictions; such a result would show that the contrastive stage cannot identify which computational features are necessary, undermining the framework's explanatory criterion.
Extended reading notes
Core claim
The paper's central claim is that behavioral alignment justifies treating a foundation model as an explanatory cognitive model only when the evaluation embeds the model in a theory-diagnostic design: adapt the task so model and humans perform functionally equivalent problems, specify a linking hypothesis that maps model outputs to human measures, evaluate fit on theoretically diagnostic contrasts rather than aggregate scores, and compare across candidate models or manipulations to identify which computational features are necessary for the fit. Under this view, a model that simply gets high accuracy or matches average human judgments remains a behavioral proxy. The paper supports the claim by showing how each of four common linking hypotheses—similarity in representational space, surprisal, prompting, and process-trace analysis—carries different theoretical commitments, and by identifying four challenges that constrain alignment claims: theoretical underdetermination, mechanistic opacity, training and developmental mismatch, and population-level variability.
Load-bearing premise
The framework assumes that adding explicit theory, diagnostic contrasts, and contrastive model comparison can overcome underdetermination and mechanistic opacity—if the same behavior can arise from very different internal computations, even a model that passes all four stages may not reveal the mechanisms of human cognition.
Editorial extensions
If this is right
- Single-model demonstrations, however strong, establish at most that a cognitive signature is reproducible, not that the model's mechanisms explain it.
- Evaluation of alignment without diagnostic contrasts—comparing conditions that discriminate between theories—cannot adjudicate between cognitive accounts.
- Developmental alignment claims require more than matching learning curves with children; they need causal manipulations of training regime or data ordering.
- Because linking hypotheses are not neutral, results should be triangulated across multiple linking assumptions before drawing explanatory conclusions.
- The framework gives concrete shape to reporting standards for model–cognition studies: state the theory, the adaptation, the linking hypothesis, the contrast, and the comparison set.
Reading between the lines
- If the framework is right, a large share of published model–behavior benchmarks should be reinterpreted as capability or similarity studies rather than cognitive models, unless they already include diagnostic contrasts and model comparison.
- A testable extension: reanalyzing landmark alignment results under alternative linking hypotheses should change the apparent fit, and the direction of change would reveal which theoretical commitment is doing the work.
- The framework implies that the field's next bottleneck is not larger models but better theory: designing tasks whose contrasts discriminate between computational mechanisms, and reporting null or negative contrasts as informative.
- The four-stage structure could generalize to other 'black box' scientific models, wherever mere behavioral fit risks being mistaken for mechanistic insight.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a perspective article that asks under what conditions behavioral alignment between foundation models (FMs) and human performance can justify treating FMs as explanatory models of cognition. It proposes a four-stage inferential framework: (1) adapting human tasks to model-compatible formats, (2) specifying linking hypotheses that map model outputs to human measures, (3) evaluating behavioral correspondence with attention to diagnostic contrasts, and (4) comparing across candidate models or manipulations. The authors argue that behavioral fit alone is insufficient and that meaningful alignment requires explicit theoretical commitments, theory-diagnostic tasks, and contrastive evaluation. They discuss four linking hypotheses (similarity, surprisal, prompting, process-trace) and four challenges (theoretical underdetermination, mechanistic opacity, training/developmental mismatch, and population variability), then distill the framework into five research guidelines, using the digit-matrix analogical reasoning task as a running example.
Significance. If the framework is adopted, it would provide a common vocabulary and set of standards for a rapidly growing literature that evaluates FMs against human and developmental data. The paper's main contributions are conceptual: it separates task adaptation from linking hypotheses and evaluation, emphasizes diagnostic contrasts over aggregate fit, and insists that alignment is a relation between model, task, linking hypothesis, and theoretical claim rather than a property of the model alone. The treatment is careful and self-consciously hedged: the authors acknowledge multiple realizability, unfaithful chain-of-thought traces, and the correlational status of developmental correspondences. The paper also offers concrete, actionable guidelines and a running example that makes the abstract stages easy to follow. No new empirical validation is provided, but for a perspective article this is appropriate; the value lies in organizing and constraining future practice.
major comments (2)
- [Section 2, Stage 4; Guideline 4] Stage 4 is described as identifying computational features that are 'necessary' to reproduce a behavioral signature, and Guideline 4 repeats this language. Finite contrastive comparisons over a few architectures, scales, ablations, or training regimes can establish at most that a feature is necessary within that particular model family and manipulation set; because of multiple realizability, which the paper itself acknowledges in Section 4.2, they cannot establish that the feature is necessary for any computational account of the behavior, let alone that it corresponds to a human mechanism. I recommend rewording these passages to say 'necessary within the class of models and manipulations under consideration' and adding a sentence that unconditional mechanistic necessity would require additional theoretical constraints beyond contrastive evaluation. This is a local but important precision issue, because the abstract and conclusion also use 'necessary' when describing what the framework can illuminate.
- [Section 3, intro and Section 3.2] The paper reviews four linking hypotheses and explains their commitments, but it does not give the reader guidance for choosing among them beyond saying that appropriateness is 'empirical and task-dependent' in the surprisal section. Since the framework makes the linking hypothesis the central theoretical commitment, I would welcome a brief selection principle, for example, prefer the most proximal mapping that preserves the theoretical construct of interest, and triangulate across at least two linking hypotheses when possible. This would make the framework more actionable.
minor comments (4)
- [References] The reference list contains several formatting artifacts from LaTeX source, such as 'Y u', 'V arma', 'F orty-third', and 'ET AL .' in headings; these should be cleaned before publication.
- [Section 3.3] The term 'proximal' linking hypothesis is used without a definition; a brief gloss (e.g., mapping inputs and outputs directly rather than through internal representations) would help readers who are not familiar with the distal/proximal distinction.
- [Section 5, Guideline 2] The guideline notes that adaptation artifacts can lower observed alignment, but it is equally possible for surface cues to inflate alignment; adding this symmetric warning would make the point more complete.
- [Section 4.4] The discussion of persona-based prompting cites conflicting findings, but does not specify which findings conflict or how a reader should interpret them; one or two concrete examples would strengthen the caution.
Circularity Check
No significant circularity: the paper argues normatively for a four-stage framework using underdetermination and mechanistic opacity as premises, and its self-citations are illustrative rather than load-bearing.
full rationale
The paper is a perspective piece that proposes a methodological framework rather than deriving empirical predictions from fitted parameters or formal equations. Its central claim—that behavioral alignment is scientifically meaningful only when embedded in explicit theoretical commitments, theory-diagnostic tasks, and contrastive evaluation—is argued from conceptual premises: theoretical underdetermination (Section 4.1), mechanistic opacity and multiple realizability (Section 4.2), training/developmental mismatch (Section 4.3), and variability (Section 4.4). The framework is explicitly introduced as a systematization of prior proposals (“Our framework systematizes such proposals” and “we distill and expand guidance from prior commentaries”), which is a legitimate synthesis rather than a disguised derivation. The paper contains many self-citations (e.g., [36], [44], [45], [56], [88], [89], [102], [105], [28], [13], [15]), but they are used as examples of existing FM alignment studies or as background for domain-specific claims, not as premises that force the framework’s conclusions. No equation is equated to another by construction, no fitted parameter is relabeled as a prediction, and no uniqueness theorem or ansatz is imported from the authors’ prior work to forbid alternatives. The skeptical concern that Stage 4’s contrastive comparisons cannot fully establish necessity of computational features under multiple realizability is a substantive validity/cogency limitation of the framework, and the paper itself acknowledges it (“two systems might produce similar outputs while relying on different internal computations”); but this is not a circularity, because the framework’s claim does not reduce to its own definition or to a self-citation chain. The framework is a proposal about how to interpret alignment evidence, and its force depends on the strength of the methodological argument, not on an input-output equivalence. Therefore no circular step is identifiable, and the appropriate score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Model outputs can be meaningfully mapped to human behavioral measures through a linking hypothesis.
- domain assumption Task adaptations can preserve the diagnostic structure of the original human experiment.
- domain assumption Contrastive comparisons and mechanistic interpretability can in principle identify which computational features are necessary for a behavioral signature.
Cite this review
Pith. "Pith review of On the use of foundation models in cognitive science." pith.science (2026). https://pith.science/paper/6ZHX24EP
@misc{pith2026260807812,
author = {Pith},
title = {Pith review of: On the use of foundation models in cognitive science},
year = {2026},
howpublished = {\url{https://pith.science/paper/6ZHX24EP}},
note = {Machine review of arXiv:2608.07812}
}
read the original abstract
A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations of their correspondence to adult performance across a range of cognitive domains, as well as whether aspects of model training track children's cognitive development. However, using FMs as candidate cognitive models poses significant methodological and conceptual challenges. A key question underlies this effort: under what conditions does behavioral alignment justify treating FMs as explanatory models of cognition? In this paper, we articulate a four-stage inferential framework for evaluating FMs as cognitive and developmental models: adapting human experimental tasks to model-compatible formats, specifying linking hypotheses that map model outputs to human measures, evaluating behavioral correspondence, and comparing across candidate models or manipulations. We clarify the role of linking hypotheses in mapping model outputs to human behavioral measures, identify challenges that constrain alignment claims, and propose principles for theory-driven and comparative evaluation. Throughout, we argue that behavioral fit alone is insufficient. Alignment becomes scientifically meaningful only when embedded within explicit theoretical commitments, theory-diagnostic tasks, and systematic contrastive evaluation across candidate models.
Reference graph
Works this paper leans on
-
[1]
Using large language models to simulate multiple humans and replicate human subject studies
Gati V Aher, Rosa I Arriaga, and Adam Tauman Kalai. Using large language models to simulate multiple humans and replicate human subject studies. In International Conference on Machine Learning , pages 337–371. PMLR, 2023
2023
-
[2]
Large language models for mathematical reasoning: Progresses and challenges
Janice Ahn, Rishu V erma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin. Large language models for mathematical reasoning: Progresses and challenges. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop , pages 225–237, 2024
2024
-
[3]
How can the human mind occur in the physical universe? Oxford University Press, 2009
John R Anderson. How can the human mind occur in the physical universe? Oxford University Press, 2009
2009
-
[4]
Claude [large language model]
Anthropic. Claude [large language model]. https://www.anthropic.com, 2026
2026
-
[5]
Transformer networks of human conceptual knowledge
Sudeep Bhatia and Russell Richie. Transformer networks of human conceptual knowledge. Psychological Review, 2022
2022
-
[6]
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle OBrien, Eric Hallahan, Moham- mad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning , pages 2397–2430. PMLR, 2023
2023
-
[7]
Using cognitive psychology to understand gpt-3
Marcel Binz and Eric Schulz. Using cognitive psychology to understand gpt-3. Proceedings of the National Academy of Sciences, 120(6):e2218523120, 2023
2023
-
[8]
A foundation model to predict and capture human cognition
Marcel Binz, Elif Akata, Matthias Bethge, Franziska Brändle, Fred Callaway, Julian Coda-Forno, Peter Dayan, Can Demircan, Maria K Eckstein, Noémi Éltet ˝o, et al. A foundation model to predict and capture human cognition. Nature, pages 1–8, 2025. On the use of Foundation models in cognitive science 9
2025
Show all 112 references
-
[9]
Using bayesian regression to test hypotheses about re- lationships between parameters and covariates in cognitive models
Udo Boehm, Helen Steingroever, and Eric-Jan Wagenmakers. Using bayesian regression to test hypotheses about re- lationships between parameters and covariates in cognitive models. Behavior research methods , 50(3):1248–1269, 2018
2018
-
[10]
Antagonistic ai
Alice Cai, Ian Arawjo, and Elena L Glassman. Antagonistic ai. arXiv preprint arXiv:2402.07350, 2024
2024 arXiv
-
[11]
What one intelligence test measures: a theoretical account of the processing in the raven progressive matrices test
Patricia A Carpenter, Marcel A Just, and Peter Shell. What one intelligence test measures: a theoretical account of the processing in the raven progressive matrices test. Psychological review, 97(3):404, 1990
1990
-
[12]
Word acquisition in neural language models
Tyler A Chang and Benjamin K Bergen. Word acquisition in neural language models. Transactions of the Association for Computational Linguistics, 10:1–16, 2022
2022
-
[13]
Hu, Jing Liu, Jaap Jumelet, Tal Linzen, Aaron Mueller, Candace Ross, Raj Sanjay Shah, Alex Warstadt, Ethan Gotlieb Wilcox, and Adina Williams
Lucas Charpentier, Leshem Choshen, Ryan Cotterell, Mustafa Omer Gul, Michael Y . Hu, Jing Liu, Jaap Jumelet, Tal Linzen, Aaron Mueller, Candace Ross, Raj Sanjay Shah, Alex Warstadt, Ethan Gotlieb Wilcox, and Adina Williams. Findings of the third BabyLM challenge: Accelerating ...
2025
-
[14]
Think deep, not just long: Measuring llm reasoning effort via deep-thinking tokens
Wei-Lin Chen, Liqian Peng, Tian Tan, Chao Zhao, Jianhang Chen, Ziqian Lin, Alec Go, and Y u Meng. Think deep, not just long: Measuring llm reasoning effort via deep-thinking tokens. In F orty-third International Conference on Machine Learning, 2026
2026
-
[15]
Babylm turns 4: Call for papers for the 2026 babylm workshop
Leshem Choshen, Ryan Cotterell, Mustafa Omer Gul, Jaap Jumelet, Tal Linzen, Aaron Mueller, Suchir Salhan, Raj San- jay Shah, Alex Warstadt, and Ethan Gotlieb Wilcox. Babylm turns 4: Call for papers for the 2026 babylm workshop. arXiv preprint arXiv:2602.20092, 2026
2026
-
[16]
Deep neural networks as scientific models
Radoslaw M Cichy and Daniel Kaiser. Deep neural networks as scientific models. Trends in cognitive sciences , 23(4): 305–317, 2019
2019
-
[17]
Cogbench: a large language model walks into a psychology lab
Julian Coda-Forno, Marcel Binz, Jane X Wang, and Eric Schulz. Cogbench: a large language model walks into a psychology lab. arXiv preprint arXiv:2402.18225, 2024
2024 arXiv
-
[18]
Decoding answers before chain-of-thought: Evidence from pre-cot probes and activation steering
Kyle Cox, Darius Kianersi, and Adrià Garriga-Alonso. Decoding answers before chain-of-thought: Evidence from pre-cot probes and activation steering. In Mechanistic Interpretability Workshop at ICML 2026 , 2026
2026
-
[19]
Ai surrogates and illusions of generalizability in cognitive science
MJ Crockett and Lisa Messeri. Ai surrogates and illusions of generalizability in cognitive science. Trends in Cognitive Sciences, 2025
2025
-
[20]
The two disciplines of scientific psychology
Lee Joseph Cronbach. The two disciplines of scientific psychology. American Psychologist, 12:671–684, 1957. URL https://api.semanticscholar.org/CorpusID:144287695
1957
-
[21]
The limitations of large language models for understanding human language and cognition
Christine Cuskley, Rebecca Woods, and Molly Flaherty. The limitations of large language models for understanding human language and cognition. Open Mind, 8:1058–1083, 2024
2024
-
[22]
The cost of thinking is similar between large reasoning models and humans
Andrea Gregor de V arda, Ferdinando Pio DElia, Hope Kean, Andrew Lampinen, and Evelina Fedorenko. The cost of thinking is similar between large reasoning models and humans. Proceedings of the National Academy of Sciences , 122 (47):e2520077122, 2025
2025
-
[23]
Systematic testing of three language models reveals low language accuracy, absence of response stability, and a yes-response bias
Vittoria Dentella, Fritz Günther, and Evelina Leivada. Systematic testing of three language models reveals low language accuracy, absence of response stability, and a yes-response bias. Proceedings of the National Academy of Sciences , 120 (51):e2309583120, 2023
2023
-
[24]
Hlb: Benchmarking llms’ humanlikeness in language use
Xufeng Duan, Bei Xiao, Xuemei Tang, and Zhenguang G Cai. Hlb: Benchmarking llms’ humanlikeness in language use. arXiv preprint arXiv:2409.15890, 2024
2024 arXiv
-
[25]
Linnea Evanson, Y air Lakretz, and Jean-Rémi King. Language acquisition: do children and language models follow similar learning stages? In Findings of the Association for Computational Linguistics: ACL 2023 , pages 12205–12218, 2023
2023
-
[26]
Costa-jussà
Javier Ferrando, Gabriele Sarti, Arianna Bisazza, and Marta R. Costa-jussà. A Primer on the Inner Workings of Transformer-based Language Models, October 2024. URL http://arxiv.org/abs/2405.00208
2024 arXiv
-
[27]
A distributional perspective on word learning in neural language mod- els
Filippo Ficarra, Ryan Cotterell, and Alex Warstadt. A distributional perspective on word learning in neural language mod- els. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technolo...
2025
-
[28]
Bridging the data gap between children and large language models
Michael C Frank. Bridging the data gap between children and large language models. Trends in Cognitive Sciences , 2023. 10 Shah ET AL
2023
-
[29]
Cognitive modeling using artificial intelligence
Michael C Frank and Noah D Goodman. Cognitive modeling using artificial intelligence. Annual Review of Psychology, 77, 2025
2025
-
[30]
Individual differences in artificial neural networks capture individual differences in human behavior
Herrick Fung, N Apurva Ratan Murty, and Dobromir Rahnev. Individual differences in artificial neural networks capture individual differences in human behavior. bioRxiv, pages 2026–02, 2026
2026
-
[31]
Relational reasoning and generalization using nonsymbolic neural networks
Atticus Geiger, Alexandra Carstensen, Michael C Frank, and Christopher Potts. Relational reasoning and generalization using nonsymbolic neural networks. Psychological Review, 130(2):308, 2023
2023
-
[32]
What have we learned about artificial intelligence from studying the brain? Biological cybernetics, 118(1):1–5, 2024
Samuel J Gershman. What have we learned about artificial intelligence from studying the brain? Biological cybernetics, 118(1):1–5, 2024
2024
-
[33]
Gemini 3.1 pro
Google. Gemini 3.1 pro. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/ ,
-
[34]
Olmo: Accelerating the science of language models
Dirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. Olmo: Accelerating the science of language models. In Proceedings of the 62nd Annual Meeting of the Association for Computatio...
2024
-
[35]
On logical inference over brains, behaviour, and artificial neural networks
Olivia Guest and Andrea E Martin. On logical inference over brains, behaviour, and artificial neural networks. Computational Brain & Behavior , 6(2):213–227, 2023
2023
-
[36]
Understanding graphical perception in data visualization through vision-language models
Grace Guo, Jenna Kang, Raj Sanjay Shah, Hanspeter Pfister, and Sashank V arma. Understanding graphical perception in data visualization through vision-language models. In NeurIPS 2024 Workshop on Behavioral Machine Learning , 2024
2024
-
[37]
Individual differences in non-verbal number acuity correlate with maths achievement
Justin Halberda, Michèle MM Mazzocco, and Lisa Feigenson. Individual differences in non-verbal number acuity correlate with maths achievement. Nature, 455(7213):665–668, 2008
2008
-
[38]
A probabilistic Earley parser as a psycholinguistic model
John Hale. A probabilistic Earley parser as a psycholinguistic model. In Second Meeting of the North American Chapter of the Association for Computational Linguistics , 2001. URL https://aclanthology.org/N01-1021
2001
-
[39]
Artificial neural network language models align neurally and behaviorally with humans even after a developmentally realistic amount of training
Eghbal A Hosseini, Martin Schrimpf, Yian Zhang, Samuel Bowman, Noga Zaslavsky, and Evelina Fedorenko. Artificial neural network language models align neurally and behaviorally with humans even after a developmentally realistic amount of training. BioRxiv, pages 2022–10, 2022
2022
-
[40]
Auxiliary task demands mask the capabilities of smaller language models
Jennifer Hu and Michael Frank. Auxiliary task demands mask the capabilities of smaller language models. In First Conference on Language Modeling, 2024
2024
-
[41]
Language models align with human judg- ments on key grammatical constructions
Jennifer Hu, Kyle Mahowald, Gary Lupyan, Anna Ivanova, and Roger Levy. Language models align with human judg- ments on key grammatical constructions. Proceedings of the National Academy of Sciences , 121(36):e2400917121, 2024
2024
-
[42]
In-context analogical reasoning with pre-trained language models
Xiaoyang Hu, Shane Storks, Richard L Lewis, and Joyce Chai. In-context analogical reasoning with pre-trained language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (V olume 1: Long Papers), pages 1953–1969, 2023
1953
-
[43]
thinking traces in large reasoning models: Cognitive cost or performative scaffolding? Proceedings of the National Academy of Sciences , 123(17):e2604554123, 2026
Y ueqing Hu. thinking traces in large reasoning models: Cognitive cost or performative scaffolding? Proceedings of the National Academy of Sciences , 123(17):e2604554123, 2026
2026
-
[44]
The representational geometry of number
Zhimin Hu, Lanhao Niu, and Sashank V arma. The representational geometry of number. arXiv preprint arXiv:2602.06843, 2026
2026
-
[45]
Are more tokens rational? inference-time scaling in language models as adaptive resource rationality
Zhimin Hu, Riya Roshan, and Sashank V arma. Are more tokens rational? inference-time scaling in language models as adaptive resource rationality. arXiv preprint arXiv:2602.10329, 2026
2026
-
[46]
How to evaluate the cognitive abilities of llms
Anna A Ivanova. How to evaluate the cognitive abilities of llms. Nature Human Behaviour, pages 1–4, 2025
2025
-
[47]
How many instructions can llms follow at once? In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025
Daniel Jaroslawicz, Brendan Whiting, Parth Shah, and Karime Maamari. How many instructions can llms follow at once? In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025
2025
-
[48]
Llms can find mathematical reasoning mis- takes by pedagogical chain-of-thought
Zhuoxuan Jiang, Haoyuan Peng, Shanshan Feng, Fan Li, and Dongsheng Li. Llms can find mathematical reasoning mis- takes by pedagogical chain-of-thought. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, pages 3439–3447, 2024
2024
-
[49]
The organization of thinking: What functional brain imaging reveals about the neuroarchitecture of complex cognition
Marcel Adam Just and Sashank V arma. The organization of thinking: What functional brain imaging reveals about the neuroarchitecture of complex cognition. Cognitive, Affective, & Behavioral Neuroscience, 7(3):153–191, 2007
2007
-
[50]
Interpretability of artificial neural network models in artificial intelligence versus neuroscience
Kohitij Kar, Simon Kornblith, and Evelina Fedorenko. Interpretability of artificial neural network models in artificial intelligence versus neuroscience. Nature Machine Intelligence, 4(12):1065–1067, 2022
2022
-
[51]
Psychometric predictive power of large language models
Tatsuki Kuribayashi, Y ohei Oseki, and Timothy Baldwin. Psychometric predictive power of large language models. In Findings of the Association for Computational Linguistics: NAACL 2024 , pages 1983–2005, 2024. On the use of Foundation models in cognitive science 11
2024
-
[52]
Large language models are human-like internally
Tatsuki Kuribayashi, Y ohei Oseki, Souhaib Ben Taieb, Kentaro Inui, and Timothy Baldwin. Large language models are human-like internally. Transactions of the Association for Computational Linguistics , 13:1743–1766, 2025
2025
-
[53]
Language models, like humans, show content effects on reasoning tasks
Andrew K Lampinen, Ishita Dasgupta, Stephanie CY Chan, Hannah R Sheahan, Antonia Creswell, Dharshan Kumaran, James L McClelland, and Felix Hill. Language models, like humans, show content effects on reasoning tasks. PNAS nexus, 3(7):pgae233, 2024
2024
-
[54]
Representation biases: will we achieve complete understanding by analyzing representations? arXiv preprint arXiv:2507.22216, 2025
Andrew Kyle Lampinen, Stephanie CY Chan, Y uxuan Li, and Katherine Hermann. Representation biases: will we achieve complete understanding by analyzing representations? arXiv preprint arXiv:2507.22216, 2025
2025 arXiv
-
[55]
Expectation-based syntactic comprehension
Roger Levy. Expectation-based syntactic comprehension. Cognition, 106(3):1126–1177, 2008. ISSN 0010-0277. . URL https://www.sciencedirect.com/science/article/pii/S0010027707001436
2008
-
[56]
Incremental comprehension of garden-path sentences by large language models: Semantic interpretation, syntactic re-analysis, and attention
Andrew Li, Xianle Feng, Siddhant Narang, Austin Peng, Tianle Cai, Raj Sanjay Shah, and Sashank V arma. Incremental comprehension of garden-path sentences by large language models: Semantic interpretation, syntactic re-analysis, and attention. In Proceedings of the Annual Meeti...
2024
-
[57]
Cogmath: Assessing llms’ authentic mathematical ability from a human cognitive perspective
Jiayu Liu, Zhenya Huang, Wei Dai, Cheng Cheng, Jinze Wu, Jing Sha, Song Li, Qi Liu, Shijin Wang, and Enhong Chen. Cogmath: Assessing llms’ authentic mathematical ability from a human cognitive perspective. In F orty-second International Conference on Machine Learning , 2025
2025
-
[58]
Llm360: Towards fully transparent open-source llms
Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Y uqi Wang, Suqi Sun, Omkar Pangarkar, et al. Llm360: Towards fully transparent open-source llms. In First Conference on Language Modeling, 2023
2023
-
[59]
Response times: Their role in inferring elementary mental organization
R Duncan Luce. Response times: Their role in inferring elementary mental organization . Oxford University Press, 1991
1991
-
[60]
Rethinking thinking tokens: Llms as improvement operators
Lovish Madaan, Aniket Didolkar, Suchin Gururangan, John Quan, Ruan Silva, Ruslan Salakhutdinov, Manzil Zaheer, Sanjeev Arora, and Anirudh Goyal. Rethinking thinking tokens: Llms as improvement operators. arXiv preprint arXiv:2510.01123, 2025
-
[61]
Dissociating language and thought in large language models
Kyle Mahowald, Anna A Ivanova, Idan A Blank, Nancy Kanwisher, Joshua B Tenenbaum, and Evelina Fedorenko. Dissociating language and thought in large language models. Trends in Cognitive Sciences, 2024
2024
-
[62]
Vision: A computational investigation into the human representation and processing of visual information
David Marr. Vision: A computational investigation into the human representation and processing of visual information . MIT press, 2010
2010
-
[63]
Cognitive assessment of language models
Daniel McDuff, David Munday, Xin Liu, and Isaac Galatzer-Levy. Cognitive assessment of language models. In ICML 2024 Workshop on LLMs and Cognition , 2024
2024
-
[64]
How can deep neural networks inform theory in psychological science? 2023
Sam Whitman McGrath, Jacob Russin, Ellie Pavlick, and Roman Feiman. How can deep neural networks inform theory in psychological science? 2023
2023
-
[65]
N-gram-like language models predict reading time best
James A Michaelov and Roger P Levy. N-gram-like language models predict reading time best. arXiv preprint arXiv:2603.09872, 2026
2026
-
[66]
Large language models are able to downplay their cognitive abilities to fit the persona they simulate
Jiˇrí Miliˇcka, Anna Marklová, Klára V anSlambrouck, Eva Pospíšilová, Jana Šimsová, Samuel Harvan, and Ondˇrej Drobil. Large language models are able to downplay their cognitive abilities to fit the persona they simulate. Plos one , 19(3): e0298522, 2024
2024
-
[67]
Interventionist methods for interpreting deep neural networks
Raphaël Millière and Cameron Buckner. Interventionist methods for interpreting deep neural networks. In Neurocogni- tive F oundations of Mind. Routledge, 1st edition, 2025
2025
-
[68]
Do language models learn typicality judgments from text? In Proceedings of the Annual Meeting of the Cognitive Science Society , volume 43, 2021
Kanishka Misra, Allyson Ettinger, and Julia Rayz. Do language models learn typicality judgments from text? In Proceedings of the Annual Meeting of the Cognitive Science Society , volume 43, 2021
2021
-
[69]
Moyer and Thomas K
Robert S. Moyer and Thomas K. Landauer. Time required for judgements of numerical inequality. Nature, 215(5109): 1519–1520, 1967
1967
-
[70]
To model human linguistic prediction, make llms less superhuman
Byung-Doh Oh and Tal Linzen. To model human linguistic prediction, make llms less superhuman. arXiv preprint arXiv:2510.05141, 2025
2025 arXiv
-
[71]
Transformer-based language model surprisal predicts human reading times best with about two billion training tokens
Byung-Doh Oh and William Schuler. Transformer-based language model surprisal predicts human reading times best with about two billion training tokens. In Findings of the association for computational linguistics: EMNLP 2023 , pages 1915–1921, 2023
2023
-
[72]
Desmond Ong. Gpt-ology, computational models, silicon sampling: How should we think about llms in cognitive science? In Proceedings of the Annual Meeting of the Cognitive Science Society , volume 46, 2024
2024
-
[73]
OpenAI. Gpt-5.6. https://openai.com/index/gpt-5-6/, 2026. Accessed: 2026-07-16
2026
-
[74]
So- cial simulacra: Creating populated prototypes for social computing systems
Joon Sung Park, Lindsay Popowski, Carrie Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. So- cial simulacra: Creating populated prototypes for social computing systems. In Proceedings of the 35th Annual ACM 12 Shah ET AL . Symposium on User Interface Softwar...
2022
-
[75]
Temporal aspects of digit and letter inequality judgments
John M Parkman. Temporal aspects of digit and letter inequality judgments. Journal of experimental psychology, 91(2): 191, 1971
1971
-
[76]
Mapping language models to grounded conceptual spaces
Roma Patel and Ellie Pavlick. Mapping language models to grounded conceptual spaces. In International conference on learning representations, 2021
2021
-
[77]
Modern language models refute chomskys approach to language
Steven Piantadosi. Modern language models refute chomskys approach to language. Lingbuzz Preprint, lingbuzz, 7180, 2023
2023
-
[78]
Why concepts are (probably) vectors
Steven T Piantadosi, Dyana CY Muller, Joshua S Rule, Karthikeya Kaushik, Mark Gorenstein, Elena R Leib, and Emily Sanford. Why concepts are (probably) vectors. Trends in Cognitive Sciences, 2024
2024
-
[79]
Conjectures and refutations: The growth of scientific knowledge
Karl Popper. Conjectures and refutations: The growth of scientific knowledge . routledge, 2014
2014
-
[80]
A controlled reevaluation of coreference resolution models
Ian Porada, Xiyuan Zou, and Jackie Chi Kit Cheung. A controlled reevaluation of coreference resolution models. In Pro- ceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 256–263, 2024
2024
-
[81]
Frank, and Gary Lupyan
Eva Portelance, Y uguang Duan, Michael C. Frank, and Gary Lupyan. Predicting age of acquisition for children’s early vocabulary in five languages using language model surprisal. Cognitive science , 47 9:e13334, 2023. URL https://api. semanticscholar.org/CorpusID:261696384
2023
-
[82]
Psychological predicates
Hilary Putnam. Psychological predicates. In W. H. Capitan and D. D. Merrill, editors, Art, Mind, and Religion , pages 37–48. University of Pittsburgh Press, 1967
1967
-
[83]
Does spatial cognition emerge in frontier models? In The Thirteenth International Conference on Learning Representations , 2024
Santhosh Kumar Ramakrishnan, Erik Wijmans, Philipp Kraehenbuehl, and Vladlen Koltun. Does spatial cognition emerge in frontier models? In The Thirteenth International Conference on Learning Representations , 2024
2024
-
[84]
Similarity judgment within and across categories: A comprehensive model comparison
Russell Richie and Sudeep Bhatia. Similarity judgment within and across categories: A comprehensive model comparison. Cognitive science, 45(8):e13030, 2021
2021
-
[85]
Perturbation: A simple and efficient adversarial tracer for representation learning in language models
Joshua Rozner and Cory Shain. Perturbation: A simple and efficient adversarial tracer for representation learning in language models. arXiv preprint arXiv:2603.23821, 2026
2026
-
[86]
Pdp models and general issues in cognitive science
David E Rumelhart and James L McClelland. Pdp models and general issues in cognitive science. In Parallel distributed processing: Explorations in the microstructure of cognition, V ol. 1: F oundations, pages 110–146. 1986
1986
-
[87]
In-context impersonation reveals large language models’ strengths and biases
Leonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz, and Zeynep Akata. In-context impersonation reveals large language models’ strengths and biases. Advances in Neural Information Processing Systems , 36, 2024
2024
-
[88]
Development of cognitive intelligence in pre-trained language models
Raj Shah, Khushi Bhardwaj, and Sashank V arma. Development of cognitive intelligence in pre-trained language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages 9632–9657, 2024
2024
-
[89]
Numeric magnitude comparison effects in large language models, 2023
Raj Sanjay Shah, Vijay Marupudi, Reba Koenen, Khushi Bhardwaj, and Sashank V arma. Numeric magnitude comparison effects in large language models, 2023
2023
-
[90]
Word frequency and predictability dissociate in naturalistic reading
Cory Shain. Word frequency and predictability dissociate in naturalistic reading. Open Mind, 8:177–201, 2024
2024
-
[91]
Open problems in mechanistic interpretability
Lee Sharkey, Bilal Chughtai, Joshua Batson, Jack Lindsey, Jeffrey Wu, Lucius Bushnaq, Nicholas Goldowsky-Dill, Stefan Heimersheim, Alejandro Ortega, Joseph Isaac Bloom, et al. Open problems in mechanistic interpretability. Transactions on Machine Learning Research, 2025
2025
-
[92]
The bitter lesson
Rich Sutton. The bitter lesson. 2019
2019
-
[93]
Devbench: A multimodal developmental benchmark for language learning
Alvin W Tan, Sunny Y u, Bria Long, Wanjing Anya, Tonya Murray, Rebecca D Silverman, Jason D Y eatman, and Michael C Frank. Devbench: A multimodal developmental benchmark for language learning. Advances in Neural Information Processing Systems, 37:77445–77467, 2024
2024
-
[94]
Numerosity discrimination in deep neural networks: Initial competence, developmental refinement and experience statistics
Alberto Testolin, Will Y Zou, and James L McClelland. Numerosity discrimination in deep neural networks: Initial competence, developmental refinement and experience statistics. Developmental science, 23(5):e12940, 2020
2020
-
[95]
Two tales of persona in llms: A survey of role-playing and personalization
Y u-Min Tseng, Y u-Chao Huang, Teng-Y un Hsiao, Wei-Lin Chen, Chao-Wei Huang, Y u Meng, and Y un-Nung Chen. Two tales of persona in llms: A survey of role-playing and personalization. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 16612–16631, 2024
2024
-
[96]
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. Advances in Neural Information Processing Systems , 36:74952– 74965, 2023
2023
-
[97]
Large language models fail on trivial alterations to theory-of-mind tasks
Tomer Ullman. Large language models fail on trivial alterations to theory-of-mind tasks. arXiv preprint arXiv:2302.08399, 2023. On the use of Foundation models in cognitive science 13
2023 arXiv
-
[98]
Individual differences as a crucible in theory construction
Benton J Underwood. Individual differences as a crucible in theory construction. American Psychologist , 30(2):128, 1975
1975
-
[99]
Reclaiming ai as a theoretical tool for cognitive science
Iris V an Rooij, Olivia Guest, Federico Adolfi, Ronald De Haan, Antonina Kolokolova, and Patricia Rich. Reclaiming ai as a theoretical tool for cognitive science. Computational Brain & Behavior , 7(4):616–636, 2024
2024
-
[100]
Correlations without causa- tion do not support claims of human–llm reasoning alignment
Ivan I V ankov, Federico Adolfi, Rachel F Heaton, Guillermo Puebla, and Jeffrey S Bowers. Correlations without causa- tion do not support claims of human–llm reasoning alignment. Proceedings of the National Academy of Sciences , 123 (12):e2536362123, 2026
2026
-
[101]
Capturing human cognitive styles with language: Towards an experimental evaluation paradigm
V asudha V aradarajan, Syeda Mahwish, Xiaoran Liu, Julia Buffolino, Christian C Luhmann, Ryan Boyd, and H Andrew Schwartz. Capturing human cognitive styles with language: Towards an experimental evaluation paradigm. In Proceed- ings of the 2025 Conference of the Nations of the...
2025
-
[102]
How well do deep learning models capture human concepts? the case of the typicality effect
Siddhartha V emuri, Raj Sanjay Shah, and Sashank V arma. How well do deep learning models capture human concepts? the case of the typicality effect. In Proceedings of the Annual Meeting of the Cognitive Science Society , volume 46, 2024
2024
-
[103]
Is a picture worth a thousand words? delving into spatial reasoning for vision language models
Jiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet, Xin Wang, Sharon Li, and Neel Joshi. Is a picture worth a thousand words? delving into spatial reasoning for vision language models. Advances in Neural Information Processing Systems, 37:75392–75421, 2024
2024
-
[104]
Coglm: Tracking cognitive development of large language models
Xinglin Wang, Peiwen Y uan, Shaoxiong Feng, Yiwei Li, Boyuan Pan, Heda Wang, Y ao Hu, and Kan Li. Coglm: Tracking cognitive development of large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational L...
2025
-
[105]
What artificial neural networks can tell us about human language ac- quisition
Alex Warstadt and Samuel R Bowman. What artificial neural networks can tell us about human language ac- quisition. In Shalom Lappin and Jean-Philippe Bernardy, editors, Algebraic Structures in Natural Language , pages 17–60. CRC Press, 2022. URL https://www.taylorfrancis.com/ch...
2022 doi
-
[106]
Blimp: The benchmark of linguistic minimal pairs for english
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R Bowman. Blimp: The benchmark of linguistic minimal pairs for english. Transactions of the Association for Computational Linguistics, 8:377–392, 2020
2020
-
[107]
Emergent analogical reasoning in large language models
Taylor Webb, Keith J Holyoak, and Hongjing Lu. Emergent analogical reasoning in large language models. Nature Human Behaviour, 7(9):1526–1541, 2023
2023
-
[108]
Evidence from counterfactual tasks supports emergent analogical reasoning in large language models
Taylor W Webb, Keith J Holyoak, and Hongjing Lu. Evidence from counterfactual tasks supports emergent analogical reasoning in large language models. PNAS nexus, 4(5):pgaf135, 2025
2025
-
[109]
On the need to improve the way individual differences in cognitive function are measured with reaction time tasks
Corey N White and Kiah N Kitchen. On the need to improve the way individual differences in cognitive function are measured with reaction time tasks. Current directions in psychological science , 31(3):223–230, 2022
2022
-
[110]
In- context learning strategies emerge rationally
Daniel Wurgaft, Ekdeep Singh Lubana, Core Francisco Park, Hidenori Tanaka, Gautam Reddy, and Noah Goodman. In- context learning strategies emerge rationally. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025
2025
-
[111]
Choosing prediction over explanation in psychology: Lessons from machine learning
Tal Y arkoni and Jacob Westfall. Choosing prediction over explanation in psychology: Lessons from machine learning. Perspectives on Psychological Science, 12(6):1100–1122, 2017
2017
-
[2026]
Accessed: 2026-05-09
2026
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.