REVIEW 2 major objections 10 minor 1 cited by
The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories
T0 review · 2 major / 10 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Pre-trained language models can serve as scientific theories of cognition when their performance profiles match human behavior, even without mechanistic explanation.
desk verdict A clearly written position paper that organizes known pitfalls into a usable framework; the load-bearing sufficiency criterion is asserted rather than demonstrated, but the paper is honest about that. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the three-stage mapping between humans and models: (1) adapting human experimental stimuli and task instructions to the model's text-input modality; (2) a linking hypothesis that converts model outputs into human-like performance indices; and (3) a goodness-of-fit comparison against human behavioral data. The paper examines three families of linking hypotheses—similarity computations in latent space, surprisal (negative log probability) values, and direct prompting of the model's response distribution—and argues that no single mapping is universally valid; each must be empirically validated per task. The sufficiency criterion is the engine of the argument: jointly, a reasonable adaptation, a valid linking hypothesis, and a matching performance profile justify using the model to predict human behavior, and this justifies treating prediction as the primary scientific criterion over mechanistic explanation.
What would settle it
Run a model that passes a broad battery of human-aligned cognitive tests, then present it with novel stimuli that preserve the same task structure but fall outside its training distribution while still being natural for humans; if human behavior is predicted substantially worse by the model than by a mechanism-based account, the sufficiency criterion is undercut. A second, sharper falsifier: construct two stimuli that are equally similar in the model's latent space but are judged at very different speeds by humans, contradicting the similarity-based linking hypothesis on a task where the paper's examples rely on that mapping.
Extended reading notes
Core claim
The paper's core claim is a sufficiency argument: a pre-trained language model becomes a usable cognitive or developmental science theory when, after appropriate adaptation of stimuli, a well-chosen linking hypothesis, and a goodness-of-fit comparison, its performance profile matches human performance on the relevant tasks. This holds even though the model is architecturally different, trained on different data, and not interpretable at the level of human mechanisms; for prediction of human behavior, alignment of functional form is the key requirement. From that criterion, the paper derives two utilities—generalizability across tasks and the ability to generate novel candidate theories—and it organizes known hazards into a new taxonomy: errors of commission (bad assumptions in the mapping, opaque mechanisms, aggregate training data) and errors of omission (ignoring psychometric intercorrelations, neural correlates, and developmental progression). The paper is a review and position statement, not an empirical demonstration, but its contribution is the explicit framework for when and how a PLM's output can be treated as evidence about human thinking.
Load-bearing premise
The load-bearing premise is that a match in functional performance profiles between a language model and humans is sufficient for the model to be used as a cognitive science theory, without requiring correspondence at the level of mechanisms or processes.
Editorial extensions
If this is right
- A PLM that matches human performance profiles across enough tasks can be used to predict human behavior on a new task without collecting new human data, provided the same linking-hypothesis validation holds.
- The sufficiency criterion licenses using PLMs as discovery engines, generating candidate cognitive theories from model behavior on novel stimuli, analogous to screening hypotheses in drug discovery.
- Researchers must validate the linking hypothesis separately for each task, since similarity, surprisal, and prompting each have distinct failure modes and context sensitivities.
- Developmental claims should be tested with intermediate training checkpoints, not only final models, so that the sequence of skill acquisition can be compared with children's developmental trajectories.
- Cognitive evaluations of PLMs should measure cross-task correlations in addition to single-task accuracy, to capture psychometric structure rather than modular performance.
Reading between the lines
- If functional alignment is sufficient, model architecture becomes scientifically irrelevant to the validity of a cognitive claim; two structurally different PLMs that both match human profiles would be equally valid theories, a claim that could be tested by comparing diverse model families on the same battery.
- The commission/omission taxonomy implies a concrete operational test: report the full correlation matrix of PLM performance across cognitive tasks and compare it with human psychometric correlation matrices; the paper discusses this as a criterion but does not run it.
- The emphasis on prediction suggests a practical triage rule for the field: when surprisal and prompting disagree, the better predictor of human reaction times or judgments should win, a rule the paper gestures at but does not formalize.
- The developmental argument predicts that curriculum order, not just final data volume, drives whether model checkpoints track child development; this is testable by training on developmentally ordered corpora and checking age-of-acquisition alignment, which the authors propose as future work.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that pre-trained language models (PLMs) can serve as credible theories in cognitive and developmental science when their performance profiles match human performance under a three-stage mapping: adaptation of human stimuli to text-compatible inputs, a linking hypothesis from model outputs to human behavioral measures, and comparison of model and human performance. The authors advocate a sufficiency criterion under which, given reasonable assumptions about these stages, matching functional profiles justifies using PLMs to predict human behavior, and they explicitly prioritize prediction over explanation. The paper reviews pitfalls, organizing them into pitfalls of commission (distal linking hypotheses, opacity and limited interpretability, training-data and stimulus mismatches, non-determinism) and pitfalls of omission (architectural differences, neglect of psychometric correlations, developmental trajectory issues, missing neuroscience context). It critically examines three families of linking hypotheses, namely latent-space similarity, surprisal, and prompting, and then enumerates two sets of criteria: appropriateness criteria (multi-experiment triangulation, multi-method interpretation, path-dependency testing, tuning controls, documented linking assumptions, cross-task correlations, embodiment) and development criteria (developmentally plausible corpora, tuning on core cognitive tasks).
Significance. The paper's main value is organizational and methodological. The commission/omission taxonomy, the three-stage mapping framework, and the explicit criteria give researchers a concrete checklist for designing and interpreting PLM-human alignment studies, and the review of linking hypotheses is grounded in specific failure cases (tokenization mismatches in Section 4, surprisal's underestimated garden-path costs, prompt-format sensitivity). The recommendation to evaluate cross-task correlation patterns (Section 5.1) is a falsifiable, implementable proposal, and the parallel between functional mapping in neuroscience and mechanistic interpretability is generative. The authors deserve credit for an unusually transparent limitations section (Section 7), which correctly flags that the criteria are asserted rather than demonstrated. The principal weakness is that the sufficiency criterion is under-specified on validation: as written it licenses prediction from benchmark fit, and the paper never defines what would count as successful or failed prediction.
major comments (2)
- [Section 2 (boxed Sufficiency criterion); Sections 5.1 and 7] The boxed sufficiency criterion states that, under reasonable assumptions about the three stages, PLMs are justified in predicting human behavior once their performance profiles match those of humans, and the paper adds that prediction is more important than explanation. As stated, this licenses prediction on the strength of fitting a finite set of benchmark tasks, which is calibration rather than prediction: a model can match a human performance profile on sampled tasks for non-cognitive reasons, as the paper's own examples show (Section 4's tokenization mismatch for 'Nine', prompt-format sensitivity, and Section 5.1's mention of pre-training data contamination). The Section 5.1 recommendations to 'Design multiple experiments' and to 'Establish task correlations' are steps toward stronger evidence, but the cross-task correlational benchmark is still a fit-based check rather than an explicit out-of-sample test. I recommend amending the sufficiency criterion to require preregistered predictions on held-out stimuli or new task framings before a PLM is credited with predictive utility, and to define what outcome would count as falsification. This matters because the paper itself states in Section 7, item 3, that no experiments are offered; without an explicit validation condition, the central claim remains an assertion rather than an operationalized proposal.
- [Section 5.2 (second bullet); Section 5.1 (Control for tuning techniques)] There is a tension between Section 5.1's 'Control for tuning techniques' criterion, which warns that tuning objectives (instruction tuning, RLHF) can make model behavior reflect the tuning goals rather than representational fidelity, and the Section 5.2 suggestion to preference-tune PLMs on 'core' cognitive tasks such as typicality experiments and then evaluate on a broader set of tasks. If the model is tuned on a subset of the very benchmarks used to establish cognitive alignment, alignment on the remaining tasks may partly reflect the tuning objective rather than the model's natural inductive biases, which is precisely the confound the paper warns about elsewhere. The recommendation should specify how the tuning stage is to be controlled (for example, held-out task families or pretraining-only baselines), or it should be framed as a hypothesis to be tested rather than as a practice to adopt without caveats.
minor comments (10)
- [Section 2, boxed sufficiency criterion] In the three-stage list inside the boxed sufficiency criterion, the third stage is numbered '(2)', duplicating the second stage; it should be numbered '(3)'.
- [Table 1] The contribution header cells contain garbled icon-like characters (a trophy symbol and a malformed 'exclamation-triangle' token) that should be replaced with clean text or removed.
- [Section 2] The sentence 'These three stages are depicted in Figure 1)' contains a stray closing parenthesis after the figure reference.
- [Section 4, Surprisal values] The term 'surprisas' should be 'surprisal' or 'surprisal values'.
- [Section 4, Surprisal values] 'Shain (2024) use PLMs to demonstrate strong surprisal predictability estimates' has a subject-verb agreement error; it should be 'uses'.
- [Section 3, Challenges of Development] 'its progression over-development' should read 'over development' without the hyphen.
- [Section 3] The heading styles 'Pitfalls of the commission' and 'Pitfalls of Commission' are inconsistent; one style should be used throughout.
- [Section 3, Pitfalls of Commission] The assertion that mechanistic interpretability is 'the wrong level of analysis' for capturing the emergent, contextual, and symbolic aspects of human thought is made without citation; the authors should either cite supporting arguments or frame the claim as a contested position.
- [Sections 1 and 3] The capitalization 'PLMS' appears alongside 'PLMs' in several places; standardize to 'PLMs'.
- [Section 4, Similarity computations] In the typicality exposition, the sentence 'the typicality of an exemplar is commonly defined as the proportion of humans that produce it' conflates the human production measure with the model-based similarity estimate; clarify that production frequency is the human measure and similarity-to-prototype is the linking hypothesis.
Circularity Check
No significant circularity: the paper is a theoretical position piece whose recommendations are explicit arguments, not derivations reduced to fitted inputs or to load-bearing self-citations.
full rationale
The paper contains no equations, fitted parameters, trained models, or quantitative derivations, and Section 7 explicitly states: "our work is theoretical and does not conduct experiments or offer empirical evidence of performance comparisons or other quantitative measures." The central "Sufficiency criterion" is an explicit argument that if the three stated stages of adaptation, linking hypothesis, and comparison are reasonable and performance profiles match, then PLMs may be used to predict human behavior; this is a normative proposal with stated assumptions, not a result derived from itself. The authors' self-citations (Shah et al. 2023; Shah et al. 2024; Vemuri et al. 2024; Bhardwaj et al. 2024; Li et al. 2024; Guo et al. 2024; Sharma et al. 2024) appear as illustrative examples of PLM-human alignment research, and the paper's recommended criteria do not depend on the truth of those specific empirical results. No uniqueness theorem is imported from the authors' prior work, no ansatz is smuggled in via a self-citation, and no known result is merely renamed. The concern that the sufficiency criterion may need stronger out-of-sample validation is a scientific correctness risk, not a circularity: the criterion is asserted as an argument rather than claimed to follow from its own conclusion.
Assumptions & free parameters
assumptions (4)
- domain assumption The sufficiency criterion: if model and human performance profiles match under reasonable linking hypotheses, the PLM can be used to predict human behavior.
- domain assumption The three-stage mapping (stimulus adaptation, linking hypothesis, comparison to human performance) is the correct decomposition of PLM-human comparisons.
- domain assumption Prediction is more important than explanation when applying machine learning models to cognitive phenomena.
- domain assumption Marr's three levels of analysis provide a valid framework for comparing brains and PLMs.
Cite this review
Pith. "Pith review of The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories." pith.science (2026). https://pith.science/paper/G7CVWVL7
@misc{pith2026250112651,
author = {Pith},
title = {Pith review of: The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories},
year = {2026},
howpublished = {\url{https://pith.science/paper/G7CVWVL7}},
note = {Machine review of arXiv:2501.12651}
}
read the original abstract
Many studies have evaluated the cognitive alignment of Pre-trained Language Models (PLMs), i.e., their correspondence to adult performance across a range of cognitive domains. Recently, the focus has expanded to the developmental alignment of these models: identifying phases during training where improvements in model performance track improvements in children's thinking over development. However, there are many challenges to the use of PLMs as cognitive science theories, including different architectures, different training data modalities and scales, and limited model interpretability. In this paper, we distill lessons learned from treating PLMs, not as engineering artifacts but as cognitive science and developmental science models. We review assumptions used by researchers to map measures of PLM performance to measures of human performance. We identify potential pitfalls of this approach to understanding human thinking, and we end by enumerating criteria for using PLMs as credible accounts of cognition and cognitive development.
Figures
Forward citations
Cited by 1 Pith paper
-
Human-Like Anaphor Resolution in Large Language Models
Some open-weight LLMs show human-like sensitivity to distance and discourse prominence in anaphor resolution, but weaker sensitivity to semantic interference.
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Gati V Aher, Rosa I Arriaga, and Adam Tauman Kalai. 2023. Using large language models to simulate multiple humans and replicate human subject studies. In International Conference on Machine Learning, pages 337--371. PMLR
2023
-
[4]
Janice Ahn, Rishu Verma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin. 2024. Large language models for mathematical reasoning: Progresses and challenges. arXiv preprint arXiv:2402.00157
arXiv 2024
-
[5]
John R Anderson. 2009. How can the human mind occur in the physical universe? Oxford University Press
2009
-
[6]
Khai Loong Aw, Syrielle Montariol, Badr AlKhamissi, Martin Schrimpf, and Antoine Bosselut. 2023. Instruction-tuning aligns llms to the human brain. In First Conference on Language Modeling
2023
-
[7]
Catarina G Belem, Markelle Kelly, Mark Steyvers, Sameer Singh, and Padhraic Smyth. 2024. Perceptions of linguistic uncertainty by language models and humans. arXiv preprint arXiv:2407.15814
arXiv 2024
-
[8]
Yoshua Bengio, J \'e r \^o me Louradour, Ronan Collobert, and Jason Weston. 2009. Curriculum learning. In Proceedings of the 26th annual international conference on machine learning, pages 41--48
2009
Show all 112 references
-
[9]
Khushi Bhardwaj, Raj Sanjay Shah, and Sashank Varma. 2024. https://arxiv.org/abs/2311.04666 Pre-training llms using human-like development data corpus . Preprint, arXiv:2311.04666
2024 arXiv
-
[10]
Sudeep Bhatia and Russell Richie. 2022. Transformer networks of human conceptual knowledge. Psychological Review
2022
-
[11]
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. 2023. Pythia: A suite for analyzing large language models across training and scaling. In ...
2023
-
[12]
Marcel Binz, Elif Akata, Matthias Bethge, Franziska Br \"a ndle, Fred Callaway, Julian Coda-Forno, Peter Dayan, Can Demircan, Maria K Eckstein, No \'e mi \'E ltet o , et al. 2024. Centaur: a foundation model of human cognition. arXiv preprint arXiv:2410.20268
2024 arXiv
-
[13]
Abeba Birhane and Marek McGann. 2024. Large models of what? mistaking engineering achievements for human linguistic agency. Language Sciences, 106:101672
2024
-
[14]
Leo Breiman. 2003. Statistical modeling: The two cultures. Quality control and applied statistics, 48(1):81--82
2003
-
[15]
Danilo Bzdok, Andrew Thieme, Oleksiy Levkovskyy, Paul Wren, Thomas Ray, and Siva Reddy. 2024. Data science opportunities of large language models for neuroscience and biomedicine. Neuron, 112(5):698--717
2024
-
[16]
Alice Cai, Ian Arawjo, and Elena L Glassman. 2024. Antagonistic ai. arXiv preprint arXiv:2402.07350
2024 arXiv
-
[17]
Chen-Chi Chang, Ching-Yuan Chen, Hung-Shin Lee, and Chih-Cheng Lee. 2024. Benchmarking cognitive domains for llms: Insights from taiwanese hakka culture. arXiv preprint arXiv:2409.01556
2024 arXiv
-
[18]
Anthony Chemero. 2023. Llms differ from human cognition because they are not embodied. Nature Human Behaviour, 7(11):1828--1829
2023
-
[19]
Julian Coda-Forno, Marcel Binz, Jane X Wang, and Eric Schulz. 2024. Cogbench: a large language model walks into a psychology lab. arXiv preprint arXiv:2402.18225
2024 arXiv
-
[20]
Lee Joseph Cronbach. 1957. https://api.semanticscholar.org/CorpusID:144287695 The two disciplines of scientific psychology. American Psychologist, 12:671--684
1957
-
[21]
Christine Cuskley, Rebecca Woods, and Molly Flaherty. 2024. The limitations of large language models for understanding human language and cognition. Open Mind, 8:1058--1083
2024
-
[22]
Dorottya Demszky, Diyi Yang, David S Yeager, Christopher J Bryan, Margarett Clapper, Susannah Chandhok, Johannes C Eichstaedt, Cameron Hecht, Jeremy Jamieson, Meghann Johnson, et al. 2023. Using large language models in psychology. Nature Reviews Psychology, 2(11):688--701
2023
-
[23]
Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, Eric P Xing, and Zhiting Hu. 2022. Rlprompt: Optimizing discrete text prompts with reinforcement learning. arXiv preprint arXiv:2205.12548
2022 arXiv
-
[24]
Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray. 2023. Can ai language models replace human participants? Trends in Cognitive Sciences, 27(7):597--600
2023
-
[25]
Yijiang River Dong, Tiancheng Hu, and Nigel Collier. 2024. Can llm be a personalized judge? arXiv preprint arXiv:2406.11657
2024 arXiv
-
[26]
Xufeng Duan, Bei Xiao, Xuemei Tang, and Zhenguang G Cai. 2024. Hlb: Benchmarking llms' humanlikeness in language use. arXiv preprint arXiv:2409.15890
2024 arXiv
-
[27]
Jeffrey L Elman. 1996. Rethinking innateness: A connectionist perspective on development, volume 10. MIT press
1996
-
[28]
Linnea Evanson, Yair Lakretz, and Jean-R \'e mi King. 2023. Language acquisition: do children and language models follow similar learning stages? arXiv preprint arXiv:2306.03586
2023 arXiv
-
[29]
Michael C Frank. 2023. Bridging the data gap between children and large language models. Trends in Cognitive Sciences
2023
-
[30]
Jakob Gawlikowski, Cedrique Rovile Njieutcheu Tassi, Mohsin Ali, Jongseok Lee, Matthias Humt, Jianxiang Feng, Anna Kruspe, Rudolph Triebel, Peter Jung, Ribana Roscher, et al. 2023. A survey of uncertainty in deep neural networks. Artificial Intelligence Review, 56(Suppl 1):1513--1589
2023
-
[31]
Gemini Team . 2023. https://arxiv.org/abs/2312.11805 Gemini: A family of highly capable multimodal models . Preprint, arXiv:2312.11805
2023 arXiv
-
[32]
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, A. Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, T...
2024
-
[33]
Grace Guo, Jenna Kang, Raj Sanjay Shah, Hanspeter Pfister, and Sashank Varma. 2024. Understanding graphical perception in data visualization through vision-language models. In NeurIPS 2024 Workshop on Behavioral Machine Learning
2024
-
[34]
Justin Halberda, Mich \`e le MM Mazzocco, and Lisa Feigenson. 2008. Individual differences in non-verbal number acuity correlate with maths achievement. Nature, 455(7213):665--668
2008
-
[35]
John Hale. 2001. https://aclanthology.org/N01-1021 A probabilistic E arley parser as a psycholinguistic model . In Second Meeting of the North A merican Chapter of the Association for Computational Linguistics
2001
-
[36]
a m \"a l \
Perttu H \"a m \"a l \"a inen, Mikke Tavast, and Anton Kunnari. 2023. Evaluating large language models in generating synthetic hci research data: a case study. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1--19
2023
-
[37]
Mathew Hardy, Ilia Sucholutsky, Bill Thompson, and Tom Griffiths. 2023. Large language models meet cognitive science: Llms as tools, models, and participants. In Proceedings of the annual meeting of the cognitive science society, volume 45
2023
-
[38]
Eghbal A Hosseini, Martin Schrimpf, Yian Zhang, Samuel Bowman, Noga Zaslavsky, and Evelina Fedorenko. 2022. Artificial neural network language models align neurally and behaviorally with humans even after a developmentally realistic amount of training. BioRxiv, pages 2022--10
2022
-
[39]
Jennifer Hu and Michael C Frank. 2024. Auxiliary task demands mask the capabilities of smaller language models. arXiv preprint arXiv:2404.02418
2024 arXiv
-
[40]
Jennifer Hu, Kyle Mahowald, Gary Lupyan, Anna Ivanova, and Roger Levy. 2024 a . Language models align with human judgments on key grammatical constructions. Proceedings of the National Academy of Sciences, 121(36):e2400917121
2024
-
[41]
Michael Y Hu, Aaron Mueller, Candace Ross, Adina Williams, Tal Linzen, Chengxu Zhuang, Ryan Cotterell, Leshem Choshen, Alex Warstadt, and Ethan Gotlieb Wilcox. 2024 b . Findings of the second babylm challenge: Sample-efficient pretraining on developmentally plausible corpora. ...
2024 arXiv
-
[42]
Xiaoyang Hu, Shane Storks, Richard L Lewis, and Joyce Chai. 2023. In-context analogical reasoning with pre-trained language models. arXiv preprint arXiv:2305.17626
2023 arXiv
-
[43]
Kuan-Jung Huang, Suhas Arehalli, Mari Kugemoto, Christian Muxica, Grusha Prasad, Brian Dillon, and Tal Linzen. 2024. Large-scale benchmark yields no evidence that language model surprisal explains syntactic disambiguation difficulty. Journal of Memory and Language, 137:104510
2024
-
[44]
Huebner, Elior Sulem, Fisher Cynthia, and Dan Roth
Philip A. Huebner, Elior Sulem, Fisher Cynthia, and Dan Roth. 2021. https://doi.org/10.18653/v1/2021.conll-1.49 B aby BERT a: Learning more grammar with small-scale child-directed language . In Proceedings of the 25th Conference on Computational Natural Language Learning, page...
2021 doi
-
[45]
Anna A Ivanova. 2023. Running cognitive evaluations on large language models: The do's and the don'ts. arXiv preprint arXiv:2312.01276
2023 arXiv
-
[46]
Anna A Ivanova, Aalok Sathe, Benjamin Lipkin, Evelina Fedorenko, and Jacob Andreas. 2024 a . Log probability scores provide a closer match to human plausibility judgments than prompt-based evaluations
2024
-
[47]
Anna A Ivanova, Aalok Sathe, Benjamin Lipkin, Unnathi Kumar, Setayesh Radkani, Thomas H Clark, Carina Kauf, Jennifer Hu, RT Pramod, Gabriel Grand, et al. 2024 b . Elements of world knowledge (ewok): A cognition-inspired framework for evaluating basic world knowledge in languag...
2024 arXiv
-
[48]
Zhuoxuan Jiang, Haoyuan Peng, Shanshan Feng, Fan Li, and Dongsheng Li. 2024. Llms can find mathematical reasoning mistakes by pedagogical chain-of-thought. arXiv preprint arXiv:2405.06705
2024 arXiv
-
[49]
Kohitij Kar, Simon Kornblith, and Evelina Fedorenko. 2022. Interpretability of artificial neural network models in artificial intelligence versus neuroscience. Nature Machine Intelligence, 4(12):1065--1067
2022
-
[50]
Carina Kauf, Anna A Ivanova, Giulia Rambelli, Emmanuele Chersoni, Jingyuan Selena She, Zawad Chowdhury, Evelina Fedorenko, and Alessandro Lenci. 2023. Event knowledge in large language models: the gap between the impossible and the unlikely. Cognitive Science, 47(11):e13386
2023
-
[51]
Eliza Kosoy, Emily Rose Reagan, Leslie Lai, Alison Gopnik, and Danielle Krettek Cobb. 2023. Comparing machines and children: Using developmental psychology experiments to assess the strengths and weaknesses of lamda responses. arXiv preprint arXiv:2305.11243
2023 arXiv
-
[52]
Thomas L Griffiths, Charles Kemp, and Joshua B Tenenbaum. 2008. Bayesian models of cognition
2008
-
[53]
Roger Levy. 2008. https://doi.org/10.1016/j.cognition.2007.05.006 Expectation-based syntactic comprehension . Cognition, 106(3):1126--1177
2008 doi
-
[54]
Andrew Li, Xianle Feng, Siddhant Narang, Austin Peng, Tianle Cai, Raj Sanjay Shah, and Sashank Varma. 2024. Incremental comprehension of garden-path sentences by large language models: Semantic interpretation, syntactic re-analysis, and attention
2024
-
[55]
Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Yuqi Wang, Suqi Sun, Omkar Pangarkar, et al. 2023. Llm360: Towards fully transparent open-source llms. arXiv preprint arXiv:2312.06550
2023 arXiv
-
[56]
Kyle Mahowald, Anna A Ivanova, Idan A Blank, Nancy Kanwisher, Joshua B Tenenbaum, and Evelina Fedorenko. 2024. Dissociating language and thought in large language models. Trends in Cognitive Sciences
2024
-
[57]
age effects in second language acquisition: Expanding the emergentist account
Viorica Marian. 2023. Studying second language acquisition in the age of large language models: Unlocking the mysteries of language and learning, a commentary on “age effects in second language acquisition: Expanding the emergentist account” by catherine l. caldwell-harris and...
2023
-
[58]
David Marr. 2010. Vision: A computational investigation into the human representation and processing of visual information. MIT press
2010
-
[59]
Sam Whitman McGrath, Jacob Russin, Ellie Pavlick, and Roman Feiman. 2023. How can deep neural networks inform theory in psychological science?
2023
-
[60]
Sean McGrath, Parth Mehta, Alexandra Zytek, Isaac Lage, and Himabindu Lakkaraju. 2020. When does uncertainty matter?: Understanding the impact of predictive uncertainty in ml assisted decision making. arXiv preprint arXiv:2011.06167
2020 arXiv
-
[61]
Ji r \' Mili c ka, Anna Marklov \'a , Kl \'a ra VanSlambrouck, Eva Posp \' s ilov \'a , Jana S imsov \'a , Samuel Harvan, and Ond r ej Drobil. 2024. Large language models are able to downplay their cognitive abilities to fit the persona they simulate. Plos one, 19(3):e0298522
2024
-
[62]
Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2024. Large language models: A survey. arXiv preprint arXiv:2402.06196
2024 arXiv
-
[63]
Kanishka Misra, Allyson Ettinger, and Julia Taylor Rayz. 2021. Do language models learn typicality judgments from text? arXiv preprint arXiv:2105.02987
2021 arXiv
-
[64]
Moyer and Thomas K
Robert S. Moyer and Thomas K. Landauer. 1967. Time required for judgements of numerical inequality. Nature, 215(5109):1519--1520
1967
-
[65]
Qian Niu, Junyu Liu, Ziqian Bi, Pohsun Feng, Benji Peng, and Keyu Chen. 2024. Large language models and cognitive science: A comprehensive review of similarities, differences, and challenges. arXiv preprint arXiv:2409.02387
2024
-
[66]
Desmond C Ong. 2024. Gpt-ology, computational models, silicon sampling: How should we think about llms in cognitive science? arXiv preprint arXiv:2406.09464
2024 arXiv
-
[67]
OpenAI. 2023. https://arxiv.org/abs/2303.08774 Gpt-4 technical report . Preprint, arXiv:2303.08774
2023 arXiv
-
[68]
Joon Sung Park, Lindsay Popowski, Carrie Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2022. Social simulacra: Creating populated prototypes for social computing systems. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Techno...
2022
-
[69]
John M Parkman. 1971. Temporal aspects of digit and letter inequality judgments. Journal of experimental psychology, 91(2):191
1971
-
[70]
Roma Patel and Ellie Pavlick. 2021. Mapping language models to grounded conceptual spaces. In International conference on learning representations
2021
-
[71]
Steven Piantadosi. 2023. Modern language models refute chomsky’s approach to language. Lingbuzz Preprint, lingbuzz, 7180
2023
-
[72]
Steven T Piantadosi, Dyana CY Muller, Joshua S Rule, Karthikeya Kaushik, Mark Gorenstein, Elena R Leib, and Emily Sanford. 2024. Why concepts are (probably) vectors. Trends in Cognitive Sciences
2024
-
[73]
Frank, and Gary Lupyan
Eva Portelance, Yuguang Duan, Michael C. Frank, and Gary Lupyan. 2023. https://api.semanticscholar.org/CorpusID:261696384 Predicting age of acquisition for children's early vocabulary in five languages using language model surprisal . Cognitive science, 47 9:e13334
2023
-
[74]
Yujin Potter, Shiyang Lai, Junsol Kim, James Evans, and Dawn Song. 2024. Hidden persuaders: Llms' political leaning and their influence on voters. arXiv preprint arXiv:2410.24190
2024 arXiv
-
[75]
Santhosh Kumar Ramakrishnan, Erik Wijmans, Philipp Kraehenbuehl, and Vladlen Koltun. 2024. Does spatial cognition emerge in frontier models? arXiv preprint arXiv:2410.06468
2024 arXiv
-
[76]
Giulia Rambelli, Emmanuele Chersoni, Davide Testa, Philippe Blache, and Alessandro Lenci. 2024. Neural generative models and the parallel architecture of language: A critical review and outlook. Topics in cognitive science
2024
-
[77]
Eleanor Rosch. 1975. Cognitive representations of semantic categories. Journal of Experimental Psychology: General, 104(3):192
1975
-
[78]
Leonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz, and Zeynep Akata. 2024. In-context impersonation reveals large language models' strengths and biases. Advances in Neural Information Processing Systems, 36
2024
-
[79]
Vinay Samuel, Henry Peng Zou, Yue Zhou, Shreyas Chaudhari, Ashwin Kalyan, Tanmay Rajpurohit, Ameet Deshpande, Karthik Narasimhan, and Vishvak Murahari. 2024. Personagym: Evaluating persona agents and llms. arXiv preprint arXiv:2407.18416
2024 arXiv
-
[80]
W Joel Schneider and Kevin S McGrew. 2012. The cattell-horn-carroll model of intelligence
2012
-
[81]
Andreas Schuller, Doris Janssen, Julian Blumenr \"o ther, Theresa Maria Probst, Michael Schmidt, and Chandan Kumar. 2024. Generating personas using llms and assessing their viability. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, pages 1--7
2024
-
[82]
Raj Shah, Khushi Bhardwaj, and Sashank Varma. 2024. Development of cognitive intelligence in pre-trained language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 9632--9657
2024
-
[83]
Raj Sanjay Shah, Vijay Marupudi, Reba Koenen, Khushi Bhardwaj, and Sashank Varma. 2023. https://arxiv.org/abs/2305.10782 Numeric magnitude comparison effects in large language models . Preprint, arXiv:2305.10782
2023 arXiv
-
[84]
Cory Shain. 2024. Word frequency and predictability dissociate in naturalistic reading. Open Mind, 8:177--201
2024
-
[85]
Mihir Sharma, Ryan Ding, Raj Sanjay Shah, and Sashank Varma. 2024. Monolingual and bilingual language acquisition in language models
2024
-
[86]
Roger N Shepard and Peter Podgorny. 1978. Cognitive processes that resemble perceptual processes. In Handbook of learning and cognitive processes: Vol. 5. Human information processing, pages 189--237. Erlbaum Hillsdale, NJ
1978
-
[87]
Linda Smith and Michael Gasser. 2005. The development of embodied cognition: Six lessons from babies. Artificial life, 11(1-2):13--29
2005
-
[88]
Richard E Snow, Patrick C Kyllonen, Brachia Marshalek, et al. 1984. The topography of ability and learning correlations. Advances in the psychology of human intelligence, 2(S 47):103
1984
-
[89]
Alvin Wei Ming Tan, Sunny Yu, Bria Long, Wanjing Anya Ma, Tonya Murray, Rebecca D Silverman, Jason D Yeatman, and Michael C Frank. 2024. Devbench: A multimodal developmental benchmark for language learning. arXiv preprint arXiv:2406.10215
2024 arXiv
-
[90]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. https://arxiv.org/abs/2302.13971 Llama:...
2023 arXiv
-
[91]
Yu-Min Tseng, Yu-Chao Huang, Teng-Yun Hsiao, Yu-Ching Hsu, Jia-Yin Foo, Chao-Wei Huang, and Yun-Nung Chen. 2024. Two tales of persona in llms: A survey of role-playing and personalization. arXiv preprint arXiv:2406.01171
2024 arXiv
-
[92]
Peter D. Turney. 2013. https://doi.org/10.1162/tacl_a_00233 Distributional semantics beyond words: Supervised learning of analogy and paraphrase . Transactions of the Association for Computational Linguistics, 1:353--366
2013 doi
-
[93]
Peter D Turney and Michael L Littman. 2005. Corpus-based learning of analogies and semantic relations. Machine Learning, 60:251--278
2005
-
[94]
Marten Van Schijndel and Tal Linzen. 2021. Single-stage prediction models do not explain the magnitude of syntactic disambiguation difficulty. Cognitive science, 45(6):e12988
2021
-
[95]
Sashank Varma. 2011. Criteria for the design and evaluation of cognitive architectures. Cognitive science, 35(7):1329--1351
2011
-
[96]
Manoj Chowdary Vattikuti. 2024. Improving drug discovery and development using ai: Opportunities and challenges. Research-gate journal, 10(10)
2024
-
[97]
H \'e ctor Javier V \'a zquez Mart \' nez. 2021. https://doi.org/10.18653/v1/2021.blackboxnlp-1.38 The acceptability delta criterion: Testing knowledge of language using the gradience of sentence acceptability . In Proceedings of the Fourth BlackboxNLP Workshop on Analyzing an...
2021 doi
-
[98]
H \'e ctor Javier V \'a zquez Mart \' nez, Annika Heuser, Charles Yang, and Jordan Kodner. 2023. https://doi.org/10.18653/v1/2023.genbench-1.4 Evaluating neural language models as cognitive models of language acquisition . In Proceedings of the 1st GenBench Workshop on (Benchm...
2023 doi
-
[99]
Siddhartha K Vemuri, Raj Sanjay Shah, and Sashank Varma. 2024. How well do deep learning models capture human concepts? the case of the typicality effect
2024
-
[100]
Xinglin Wang, Peiwen Yuan, Shaoxiong Feng, Yiwei Li, Boyuan Pan, Heda Wang, Yao Hu, and Kan Li. 2024. Coglm: Tracking cognitive development of large language models. arXiv preprint arXiv:2408.09150
2024 arXiv
-
[101]
Alex Warstadt and Samuel R. Bowman. 2024. https://arxiv.org/abs/2208.07998 What artificial neural networks can tell us about human language acquisition . Preprint, arXiv:2208.07998
2024 arXiv
-
[102]
Alex Warstadt, Aaron Mueller, Leshem Choshen, Ethan Wilcox, Chengxu Zhuang, Juan Ciro, Rafael Mosquera, Bhargavi Paranjabe, Adina Williams, Tal Linzen, and Ryan Cotterell, editors. 2023. https://aclanthology.org/2023.conll-babylm.0 Proceedings of the BabyLM Challenge at the 27...
2023
-
[103]
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R Bowman. 2020. Blimp: The benchmark of linguistic minimal pairs for english. Transactions of the Association for Computational Linguistics, 8:377--392
2020
-
[104]
Taylor Webb, Keith J Holyoak, and Hongjing Lu. 2023. Emergent analogical reasoning in large language models. Nature Human Behaviour, 7(9):1526--1541
2023
-
[105]
Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022. Emergent abilities of large language mod...
2022 arXiv
-
[106]
Benjue Weng. 2024. Navigating the landscape of large language models: A comprehensive review and analysis of paradigms and fine-tuning strategies. arXiv preprint arXiv:2404.09022
2024 arXiv
-
[107]
Haoran Yang, Yumeng Zhang, Jiaqi Xu, Hongyuan Lu, Pheng Ann Heng, and Wai Lam. 2024. Unveiling the generalization power of fine-tuned large language models. arXiv preprint arXiv:2403.09162
2024 arXiv
-
[108]
Tal Yarkoni and Jacob Westfall. 2017. https://api.semanticscholar.org/CorpusID:25324374 Choosing prediction over explanation in psychology: Lessons from machine learning . Perspectives on Psychological Science, 12:1100 -- 1122
2017
-
[109]
Michihiro Yasunaga, Xinyun Chen, Yujia Li, Panupong Pasupat, Jure Leskovec, Percy Liang, Ed H Chi, and Denny Zhou. 2023. Large language models as analogical reasoners. arXiv preprint arXiv:2310.01714
2023 arXiv
-
[110]
Eunice Yiu, Maan Qraitem, Charlie Wong, Anisa Noor Majhi, Yutong Bai, Shiry Ginosar, Alison Gopnik, and Kate Saenko. 2024. Kiva: Kid-inspired visual analogies for testing large multimodal models. arXiv preprint arXiv:2407.17773
2024
-
[111]
Pardis Sadat Zahraei and Ali Emami. 2024. Wsc+: Enhancing the winograd schema challenge using tree-of-experts. arXiv preprint arXiv:2401.17703
2024 arXiv
-
[112]
Yan Zhuang, Qi Liu, Yuting Ning, Weizhe Huang, Rui Lv, Zhenya Huang, Guanhao Zhao, Zheng Zhang, Qingyang Mao, Shijin Wang, et al. 2023. Efficiently measuring the cognitive ability of llms: An adaptive testing perspective. arXiv preprint arXiv:2306.10512
2023 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.