Pith. sign in

REVIEW 2 major objections 10 minor 1 cited by

The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories

T0 review · 2 major / 10 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Pre-trained language models can serve as scientific theories of cognition when their performance profiles match human behavior, even without mechanistic explanation.

desk verdict A clearly written position paper that organizes known pitfalls into a usable framework; the load-bearing sufficiency criterion is asserted rather than demonstrated, but the paper is honest about that. read the letter →

arxiv 2501.12651 v1 pith:G7CVWVL7 submitted 2025-01-22 cs.CL cs.AI

classification cs.CLcs.AI
keywords pre-trainedlanguagemodelscognitivesciencetheoriesalignmentdevelopmentallinkinghypothesespredictionversusexplanationsufficiencycriterionpsychometrics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Pre-trained language models, trained to predict text rather than to model minds, can nonetheless serve as scientific theories of human cognition and development, this paper argues. The condition is functional alignment: once a human experimental task is adapted to a model's text input, model outputs are mapped to human behavioral measures through an explicit linking hypothesis, and the comparison shows that the model's performance profile matches humans' profile, then the model is justified for predicting human behavior on those stimuli. The authors deliberately rank prediction over explanation, treating a match in functional form as sufficient even when the model's internal mechanisms differ from human mental processes. They then classify the ways this program goes wrong—pitfalls of commission, such as distal linking hypotheses, and pitfalls of omission, such as ignoring psychometric correlations and developmental trajectories—and close with concrete criteria for making such models credible in cognitive and developmental science.

What carries the argument

The central object is the three-stage mapping between humans and models: (1) adapting human experimental stimuli and task instructions to the model's text-input modality; (2) a linking hypothesis that converts model outputs into human-like performance indices; and (3) a goodness-of-fit comparison against human behavioral data. The paper examines three families of linking hypotheses—similarity computations in latent space, surprisal (negative log probability) values, and direct prompting of the model's response distribution—and argues that no single mapping is universally valid; each must be empirically validated per task. The sufficiency criterion is the engine of the argument: jointly, a reasonable adaptation, a valid linking hypothesis, and a matching performance profile justify using the model to predict human behavior, and this justifies treating prediction as the primary scientific criterion over mechanistic explanation.

What would settle it

Run a model that passes a broad battery of human-aligned cognitive tests, then present it with novel stimuli that preserve the same task structure but fall outside its training distribution while still being natural for humans; if human behavior is predicted substantially worse by the model than by a mechanism-based account, the sufficiency criterion is undercut. A second, sharper falsifier: construct two stimuli that are equally similar in the model's latent space but are judged at very different speeds by humans, contradicting the similarity-based linking hypothesis on a task where the paper's examples rely on that mapping.

Watch

Extended reading notes

Core claim

The paper's core claim is a sufficiency argument: a pre-trained language model becomes a usable cognitive or developmental science theory when, after appropriate adaptation of stimuli, a well-chosen linking hypothesis, and a goodness-of-fit comparison, its performance profile matches human performance on the relevant tasks. This holds even though the model is architecturally different, trained on different data, and not interpretable at the level of human mechanisms; for prediction of human behavior, alignment of functional form is the key requirement. From that criterion, the paper derives two utilities—generalizability across tasks and the ability to generate novel candidate theories—and it organizes known hazards into a new taxonomy: errors of commission (bad assumptions in the mapping, opaque mechanisms, aggregate training data) and errors of omission (ignoring psychometric intercorrelations, neural correlates, and developmental progression). The paper is a review and position statement, not an empirical demonstration, but its contribution is the explicit framework for when and how a PLM's output can be treated as evidence about human thinking.

Load-bearing premise

The load-bearing premise is that a match in functional performance profiles between a language model and humans is sufficient for the model to be used as a cognitive science theory, without requiring correspondence at the level of mechanisms or processes.

Editorial extensions

If this is right

  • A PLM that matches human performance profiles across enough tasks can be used to predict human behavior on a new task without collecting new human data, provided the same linking-hypothesis validation holds.
  • The sufficiency criterion licenses using PLMs as discovery engines, generating candidate cognitive theories from model behavior on novel stimuli, analogous to screening hypotheses in drug discovery.
  • Researchers must validate the linking hypothesis separately for each task, since similarity, surprisal, and prompting each have distinct failure modes and context sensitivities.
  • Developmental claims should be tested with intermediate training checkpoints, not only final models, so that the sequence of skill acquisition can be compared with children's developmental trajectories.
  • Cognitive evaluations of PLMs should measure cross-task correlations in addition to single-task accuracy, to capture psychometric structure rather than modular performance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If functional alignment is sufficient, model architecture becomes scientifically irrelevant to the validity of a cognitive claim; two structurally different PLMs that both match human profiles would be equally valid theories, a claim that could be tested by comparing diverse model families on the same battery.
  • The commission/omission taxonomy implies a concrete operational test: report the full correlation matrix of PLM performance across cognitive tasks and compare it with human psychometric correlation matrices; the paper discusses this as a criterion but does not run it.
  • The emphasis on prediction suggests a practical triage rule for the field: when surprisal and prompting disagree, the better predictor of human reaction times or judgments should win, a rule the paper gestures at but does not formalize.
  • The developmental argument predicts that curriculum order, not just final data volume, drives whether model checkpoints track child development; this is testable by training on developmentally ordered corpora and checking age-of-acquisition alignment, which the authors propose as future work.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 10 minor

Summary. This position paper argues that pre-trained language models (PLMs) can serve as credible theories in cognitive and developmental science when their performance profiles match human performance under a three-stage mapping: adaptation of human stimuli to text-compatible inputs, a linking hypothesis from model outputs to human behavioral measures, and comparison of model and human performance. The authors advocate a sufficiency criterion under which, given reasonable assumptions about these stages, matching functional profiles justifies using PLMs to predict human behavior, and they explicitly prioritize prediction over explanation. The paper reviews pitfalls, organizing them into pitfalls of commission (distal linking hypotheses, opacity and limited interpretability, training-data and stimulus mismatches, non-determinism) and pitfalls of omission (architectural differences, neglect of psychometric correlations, developmental trajectory issues, missing neuroscience context). It critically examines three families of linking hypotheses, namely latent-space similarity, surprisal, and prompting, and then enumerates two sets of criteria: appropriateness criteria (multi-experiment triangulation, multi-method interpretation, path-dependency testing, tuning controls, documented linking assumptions, cross-task correlations, embodiment) and development criteria (developmentally plausible corpora, tuning on core cognitive tasks).

Significance. The paper's main value is organizational and methodological. The commission/omission taxonomy, the three-stage mapping framework, and the explicit criteria give researchers a concrete checklist for designing and interpreting PLM-human alignment studies, and the review of linking hypotheses is grounded in specific failure cases (tokenization mismatches in Section 4, surprisal's underestimated garden-path costs, prompt-format sensitivity). The recommendation to evaluate cross-task correlation patterns (Section 5.1) is a falsifiable, implementable proposal, and the parallel between functional mapping in neuroscience and mechanistic interpretability is generative. The authors deserve credit for an unusually transparent limitations section (Section 7), which correctly flags that the criteria are asserted rather than demonstrated. The principal weakness is that the sufficiency criterion is under-specified on validation: as written it licenses prediction from benchmark fit, and the paper never defines what would count as successful or failed prediction.

major comments (2)
  1. [Section 2 (boxed Sufficiency criterion); Sections 5.1 and 7] The boxed sufficiency criterion states that, under reasonable assumptions about the three stages, PLMs are justified in predicting human behavior once their performance profiles match those of humans, and the paper adds that prediction is more important than explanation. As stated, this licenses prediction on the strength of fitting a finite set of benchmark tasks, which is calibration rather than prediction: a model can match a human performance profile on sampled tasks for non-cognitive reasons, as the paper's own examples show (Section 4's tokenization mismatch for 'Nine', prompt-format sensitivity, and Section 5.1's mention of pre-training data contamination). The Section 5.1 recommendations to 'Design multiple experiments' and to 'Establish task correlations' are steps toward stronger evidence, but the cross-task correlational benchmark is still a fit-based check rather than an explicit out-of-sample test. I recommend amending the sufficiency criterion to require preregistered predictions on held-out stimuli or new task framings before a PLM is credited with predictive utility, and to define what outcome would count as falsification. This matters because the paper itself states in Section 7, item 3, that no experiments are offered; without an explicit validation condition, the central claim remains an assertion rather than an operationalized proposal.
  2. [Section 5.2 (second bullet); Section 5.1 (Control for tuning techniques)] There is a tension between Section 5.1's 'Control for tuning techniques' criterion, which warns that tuning objectives (instruction tuning, RLHF) can make model behavior reflect the tuning goals rather than representational fidelity, and the Section 5.2 suggestion to preference-tune PLMs on 'core' cognitive tasks such as typicality experiments and then evaluate on a broader set of tasks. If the model is tuned on a subset of the very benchmarks used to establish cognitive alignment, alignment on the remaining tasks may partly reflect the tuning objective rather than the model's natural inductive biases, which is precisely the confound the paper warns about elsewhere. The recommendation should specify how the tuning stage is to be controlled (for example, held-out task families or pretraining-only baselines), or it should be framed as a hypothesis to be tested rather than as a practice to adopt without caveats.
minor comments (10)
  1. [Section 2, boxed sufficiency criterion] In the three-stage list inside the boxed sufficiency criterion, the third stage is numbered '(2)', duplicating the second stage; it should be numbered '(3)'.
  2. [Table 1] The contribution header cells contain garbled icon-like characters (a trophy symbol and a malformed 'exclamation-triangle' token) that should be replaced with clean text or removed.
  3. [Section 2] The sentence 'These three stages are depicted in Figure 1)' contains a stray closing parenthesis after the figure reference.
  4. [Section 4, Surprisal values] The term 'surprisas' should be 'surprisal' or 'surprisal values'.
  5. [Section 4, Surprisal values] 'Shain (2024) use PLMs to demonstrate strong surprisal predictability estimates' has a subject-verb agreement error; it should be 'uses'.
  6. [Section 3, Challenges of Development] 'its progression over-development' should read 'over development' without the hyphen.
  7. [Section 3] The heading styles 'Pitfalls of the commission' and 'Pitfalls of Commission' are inconsistent; one style should be used throughout.
  8. [Section 3, Pitfalls of Commission] The assertion that mechanistic interpretability is 'the wrong level of analysis' for capturing the emergent, contextual, and symbolic aspects of human thought is made without citation; the authors should either cite supporting arguments or frame the claim as a contested position.
  9. [Sections 1 and 3] The capitalization 'PLMS' appears alongside 'PLMs' in several places; standardize to 'PLMs'.
  10. [Section 4, Similarity computations] In the typicality exposition, the sentence 'the typicality of an exemplar is commonly defined as the proportion of humans that produce it' conflates the human production measure with the model-based similarity estimate; clarify that production frequency is the human measure and similarity-to-prototype is the linking hypothesis.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a theoretical position piece whose recommendations are explicit arguments, not derivations reduced to fitted inputs or to load-bearing self-citations.

full rationale

The paper contains no equations, fitted parameters, trained models, or quantitative derivations, and Section 7 explicitly states: "our work is theoretical and does not conduct experiments or offer empirical evidence of performance comparisons or other quantitative measures." The central "Sufficiency criterion" is an explicit argument that if the three stated stages of adaptation, linking hypothesis, and comparison are reasonable and performance profiles match, then PLMs may be used to predict human behavior; this is a normative proposal with stated assumptions, not a result derived from itself. The authors' self-citations (Shah et al. 2023; Shah et al. 2024; Vemuri et al. 2024; Bhardwaj et al. 2024; Li et al. 2024; Guo et al. 2024; Sharma et al. 2024) appear as illustrative examples of PLM-human alignment research, and the paper's recommended criteria do not depend on the truth of those specific empirical results. No uniqueness theorem is imported from the authors' prior work, no ansatz is smuggled in via a self-citation, and no known result is merely renamed. The concern that the sufficiency criterion may need stronger out-of-sample validation is a scientific correctness risk, not a circularity: the criterion is asserted as an argument rather than claimed to follow from its own conclusion.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central argument rests on a small number of explicit value judgments and framing assumptions, not on fitted parameters or invented entities. The three main axioms are the sufficiency criterion, the three-stage decomposition, and the priority of prediction over explanation. These are stated in Section 2 and used throughout.

assumptions (4)
  • domain assumption The sufficiency criterion: if model and human performance profiles match under reasonable linking hypotheses, the PLM can be used to predict human behavior.
    Section 2 (boxed 'Sufficiency criterion'); this is the core philosophical premise that justifies the paper's recommendations.
  • domain assumption The three-stage mapping (stimulus adaptation, linking hypothesis, comparison to human performance) is the correct decomposition of PLM-human comparisons.
    Section 2 and Figure 1 structure the entire analysis; if this decomposition is wrong, the pitfalls and criteria may be mis-framed.
  • domain assumption Prediction is more important than explanation when applying machine learning models to cognitive phenomena.
    Section 2, citing Breiman (2003) and Yarkoni and Westfall (2017); this value judgment drives the emphasis on functional alignment and underlies the criteria in Section 5.
  • domain assumption Marr's three levels of analysis provide a valid framework for comparing brains and PLMs.
    Section 3 (Pitfalls of Omission), citing Marr (2010); used to locate architectural differences at the implementational level.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories." pith.science (2026). https://pith.science/paper/G7CVWVL7

@misc{pith2026250112651,
  author       = {Pith},
  title        = {Pith review of: The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/G7CVWVL7}},
  note         = {Machine review of arXiv:2501.12651}
}
read the original abstract

Many studies have evaluated the cognitive alignment of Pre-trained Language Models (PLMs), i.e., their correspondence to adult performance across a range of cognitive domains. Recently, the focus has expanded to the developmental alignment of these models: identifying phases during training where improvements in model performance track improvements in children's thinking over development. However, there are many challenges to the use of PLMs as cognitive science theories, including different architectures, different training data modalities and scales, and limited model interpretability. In this paper, we distill lessons learned from treating PLMs, not as engineering artifacts but as cognitive science and developmental science models. We review assumptions used by researchers to map measures of PLM performance to measures of human performance. We identify potential pitfalls of this approach to understanding human thinking, and we end by enumerating criteria for using PLMs as credible accounts of cognition and cognitive development.

Figures

Figures reproduced from arXiv: 2501.12651 by the authors.

Figure 1
Figure 1. The three-stage mapping between human data and model performance when establishing the sufficiency [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Human-Like Anaphor Resolution in Large Language Models

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Some open-weight LLMs show human-like sensitivity to distance and discourse prominence in anaphor resolution, but weaker sensitivity to semantic interference.

Reference graph

Works this paper leans on

112 extracted references · 39 canonical work pages · cited by 1 Pith paper

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Gati V Aher, Rosa I Arriaga, and Adam Tauman Kalai. 2023. Using large language models to simulate multiple humans and replicate human subject studies. In International Conference on Machine Learning, pages 337--371. PMLR

  4. [4]

    Janice Ahn, Rishu Verma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin. 2024. Large language models for mathematical reasoning: Progresses and challenges. arXiv preprint arXiv:2402.00157

  5. [5]

    John R Anderson. 2009. How can the human mind occur in the physical universe? Oxford University Press

  6. [6]

    Khai Loong Aw, Syrielle Montariol, Badr AlKhamissi, Martin Schrimpf, and Antoine Bosselut. 2023. Instruction-tuning aligns llms to the human brain. In First Conference on Language Modeling

  7. [7]

    Catarina G Belem, Markelle Kelly, Mark Steyvers, Sameer Singh, and Padhraic Smyth. 2024. Perceptions of linguistic uncertainty by language models and humans. arXiv preprint arXiv:2407.15814

  8. [8]

    Yoshua Bengio, J \'e r \^o me Louradour, Ronan Collobert, and Jason Weston. 2009. Curriculum learning. In Proceedings of the 26th annual international conference on machine learning, pages 41--48

Show all 112 references
  1. [9]

    Khushi Bhardwaj, Raj Sanjay Shah, and Sashank Varma. 2024. https://arxiv.org/abs/2311.04666 Pre-training llms using human-like development data corpus . Preprint, arXiv:2311.04666

  2. [10]

    Sudeep Bhatia and Russell Richie. 2022. Transformer networks of human conceptual knowledge. Psychological Review

  3. [11]

    Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. 2023. Pythia: A suite for analyzing large language models across training and scaling. In ...

  4. [12]

    Marcel Binz, Elif Akata, Matthias Bethge, Franziska Br \"a ndle, Fred Callaway, Julian Coda-Forno, Peter Dayan, Can Demircan, Maria K Eckstein, No \'e mi \'E ltet o , et al. 2024. Centaur: a foundation model of human cognition. arXiv preprint arXiv:2410.20268

  5. [13]

    Abeba Birhane and Marek McGann. 2024. Large models of what? mistaking engineering achievements for human linguistic agency. Language Sciences, 106:101672

  6. [14]

    Leo Breiman. 2003. Statistical modeling: The two cultures. Quality control and applied statistics, 48(1):81--82

  7. [15]

    Danilo Bzdok, Andrew Thieme, Oleksiy Levkovskyy, Paul Wren, Thomas Ray, and Siva Reddy. 2024. Data science opportunities of large language models for neuroscience and biomedicine. Neuron, 112(5):698--717

  8. [16]

    Alice Cai, Ian Arawjo, and Elena L Glassman. 2024. Antagonistic ai. arXiv preprint arXiv:2402.07350

  9. [17]

    Chen-Chi Chang, Ching-Yuan Chen, Hung-Shin Lee, and Chih-Cheng Lee. 2024. Benchmarking cognitive domains for llms: Insights from taiwanese hakka culture. arXiv preprint arXiv:2409.01556

  10. [18]

    Anthony Chemero. 2023. Llms differ from human cognition because they are not embodied. Nature Human Behaviour, 7(11):1828--1829

  11. [19]

    Julian Coda-Forno, Marcel Binz, Jane X Wang, and Eric Schulz. 2024. Cogbench: a large language model walks into a psychology lab. arXiv preprint arXiv:2402.18225

  12. [20]

    Lee Joseph Cronbach. 1957. https://api.semanticscholar.org/CorpusID:144287695 The two disciplines of scientific psychology. American Psychologist, 12:671--684

  13. [21]

    Christine Cuskley, Rebecca Woods, and Molly Flaherty. 2024. The limitations of large language models for understanding human language and cognition. Open Mind, 8:1058--1083

  14. [22]

    Dorottya Demszky, Diyi Yang, David S Yeager, Christopher J Bryan, Margarett Clapper, Susannah Chandhok, Johannes C Eichstaedt, Cameron Hecht, Jeremy Jamieson, Meghann Johnson, et al. 2023. Using large language models in psychology. Nature Reviews Psychology, 2(11):688--701

  15. [23]

    Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, Eric P Xing, and Zhiting Hu. 2022. Rlprompt: Optimizing discrete text prompts with reinforcement learning. arXiv preprint arXiv:2205.12548

  16. [24]

    Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray. 2023. Can ai language models replace human participants? Trends in Cognitive Sciences, 27(7):597--600

  17. [25]

    Yijiang River Dong, Tiancheng Hu, and Nigel Collier. 2024. Can llm be a personalized judge? arXiv preprint arXiv:2406.11657

  18. [26]

    Xufeng Duan, Bei Xiao, Xuemei Tang, and Zhenguang G Cai. 2024. Hlb: Benchmarking llms' humanlikeness in language use. arXiv preprint arXiv:2409.15890

  19. [27]

    Jeffrey L Elman. 1996. Rethinking innateness: A connectionist perspective on development, volume 10. MIT press

  20. [28]

    Linnea Evanson, Yair Lakretz, and Jean-R \'e mi King. 2023. Language acquisition: do children and language models follow similar learning stages? arXiv preprint arXiv:2306.03586

  21. [29]

    Michael C Frank. 2023. Bridging the data gap between children and large language models. Trends in Cognitive Sciences

  22. [30]

    Jakob Gawlikowski, Cedrique Rovile Njieutcheu Tassi, Mohsin Ali, Jongseok Lee, Matthias Humt, Jianxiang Feng, Anna Kruspe, Rudolph Triebel, Peter Jung, Ribana Roscher, et al. 2023. A survey of uncertainty in deep neural networks. Artificial Intelligence Review, 56(Suppl 1):1513--1589

  23. [31]

    Gemini Team . 2023. https://arxiv.org/abs/2312.11805 Gemini: A family of highly capable multimodal models . Preprint, arXiv:2312.11805

  24. [32]

    Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, A. Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, T...

  25. [33]

    Grace Guo, Jenna Kang, Raj Sanjay Shah, Hanspeter Pfister, and Sashank Varma. 2024. Understanding graphical perception in data visualization through vision-language models. In NeurIPS 2024 Workshop on Behavioral Machine Learning

  26. [34]

    Justin Halberda, Mich \`e le MM Mazzocco, and Lisa Feigenson. 2008. Individual differences in non-verbal number acuity correlate with maths achievement. Nature, 455(7213):665--668

  27. [35]

    John Hale. 2001. https://aclanthology.org/N01-1021 A probabilistic E arley parser as a psycholinguistic model . In Second Meeting of the North A merican Chapter of the Association for Computational Linguistics

  28. [36]

    a m \"a l \

    Perttu H \"a m \"a l \"a inen, Mikke Tavast, and Anton Kunnari. 2023. Evaluating large language models in generating synthetic hci research data: a case study. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1--19

  29. [37]

    Mathew Hardy, Ilia Sucholutsky, Bill Thompson, and Tom Griffiths. 2023. Large language models meet cognitive science: Llms as tools, models, and participants. In Proceedings of the annual meeting of the cognitive science society, volume 45

  30. [38]

    Eghbal A Hosseini, Martin Schrimpf, Yian Zhang, Samuel Bowman, Noga Zaslavsky, and Evelina Fedorenko. 2022. Artificial neural network language models align neurally and behaviorally with humans even after a developmentally realistic amount of training. BioRxiv, pages 2022--10

  31. [39]

    Jennifer Hu and Michael C Frank. 2024. Auxiliary task demands mask the capabilities of smaller language models. arXiv preprint arXiv:2404.02418

  32. [40]

    Jennifer Hu, Kyle Mahowald, Gary Lupyan, Anna Ivanova, and Roger Levy. 2024 a . Language models align with human judgments on key grammatical constructions. Proceedings of the National Academy of Sciences, 121(36):e2400917121

  33. [41]

    Michael Y Hu, Aaron Mueller, Candace Ross, Adina Williams, Tal Linzen, Chengxu Zhuang, Ryan Cotterell, Leshem Choshen, Alex Warstadt, and Ethan Gotlieb Wilcox. 2024 b . Findings of the second babylm challenge: Sample-efficient pretraining on developmentally plausible corpora. ...

  34. [42]

    Xiaoyang Hu, Shane Storks, Richard L Lewis, and Joyce Chai. 2023. In-context analogical reasoning with pre-trained language models. arXiv preprint arXiv:2305.17626

  35. [43]

    Kuan-Jung Huang, Suhas Arehalli, Mari Kugemoto, Christian Muxica, Grusha Prasad, Brian Dillon, and Tal Linzen. 2024. Large-scale benchmark yields no evidence that language model surprisal explains syntactic disambiguation difficulty. Journal of Memory and Language, 137:104510

  36. [44]

    Huebner, Elior Sulem, Fisher Cynthia, and Dan Roth

    Philip A. Huebner, Elior Sulem, Fisher Cynthia, and Dan Roth. 2021. https://doi.org/10.18653/v1/2021.conll-1.49 B aby BERT a: Learning more grammar with small-scale child-directed language . In Proceedings of the 25th Conference on Computational Natural Language Learning, page...

  37. [45]

    Anna A Ivanova. 2023. Running cognitive evaluations on large language models: The do's and the don'ts. arXiv preprint arXiv:2312.01276

  38. [46]

    Anna A Ivanova, Aalok Sathe, Benjamin Lipkin, Evelina Fedorenko, and Jacob Andreas. 2024 a . Log probability scores provide a closer match to human plausibility judgments than prompt-based evaluations

  39. [47]

    Anna A Ivanova, Aalok Sathe, Benjamin Lipkin, Unnathi Kumar, Setayesh Radkani, Thomas H Clark, Carina Kauf, Jennifer Hu, RT Pramod, Gabriel Grand, et al. 2024 b . Elements of world knowledge (ewok): A cognition-inspired framework for evaluating basic world knowledge in languag...

  40. [48]

    Zhuoxuan Jiang, Haoyuan Peng, Shanshan Feng, Fan Li, and Dongsheng Li. 2024. Llms can find mathematical reasoning mistakes by pedagogical chain-of-thought. arXiv preprint arXiv:2405.06705

  41. [49]

    Kohitij Kar, Simon Kornblith, and Evelina Fedorenko. 2022. Interpretability of artificial neural network models in artificial intelligence versus neuroscience. Nature Machine Intelligence, 4(12):1065--1067

  42. [50]

    Carina Kauf, Anna A Ivanova, Giulia Rambelli, Emmanuele Chersoni, Jingyuan Selena She, Zawad Chowdhury, Evelina Fedorenko, and Alessandro Lenci. 2023. Event knowledge in large language models: the gap between the impossible and the unlikely. Cognitive Science, 47(11):e13386

  43. [51]

    Eliza Kosoy, Emily Rose Reagan, Leslie Lai, Alison Gopnik, and Danielle Krettek Cobb. 2023. Comparing machines and children: Using developmental psychology experiments to assess the strengths and weaknesses of lamda responses. arXiv preprint arXiv:2305.11243

  44. [52]

    Thomas L Griffiths, Charles Kemp, and Joshua B Tenenbaum. 2008. Bayesian models of cognition

  45. [53]

    Roger Levy. 2008. https://doi.org/10.1016/j.cognition.2007.05.006 Expectation-based syntactic comprehension . Cognition, 106(3):1126--1177

  46. [54]

    Andrew Li, Xianle Feng, Siddhant Narang, Austin Peng, Tianle Cai, Raj Sanjay Shah, and Sashank Varma. 2024. Incremental comprehension of garden-path sentences by large language models: Semantic interpretation, syntactic re-analysis, and attention

  47. [55]

    Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Yuqi Wang, Suqi Sun, Omkar Pangarkar, et al. 2023. Llm360: Towards fully transparent open-source llms. arXiv preprint arXiv:2312.06550

  48. [56]

    Kyle Mahowald, Anna A Ivanova, Idan A Blank, Nancy Kanwisher, Joshua B Tenenbaum, and Evelina Fedorenko. 2024. Dissociating language and thought in large language models. Trends in Cognitive Sciences

  49. [57]

    age effects in second language acquisition: Expanding the emergentist account

    Viorica Marian. 2023. Studying second language acquisition in the age of large language models: Unlocking the mysteries of language and learning, a commentary on “age effects in second language acquisition: Expanding the emergentist account” by catherine l. caldwell-harris and...

  50. [58]

    David Marr. 2010. Vision: A computational investigation into the human representation and processing of visual information. MIT press

  51. [59]

    Sam Whitman McGrath, Jacob Russin, Ellie Pavlick, and Roman Feiman. 2023. How can deep neural networks inform theory in psychological science?

  52. [60]

    Sean McGrath, Parth Mehta, Alexandra Zytek, Isaac Lage, and Himabindu Lakkaraju. 2020. When does uncertainty matter?: Understanding the impact of predictive uncertainty in ml assisted decision making. arXiv preprint arXiv:2011.06167

  53. [61]

    Ji r \' Mili c ka, Anna Marklov \'a , Kl \'a ra VanSlambrouck, Eva Posp \' s ilov \'a , Jana S imsov \'a , Samuel Harvan, and Ond r ej Drobil. 2024. Large language models are able to downplay their cognitive abilities to fit the persona they simulate. Plos one, 19(3):e0298522

  54. [62]

    Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2024. Large language models: A survey. arXiv preprint arXiv:2402.06196

  55. [63]

    Kanishka Misra, Allyson Ettinger, and Julia Taylor Rayz. 2021. Do language models learn typicality judgments from text? arXiv preprint arXiv:2105.02987

  56. [64]

    Moyer and Thomas K

    Robert S. Moyer and Thomas K. Landauer. 1967. Time required for judgements of numerical inequality. Nature, 215(5109):1519--1520

  57. [65]

    Qian Niu, Junyu Liu, Ziqian Bi, Pohsun Feng, Benji Peng, and Keyu Chen. 2024. Large language models and cognitive science: A comprehensive review of similarities, differences, and challenges. arXiv preprint arXiv:2409.02387

  58. [66]

    Desmond C Ong. 2024. Gpt-ology, computational models, silicon sampling: How should we think about llms in cognitive science? arXiv preprint arXiv:2406.09464

  59. [67]

    OpenAI. 2023. https://arxiv.org/abs/2303.08774 Gpt-4 technical report . Preprint, arXiv:2303.08774

  60. [68]

    Joon Sung Park, Lindsay Popowski, Carrie Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2022. Social simulacra: Creating populated prototypes for social computing systems. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Techno...

  61. [69]

    John M Parkman. 1971. Temporal aspects of digit and letter inequality judgments. Journal of experimental psychology, 91(2):191

  62. [70]

    Roma Patel and Ellie Pavlick. 2021. Mapping language models to grounded conceptual spaces. In International conference on learning representations

  63. [71]

    Steven Piantadosi. 2023. Modern language models refute chomsky’s approach to language. Lingbuzz Preprint, lingbuzz, 7180

  64. [72]

    Steven T Piantadosi, Dyana CY Muller, Joshua S Rule, Karthikeya Kaushik, Mark Gorenstein, Elena R Leib, and Emily Sanford. 2024. Why concepts are (probably) vectors. Trends in Cognitive Sciences

  65. [73]

    Frank, and Gary Lupyan

    Eva Portelance, Yuguang Duan, Michael C. Frank, and Gary Lupyan. 2023. https://api.semanticscholar.org/CorpusID:261696384 Predicting age of acquisition for children's early vocabulary in five languages using language model surprisal . Cognitive science, 47 9:e13334

  66. [74]

    Yujin Potter, Shiyang Lai, Junsol Kim, James Evans, and Dawn Song. 2024. Hidden persuaders: Llms' political leaning and their influence on voters. arXiv preprint arXiv:2410.24190

  67. [75]

    Santhosh Kumar Ramakrishnan, Erik Wijmans, Philipp Kraehenbuehl, and Vladlen Koltun. 2024. Does spatial cognition emerge in frontier models? arXiv preprint arXiv:2410.06468

  68. [76]

    Giulia Rambelli, Emmanuele Chersoni, Davide Testa, Philippe Blache, and Alessandro Lenci. 2024. Neural generative models and the parallel architecture of language: A critical review and outlook. Topics in cognitive science

  69. [77]

    Eleanor Rosch. 1975. Cognitive representations of semantic categories. Journal of Experimental Psychology: General, 104(3):192

  70. [78]

    Leonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz, and Zeynep Akata. 2024. In-context impersonation reveals large language models' strengths and biases. Advances in Neural Information Processing Systems, 36

  71. [79]

    Vinay Samuel, Henry Peng Zou, Yue Zhou, Shreyas Chaudhari, Ashwin Kalyan, Tanmay Rajpurohit, Ameet Deshpande, Karthik Narasimhan, and Vishvak Murahari. 2024. Personagym: Evaluating persona agents and llms. arXiv preprint arXiv:2407.18416

  72. [80]

    W Joel Schneider and Kevin S McGrew. 2012. The cattell-horn-carroll model of intelligence

  73. [81]

    Andreas Schuller, Doris Janssen, Julian Blumenr \"o ther, Theresa Maria Probst, Michael Schmidt, and Chandan Kumar. 2024. Generating personas using llms and assessing their viability. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, pages 1--7

  74. [82]

    Raj Shah, Khushi Bhardwaj, and Sashank Varma. 2024. Development of cognitive intelligence in pre-trained language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 9632--9657

  75. [83]

    Raj Sanjay Shah, Vijay Marupudi, Reba Koenen, Khushi Bhardwaj, and Sashank Varma. 2023. https://arxiv.org/abs/2305.10782 Numeric magnitude comparison effects in large language models . Preprint, arXiv:2305.10782

  76. [84]

    Cory Shain. 2024. Word frequency and predictability dissociate in naturalistic reading. Open Mind, 8:177--201

  77. [85]

    Mihir Sharma, Ryan Ding, Raj Sanjay Shah, and Sashank Varma. 2024. Monolingual and bilingual language acquisition in language models

  78. [86]

    Roger N Shepard and Peter Podgorny. 1978. Cognitive processes that resemble perceptual processes. In Handbook of learning and cognitive processes: Vol. 5. Human information processing, pages 189--237. Erlbaum Hillsdale, NJ

  79. [87]

    Linda Smith and Michael Gasser. 2005. The development of embodied cognition: Six lessons from babies. Artificial life, 11(1-2):13--29

  80. [88]

    Richard E Snow, Patrick C Kyllonen, Brachia Marshalek, et al. 1984. The topography of ability and learning correlations. Advances in the psychology of human intelligence, 2(S 47):103

  81. [89]

    Alvin Wei Ming Tan, Sunny Yu, Bria Long, Wanjing Anya Ma, Tonya Murray, Rebecca D Silverman, Jason D Yeatman, and Michael C Frank. 2024. Devbench: A multimodal developmental benchmark for language learning. arXiv preprint arXiv:2406.10215

  82. [90]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. https://arxiv.org/abs/2302.13971 Llama:...

  83. [91]

    Yu-Min Tseng, Yu-Chao Huang, Teng-Yun Hsiao, Yu-Ching Hsu, Jia-Yin Foo, Chao-Wei Huang, and Yun-Nung Chen. 2024. Two tales of persona in llms: A survey of role-playing and personalization. arXiv preprint arXiv:2406.01171

  84. [92]

    Peter D. Turney. 2013. https://doi.org/10.1162/tacl_a_00233 Distributional semantics beyond words: Supervised learning of analogy and paraphrase . Transactions of the Association for Computational Linguistics, 1:353--366

  85. [93]

    Peter D Turney and Michael L Littman. 2005. Corpus-based learning of analogies and semantic relations. Machine Learning, 60:251--278

  86. [94]

    Marten Van Schijndel and Tal Linzen. 2021. Single-stage prediction models do not explain the magnitude of syntactic disambiguation difficulty. Cognitive science, 45(6):e12988

  87. [95]

    Sashank Varma. 2011. Criteria for the design and evaluation of cognitive architectures. Cognitive science, 35(7):1329--1351

  88. [96]

    Manoj Chowdary Vattikuti. 2024. Improving drug discovery and development using ai: Opportunities and challenges. Research-gate journal, 10(10)

  89. [97]

    H \'e ctor Javier V \'a zquez Mart \' nez. 2021. https://doi.org/10.18653/v1/2021.blackboxnlp-1.38 The acceptability delta criterion: Testing knowledge of language using the gradience of sentence acceptability . In Proceedings of the Fourth BlackboxNLP Workshop on Analyzing an...

  90. [98]

    H \'e ctor Javier V \'a zquez Mart \' nez, Annika Heuser, Charles Yang, and Jordan Kodner. 2023. https://doi.org/10.18653/v1/2023.genbench-1.4 Evaluating neural language models as cognitive models of language acquisition . In Proceedings of the 1st GenBench Workshop on (Benchm...

  91. [99]

    Siddhartha K Vemuri, Raj Sanjay Shah, and Sashank Varma. 2024. How well do deep learning models capture human concepts? the case of the typicality effect

  92. [100]

    Xinglin Wang, Peiwen Yuan, Shaoxiong Feng, Yiwei Li, Boyuan Pan, Heda Wang, Yao Hu, and Kan Li. 2024. Coglm: Tracking cognitive development of large language models. arXiv preprint arXiv:2408.09150

  93. [101]

    Alex Warstadt and Samuel R. Bowman. 2024. https://arxiv.org/abs/2208.07998 What artificial neural networks can tell us about human language acquisition . Preprint, arXiv:2208.07998

  94. [102]

    Alex Warstadt, Aaron Mueller, Leshem Choshen, Ethan Wilcox, Chengxu Zhuang, Juan Ciro, Rafael Mosquera, Bhargavi Paranjabe, Adina Williams, Tal Linzen, and Ryan Cotterell, editors. 2023. https://aclanthology.org/2023.conll-babylm.0 Proceedings of the BabyLM Challenge at the 27...

  95. [103]

    Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R Bowman. 2020. Blimp: The benchmark of linguistic minimal pairs for english. Transactions of the Association for Computational Linguistics, 8:377--392

  96. [104]

    Taylor Webb, Keith J Holyoak, and Hongjing Lu. 2023. Emergent analogical reasoning in large language models. Nature Human Behaviour, 7(9):1526--1541

  97. [105]

    Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

    Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022. Emergent abilities of large language mod...

  98. [106]

    Benjue Weng. 2024. Navigating the landscape of large language models: A comprehensive review and analysis of paradigms and fine-tuning strategies. arXiv preprint arXiv:2404.09022

  99. [107]

    Haoran Yang, Yumeng Zhang, Jiaqi Xu, Hongyuan Lu, Pheng Ann Heng, and Wai Lam. 2024. Unveiling the generalization power of fine-tuned large language models. arXiv preprint arXiv:2403.09162

  100. [108]

    Tal Yarkoni and Jacob Westfall. 2017. https://api.semanticscholar.org/CorpusID:25324374 Choosing prediction over explanation in psychology: Lessons from machine learning . Perspectives on Psychological Science, 12:1100 -- 1122

  101. [109]

    Michihiro Yasunaga, Xinyun Chen, Yujia Li, Panupong Pasupat, Jure Leskovec, Percy Liang, Ed H Chi, and Denny Zhou. 2023. Large language models as analogical reasoners. arXiv preprint arXiv:2310.01714

  102. [110]

    Eunice Yiu, Maan Qraitem, Charlie Wong, Anisa Noor Majhi, Yutong Bai, Shiry Ginosar, Alison Gopnik, and Kate Saenko. 2024. Kiva: Kid-inspired visual analogies for testing large multimodal models. arXiv preprint arXiv:2407.17773

  103. [111]

    Pardis Sadat Zahraei and Ali Emami. 2024. Wsc+: Enhancing the winograd schema challenge using tree-of-experts. arXiv preprint arXiv:2401.17703

  104. [112]

    Yan Zhuang, Qi Liu, Yuting Ning, Weizhe Huang, Rui Lv, Zhenya Huang, Guanhao Zhao, Zheng Zhang, Qingyang Mao, Shijin Wang, et al. 2023. Efficiently measuring the cognitive ability of llms: An adaptive testing perspective. arXiv preprint arXiv:2306.10512

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.