REVIEW 3 major objections 6 minor 35 references
When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Both humans and LLMs misread garden-path sentences for the same three reasons.
desk verdict A systematic, well-run comparison of garden-path comprehension in humans and many LLMs with honest human data, but the paper overclaims by-condition Spearman correlations and should report a marginal syntactic effect as non-significant. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the object/subject garden-path sentence, a temporarily ambiguous construction of the form "While [embedded verb] [NP] [main verb] ..." whose first parse (NP as object of embedded verb) must be revised. The argument works by manipulating three features—clause order (garden-path vs. non-garden-path), plausibility of the NP as a direct object, and verb type (optionally transitive vs. reflexive/unaccusative)—and measuring, on identical items, human accuracy and the probability LLMs assign to the correct answer token. The comparison metric is the rank correlation between human item accuracy and LLM answer probability, with Spearman correlation across the six experimental conditions as a secondary check.
What would settle it
Run the same comprehension questions with a three-way forced choice (Yes / No / Not necessarily) on the same garden-path and non-garden-path items. If participants select 'Yes' or 'No' at similar rates across structures, or if accuracy computed with 'Yes' counted as correct matches non-GP accuracy, the reported garden-path deficit is an artifact of the binary scoring rule.
Extended reading notes
Core claim
The paper's central discovery is that object/subject garden-path sentences—where a noun phrase is temporarily attachable as the object of an embedded verb—produce the same pattern of comprehension errors in humans and in large language models. Using a single-trial comprehension task with questions like "Did the man hunt the deer?", the authors show that accuracy falls when the sentence requires syntactic reanalysis, when the noun is a plausible object for the verb, and when the verb is optionally transitive rather than reflexive or unaccusative. For LLMs, the same manipulations move answer probabilities in the same direction, and the rank correlation of item difficulty with human accuracy rises with model scale (the strongest model reaches 78% accuracy and the highest correlations). The authors take this as evidence that LLMs and humans share underlying sensitivity to the same syntactic and semantic pressures, and they corroborate the finding with a paraphrasing task and image generation, where the same misinterpretations appear.
Load-bearing premise
The entire error pattern rests on the labeling rule that a 'yes' answer to 'Did the man hunt the deer?' is wrong for the sentence 'While the man hunted the deer ran into the woods,' because the authors treat the sentence as not entailing that the man hunted the deer; if a permissive reading that allows 'hunt the deer' were treated as correct, the reported human and model accuracy gaps would shrink or disappear.
Editorial extensions
If this is right
- If the claim holds, LLM performance on comprehension questions can serve as a behavioral proxy for human garden-path processing, letting psycholinguists run large-scale, cheap replications of reading studies.
- Model scaling (size and pretraining tokens) should continue to increase alignment with human error patterns, so the correlation becomes a usable benchmark for whether a new model is becoming more 'human-like' in this specific sense.
- The three-factor account—syntax, plausibility, verb transitivity—should predict error rates on new garden-path items, not just the sets tested here; a new set of sentences varying these factors should reproduce the same ordering.
- Since the effect appears in comprehension, paraphrasing, and image generation, the misinterpretation is not an artifact of the question format; any model that understands a GP sentence correctly should also avoid the error in open-ended generation.
Reading between the lines
- An unresolved question is whether the reported 'misinterpretation' is partly a task artifact: the paper scores 'yes' as wrong for 'Did the man hunt the deer?' even though the sentence permits that reading. Reanalyzing the same data with a permissive scoring rule—counting 'yes' as acceptable—would show how much of the human-model error pattern survives.
- The paper groups all optionally transitive verbs together despite finding only a weak correlation (≤0.19) between transitivity bias and accuracy; a finer-grained analysis separating high- and low-bias verbs might reveal that verb-specific statistics, not a categorical distinction, drive the effect.
- If LLM errors mirror human errors because both reflect shallow distributional plausibility, then making the implausible condition more extreme (e.g., 'the rhino ran into the woods' after 'hunted') should push both humans and models toward ceiling; a graded plausibility curve could separate the syntactic and semantic contributions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper compares humans and a broad suite of LLMs on object/subject garden-path sentences using the same comprehension-question task. It manipulates three factors—GP vs. non-GP syntax, plausible vs. implausible direct-object readings, and optionally transitive vs. reflexive/unaccusative verbs—and measures accuracy for humans (Prolific, speeded word-by-word presentation) and token probabilities for LLMs. It also reports paraphrase and text-to-image tasks for LLMs. The main empirical claims are that all three hypothesized factors affect humans and many LLMs, that stronger models correlate better with human item difficulty, and that the two auxiliary tasks reproduce the same error pattern.
Significance. If the pattern holds, this is a useful contribution to the emerging comparison between human and LLM sentence processing: it uses a same-task design, covers multiple model families and checkpoints, and triangulates with paraphrasing and image generation. The human experiment is carefully controlled (single-trial design, native speakers, GLMM analyses), and the full sentence materials are provided in Appendix B. The main limitations are the absence of confidence intervals for the rank correlations, the unvalidated binary scoring rule for the comprehension questions (acknowledged in Section 2), and the treatment of timeouts as incorrect. Still, the central observation—garden-path difficulty is not unique to humans—is credible and should be of interest to both psycholinguists and NLP researchers.
major comments (3)
- [Section 5, Figure 5 and Table 4] The claim that 'all models show a high Spearman rank correlation with human data' is contradicted by the condition-level numbers in Table 4. For example, Olmo-7B-Tokens-8B has accuracies [0.655, 0.665, 0.654, 0.663, 0.649, 0.657] across the six conditions; against the human condition means this yields Spearman rho of about 0.09, not a high correlation, and Olmo-7B-Tokens-111B is essentially flat across conditions, making its rank correlation near zero or undefined. The paper should report the per-model Spearman values (with a note on tied ranks) and restrict the claim to the models that actually show high correlation, or provide a statistical summary with confidence intervals. This is load-bearing because the 'stronger models are more similar to humans' narrative depends on the reliability of these correlations.
- [Section 2, scoring of comprehension questions] The paper states that for questions like 'Did the man hunt the deer?' the logically accurate answer is 'Not necessarily', but then scores 'yes' as wrong and 'no' as right. All human and LLM accuracy numbers, condition means, and rank correlations inherit this coding. This is the standard convention in the garden-path literature, so it is not an internal inconsistency, but it is a validity limitation: some 'no' responses may reflect uncertainty avoidance rather than successful reanalysis, and some 'yes' responses may reflect a pragmatically licensed inference rather than a lingering misparse. The authors should report a sensitivity analysis (e.g., excluding timeouts, or re-scoring with a permissive rule) and should also state whether any human data were collected on a three-option or paraphrase version to validate the forced binary coding.
- [Section 3.1, Procedure] The procedure marks a response as incorrect if it is not given within five seconds after the question. Because garden-path conditions are likely to slow response times, this rule can inflate the human GP deficit purely through a speed-accuracy trade-off. The paper does not report timeout rates or model them statistically. Please report the proportion of timed-out trials by condition and re-run the main GLMMs excluding or modeling timeouts, to show that the human effects are not an artifact of the response deadline.
minor comments (6)
- [Section 6 heading and Figure 4 caption] There are typos: 'paraphraing' should be 'paraphrasing', and 'Gloabal' should be 'Global'.
- [Section 3.1, plausibility pretesting] The plausibility manipulation was selected using GPT-4 ratings, and GPT-family models are then evaluated on those same items. This does not by itself invalidate the results because the effect also appears in humans and non-GPT models, but the paper should discuss the possible circularity and ideally show that the plausibility effect survives when the items were selected by human norms (or by a non-GPT model).
- [Section 5, correlation methodology] Spearman correlations based on six condition means are coarse and unstable; the paper should report the actual correlation values, their uncertainty, and the handling of tied ranks, especially for near-flat models like the early OLMo checkpoints.
- [Figure 1] The figure caption does not explicitly define the order of the six condition bars; please add a legend or a sentence such as 'conditions are GP/non-GP for plausible, implausible, and reflexive verb types'.
- [Section 6.1] The sentence 'The man hunted the child. seem to be too out of distribution for our LLMs to generate' contains a grammatical error and should be rephrased; also, the format metric description says 'consistes' instead of 'consists'.
- [General] The paper does not state whether the experimental materials, human response data, and model output probabilities will be released; providing these would strengthen reproducibility.
Circularity Check
No significant circularity: the paper is an empirical human-LLM comparison whose central effects are supported by independent human data and multiple non-GPT model families.
full rationale
The paper does not derive its conclusions from its own inputs by construction. The hypotheses are drawn from prior psycholinguistic work (Christianson et al. 2001; Patson et al. 2009) and tested on both human participants and a wide array of LLM families. The only self-citation, Amouyal et al. (2024), is used in Section 3.1 to motivate using GPT-4 to rate stimulus plausibility for the plausibility manipulation. This is not load-bearing circularity: the plausibility effect is independently observed in human responses (p = 4.11e-16) and appears across GPT, Llama, Qwen, Gemma, and OLMo models, not just in GPT models whose ratings selected the stimuli. The plausibility ratings are not a fitted parameter later relabeled as a prediction; they are a stimulus-selection tool. The Section 5 correlations compare independently measured human accuracies with LLM answer-token probabilities, with no equation that equates the two by definition. The binary scoring convention in Section 2 ('we consider "yes" to be a wrong answer here, whereas "no" is considered the right answer') is a coding assumption inherited from the garden-path literature and is not a circular derivation, especially since converging evidence comes from paraphrase and text-to-image tasks. No uniqueness theorem from the authors is invoked, no ansatz is smuggled in via self-citation, and no known result is merely renamed as a new organization. The limitations section honestly notes the absence of reading-time and eye-gaze data, but that is a completeness concern, not a circularity concern.
Assumptions & free parameters
assumptions (4)
- domain assumption The correct answer to the GP comprehension question is 'No' even though the sentence permits a reading where the answer is 'Yes' (i.e., the intended final parse is the reference).
- domain assumption GPT-4 plausibility ratings (1-7 scale) are a valid proxy for human plausibility for stimulus selection.
- domain assumption Average next-token probability of the correct answer under few-shot prompting reflects LLM comprehension of the sentence-question pair.
- domain assumption Word-by-word presentation at 400 ms per word engages the same reading-comprehension processes as natural reading.
Cite this review
Pith. "Pith review of When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models." pith.science (2026). https://pith.science/paper/DUCJR4A3
@misc{pith2026250209307,
author = {Pith},
title = {Pith review of: When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models},
year = {2026},
howpublished = {\url{https://pith.science/paper/DUCJR4A3}},
note = {Machine review of arXiv:2502.09307}
}
read the original abstract
Modern Large Language Models (LLMs) have shown human-like abilities in many language tasks, sparking interest in comparing LLMs' and humans' language processing. In this paper, we conduct a detailed comparison of the two on a sentence comprehension task using garden-path constructions, which are notoriously challenging for humans. Based on psycholinguistic research, we formulate hypotheses on why garden-path sentences are hard, and test these hypotheses on human participants and a large suite of LLMs using comprehension questions. Our findings reveal that both LLMs and humans struggle with specific syntactic complexities, with some models showing high correlation with human comprehension. To complement our findings, we test LLM comprehension of garden-path constructions with paraphrasing and text-to-image generation tasks, and find that the results mirror the sentence comprehension question results, further validating our findings on LLM understanding of these constructions.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Merouane Debbah, Etienne Goffinet, Daniel Heslow, Julien Launay, Quentin Malartic, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023. Falcon-40B : an open large language model with state-of-the-art performance
work page 2023
-
[4]
Samuel Amouyal, Aya Meltzer-Asscher, and Jonathan Berant. 2024. https://aclanthology.org/2024.findings-eacl.12 Large language models for psycholinguistic plausibility pretesting . In Findings of the Association for Computational Linguistics: EACL 2024, pages 166--181, St. Julian ' s, Malta. Association for Computational Linguistics
work page 2024
-
[5]
Suhas Arehalli, Brian Dillon, and Tal Linzen. 2022. https://api.semanticscholar.org/CorpusID:253098758 Syntactic surprisal from neural models predicts, but underestimates, human processing difficulty from syntactic ambiguities . ArXiv, abs/2210.12187
arXiv 2022
-
[6]
Charlotte Cacheteux and Jean-Rémi King. 2022. https://doi.org/10.1038/s42003-022-03036-1 An activation-based model of sentence processing as skilled memory retrieval . Nature, pages 375--419
-
[7]
Kiel Christianson, Jack Dempsey, Anna Tsiola, and Maria Goldshtein. 2022. What if they're just not that into you (or your experiment)? on motivation and psycholinguistics. In Psychology of learning and motivation, volume 76, pages 51--88. Elsevier
work page 2022
-
[8]
Kiel Christianson, Andrew Hollingworth, John F Halliwell, and Fernanda Ferreira. 2001. Thematic roles assigned along the garden path linger. Cognitive psychology, 42(4):368--407
work page 2001
Show all 35 references
-
[9]
Kiel Christianson, Carrick C Williams, Rose T Zacks, and Fernanda Ferreira. 2006. Misinterpretations of garden-path sentences by older and younger adults. Discourse Processes, 42:205--238
2006
-
[10]
Fernanda Ferreira and John M Henderson. 1990. Use of verb information in syntactic parsing: evidence from eye movements and word-by-word self-paced reading. Journal of Experimental Psychology: Learning, Memory, and Cognition, 16(4):555
1990
-
[11]
Alex B Fine, T Florian Jaeger, Thomas A Farmer, and Ting Qian. 2013. Rapid expectation adaptation during syntactic comprehension. PloS one, 8(10):e77661
2013
-
[12]
Susan M Garnsey, Neal J Pearlmutter, Elizabeth Myers, and Melanie A Lotocky. 1997. The contributions of verb bias and plausibility to the comprehension of temporarily ambiguous sentences. Journal of memory and language, 37(1):58--93
1997
-
[13]
Team Gemini. 2024. https://arxiv.org/abs/2403.05530 Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context . Preprint, arXiv:2403.05530
2024 arXiv
-
[14]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...
2024 arXiv
-
[15]
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack...
2024 arXiv
-
[16]
Michael Hanna and Aaron Mueller. 2024. https://arxiv.org/abs/2412.05353 Incremental sentence processing mechanisms in autoregressive transformer language models . Preprint, arXiv:2412.05353
2024 arXiv
-
[17]
Jennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox, and Roger Levy. 2020. https://doi.org/10.18653/v1/2020.acl-main.158 A systematic assessment of syntactic generalization in neural language models . In Proceedings of the 58th Annual Meeting of the Association for Computationa...
2020 doi
-
[18]
Tovah Irwin, Kyra Wilson, and Alec Marantz. 2023. https://api.semanticscholar.org/CorpusID:258378287 Bert shows garden path effects . In Conference of the European Chapter of the Association for Computational Linguistics
2023
-
[19]
Tatsuki Kuribayashi, Yohei Oseki, Souhaib Ben Taieb, Kentaro Inui, and Timothy Baldwin. 2025. Large language models are human-like internally. arXiv preprint arXiv:2502.01615
2025 arXiv
-
[20]
Andrew Li, Xianle Feng, Siddhant Narang, Austin Peng, Tianle Cai, Raj Sanjay Shah, and Sashank Varma. 2024. https://api.semanticscholar.org/CorpusID:270063738 Incremental comprehension of garden-path sentences by large language models: Semantic interpretation, syntactic re-ana...
2024 arXiv
-
[21]
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016. https://doi.org/10.1162/tacl_a_00115 Assessing the ability of LSTM s to learn syntax-sensitive dependencies . Transactions of the Association for Computational Linguistics, 4:521--535
2016 doi
-
[22]
OpenAI. 2023. https://arxiv.org/abs/2303.08774 Gpt-4 technical report . ArXiv preprint, abs/2303.08774
2023 arXiv
-
[23]
OpenAI. 2024. https://cdn.openai.com/papers/DALL_E_3_System_Card.pdf Dall-e 3 system card . OpenAI
2024
-
[24]
Nikole D Patson, Emily S Darowski, Nicole Moon, and Fernanda Ferreira. 2009. Lingering misinterpretations in garden-path sentences: evidence from a paraphrasing task. Journal of Experimental Psychology: Learning, Memory, and Cognition, 35(1):280
2009
-
[25]
Qwen Team . 2024. https://qwenlm.github.io/blog/qwen2.5/ Qwen2.5: A party of foundation models
2024
-
[26]
Adrielli Lopes Rego, Joshua Snell, and Martijn Meeter. 2024. https://api.semanticscholar.org/CorpusID:269527476 Language models outperform cloze predictability in a cognitive model of reading . PLOS Computational Biology, 20
2024
-
[27]
Yuqi Ren, Renren Jin, Tongxuan Zhang, and Deyi Xiong. 2024. https://api.semanticscholar.org/CorpusID:268041437 Do large language models mirror cognitive language processing? ArXiv, abs/2402.18023
2024 arXiv
-
[28]
Hosseini, Nancy Kanwisher, Joshua B
Martin Schrimpf, Idan Asher Blank, Greta Tuckute, Carina Kauf, Eghbal A. Hosseini, Nancy Kanwisher, Joshua B. Tenenbaum, and Evelina Fedorenko. 2021. https://doi.org/10.1073/pnas.2105646118 The neural architecture of language: Integrative modeling converges on predictive proce...
2021 doi
-
[29]
Kun Sun and Rong Wang. 2024. https://api.semanticscholar.org/CorpusID:268680327 Computational sentence-level metrics predicting human sentence comprehension . ArXiv, abs/2403.15822
2024 arXiv
-
[30]
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan...
2024 arXiv
-
[31]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, Aur'elien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. https://arxiv.org/abs/2302.139...
2023 arXiv
-
[32]
John C Trueswell, Michael K Tanenhaus, and Christopher Kello. 1993. Verb-specific constraints in sentence processing: separating effects of lexical preference from garden-paths. Journal of Experimental psychology: Learning, memory, and Cognition, 19(3):528
1993
-
[33]
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019. https://doi.org/10.1162/tacl_a_00290 Neural network acceptability judgments . Transactions of the Association for Computational Linguistics, 7:625--641
2019 doi
-
[34]
Ethan Gotlieb Wilcox, Jon Gauthier, Jennifer Hu, Peng Qian, and Roger Levy. 2022. Learning syntactic structures from string input. Algebraic Structures in Natural Language, page 113
2022
-
[35]
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, and 40 others. 2024. Qwen2 technical repo...
2024 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.