Pith. sign in

REVIEW 3 major objections 6 minor 35 references

When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Both humans and LLMs misread garden-path sentences for the same three reasons.

desk verdict A systematic, well-run comparison of garden-path comprehension in humans and many LLMs with honest human data, but the paper overclaims by-condition Spearman correlations and should report a marginal syntactic effect as non-significant. read the letter →

arxiv 2502.09307 v1 pith:DUCJR4A3 submitted 2025-02-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords garden-pathsentenceslanguagemodelshumansentenceprocessingcomprehensionpsycholinguisticsplausibilitytransitivityparaphrasing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether large language models fail on the same sentences that reliably trip up human readers. It tests humans and a wide range of LLMs on the same comprehension questions about garden-path sentences—sentences like "While the man hunted the deer ran into the woods," where the first clause tempts a misparse. The paper claims that three factors—the garden-path syntax itself, the semantic plausibility of the noun as a direct object, and whether the verb is transitive—drive errors in both humans and models, and that stronger models resemble humans more closely. If true, it means LLM errors on this construction are not arbitrary; they follow the same psycholinguistic pressures that shape human misreading, and LLMs could serve as a test bed for theories of human sentence processing. The claim is validated further by paraphrasing and text-to-image tasks, which show the same error pattern.

What carries the argument

The central object is the object/subject garden-path sentence, a temporarily ambiguous construction of the form "While [embedded verb] [NP] [main verb] ..." whose first parse (NP as object of embedded verb) must be revised. The argument works by manipulating three features—clause order (garden-path vs. non-garden-path), plausibility of the NP as a direct object, and verb type (optionally transitive vs. reflexive/unaccusative)—and measuring, on identical items, human accuracy and the probability LLMs assign to the correct answer token. The comparison metric is the rank correlation between human item accuracy and LLM answer probability, with Spearman correlation across the six experimental conditions as a secondary check.

What would settle it

Run the same comprehension questions with a three-way forced choice (Yes / No / Not necessarily) on the same garden-path and non-garden-path items. If participants select 'Yes' or 'No' at similar rates across structures, or if accuracy computed with 'Yes' counted as correct matches non-GP accuracy, the reported garden-path deficit is an artifact of the binary scoring rule.

Watch

Extended reading notes

Core claim

The paper's central discovery is that object/subject garden-path sentences—where a noun phrase is temporarily attachable as the object of an embedded verb—produce the same pattern of comprehension errors in humans and in large language models. Using a single-trial comprehension task with questions like "Did the man hunt the deer?", the authors show that accuracy falls when the sentence requires syntactic reanalysis, when the noun is a plausible object for the verb, and when the verb is optionally transitive rather than reflexive or unaccusative. For LLMs, the same manipulations move answer probabilities in the same direction, and the rank correlation of item difficulty with human accuracy rises with model scale (the strongest model reaches 78% accuracy and the highest correlations). The authors take this as evidence that LLMs and humans share underlying sensitivity to the same syntactic and semantic pressures, and they corroborate the finding with a paraphrasing task and image generation, where the same misinterpretations appear.

Load-bearing premise

The entire error pattern rests on the labeling rule that a 'yes' answer to 'Did the man hunt the deer?' is wrong for the sentence 'While the man hunted the deer ran into the woods,' because the authors treat the sentence as not entailing that the man hunted the deer; if a permissive reading that allows 'hunt the deer' were treated as correct, the reported human and model accuracy gaps would shrink or disappear.

Editorial extensions

If this is right

  • If the claim holds, LLM performance on comprehension questions can serve as a behavioral proxy for human garden-path processing, letting psycholinguists run large-scale, cheap replications of reading studies.
  • Model scaling (size and pretraining tokens) should continue to increase alignment with human error patterns, so the correlation becomes a usable benchmark for whether a new model is becoming more 'human-like' in this specific sense.
  • The three-factor account—syntax, plausibility, verb transitivity—should predict error rates on new garden-path items, not just the sets tested here; a new set of sentences varying these factors should reproduce the same ordering.
  • Since the effect appears in comprehension, paraphrasing, and image generation, the misinterpretation is not an artifact of the question format; any model that understands a GP sentence correctly should also avoid the error in open-ended generation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An unresolved question is whether the reported 'misinterpretation' is partly a task artifact: the paper scores 'yes' as wrong for 'Did the man hunt the deer?' even though the sentence permits that reading. Reanalyzing the same data with a permissive scoring rule—counting 'yes' as acceptable—would show how much of the human-model error pattern survives.
  • The paper groups all optionally transitive verbs together despite finding only a weak correlation (≤0.19) between transitivity bias and accuracy; a finer-grained analysis separating high- and low-bias verbs might reveal that verb-specific statistics, not a categorical distinction, drive the effect.
  • If LLM errors mirror human errors because both reflect shallow distributional plausibility, then making the implausible condition more extreme (e.g., 'the rhino ran into the woods' after 'hunted') should push both humans and models toward ceiling; a graded plausibility curve could separate the syntactic and semantic contributions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper compares humans and a broad suite of LLMs on object/subject garden-path sentences using the same comprehension-question task. It manipulates three factors—GP vs. non-GP syntax, plausible vs. implausible direct-object readings, and optionally transitive vs. reflexive/unaccusative verbs—and measures accuracy for humans (Prolific, speeded word-by-word presentation) and token probabilities for LLMs. It also reports paraphrase and text-to-image tasks for LLMs. The main empirical claims are that all three hypothesized factors affect humans and many LLMs, that stronger models correlate better with human item difficulty, and that the two auxiliary tasks reproduce the same error pattern.

Significance. If the pattern holds, this is a useful contribution to the emerging comparison between human and LLM sentence processing: it uses a same-task design, covers multiple model families and checkpoints, and triangulates with paraphrasing and image generation. The human experiment is carefully controlled (single-trial design, native speakers, GLMM analyses), and the full sentence materials are provided in Appendix B. The main limitations are the absence of confidence intervals for the rank correlations, the unvalidated binary scoring rule for the comprehension questions (acknowledged in Section 2), and the treatment of timeouts as incorrect. Still, the central observation—garden-path difficulty is not unique to humans—is credible and should be of interest to both psycholinguists and NLP researchers.

major comments (3)
  1. [Section 5, Figure 5 and Table 4] The claim that 'all models show a high Spearman rank correlation with human data' is contradicted by the condition-level numbers in Table 4. For example, Olmo-7B-Tokens-8B has accuracies [0.655, 0.665, 0.654, 0.663, 0.649, 0.657] across the six conditions; against the human condition means this yields Spearman rho of about 0.09, not a high correlation, and Olmo-7B-Tokens-111B is essentially flat across conditions, making its rank correlation near zero or undefined. The paper should report the per-model Spearman values (with a note on tied ranks) and restrict the claim to the models that actually show high correlation, or provide a statistical summary with confidence intervals. This is load-bearing because the 'stronger models are more similar to humans' narrative depends on the reliability of these correlations.
  2. [Section 2, scoring of comprehension questions] The paper states that for questions like 'Did the man hunt the deer?' the logically accurate answer is 'Not necessarily', but then scores 'yes' as wrong and 'no' as right. All human and LLM accuracy numbers, condition means, and rank correlations inherit this coding. This is the standard convention in the garden-path literature, so it is not an internal inconsistency, but it is a validity limitation: some 'no' responses may reflect uncertainty avoidance rather than successful reanalysis, and some 'yes' responses may reflect a pragmatically licensed inference rather than a lingering misparse. The authors should report a sensitivity analysis (e.g., excluding timeouts, or re-scoring with a permissive rule) and should also state whether any human data were collected on a three-option or paraphrase version to validate the forced binary coding.
  3. [Section 3.1, Procedure] The procedure marks a response as incorrect if it is not given within five seconds after the question. Because garden-path conditions are likely to slow response times, this rule can inflate the human GP deficit purely through a speed-accuracy trade-off. The paper does not report timeout rates or model them statistically. Please report the proportion of timed-out trials by condition and re-run the main GLMMs excluding or modeling timeouts, to show that the human effects are not an artifact of the response deadline.
minor comments (6)
  1. [Section 6 heading and Figure 4 caption] There are typos: 'paraphraing' should be 'paraphrasing', and 'Gloabal' should be 'Global'.
  2. [Section 3.1, plausibility pretesting] The plausibility manipulation was selected using GPT-4 ratings, and GPT-family models are then evaluated on those same items. This does not by itself invalidate the results because the effect also appears in humans and non-GPT models, but the paper should discuss the possible circularity and ideally show that the plausibility effect survives when the items were selected by human norms (or by a non-GPT model).
  3. [Section 5, correlation methodology] Spearman correlations based on six condition means are coarse and unstable; the paper should report the actual correlation values, their uncertainty, and the handling of tied ranks, especially for near-flat models like the early OLMo checkpoints.
  4. [Figure 1] The figure caption does not explicitly define the order of the six condition bars; please add a legend or a sentence such as 'conditions are GP/non-GP for plausible, implausible, and reflexive verb types'.
  5. [Section 6.1] The sentence 'The man hunted the child. seem to be too out of distribution for our LLMs to generate' contains a grammatical error and should be rephrased; also, the format metric description says 'consistes' instead of 'consists'.
  6. [General] The paper does not state whether the experimental materials, human response data, and model output probabilities will be released; providing these would strengthen reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is an empirical human-LLM comparison whose central effects are supported by independent human data and multiple non-GPT model families.

full rationale

The paper does not derive its conclusions from its own inputs by construction. The hypotheses are drawn from prior psycholinguistic work (Christianson et al. 2001; Patson et al. 2009) and tested on both human participants and a wide array of LLM families. The only self-citation, Amouyal et al. (2024), is used in Section 3.1 to motivate using GPT-4 to rate stimulus plausibility for the plausibility manipulation. This is not load-bearing circularity: the plausibility effect is independently observed in human responses (p = 4.11e-16) and appears across GPT, Llama, Qwen, Gemma, and OLMo models, not just in GPT models whose ratings selected the stimuli. The plausibility ratings are not a fitted parameter later relabeled as a prediction; they are a stimulus-selection tool. The Section 5 correlations compare independently measured human accuracies with LLM answer-token probabilities, with no equation that equates the two by definition. The binary scoring convention in Section 2 ('we consider "yes" to be a wrong answer here, whereas "no" is considered the right answer') is a coding assumption inherited from the garden-path literature and is not a circular derivation, especially since converging evidence comes from paraphrase and text-to-image tasks. No uniqueness theorem from the authors is invoked, no ansatz is smuggled in via self-citation, and no known result is merely renamed as a new organization. The limitations section honestly notes the absence of reading-time and eye-gaze data, but that is a completeness concern, not a circularity concern.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new theoretical entities. It relies on standard psycholinguistic assumptions about garden-path misanalysis and on the validity of LLM probability scores as comprehension measures. The main empirical burden is the correctness labeling of GP questions.

assumptions (4)
  • domain assumption The correct answer to the GP comprehension question is 'No' even though the sentence permits a reading where the answer is 'Yes' (i.e., the intended final parse is the reference).
    Section 2: 'we consider yes to be a wrong answer here, whereas no is considered the right answer.' This is standard in psycholinguistic GP research but is a substantive labeling choice on which all accuracy scores depend.
  • domain assumption GPT-4 plausibility ratings (1-7 scale) are a valid proxy for human plausibility for stimulus selection.
    Section 3.1: materials for the plausibility manipulation are selected using GPT-4 ratings, relying on the authors' prior work (Amouyal et al. 2024). If GPT-4 ratings diverge from human judgments, the plausible/implausible contrast could be confounded.
  • domain assumption Average next-token probability of the correct answer under few-shot prompting reflects LLM comprehension of the sentence-question pair.
    Section 4.1 describes extracting probabilities of correct/incorrect answer tokens across 8 prompts. This standard practice assumes token probabilities track model understanding; floor/ceiling effects vary by model.
  • domain assumption Word-by-word presentation at 400 ms per word engages the same reading-comprehension processes as natural reading.
    Section 3.1: humans read word-by-word (400 ms + 100 ms blank). This is a self-paced reading style common in psycholinguistics, but it differs from the full-sentence presentation given to LLMs, so the human-LM comparison assumes this format difference does not drive the results.

how reviews work

0 comments
Cite this review

Pith. "Pith review of When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models." pith.science (2026). https://pith.science/paper/DUCJR4A3

@misc{pith2026250209307,
  author       = {Pith},
  title        = {Pith review of: When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DUCJR4A3}},
  note         = {Machine review of arXiv:2502.09307}
}
read the original abstract

Modern Large Language Models (LLMs) have shown human-like abilities in many language tasks, sparking interest in comparing LLMs' and humans' language processing. In this paper, we conduct a detailed comparison of the two on a sentence comprehension task using garden-path constructions, which are notoriously challenging for humans. Based on psycholinguistic research, we formulate hypotheses on why garden-path sentences are hard, and test these hypotheses on human participants and a large suite of LLMs using comprehension questions. Our findings reveal that both LLMs and humans struggle with specific syntactic complexities, with some models showing high correlation with human comprehension. To complement our findings, we test LLM comprehension of garden-path constructions with paraphrasing and text-to-image generation tasks, and find that the results mirror the sentence comprehension question results, further validating our findings on LLM understanding of these constructions.

Figures

Figures reproduced from arXiv: 2502.09307 by the authors.

Figure 1
Figure 1. Top: The manipulations made to an example garden-path sentence along with predictions from hu￾mans and LLMs for these sentences. Bottom: human and the Gemma-2-9B average performance on the dif￾ferent experimental conditions. The behaviour of hu￾mans and Gemma-2-9B is similar. and Wang, 2024; Kuribayashi et al., 2025). While LLMs mostly succeed where humans succeed, less is known on whether LLMs fail where humans fai… view at source ↗
Figure 2
Figure 2. Dall-e-3 incorrectly generates an image where the boy washes the dog (Left) given a GP sentence, but [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Performance of models from all families on our experimental conditions. Models with an “-Inst” suffix [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (6 more)
Figure 5
Figure 5. Figure 5: Spearman rank correlation per model family [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Average paraphrase accuracy for each condition per model family [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Example image for each of our manually-assigned labels for text-to-image generation examples. From [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Proportion of images classified as correctly [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 9
Figure 9. Figure 9: Example of a prompt [PITH_FULL_IMAGE:figures/full_fig_p015_9.png]
Figure 10
Figure 10. Figure 10: Example of a paraphrase task prompt [PITH_FULL_IMAGE:figures/full_fig_p016_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

35 extracted references · 17 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Merouane Debbah, Etienne Goffinet, Daniel Heslow, Julien Launay, Quentin Malartic, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023. Falcon-40B : an open large language model with state-of-the-art performance

  4. [4]

    Samuel Amouyal, Aya Meltzer-Asscher, and Jonathan Berant. 2024. https://aclanthology.org/2024.findings-eacl.12 Large language models for psycholinguistic plausibility pretesting . In Findings of the Association for Computational Linguistics: EACL 2024, pages 166--181, St. Julian ' s, Malta. Association for Computational Linguistics

  5. [5]

    Suhas Arehalli, Brian Dillon, and Tal Linzen. 2022. https://api.semanticscholar.org/CorpusID:253098758 Syntactic surprisal from neural models predicts, but underestimates, human processing difficulty from syntactic ambiguities . ArXiv, abs/2210.12187

  6. [6]

    Charlotte Cacheteux and Jean-Rémi King. 2022. https://doi.org/10.1038/s42003-022-03036-1 An activation-based model of sentence processing as skilled memory retrieval . Nature, pages 375--419

  7. [7]

    Kiel Christianson, Jack Dempsey, Anna Tsiola, and Maria Goldshtein. 2022. What if they're just not that into you (or your experiment)? on motivation and psycholinguistics. In Psychology of learning and motivation, volume 76, pages 51--88. Elsevier

  8. [8]

    Kiel Christianson, Andrew Hollingworth, John F Halliwell, and Fernanda Ferreira. 2001. Thematic roles assigned along the garden path linger. Cognitive psychology, 42(4):368--407

Show all 35 references
  1. [9]

    Kiel Christianson, Carrick C Williams, Rose T Zacks, and Fernanda Ferreira. 2006. Misinterpretations of garden-path sentences by older and younger adults. Discourse Processes, 42:205--238

  2. [10]

    Fernanda Ferreira and John M Henderson. 1990. Use of verb information in syntactic parsing: evidence from eye movements and word-by-word self-paced reading. Journal of Experimental Psychology: Learning, Memory, and Cognition, 16(4):555

  3. [11]

    Alex B Fine, T Florian Jaeger, Thomas A Farmer, and Ting Qian. 2013. Rapid expectation adaptation during syntactic comprehension. PloS one, 8(10):e77661

  4. [12]

    Susan M Garnsey, Neal J Pearlmutter, Elizabeth Myers, and Melanie A Lotocky. 1997. The contributions of verb bias and plausibility to the comprehension of temporarily ambiguous sentences. Journal of memory and language, 37(1):58--93

  5. [13]

    Team Gemini. 2024. https://arxiv.org/abs/2403.05530 Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context . Preprint, arXiv:2403.05530

  6. [14]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...

  7. [15]

    Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack...

  8. [16]

    Michael Hanna and Aaron Mueller. 2024. https://arxiv.org/abs/2412.05353 Incremental sentence processing mechanisms in autoregressive transformer language models . Preprint, arXiv:2412.05353

  9. [17]

    Jennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox, and Roger Levy. 2020. https://doi.org/10.18653/v1/2020.acl-main.158 A systematic assessment of syntactic generalization in neural language models . In Proceedings of the 58th Annual Meeting of the Association for Computationa...

  10. [18]

    Tovah Irwin, Kyra Wilson, and Alec Marantz. 2023. https://api.semanticscholar.org/CorpusID:258378287 Bert shows garden path effects . In Conference of the European Chapter of the Association for Computational Linguistics

  11. [19]

    Tatsuki Kuribayashi, Yohei Oseki, Souhaib Ben Taieb, Kentaro Inui, and Timothy Baldwin. 2025. Large language models are human-like internally. arXiv preprint arXiv:2502.01615

  12. [20]

    Andrew Li, Xianle Feng, Siddhant Narang, Austin Peng, Tianle Cai, Raj Sanjay Shah, and Sashank Varma. 2024. https://api.semanticscholar.org/CorpusID:270063738 Incremental comprehension of garden-path sentences by large language models: Semantic interpretation, syntactic re-ana...

  13. [21]

    Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016. https://doi.org/10.1162/tacl_a_00115 Assessing the ability of LSTM s to learn syntax-sensitive dependencies . Transactions of the Association for Computational Linguistics, 4:521--535

  14. [22]

    OpenAI. 2023. https://arxiv.org/abs/2303.08774 Gpt-4 technical report . ArXiv preprint, abs/2303.08774

  15. [23]

    OpenAI. 2024. https://cdn.openai.com/papers/DALL_E_3_System_Card.pdf Dall-e 3 system card . OpenAI

  16. [24]

    Nikole D Patson, Emily S Darowski, Nicole Moon, and Fernanda Ferreira. 2009. Lingering misinterpretations in garden-path sentences: evidence from a paraphrasing task. Journal of Experimental Psychology: Learning, Memory, and Cognition, 35(1):280

  17. [25]

    Qwen Team . 2024. https://qwenlm.github.io/blog/qwen2.5/ Qwen2.5: A party of foundation models

  18. [26]

    Adrielli Lopes Rego, Joshua Snell, and Martijn Meeter. 2024. https://api.semanticscholar.org/CorpusID:269527476 Language models outperform cloze predictability in a cognitive model of reading . PLOS Computational Biology, 20

  19. [27]

    Yuqi Ren, Renren Jin, Tongxuan Zhang, and Deyi Xiong. 2024. https://api.semanticscholar.org/CorpusID:268041437 Do large language models mirror cognitive language processing? ArXiv, abs/2402.18023

  20. [28]

    Hosseini, Nancy Kanwisher, Joshua B

    Martin Schrimpf, Idan Asher Blank, Greta Tuckute, Carina Kauf, Eghbal A. Hosseini, Nancy Kanwisher, Joshua B. Tenenbaum, and Evelina Fedorenko. 2021. https://doi.org/10.1073/pnas.2105646118 The neural architecture of language: Integrative modeling converges on predictive proce...

  21. [29]

    Kun Sun and Rong Wang. 2024. https://api.semanticscholar.org/CorpusID:268680327 Computational sentence-level metrics predicting human sentence comprehension . ArXiv, abs/2403.15822

  22. [30]

    Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan...

  23. [31]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, Aur'elien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. https://arxiv.org/abs/2302.139...

  24. [32]

    John C Trueswell, Michael K Tanenhaus, and Christopher Kello. 1993. Verb-specific constraints in sentence processing: separating effects of lexical preference from garden-paths. Journal of Experimental psychology: Learning, memory, and Cognition, 19(3):528

  25. [33]

    Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019. https://doi.org/10.1162/tacl_a_00290 Neural network acceptability judgments . Transactions of the Association for Computational Linguistics, 7:625--641

  26. [34]

    Ethan Gotlieb Wilcox, Jon Gauthier, Jennifer Hu, Peng Qian, and Roger Levy. 2022. Learning syntactic structures from string input. Algebraic Structures in Natural Language, page 113

  27. [35]

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, and 40 others. 2024. Qwen2 technical repo...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.