REVIEW 3 major objections 5 minor 26 references
Human-Like Anaphor Resolution in Large Language Models
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Large language models show human-like sensitivity to discourse prominence and distance in anaphor resolution, but largely miss semantic interference effects.
desk verdict A transparent, novel evaluation of six classic cognitive factors in anaphor resolution across five LLMs, but the central claim of selective alignment rests on descriptive bar-chart patterns with no inferential statistics; worth peer review as a promising descriptive study that needs statistical backing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the surprisal linking hypothesis—the assumption that a model's $-\log_2 p$ for the anaphor token indexes processing difficulty the way reading time does for humans—combined with comprehension questions scored by an automatic judge. These two measures are applied to factor-manipulated texts drawn from prior cognitive science experiments, so each text's four versions isolate one accessibility factor at a time. What carries the argument is the pattern across conditions: if models show the same ordinal ordering as human reading-time and accuracy findings, they exhibit cognitive alignment.
What would settle it
Collect human ratings on the same model-generated comprehension answers for all text versions; if the condition ordering in accuracy does not reproduce for any model, the accuracy-side evidence for cognitive alignment collapses.
Extended reading notes
Core claim
Using materials from earlier studies of anaphor resolution, the paper manipulates six factors orthogonally: antecedent topicality, sentential distance, spatial distance, temporal duration, semantic overlap between anaphor and antecedent, and semantic interference from a non-antecedent distractor. For each of five open-weight LLMs, it measures mean surprisal on the anaphor and accuracy on a comprehension question about the antecedent. The predicted human pattern is that resolution is fastest and most accurate when the antecedent is topicalized, close in surface or situational distance, semantically similar to the anaphor, and free of competing typical distractors. The results show that several models track the topicality and sentential-distance pattern in surprisal, that two models track spatial and temporal distance, and that only isolated models show the predicted semantic patterns. The authors conclude that some LLMs, especially Mistral-7B and GPT-2-XL, approximate human anaphor resolution on prominence and distance factors but diverge on semantic interference.
Load-bearing premise
The reported comprehension-accuracy results depend on an automatic judge scoring model responses correctly enough to preserve the ordering across conditions, and the paper itself flags that this judge can differ from human judgments.
Editorial extensions
If this is right
- Surprisal at the anaphor can serve as a process-level measure for discourse phenomena, not just word- and sentence-level effects.
- Models like Mistral-7B and GPT-2-XL become candidate tools for generating behavioral predictions about human anaphor resolution.
- Benchmarks for coreference and anaphor resolution should manipulate the cognitive factors that drive human resolution, since raw accuracy alone cannot distinguish human-like processing from shallow heuristics.
- Semantic interference is a reliable diagnostic dimension on which current LLMs diverge from human comprehension.
- The documented recency and ceiling effects in larger models can mask distance-based effects, so model size alone does not determine cognitive alignment.
Reading between the lines
- An extension the paper does not run: replacing the automatic judge with human raters on the same model outputs would show whether the reported accuracy orderings are artifacts of judge bias.
- If the absence of semantic interference is genuine, it points to a structural difference—current LLMs lack the competitive memory retrieval that cognitive theories posit—which would recommend memory-augmented architectures, a prediction the paper does not make.
- The small number of texts makes the conclusions descriptive; a larger-item replication with inferential statistics is the natural next step, which the authors explicitly defer.
- The asymmetry between prominence and distance effects on the one hand and semantic effects on the other suggests a possible ordering of difficulty for cognitive alignment in models: surface accessibility first, situation-model distance second, semantic competition last.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper evaluates whether five open-weight LLMs (GPT-2-XL, Llama-3.1-8B, Pythia-12B, Mistral-7B, Mistral-24B) are sensitive to six factors that cognitive science has shown to affect human anaphor resolution: antecedent topicality, sentential distance, spatial distance, temporal duration, semantic similarity, and semantic interference. Three experiments use published human materials, and each compares four text versions with model surprisal at the anaphor as a reading-time proxy and model accuracy on a comprehension question as a success measure, with accuracy scored by an LLM-as-a-judge. The paper reports qualitative ordinal patterns and concludes that LLMs show selective cognitive alignment: more consistent human-like sensitivity to discourse-prominence and distance-based factors than to semantic interference, suggesting that LLMs can serve as partial cognitive models of discourse processing while localizing some divergences.
Significance. If the reported selective-alignment pattern were statistically established, this would be a useful contribution to the cognitive-science evaluation of LLMs: it moves beyond coreference-accuracy benchmarks, uses classic human experimental materials, tests predictions derived from external cognitive science findings rather than fitting model parameters, and publicly releases code and data. The asymmetry between distance/discourse factors and semantic-interference factors is also a concrete, falsifiable claim about where current LLM architectures diverge from human discourse comprehension. However, the current evidentiary basis is almost entirely descriptive, so the significance of the contribution is conditional on the missing inferential support and on validation of the automatic judge.
major comments (3)
- [General Discussion, limitations; Figures 1-6] The central claim of selective cognitive alignment rests on ordinal patterns of descriptive means computed over 16 texts in Experiment 1 and 19 texts in Experiments 2 and 3. The General Discussion itself states that the item count 'precluded running statistical analyses' and describes the trends as 'informal.' With five models x two measures x three experiments, there are many opportunities for chance to yield the predicted A/D ordering or a full predicted pattern, and several error bars visibly overlap, especially in the accuracy panels (e.g., Figures 2, 4, and 6). The surprisal results, which carry the headline generalization, are reported without any reliability estimate. This is not an internal inconsistency, but it means the empirical claim that 'some LLMs exhibit human-like sensitivity' is currently unestablished. The released data would permit item-level mixed-effects models or bootstrap confidence intervals for each model-by-factor contrast; these should be added, with appropriate correction for multiple comparisons.
- [Experiment 1, Comprehension Question Answering; General Discussion limitations] All comprehension-accuracy results pass through an automatic judge (Gemini-2.5-flash-preview) that the authors acknowledge 'introduces known biases and can be different from human judgments.' No human-agreement sample, judge reliability statistic, or error analysis is reported. Because accuracy is one of only two dependent measures and is used in every experiment to support the selective-alignment narrative, the paper needs at least a validation subset scored by human annotators, or an error analysis showing that judge disagreements do not systematically favor the predicted conditions.
- [Table 1 and Results sections] The summary table does not state what count as 'human-like' for each cell, and it is not always consistent with the text. In Experiment 1, the text reports that GPT-2-XL, Llama-3.1-8B, Pythia-12B, and Mistral-7B showed the predicted surprisal pattern, whereas Table 1 codes 'All models'; in Experiment 2, the text identifies only Llama-3.1-8B and Mistral-7B as showing the predicted pattern, while Table 1's 'All models except Pythia-12B' implies additional models did so. Similar ambiguities appear in the accuracy columns, where 'None/weak effects' coexists with 'all of the other models correctly order versions A and D.' Since the main conclusion about selective alignment is a meta-level summary of these table cells, the coding criteria need to be explicit and applied consistently.
minor comments (5)
- [Experiment 1, Procedure and Dependent Measures] The normalization equation is unnumbered and its notation is loose; please label it as Equation (1), define whether the min and max are taken over the four version-level surprisals for each text, and justify min-max normalization over z-scoring, especially given that it can change the relative weighting of texts with different within-text variance.
- [General Discussion, Table 1] The model name 'Pythia' is used instead of 'Pythia-12B', and the paper alternates between 'LLaMa' and 'Llama'; please harmonize names across the text, tables, and figures.
- [Experiment 3, Design and Materials] The description says the base texts are 'the version D (far spatial distance, long temporal duration) texts from Experiment 2'; please clarify whether the surrounding discourse and the anaphor tokens are identical to Experiment 2 or were modified, and how the semantic manipulations were crossed with the pre-existing spatial/temporal properties.
- [Experiment 1, Comprehension Question Answering] The date-stamped judge name 'Gemini-2.5-flash-preview-09-2025' will age quickly; please report the exact model version used and consider adding a note about reproducibility as the judge model changes.
- [Figures 1-6] The figures report only means and standard errors; adding item-level points or box plots would make the degree of overlap and the ordinal claims much easier to assess visually.
Circularity Check
No significant circularity: the study is an empirical evaluation with predictions imported from external cognitive science results, and no fitted parameter is renamed as a prediction.
full rationale
This paper is an empirical evaluation rather than a derivation, so the main circularity patterns do not apply. The predictions (e.g., that surprisal should be lowest for version A, highest for version D, and intermediate for B and C) are taken from published cognitive science findings such as O'Brien (1987), Clark and Sengul (1979), Morrow et al. (1987), Anderson et al. (1983), and Varma and Janssen (2019). No model parameter is fitted to the experimental data and then called a prediction; the only transformation of the raw surprisal values is a min-max normalization within each text, which is explicitly defined and does not encode the predicted ordering. The comprehension accuracy scores are produced by an external judge, Gemini-2.5-flash-preview, which is not one of the five evaluated LLMs, so there is no self-referential loop in which the tested model scores itself. The paper's self-citations (e.g., Shah & Varma, 2025; Li et al., 2024; Varma & Janssen, 2019) are used to motivate the cognitive-modeling framing and to supply stimulus materials, not to justify the empirical outcome. The paper itself candidly notes its largest limitation: the small number of items precluded inferential statistics, and it describes the reported results as cautiously interpreted descriptive trends. That is a statistical robustness concern, not a circularity concern. The central claim is therefore self-contained against external benchmarks and does not reduce to its own inputs.
Assumptions & free parameters
assumptions (3)
- domain assumption The standard linking hypothesis that higher model surprisal corresponds to longer human reading time is valid.
- domain assumption The six cognitive factors manipulate human anaphor resolution as described in the cited literature.
- domain assumption Gemini-2.5-flash-preview produces valid and reliable binary accuracy judgments for model answers.
Cite this review
Pith. "Pith review of Human-Like Anaphor Resolution in Large Language Models." pith.science (2026). https://pith.science/paper/MG27FO3S
@misc{pith2026260805630,
author = {Pith},
title = {Pith review of: Human-Like Anaphor Resolution in Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/MG27FO3S}},
note = {Machine review of arXiv:2608.05630}
}
read the original abstract
Anaphors are expressions that refer to other expressions, called antecedents. The process of connecting the two is called resolution. Cognitive science has identified multiple factors that affect the speed and success of anaphor resolution, including discourse structure, situation-model properties, and semantic factors. Here, we investigate whether these factors also affect anaphor resolution in five Large Language Models (LLMs) with open weights: GPT-2-XL, Llama-3.1-8B, Pythia-12B, Mistral-7B, and Mistral-24B. To model processing difficulty, we adopt the standard linking hypothesis that relates human reading times to model surprisal at the anaphor. As a second behavioral measure, we compare model accuracy to human accuracy on comprehension questions probing the antecedents of anaphors. The results show selective cognitive alignment: some LLMs exhibit human-like sensitivity to discourse prominence and distance-based factors in anaphor resolution, while showing weaker or absent sensitivity to semantic interference effects. These findings delimit the conditions under which LLMs approximate human anaphor resolution.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Anderson, A., Garrod, S. C., & Sanford, A. J. (1983). The accessibility of pronominal antecedents as a function of episode shifts in narrative text.The Quarterly Journal of Experimental Psychology A,35, 427–440. BabyLMOrganizers.(2025,November).Findingsofthethird BabyLM challenge: Accelerating language modeling re- search with cognitively plausible data. ...
work page 1983
-
[2]
Jumelet, T. Linzen, A. Mueller, C. Ross, R. S. Shah, A. Warstadt,E.G.Wilcox,&A.Williams(Eds.),Proceedings ofthefirstbabylmworkshop(pp.399–420).Associationfor Computational Linguistics. https://doi.org/10.18653/v1/ 2025.babylm-main.28
doi:10.18653/v1/ 2025
-
[3]
Cambria, E. (2025). Semantics processing. Springer
work page 2025
-
[4]
Clark, H. H., & Sengul, C. J. (1979). In search of referents for nouns and pronouns.Memory & Cognition,7, 35–41. https://doi.org/10.3758/BF03196932
-
[5]
Corbett, A. T. (1984). Prenominal adjectives and the disam- biguation of anaphoric nouns.Journal of Verbal Learning and Verbal Behavior,23, 683–695
work page 1984
-
[6]
Daneman, M., & Carpenter, P. A. (1980). Individual differ- ences in working memory and reading.Journal of Verbal Learning and Verbal Behavior,19, 450–466
work page 1980
-
[7]
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). Bert:Pre-trainingofdeepbidirectionaltransformersforlan- guage understanding
work page 2019
-
[8]
Gan, Y., Poesio, M., & Yu, J. (2024, May). Assessing the ca- pabilitiesoflargelanguagemodelsincoreference:Aneval- uation. In N. Calzolari, M.-Y. Kan, V. Hoste, A. Lenci, S. Sakti,&N.Xue(Eds.),Proceedingsofthe2024jointinter- nationalconferenceoncomputationallinguistics,language resources and evaluation (lrec-coling 2024)(pp. 1645– 1665). ELRA; ICCL. https:...
work page 2024
Show all 26 references
-
[9]
Garrod, S., & Sanford, A. J. (1977). Interpreting anaphoric relations: The integration of semantic information while reading.Journal of Verbal Learning and Verbal Behavior, 16, 77–90
1977
- [10]
-
[11]
Hale, J. (2001). A probabilistic earley parser as a psycholin- guistic model.Proceedings of the 2nd Meeting of NAACL. Hu,J.,Chen,S.Y.,&Levy,R.(2020).Acloserlookattheper- formanceofneurallanguagemodelsonreflexiveanaphorli- censing.ProceedingsoftheSocietyforComputationinLin- gui...
2001
-
[12]
Lee, S.-H., & Schuster, S. (2022). Can language models cap- ture syntactic associations without surface cues? a case studyofreflexiveanaphorlicensinginEnglishcontrolcon- structions.Proceedings of the Society for Computation in Linguistics 2022, 206–211. https://aclanthology.or...
2022
- [13]
-
[14]
Manikantan, K., Tapaswi, M., Gandhi, V., & Toshniwal, S. (2024). Identifyme: A challenging long-context mention resolution benchmark
2024
-
[15]
S., & Magliano, J
McNamara, D. S., & Magliano, J. (2009). Toward a compre- hensive model of comprehension [In B. Ross (Ed.), The psychology of learning and motivation]. MistralAITeam.(2025,January).Mistralsmall3[Accessed: 2026-01-30]. https://mistral.ai/news/mistral-small-3
2009
-
[16]
G., Greenspan, S
Morrow, D. G., Greenspan, S. L., & Bower, G. H. (1987). Accessibilityandsituationmodelsinnarrativecomprehen- sion.Journal of Memory and Language,26, 165–187. Novák, M., et al. (2025). Findings of the fourth shared task on multilingual coreference resolution.Proceedings of the ...
1987 doi
-
[17]
Learning, Memory, and Cognition,13, 278–290. OpenAI. (2025, August). Introducing gpt-5 [Accessed: 2026- 01-30]. https://openai.com/index/introducing-gpt-5/ Pandit,O.,&Hou,Y.(2021).Probingforbridginginferencein transformerlanguagemodels.Proceedingsofthe2021Con- ference of the N...
2021 doi
-
[18]
T., et al
Piantadosi, S. T., et al. (2024). Why concepts are (probably) vectors.Trends in Cognitive Sciences,28, 844–856
2024
-
[19]
Sutskever, I. (2019). Language models are unsupervised multitask learners
2019
-
[20]
Rinck, M., & Bower, G. H. (1995). Anaphora resolution and the focus of attention in situation models.Journal of Mem- ory and Language,34, 110–131
1995
-
[21]
Boyes-Braem, P. (1976). Basic objects and natural cate- gories.Cognitive Psychology,9, 382–440. Shah,R.S.,&Varma,S.(2025).Thepotentialandthepitfalls of using pre-trained language models as cognitive science theories. https://arxiv.org/abs/2501.12651
1976 arXiv
-
[22]
Talukdar, C., & Rahman, M. (2025). Coreference resolution in machine learning: A survey.2025 IEEE Guwahati Sub- section Conference (GCON). https://doi.org/10.1109/ GCON65540.2025.11173320
2025
-
[23]
N., & Radvansky, G
Thompson, A. N., & Radvansky, G. A. (2016). Event bound- ariesandanaphoricreference.PsychonomicBulletin&Re- view,23, 849–856. vanDijk,T.A.,&Kintsch,W.(1983).Strategiesofdiscourse comprehension. Academic Press
2016
-
[24]
Varma, S., & Janssen, A. (2019). The structure of situation models as revealed by anaphor resolution.Language Sci- ences,72, 104–115
2019
- [25]
-
[26]
Zwaan, R. A. (1996). Processing narrative time shifts.Jour- nal of Experimental Psychology: Learning, Memory, and Cognition,22, 1196–1207
1996
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.