REVIEW 3 major objections 6 minor 1 cited by
KRISTEVA: Close Reading as a Novel Task for Benchmarking Interpretive Reasoning
T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read KRISTEVA, the first close-reading benchmark for LLMs, finds that best models score 49.7–69.7 percent while experienced human readers lead on 10 of 11 tasks.
desk verdict KRISTEVA is a genuinely new and useful close-reading benchmark, but its main 'humans beat LLMs on 10/11 tasks' claim rests on a three-evaluator per-task-max baseline and needs a stronger human study. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is KRISTEVA's eleven task types, arranged in three ascending clusters that approximate the stages of close reading. The first cluster asks models to extract stylistic features—detect the device, locate it, name its elements, infer its purpose and significance. The second asks them to retrieve and rank relevant external contexts from parametric knowledge. The third requires multi-hop reasoning that connects a specific stylistic feature to a specific external context and articulates the connection. The benchmark's ground truth comes from instructor-graded student essays, and its wrong answers are generated by a reasoning LLM to be plausible but weaker interpretations; the claim that accuracy measures interpretive reasoning depends on that ordering of answer quality.
What would settle it
Take a random sample of 200 KRISTEVA questions and ask experienced close readers to rank the four options without knowing the keyed answer. If in a substantial share of items the LLM-generated distractor is ranked as more reasonable than the keyed answer, then the accuracy gap between humans and models reflects option quality, not close-reading ability.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that close reading can be operationalized as a structured, multiple-choice evaluation, and that under that operationalization LLMs show partial but incomplete competence. The strongest model reaches 69.7% overall and 64.3% on reasoning-heavy questions, while the best human evaluator outperforms the best model on 8 of 11 tasks, and human evaluators collectively match or beat every model on 10 of 11 tasks. This gap survives even though the human baseline is likely conservative, since evaluators had to adapt to an MCQ format far removed from open-ended close reading. The benchmark therefore establishes a measurable distance between machine and human interpretive skill and provides a reusable scaffold for closing it.
Load-bearing premise
The benchmark is valid only if the correct answers taken from high-scoring student essays remain more defensible than the LLM-generated wrong choices, so that picking the keyed answer reflects interpretive skill rather than an artifact of how the options were written.
Editorial extensions
If this is right
- A standardized evaluation now exists for a form of reasoning that relies on relative plausibility rather than a uniquely correct answer, so future work can compare models and track progress on interpretive judgment.
- Because the tasks combine figurative-language understanding with multi-hop reasoning, gains on KRISTEVA should transfer to both NLP areas rather than to a single narrow skill.
- The 10-of-11 human advantage indicates that current LLM pretraining and instruction tuning leave room for improvement in text-grounded aesthetic and interpretive judgment.
- The 49-essay, 1,331-question pipeline shows that routine college classroom writing can become a scalable source of benchmark data.
Reading between the lines
- If the benchmark's validity assumption holds, the task ordering implies a diagnostic profile: models' weakest points are significance ranking and feature-context reasoning, suggesting that connecting form to outside knowledge—not detecting devices—is the current bottleneck.
- A testable extension would convert KRISTEVA to free-response format and score it with expert rubrics; the paper names this as future work, and it would reveal how much of the human-model gap is an artifact of multiple-choice options.
- The same question-generation pipeline could be applied to other humanities disciplines—history, philosophy, art criticism—that teach evidence-based interpretation, giving NLP a family of benchmarks for judgment rather than fact retrieval.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces KRISTEVA, a new benchmark of 1,331 multiple-choice questions for close-reading interpretation, built from 49 student essays from a college literature course and organized into 11 task types following the CRIT pedagogical framework. The tasks span stylistic feature extraction (Q1–Q6), external context retrieval (Q7–Q9), and multi-hop feature–context reasoning (Q10–Q11). The dataset is constructed by using GPT-4o to extract structured features and answers from the essays and o1-preview to generate distractors for 7 of the 11 task types. The authors evaluate 19 LLMs in a zero-shot setting and report a best model (Phi-4) at 69.7% overall accuracy, compared with a human baseline from three PhD-student evaluators, and claim that the best LLM trails experienced human readers on 10 of 11 tasks.
Significance. If the benchmark is valid, it fills a real gap: existing multi-discipline benchmarks largely omit literature and interpretive reasoning, despite close reading being a core college-level critical-thinking skill. The paper's strengths include a publicly available dataset, a detailed and reproducible construction pipeline, a transparent evaluation harness, and a task decomposition that connects figurative-language understanding with multi-hop reading comprehension in a novel way. The task structure itself, grounded in an actual pedagogical framework, is a useful contribution independent of the human-model comparison. However, the central human-superiority claim and the construct validity of the gold labels both rest on assumptions that the current manuscript does not adequately support.
major comments (3)
- [§4.2, Table 2] The human baseline is too small and the '10 out of 11' claim is not statistically supported. The sentence 'which accounts for percentage of the dataset' is missing the actual number, and each task is answered by at most three evaluators, with per-task scores such as Q5 (75/0/0) and Q8 (0/66.7/28.6) showing extreme instability (average pairwise standard deviation 29.3 vs. 5.47 for LLMs). Table 2's 'top-line human performance' takes the maximum over evaluators per task, so the claim that the best LLM trails humans on 10/11 tasks compares best-of-three humans against best-of-many models per task; the human weighted average (65.6) is actually below Phi-4 (69.7). Please report the exact per-task sample sizes, add confidence intervals or significance tests, and either compare against a single pre-registered human aggregator or soften the claim to reflect the uncertainty.
- [§3.2.1–3.2.2, Appendix C.1] The gold answers and distractors are both produced by LLMs, which threatens construct validity in a way that affects the headline accuracy numbers. GPT-4o performs the structured extraction from which questions and answers are built, and o1-preview generates all three distractors for 1,178 of the questions (7 of 11 question types). Because the evaluated models come from the same model lineages, the benchmark may partly measure how well models second-guess distractor-writing conventions rather than close-reading ability. The instructor check described in §3.1.2 only verifies that less-reasonable distractors could in principle be generated; it does not validate the final items. Please provide evidence that the correct options are not systematically distinguishable from the distractors (e.g., human judges cannot identify the gold answer above chance given the passage and choices, or a perturbation/format-control analysis), or report results on a subset with human-authored distractors.
- [§5, §6] The claim that models 'trail' humans on 10/11 tasks is also internally at odds with the aggregate human performance in Table 2. On Q5 and Q8, for instance, the human weighted averages are 28.8 and 25.1, far below the best LLM scores of 62.3 and 35.7, respectively; these are only converted into human 'wins' by taking the maximum over three evaluators, with a single evaluator contributing the high score in each case. The paper acknowledges variability in §6 but does not adjust the headline claim. The conclusion 'LLMs still lag behind human performance' should be qualified to 'lag behind the best of three human evaluators' unless a more robust human aggregate or statistical comparison is provided.
minor comments (6)
- [§4.2] The sentence 'which accounts for percentage of the dataset' appears to be missing the numeric value; please fix the typo and report the actual fraction of the 1,331 questions used for the human baseline.
- [Table 3, Q9 row] The full-question template for Q9 repeats the Q6 wording about 'the significance of this device' and 'its effects on the reader'; for a context-significance task it should refer to the context (e.g., 'the significance of this contextual information') rather than to a device.
- [Appendix B.2] The phrase 'boarder circumstances' should be 'broader circumstances'.
- [§4.1] The manuscript says 'We then extracte the answer'; this should be 'extracted'.
- [References] The reference to Comsa et al. (2022) lists the author as 'Iulia Coms, a'; the author name should be 'Iulia Comsa'.
- [Figure 2] Figure 2 is informative but the arrow conventions are not explained; a note in the caption clarifying that the three clusters correspond to the three progressive difficulty groups would improve readability.
Circularity Check
No significant circularity: the benchmark's correct answers are instructor-graded student-essay content, and the LLM-generated-distractor concern is a validity/contamination risk, not a definitional reduction.
full rationale
The benchmark's ground truth is not defined by the models being evaluated. Correct answers come from student close-reading essays filtered by the course instructor's grade and then manually checked by the instructor to ensure that three plausible distractors could be generated (Section 3.1.2). GPT-4o and o1-preview are used as construction tools for structured extraction and distractor generation (Section 3.2), but no fitted parameter is later renamed as a prediction, and no equation in the paper defines model accuracy in terms of the human baseline or vice versa. The headline human-vs-model comparison is an empirical measurement, not an identity forced by the construction pipeline. The paper explicitly flags the LLM-generated-distractor issue in Limitations 8.2: 'its MCQs contain LLM-generated distractors, which could bias the benchmark towards being more LLM-solvable.' This is a threat to construct validity and potential contamination, not circularity, because the correct answer was fixed by the human essay and instructor check before distractors were generated, and the scoring rule compares the model's choice against that fixed answer. The only self-citation (Sui et al., 2024) appears in a future-work aside about confabulation guardrails and is not load-bearing for the benchmark's central claims. No uniqueness theorem, ansatz-via-citation, or renaming of a known result carries the argument. The incomplete sentence in Section 4.2 ('which accounts for percentage of the dataset') and the lack of reported confidence intervals are statistical-reporting weaknesses, not circular steps.
Assumptions & free parameters
free parameters (3)
- minimum essay grade threshold =
80%
- number of distractors per MCQ =
3
- number of human evaluators =
3
assumptions (4)
- domain assumption The CRIT framework is a valid operationalization of close reading for evaluation purposes
- domain assumption Instructor-assigned grades on CRIT essays are a valid proxy for the quality of close reading
- domain assumption Multiple-choice questions can validly assess interpretive reasoning that lacks a unique correct answer
- domain assumption LLM-generated distractors are less reasonable than the human-derived correct answer and do not introduce systematic artifacts
Cite this review
Pith. "Pith review of KRISTEVA: Close Reading as a Novel Task for Benchmarking Interpretive Reasoning." pith.science (2026). https://pith.science/paper/46AV54WI
@misc{pith2026250509825,
author = {Pith},
title = {Pith review of: KRISTEVA: Close Reading as a Novel Task for Benchmarking Interpretive Reasoning},
year = {2026},
howpublished = {\url{https://pith.science/paper/46AV54WI}},
note = {Machine review of arXiv:2505.09825}
}
read the original abstract
Each year, tens of millions of essays are written and graded in college-level English courses. Students are asked to analyze literary and cultural texts through a process known as close reading, in which they gather textual details to formulate evidence-based arguments. Despite being viewed as a basis for critical thinking and widely adopted as a required element of university coursework, close reading has never been evaluated on large language models (LLMs), and multi-discipline benchmarks like MMLU do not include literature as a subject. To fill this gap, we present KRISTEVA, the first close reading benchmark for evaluating interpretive reasoning, consisting of 1331 multiple-choice questions adapted from classroom data. With KRISTEVA, we propose three progressively more difficult sets of tasks to approximate different elements of the close reading process, which we use to test how well LLMs may seem to understand and reason about literary works: 1) extracting stylistic features, 2) retrieving relevant contextual information from parametric knowledge, and 3) multi-hop reasoning between style and external contexts. Our baseline results find that, while state-of-the-art LLMs possess some college-level close reading competency (accuracy 49.7% - 69.7%), their performances still trail those of experienced human evaluators on 10 out of our 11 tasks.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
Tell, Don't Show: Leveraging Language Models' Abstractive Retellings to Model Literary Themes
Retell, a method that runs LDA on language models' abstractive retellings of literary passages, produces more theme-level topics than LDA alone or direct LM topic labeling.
Reference graph
Works this paper leans on
-
[1]
For each question object in the input JSON, generate exactly three (3) distractors
-
[2]
Append those three distractors plus the correct answer (in the fourth position) to a new array "choices" within that same question object
-
[3]
Do not add any additional commentary or fields; only add "choices" to each question object, containing [distractor1, distractor2, distractor3, correctAnswer] . The way you generate the three distractors depends on question_number: Q1 Here is a snippet of a poem: {selected_passage}, selected for literary analysis from the full poem: {full_passage}. One int...
-
[4]
A framework for few-shot language model evaluation. John Guillory. 2025. On Close Reading. University of Chicago Press. Jakob Hauser, Dániel Kondor, Jenny Reddish, Majid Be- nam, Enrico Cioni, Federica Villa, James S. Bennett, Daniel Hoyer, Pieter Francois, Peter Turchin, et al
work page 2025
-
[5]
Large language models’ expert-level global history knowledge benchmark (HiST-LLM). In The Thirty-eight Conference on Neural Information Pro- cessing Systems Datasets and Benchmarks Track. N. Katherine Hayles. 2010. How we read: Close, hyper, machine. ADE Bulletin, 150(18):62–79. Jack Hessel, Ana Marasovic, Jena D. Hwang, Lillian Lee, Jeff Da, Rowan Zeller...
work page 2010
-
[7]
ARB: Advanced reasoning benchmark for large language models. arXiv:2307.13692. Matthew Sims, Jong Ho Park, and David Bamman. 2019. Literary event detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3623–3634, Florence, Italy. Asso- ciation for Computational Linguistics. Dan Sinykin and Johanna Winan...
arXiv 2019
-
[8]
Cs-Bench: A comprehensive benchmark for large language models towards computer science mastery. arXiv:2406.08587. Zayne Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang, Manya Wadhwa, Prasann Singhal, Xinyu Zhao, Xi Ye, Kyle Mahowald, and Greg Dur- rett. 2024. To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.CoR...
arXiv 2024
-
[9]
PUB: A pragmatics understanding benchmark for assessing LLMs’ pragmatics capabilities. In Find- ings of the Association for Computational Linguistics: ACL 2024, pages 12075–12097, Bangkok, Thailand. Association for Computational Linguistics. Katherine Stasaski and Marti A Hearst. 2017. Multiple choice question generation utilizing an ontology. In Proceedi...
work page 2024
Show all 17 references
-
[10]
In Proceedings of the 60th Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 5375–5388, Dublin, Ireland
IMPLI: Investigating NLI models’ perfor- mance on figurative language. In Proceedings of the 60th Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 5375–5388, Dublin, Ireland. Association for Compu- tational Linguistics. Peiqi Sui...
-
[11]
In Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 14274–14284, Bangkok, Thailand
Confabulation: The surprising value of large language model hallucinations. In Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 14274–14284, Bangkok, Thailand. Association for Computational Linguistics. Ka...
2019 arXiv
-
[14]
understanding,
ChatMusician: Understanding and generating music intrinsically with LLM. In Findings of the As- sociation for Computational Linguistics: ACL 2024, pages 6252–6271, Bangkok, Thailand. Association for Computational Linguistics. Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, R...
2024 arXiv
-
[2016]
In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2358–2367, Berlin, Germany
A thorough examination of the CNN/Daily Mail reading comprehension task. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2358–2367, Berlin, Germany. Association for Computational Linguistics. 10 Daixuan ...
2023
-
[2018]
Transactions of the Association for Computational Linguistics, 6:287– 302
Constructing datasets for multi-hop reading comprehension across documents. Transactions of the Association for Computational Linguistics, 6:287– 302. Hayden White. 1999. Figural Realism: Studies in the Mimesis Effect. Johns Hopkins University Press. Zijun Yao, Yantao Liu, Xin...
1999
-
[2020]
arXiv:2002.04326
Reclor: A reading comprehension dataset re- quiring logical reasoning. arXiv:2002.04326. Ruibin Yuan, Hanfeng Lin, Yi Wang, Zeyue Tian, Shangda Wu, Tianhao Shen, Ge Zhang, Yuhang Wu, Cong Liu, Ziya Zhou, Liumeng Xue, Ziyang Ma, Qin Liu, Tianyu Zheng, Yizhi Li, Yinghao Ma, Yimi...
2002 arXiv
-
[2022]
MiQA: A benchmark for inference on metaphorical questions. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Asso- ciation for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), p...
2021
-
[2023]
arXiv:2309.12338
Artificial intelligence and aesthetic judgment. arXiv:2309.12338. Robin Jia and Percy Liang. 2017. Adversarial exam- ples for evaluating reading comprehension systems. In Proceedings of the 2017 Conference on Empiri- cal Methods in Natural Language Processing, pages 2021–2031,...
2017 arXiv
-
[2024]
arXiv preprint arXiv:2404.06283
LLMs’ reading comprehension is affected by parametric knowledge and struggles with hypotheti- cal statements. arXiv preprint arXiv:2404.06283. Don Bialostosky. 2006. Should college English be close reading? College English, 69(2):111–116. Jiahuan Cao, Yang Liu, Yongxin Shi, Ka...
2006 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.