{"id":"e23ae249-e4af-46fb-b31a-0f4db18ced83","arxiv_id":"2505.09825","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"KRISTEVA is the first close reading benchmark for large language models, and current models still underperform experienced human readers on 10 of its 11 tasks.","lead":"KRISTEVA is a new benchmark of 1,331 multiple-choice questions that tests how well language models can perform close reading of literature, a college skill centered on analyzing style and meaning. Evaluations of 19 large language models show they lag behind experienced human readers on 10 of the benchmark's 11 task types.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 10/11 human-superiority claim rests on a three-evaluator baseline compared per-task against the best model; §4.2 omits the subset size and no statistical tests are reported.","rationale":"The reader's weakest assumption concerns the MCQ construction pipeline preserving close-reading construct. That is a genuine validity concern, but the paper provides some safeguards: instructor filtering, manual checks, and human inspection of distractors. The human baseline, by contrast, has no statistical safeguards: three evaluators, an unstated subset size, no CIs, and a best-of-three per-task comparison that maximizes human performance. Since the abstract's headline is explicitly comparative, the baseline is the least secure condition for the central claim. If the human baseline is unrepresentative or its advantage is noise, the headline conclusion fails even if the benchmark itself is valid. The benchmark remains a useful contribution, so the CONDITIONAL verdict is appropriate, but the comparative claim should be made contingent on a stronger baseline.","tokens_in":20535,"tokens_out":10895,"duration_ms":100809,"concrete_test":"Recruit at least 10 experienced close readers to answer a stratified random sample of 300 questions (e.g., 30 per task) from KRISTEVA. Report per-task mean accuracy, 95% bootstrap CIs, and the difference between the human mean and the best LLM per task. Also report the actual subset size in §4.2. If the human mean does not exceed the best model on at least 8 of 11 tasks by more than the CI, or if any single human does not reproduce the 10/11 pattern, the headline claim should be revised to a conditional or task-specific statement.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2 constructs the human baseline from three PhD-student evaluators, each answering one or two unit tests. The sentence \"which accounts for percentage of the dataset\" is missing the actual number, so the per-task human sample sizes are unknown. Table 2 reports \"top-line human performance\" as the best evaluator per task; this is not any single human. Evaluator 2, the best overall (74.7), beats Phi-4 (69.7) on only 8 of 11 tasks (Section 5), and the human weighted average (65.6) is below Phi-4. Thus the \"10 out of 11\" advantage is an artifact of taking the maximum over three evaluators per task. With n=3 and average pairwise standard deviation 29.3 (vs 5.47 for LLMs), per-task differences such as Q5 (75/0/0) and Q8 (0/66.7/28.6) are not stable; no confidence intervals or significance tests are provided. The central claim that LLMs trail experienced human readers is therefore not statistically established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces KRISTEVA, a new benchmark of 1,331 multiple-choice questions for close-reading interpretation, built from 49 student essays from a college literature course and organized into 11 task types following the CRIT pedagogical framework. The tasks span stylistic feature extraction (Q1–Q6), external context retrieval (Q7–Q9), and multi-hop feature–context reasoning (Q10–Q11). The dataset is constructed by using GPT-4o to extract structured features and answers from the essays and o1-preview to generate distractors for 7 of the 11 task types. The authors evaluate 19 LLMs in a zero-shot setting and report a best model (Phi-4) at 69.7% overall accuracy, compared with a human baseline from three PhD-student evaluators, and claim that the best LLM trails experienced human readers on 10 of 11 tasks.","tokens_in":20658,"tokens_out":4294,"duration_ms":44416,"significance":"If the benchmark is valid, it fills a real gap: existing multi-discipline benchmarks largely omit literature and interpretive reasoning, despite close reading being a core college-level critical-thinking skill. The paper's strengths include a publicly available dataset, a detailed and reproducible construction pipeline, a transparent evaluation harness, and a task decomposition that connects figurative-language understanding with multi-hop reading comprehension in a novel way. The task structure itself, grounded in an actual pedagogical framework, is a useful contribution independent of the human-model comparison. However, the central human-superiority claim and the construct validity of the gold labels both rest on assumptions that the current manuscript does not adequately support.","major_comments":[{"comment":"The human baseline is too small and the '10 out of 11' claim is not statistically supported. The sentence 'which accounts for percentage of the dataset' is missing the actual number, and each task is answered by at most three evaluators, with per-task scores such as Q5 (75/0/0) and Q8 (0/66.7/28.6) showing extreme instability (average pairwise standard deviation 29.3 vs. 5.47 for LLMs). Table 2's 'top-line human performance' takes the maximum over evaluators per task, so the claim that the best LLM trails humans on 10/11 tasks compares best-of-three humans against best-of-many models per task; the human weighted average (65.6) is actually below Phi-4 (69.7). Please report the exact per-task sample sizes, add confidence intervals or significance tests, and either compare against a single pre-registered human aggregator or soften the claim to reflect the uncertainty.","section":"§4.2, Table 2"},{"comment":"The gold answers and distractors are both produced by LLMs, which threatens construct validity in a way that affects the headline accuracy numbers. GPT-4o performs the structured extraction from which questions and answers are built, and o1-preview generates all three distractors for 1,178 of the questions (7 of 11 question types). Because the evaluated models come from the same model lineages, the benchmark may partly measure how well models second-guess distractor-writing conventions rather than close-reading ability. The instructor check described in §3.1.2 only verifies that less-reasonable distractors could in principle be generated; it does not validate the final items. Please provide evidence that the correct options are not systematically distinguishable from the distractors (e.g., human judges cannot identify the gold answer above chance given the passage and choices, or a perturbation/format-control analysis), or report results on a subset with human-authored distractors.","section":"§3.2.1–3.2.2, Appendix C.1"},{"comment":"The claim that models 'trail' humans on 10/11 tasks is also internally at odds with the aggregate human performance in Table 2. On Q5 and Q8, for instance, the human weighted averages are 28.8 and 25.1, far below the best LLM scores of 62.3 and 35.7, respectively; these are only converted into human 'wins' by taking the maximum over three evaluators, with a single evaluator contributing the high score in each case. The paper acknowledges variability in §6 but does not adjust the headline claim. The conclusion 'LLMs still lag behind human performance' should be qualified to 'lag behind the best of three human evaluators' unless a more robust human aggregate or statistical comparison is provided.","section":"§5, §6"}],"minor_comments":[{"comment":"The sentence 'which accounts for percentage of the dataset' appears to be missing the numeric value; please fix the typo and report the actual fraction of the 1,331 questions used for the human baseline.","section":"§4.2"},{"comment":"The full-question template for Q9 repeats the Q6 wording about 'the significance of this device' and 'its effects on the reader'; for a context-significance task it should refer to the context (e.g., 'the significance of this contextual information') rather than to a device.","section":"Table 3, Q9 row"},{"comment":"The phrase 'boarder circumstances' should be 'broader circumstances'.","section":"Appendix B.2"},{"comment":"The manuscript says 'We then extracte the answer'; this should be 'extracted'.","section":"§4.1"},{"comment":"The reference to Comsa et al. (2022) lists the author as 'Iulia Coms, a'; the author name should be 'Iulia Comsa'.","section":"References"},{"comment":"Figure 2 is informative but the arrow conventions are not explained; a note in the caption clarifying that the three clusters correspond to the three progressive difficulty groups would improve readability.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The central human-superiority claim is currently supported only by a maximum-over-three-evaluators baseline with no inferential statistics, and the benchmark-validity concern about LLM-generated distractors is substantive. That said, the resource is novel, the dataset is public, and the LLM results are reproducible with the evaluation harness; I believe the issues can be fixed with additional validation and a more careful framing, so major revision rather than rejection is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"KRISTEVA is a real new resource — the first public MCQ benchmark for close reading, built from classroom essays with a task taxonomy grounded in the CRIT pedagogy — but the headline that LLMs trail experienced readers on 10 of 11 tasks is not statistically established by the data in the paper.\n\nWhat's genuinely good: the task design. The 11 question types are not an ad-hoc grab bag; they map onto the Observe, Contextualize, Analyze steps of a framework actually used in university teaching. The pipeline from student essays to structured features is transparent, and the authors evaluate 19 LLMs with a standard harness. The dataset is public. That is a contribution people will build on.\n\nThe soft spots are in the human baseline and the claim built on it. The human sample is three PhD students, each answering a subset; Section 4.2 says the subset 'accounts for percentage of the dataset' — the number is missing. Table 2 reports 'top-line human performance' as the best evaluator on each task, so the 10-of-11 result is a per-task max, not any single human. Evaluator 2, the best overall at 74.7, beats Phi-4 (69.7) on 8 of 11 tasks, and the human weighted average (65.6) is below Phi-4. With n=3 and per-task scores like 75/0/0, no confidence intervals or significance tests are given. The stress-test note is right: that claim needs a bigger, better-defined human sample.\n\nThe distractor concern is real but partly self-inflicted: o1-preview wrote distractors for 7 of 11 question types, and the paper says so plainly. If those distractors are systematically distinguishable from the correct answers, model accuracy is inflated; that direction would make the qualitative human-LLM gap larger, not smaller, but it does mean the per-model numbers should be read cautiously.\n\nMinor: Q9's question text in Table 3 looks copy-pasted from Q6 ('significance of this device' for a context question). Fix it.\n\nI'd send this to serious peer review. The benchmark is worth having and the methodology can be tightened in revision. I'd cite it for the dataset and the task taxonomy, not for the human-comparison headline.","headline":"KRISTEVA is a genuinely new and useful close-reading benchmark, but its main 'humans beat LLMs on 10/11 tasks' claim rests on a three-evaluator per-task-max baseline and needs a stronger human study.","tokens_in":21292,"tokens_out":3669,"would_cite":true,"duration_ms":34342,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"KRISTEVA, the first close-reading benchmark for LLMs, finds that best models score 49.7–69.7 percent while experienced human readers lead on 10 of 11 tasks.","keywords":["close reading","interpretive reasoning","LLM benchmark","figurative language understanding","multi-hop reading comprehension","human evaluation","literary analysis","multiple-choice question generation"],"falsifier":"Take a random sample of 200 KRISTEVA questions and ask experienced close readers to rank the four options without knowing the keyed answer. If in a substantial share of items the LLM-generated distractor is ranked as more reasonable than the keyed answer, then the accuracy gap between humans and models reflects option quality, not close-reading ability.","tokens_in":20256,"feed_emoji":"📖","tokens_out":6039,"duration_ms":58854,"temperature":0.7,"pith_summary":"KRISTEVA is the first benchmark built to measure whether large language models can do close reading: gathering textual details and turning them into evidence-based interpretations of literary passages. It converts 49 high-scoring college exam essays into 1,331 multiple-choice questions organized into eleven tasks that mirror a six-step teaching heuristic. On these tasks, state-of-the-art LLMs score between 49.7% and 69.7%, and the best model still trails the strongest experienced human reader on most individual tasks. The paper's point is not that machines cannot interpret literature at all, but that interpretive reasoning—where answers are judged by plausibility rather than by a single right answer—is a distinct, testable capability that current benchmarks omit.","feed_headline":"New benchmark: AI trails expert readers on close reading","feed_subtitle":"The 1,331-question KRISTEVA test puts top LLMs at 49.7–69.7%, behind humans on 10 of 11 tasks.","key_machinery":"The load-bearing mechanism is KRISTEVA's eleven task types, arranged in three ascending clusters that approximate the stages of close reading. The first cluster asks models to extract stylistic features—detect the device, locate it, name its elements, infer its purpose and significance. The second asks them to retrieve and rank relevant external contexts from parametric knowledge. The third requires multi-hop reasoning that connects a specific stylistic feature to a specific external context and articulates the connection. The benchmark's ground truth comes from instructor-graded student essays, and its wrong answers are generated by a reasoning LLM to be plausible but weaker interpretations; the claim that accuracy measures interpretive reasoning depends on that ordering of answer quality.","core_discovery":"On the paper's own terms, the central discovery is that close reading can be operationalized as a structured, multiple-choice evaluation, and that under that operationalization LLMs show partial but incomplete competence. The strongest model reaches 69.7% overall and 64.3% on reasoning-heavy questions, while the best human evaluator outperforms the best model on 8 of 11 tasks, and human evaluators collectively match or beat every model on 10 of 11 tasks. This gap survives even though the human baseline is likely conservative, since evaluators had to adapt to an MCQ format far removed from open-ended close reading. The benchmark therefore establishes a measurable distance between machine and human interpretive skill and provides a reusable scaffold for closing it.","pith_inferences":["If the benchmark's validity assumption holds, the task ordering implies a diagnostic profile: models' weakest points are significance ranking and feature-context reasoning, suggesting that connecting form to outside knowledge—not detecting devices—is the current bottleneck.","A testable extension would convert KRISTEVA to free-response format and score it with expert rubrics; the paper names this as future work, and it would reveal how much of the human-model gap is an artifact of multiple-choice options.","The same question-generation pipeline could be applied to other humanities disciplines—history, philosophy, art criticism—that teach evidence-based interpretation, giving NLP a family of benchmarks for judgment rather than fact retrieval."],"forward_implications":["A standardized evaluation now exists for a form of reasoning that relies on relative plausibility rather than a uniquely correct answer, so future work can compare models and track progress on interpretive judgment.","Because the tasks combine figurative-language understanding with multi-hop reasoning, gains on KRISTEVA should transfer to both NLP areas rather than to a single narrow skill.","The 10-of-11 human advantage indicates that current LLM pretraining and instruction tuning leave room for improvement in text-grounded aesthetic and interpretive judgment.","The 49-essay, 1,331-question pipeline shows that routine college classroom writing can become a scalable source of benchmark data."],"supporting_citations":[{"why":"Reports positive pedagogical effects of the teaching heuristic that KRISTEVA adapts into its eleven task types.","marker":"(Bares et al., 2020)"},{"why":"Supplies the theory of stylistic features as deviations from ordinary language that grounds the feature-extraction tasks.","marker":"(Richards, 1929)"},{"why":"Frames interpretation as plausibility rather than correctness, supporting the benchmark's ground-truth model.","marker":"(Sinykin and Winant, 2025)"},{"why":"Defines figurative language understanding as reasoning, the existing task family that KRISTEVA extends to purpose, significance, and effect.","marker":"(Chakrabarty et al., 2022b)"},{"why":"Defines multi-hop reading comprehension across documents, which KRISTEVA adapts to feature-context reasoning.","marker":"(Welbl et al., 2018)"},{"why":"Provides evidence that chain-of-thought mainly helps math and logic tasks, justifying the zero-shot evaluation setting and interpretation of reasoning-model results.","marker":"(Sprague et al., 2024)"},{"why":"Provides the account of close reading as heightened attention to textual detail used to motivate the tasks.","marker":"(Guillory, 2025)"}],"fun_headline_variants":["Close reading benchmark: LLMs trail humans on 10 of 11 tasks","KRISTEVA benchmark: AI achieves 49.7-69.7% on close reading, humans ahead","New close reading test shows LLM limits: humans win 10 of 11 tasks","Interpretive reasoning gap: LLMs score 49.7-69.7% on KRISTEVA close reading","First close reading benchmark: AI partial, humans lead on most tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark is valid only if the correct answers taken from high-scoring student essays remain more defensible than the LLM-generated wrong choices, so that picking the keyed answer reflects interpretive skill rather than an artifact of how the options were written.","fun_headline_variants_meta":{"raw":{"variants":["Close reading benchmark: LLMs trail humans on 10 of 11 tasks","KRISTEVA benchmark: AI achieves 49.7-69.7% on close reading, humans ahead","New close reading test shows LLM limits: humans win 10 of 11 tasks","Interpretive reasoning gap: LLMs score 49.7-69.7% on KRISTEVA close reading","First close reading benchmark: AI partial, humans lead on most tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00023,"raw_usage":{"total_tokens":1480,"prompt_tokens":938,"completion_tokens":542,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":424}},"tokens_in":554,"tokens_out":542,"duration_ms":5348,"temperature":1.0,"reasoning_tokens":424,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:22:52.373768+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of 200 KRISTEVA questions and ask experienced close readers to rank the four options without knowing the keyed answer. If in a substantial share of items the LLM-generated distractor is ranked as more reasonable than the keyed answer, then the accuracy gap between humans and models reflects option quality, not close-reading ability.","supporting_citations":[],"review_version":1}