Pith. sign in

REVIEW 5 cited by

PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.09848 v2 pith:KHCGPTO3 submitted 2025-08-13 cs.CL cs.AI

PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts

classification cs.CL cs.AI
keywords reasoningtaskbenchmarkcomprehensionglobalhumanslong-contextnarrative
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce PRELUDE, a benchmark for evaluating long-context understanding through the task of determining whether a character's prequel story is consistent with the canonical narrative of the original book. Our task poses a stronger demand for global comprehension and deep reasoning than existing benchmarks -- as the prequels are not part of the original story, assessing their plausibility typically requires searching and integrating information that is only indirectly related. Empirically, 88% of instances require evidence from multiple parts of the narrative. Experimental results highlight the challenge of our task: in-context learning, RAG and in-domain training with state-of-the-art LLMs, and commercial DeepResearch services, lag behind humans by >15%. A further human study reveals that models often produce correct answers with flawed reasoning, leading to an over 30% gap in reasoning accuracy compared to humans. These findings underscore the substantial room for improvement in long-context understanding and reasoning.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Many-Shot CoT-ICL: Making In-Context Learning Truly Learn

    cs.CL 2026-05 conditional novelty 7.0

    Many-shot CoT-ICL functions as test-time learning when demonstrations are ordered for smooth conceptual progression rather than similarity, enabling a new selection method that improves reasoning performance.

  2. A New Role for Relevance: Guiding Corpus Interaction in Agentic Search

    cs.CL 2026-07 conditional novelty 6.0

    Ordering ripgrep by embedding relevance, seeding entry paragraphs, and reranking matches yields higher accuracy with fewer tool calls than RISE, DCI, and retrieval agents.

  3. Many-Shot CoT-ICL: Making In-Context Learning Truly Learn

    cs.CL 2026-05 unverdicted novelty 6.0

    Many-shot CoT-ICL improves when demonstrations are ordered for smooth conceptual progression, with CDS delivering up to 5.42 percentage-point gains on math tasks using 64 examples.

  4. HGMEM: Hypergraph-based Working Memory to Improve Multi-step RAG for Long-Context Complex Relational Modeling

    cs.CL 2025-12 conditional novelty 6.0

    A working memory represented as a hypergraph, whose hyperedges are updated, inserted, and progressively merged by the LLM, improves multi-step RAG on long-context sense-making benchmarks.

  5. Towards High-Level Semantic Intelligence

    cs.AI 2026-07 conditional novelty 4.0

    A survey proposing that AI's next stage should be understood as High-Level Semantic Intelligence: mastering humor, sarcasm, metaphor, empathy, persuasion, and narrative across modalities.