Pith. sign in

REVIEW 2 cited by

Needle Threading: Can LLMs Follow Threads through Near-Million-Scale Haystacks?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.05000 v2 pith:XLQO5R65 submitted 2024-11-07 cs.CL

classification cs.CL
keywords contextllmsmodelsinformationmanythreadsdifferentfind
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As the context limits of Large Language Models (LLMs) increase, the range of possible applications and downstream functions broadens. In many real-world tasks, decisions depend on details scattered across collections of often disparate documents containing mostly irrelevant information. Long-context LLMs appear well-suited to this form of complex information retrieval and reasoning, which has traditionally proven costly and time-consuming. However, although the development of longer context models has seen rapid gains in recent years, our understanding of how effectively LLMs use their context has not kept pace. To address this, we conduct a set of retrieval experiments designed to evaluate the capabilities of 17 leading LLMs, such as their ability to follow threads of information through the context window. Strikingly, we find that many models are remarkably threadsafe: capable of simultaneously following multiple threads without significant loss in performance. Still, for many models, we find the effective context limit is significantly shorter than the supported context length, with accuracy decreasing as the context window grows. Our study also highlights the important point that token counts from different tokenizers should not be directly compared -- they often correspond to substantially different numbers of written characters. We release our code and long-context experimental data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels

    cs.CL 2025-05 conditional novelty 7.0 of 10

    None of the seven frontier LLMs tested retain stable performance on novel-level summary, storyworld, and narrative-time tasks once input length exceeds roughly 64k tokens, despite advertised context windows up to 10M tokens.

  2. LLMs for Legal Subsumption in German Employment Contracts

    cs.CL 2025-07 conditional novelty 5.0 of 10

    LLMs reach 80% weighted F1 on German employment contract clause review when given lawyer-distilled examination guidelines, but lag human lawyers when reading full legal sources.

Pith tools