REVIEW 6 cited by
DOCBENCH: A Benchmark for Evaluating LLM-based Document Reading Systems
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recently, there has been a growing interest among large language model (LLM) developers in LLM-based document reading systems, which enable users to upload their own documents and pose questions related to the document contents, going beyond simple reading comprehension tasks. Consequently, these systems have been carefully designed to tackle challenges such as file parsing, metadata extraction, multi-modal information understanding and long-context reading. However, no current benchmark exists to evaluate their performance in such scenarios, where a raw file and questions are provided as input, and a corresponding response is expected as output. In this paper, we introduce DocBench, a new benchmark designed to evaluate LLM-based document reading systems. Our benchmark involves a meticulously crafted process, including the recruitment of human annotators and the generation of synthetic questions. It includes 229 real documents and 1,102 questions, spanning across five different domains and four major types of questions. We evaluate both proprietary LLM-based systems accessible via web interfaces or APIs, and a parse-then-read pipeline employing open-source LLMs. Our evaluations reveal noticeable gaps between existing LLM-based document reading systems and human performance, underscoring the challenges of developing proficient systems. To summarize, DocBench aims to establish a standardized benchmark for evaluating LLM-based document reading systems under diverse real-world scenarios, thereby guiding future advancements in this research area.
Forward citations
Cited by 6 Pith papers
-
GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents
On GDP.pdf, 100 expert-authored professional PDF tasks, seventeen frontier multimodal models pass at most 30.7% of items, and most failures come from missed footnotes, exclusions, tables, and spatial details.
-
CXXCrafter: An LLM-Based Agent for Automated C/C++ Open Source Software Building
An LLM-driven agent, CXXCrafter, automatically builds 587 of 752 C/C++ open-source projects (78%), beating default build commands (39%) and bare LLMs (32 to 38%).
-
T-GRAG: A Dynamic GraphRAG Framework for Resolving Temporal Conflicts and Redundancy in Knowledge Retrieval
A temporal GraphRAG framework that partitions knowledge graphs by timestamp and retrieves at subgraph, node, and knowledge levels outperforms RAG baselines on a new Audi annual-report QA benchmark.
-
MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs
MMGraphRAG links scene-graph entities from images to text knowledge graph entities via SpecLink, and reports accuracy gains over naive RAG and GraphRAG on multimodal document QA.
-
Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing
Across three invoice datasets, multimodal LLMs extract fields more accurately from raw images than from markdown converted by a parsing tool, with Gemini 2.5 Pro leading.
-
Evaluating Hierarchical Clinical Document Classification Using Reasoning-Based LLMs
Across 1,500 MIMIC-IV discharge summaries and 11 LLMs, no model exceeded 57% F1 on ICD-10 coding, with reasoning-labeled models slightly ahead of others, but the comparison is confounded by model differences.
Discussion (0). Continue with ORCID to comment.