Pith. sign in

REVIEW 6 cited by

DocBank: A Benchmark Dataset for Document Layout Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.01038 v3 pith:M7CEEOPA submitted 2020-06-01 cs.CL

classification cs.CL
keywords docbankdocumentlayoutanalysisdatasetdocumentsinformationmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Document layout analysis usually relies on computer vision models to understand documents while ignoring textual information that is vital to capture. Meanwhile, high quality labeled datasets with both visual and textual information are still insufficient. In this paper, we present \textbf{DocBank}, a benchmark dataset that contains 500K document pages with fine-grained token-level annotations for document layout analysis. DocBank is constructed using a simple yet effective way with weak supervision from the \LaTeX{} documents available on the arXiv.com. With DocBank, models from different modalities can be compared fairly and multi-modal approaches will be further investigated and boost the performance of document layout analysis. We build several strong baselines and manually split train/dev/test sets for evaluation. Experiment results show that models trained on DocBank accurately recognize the layout information for a variety of documents. The DocBank dataset is publicly available at \url{https://github.com/doc-analysis/DocBank}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bricker to BRACE: A Bracket Exposure RAW Dataset and Restoration Model for Flicker-Banding

    cs.CV 2026-06 unverdicted novelty 6.5 of 10

    Presents Bricker dataset and BRACE multi-frame model using frequency priors and cross-attention for flicker-banding removal in RAW screen captures, with new SFC metric.

  2. SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding

    cs.CL 2025-10 unverdicted novelty 6.0 of 10

    SARA combines natural-language snippets with semantic compression vectors in RAG to improve answer relevance, correctness, and similarity on 9 datasets across 5 LLMs.

  3. DREAM: Document Reconstruction via End-to-end Autoregressive Model

    cs.CV 2025-07 reject novelty 6.0 of 10

    A single model, DREAM, jointly predicts layout elements, coordinates, and transcriptions for document reconstruction, along with a new metric (DSM) and benchmark (DocRec1K).

  4. P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark

    cs.CL 2025-05 conditional novelty 6.0 of 10

    P2P is a multi-agent framework that automatically generates HTML-rendered academic posters from papers, backed by a 30k instruction dataset and a 121-pair evaluation benchmark.

  5. AI-Generated Lecture Slides for Improving Slide Element Detection and Retrieval

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Synthetic lecture slides generated by an LLM pipeline improve few-shot slide element detection and text-based retrieval when used as pre-training data.

  6. Hierarchical Document Parsing via Large Margin Feature Matching and Heuristics

    cs.CL 2025-02 conditional novelty 4.0 of 10

    By adding an ArcFace-style margin to a CLIP-like matching loss and applying dataset-specific greedy rules, the solution reaches 0.98904 private-leaderboard accuracy on the VRD-IU document hierarchy task.

Pith tools