REVIEW 6 cited by
DocBank: A Benchmark Dataset for Document Layout Analysis
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Document layout analysis usually relies on computer vision models to understand documents while ignoring textual information that is vital to capture. Meanwhile, high quality labeled datasets with both visual and textual information are still insufficient. In this paper, we present \textbf{DocBank}, a benchmark dataset that contains 500K document pages with fine-grained token-level annotations for document layout analysis. DocBank is constructed using a simple yet effective way with weak supervision from the \LaTeX{} documents available on the arXiv.com. With DocBank, models from different modalities can be compared fairly and multi-modal approaches will be further investigated and boost the performance of document layout analysis. We build several strong baselines and manually split train/dev/test sets for evaluation. Experiment results show that models trained on DocBank accurately recognize the layout information for a variety of documents. The DocBank dataset is publicly available at \url{https://github.com/doc-analysis/DocBank}.
Forward citations
Cited by 6 Pith papers
-
Bricker to BRACE: A Bracket Exposure RAW Dataset and Restoration Model for Flicker-Banding
Presents Bricker dataset and BRACE multi-frame model using frequency priors and cross-attention for flicker-banding removal in RAW screen captures, with new SFC metric.
-
SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding
SARA combines natural-language snippets with semantic compression vectors in RAG to improve answer relevance, correctness, and similarity on 9 datasets across 5 LLMs.
-
DREAM: Document Reconstruction via End-to-end Autoregressive Model
A single model, DREAM, jointly predicts layout elements, coordinates, and transcriptions for document reconstruction, along with a new metric (DSM) and benchmark (DocRec1K).
-
P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark
P2P is a multi-agent framework that automatically generates HTML-rendered academic posters from papers, backed by a 30k instruction dataset and a 121-pair evaluation benchmark.
-
AI-Generated Lecture Slides for Improving Slide Element Detection and Retrieval
Synthetic lecture slides generated by an LLM pipeline improve few-shot slide element detection and text-based retrieval when used as pre-training data.
-
Hierarchical Document Parsing via Large Margin Feature Matching and Heuristics
By adding an ArcFace-style margin to a CLIP-like matching loss and applying dataset-specific greedy rules, the solution reaches 0.98904 private-leaderboard accuracy on the VRD-IU document hierarchy task.
Discussion (0). Continue with ORCID to comment.