REVIEW 4 cited by
LMDX: Language Model-based Document Information Extraction and Localization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLM) have revolutionized Natural Language Processing (NLP), improving state-of-the-art and exhibiting emergent capabilities across various tasks. However, their application in extracting information from visually rich documents, which is at the core of many document processing workflows and involving the extraction of key entities from semi-structured documents, has not yet been successful. The main obstacles to adopting LLMs for this task include the absence of layout encoding within LLMs, which is critical for high quality extraction, and the lack of a grounding mechanism to localize the predicted entities within the document. In this paper, we introduce Language Model-based Document Information Extraction and Localization (LMDX), a methodology to reframe the document information extraction task for a LLM. LMDX enables extraction of singular, repeated, and hierarchical entities, both with and without training data, while providing grounding guarantees and localizing the entities within the document. Finally, we apply LMDX to the PaLM 2-S and Gemini Pro LLMs and evaluate it on VRDU and CORD benchmarks, setting a new state-of-the-art and showing how LMDX enables the creation of high quality, data-efficient parsers.
Forward citations
Cited by 4 Pith papers
-
SAIL: Sample-Centric In-Context Learning for Document Information Extraction
SAIL improves training-free document information extraction by choosing per-sample examples using entity-level text and layout similarities, outperforming prior ICL methods in F1.
-
CRAWLDoc: A Dataset for Robust Ranking of Bibliographic Documents
CRAWLDoc ranks linked web documents by embedding similarity to a paper's landing page, evaluated on a new manually labeled dataset of 600 publications from six publishers.
-
An archaeological Catalog Collection Method Based on Large Vision-Language Models
A VLM-based pipeline with localize, comprehend, and match modules collects artifact catalogs at AP up to 35.4%, though evaluation circularity weakens the claim.
-
Patchfinder: Leveraging Visual Language Models for Accurate Information Retrieval using Model Uncertainty
PatchFinder uses VLM token confidence to select patch size and the most confident patch, achieving 94% field-extraction accuracy on 190 noisy scanned well documents.
Discussion (0). Continue with ORCID to comment.