Pith. sign in

REVIEW 4 cited by

LMDX: Language Model-based Document Information Extraction and Localization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.10952 v2 pith:F443JFVM submitted 2023-09-19 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords documentextractionlmdxentitiesinformationlanguagellmsdocuments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLM) have revolutionized Natural Language Processing (NLP), improving state-of-the-art and exhibiting emergent capabilities across various tasks. However, their application in extracting information from visually rich documents, which is at the core of many document processing workflows and involving the extraction of key entities from semi-structured documents, has not yet been successful. The main obstacles to adopting LLMs for this task include the absence of layout encoding within LLMs, which is critical for high quality extraction, and the lack of a grounding mechanism to localize the predicted entities within the document. In this paper, we introduce Language Model-based Document Information Extraction and Localization (LMDX), a methodology to reframe the document information extraction task for a LLM. LMDX enables extraction of singular, repeated, and hierarchical entities, both with and without training data, while providing grounding guarantees and localizing the entities within the document. Finally, we apply LMDX to the PaLM 2-S and Gemini Pro LLMs and evaluate it on VRDU and CORD benchmarks, setting a new state-of-the-art and showing how LMDX enables the creation of high quality, data-efficient parsers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SAIL: Sample-Centric In-Context Learning for Document Information Extraction

    cs.CL 2024-12 conditional novelty 6.0 of 10

    SAIL improves training-free document information extraction by choosing per-sample examples using entity-level text and layout similarities, outperforming prior ICL methods in F1.

  2. CRAWLDoc: A Dataset for Robust Ranking of Bibliographic Documents

    cs.CL 2025-06 conditional novelty 5.0 of 10

    CRAWLDoc ranks linked web documents by embedding similarity to a paper's landing page, evaluated on a new manually labeled dataset of 600 publications from six publishers.

  3. An archaeological Catalog Collection Method Based on Large Vision-Language Models

    cs.CV 2024-12 reject novelty 5.0 of 10

    A VLM-based pipeline with localize, comprehend, and match modules collects artifact catalogs at AP up to 35.4%, though evaluation circularity weakens the claim.

  4. Patchfinder: Leveraging Visual Language Models for Accurate Information Retrieval using Model Uncertainty

    cs.CV 2024-12 conditional novelty 5.0 of 10

    PatchFinder uses VLM token confidence to select patch size and the most confident patch, achieving 94% field-extraction accuracy on 190 noisy scanned well documents.

Pith tools