Pith. sign in

REVIEW 6 cited by

LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.14740 v4 pith:3RGHD4GA submitted 2020-12-29 cs.CL

classification cs.CL
keywords layoutlmv2modelpre-trainingtasksarchitecturedocumentmulti-modaltext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Pre-training of text and layout has proved effective in a variety of visually-rich document understanding tasks due to its effective model architecture and the advantage of large-scale unlabeled scanned/digital-born documents. We propose LayoutLMv2 architecture with new pre-training tasks to model the interaction among text, layout, and image in a single multi-modal framework. Specifically, with a two-stream multi-modal Transformer encoder, LayoutLMv2 uses not only the existing masked visual-language modeling task but also the new text-image alignment and text-image matching tasks, which make it better capture the cross-modality interaction in the pre-training stage. Meanwhile, it also integrates a spatial-aware self-attention mechanism into the Transformer architecture so that the model can fully understand the relative positional relationship among different text blocks. Experiment results show that LayoutLMv2 outperforms LayoutLM by a large margin and achieves new state-of-the-art results on a wide variety of downstream visually-rich document understanding tasks, including FUNSD (0.7895 $\to$ 0.8420), CORD (0.9493 $\to$ 0.9601), SROIE (0.9524 $\to$ 0.9781), Kleister-NDA (0.8340 $\to$ 0.8520), RVL-CDIP (0.9443 $\to$ 0.9564), and DocVQA (0.7295 $\to$ 0.8672). We made our model and code publicly available at \url{https://aka.ms/layoutlmv2}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 59 citations worldwide. Full citation record

  1. SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding

    cs.CL 2025-10 unverdicted novelty 6.0 of 10

    SARA combines natural-language snippets with semantic compression vectors in RAG to improve answer relevance, correctness, and similarity on 9 datasets across 5 LLMs.

  2. Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A vision-language model fine-tuned on a new synthetic slide-animation dataset outperforms GPT-4.1 and Gemini-2.5-Pro at describing slide animations, especially on synthetic evaluation data.

  3. Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A two-stage model that analyzes page layout first, then parses text, tables, and formulas in parallel, reports state-of-the-art accuracy and speed on the benchmarks it evaluates.

  4. Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    A new benchmark and method (Spot-IT) aim to improve multimodal LLMs' ability to locate fine details in documents, with reported significant gains.

  5. Spatial ModernBERT: Spatial-Aware Transformer for Table and Key-Value Extraction in Financial Documents at Scale

    cs.CL 2025-07 reject novelty 4.0 of 10

    Spatial ModernBERT is a token-classification model that adds layout coordinates to ModernBERT to extract tables and key-value fields, with benchmark scores that fall short of the claimed state of the art.

  6. DocAnnot -- Accelerating the Creation of Key Information Extraction Datasets with GenAI-Powered Auto-annotation

    cs.IR 2026-05 conditional novelty 3.0 of 10

    DocAnnot combines an LVLM, OCR, and a spatial matching heuristic to auto-annotate KIE documents at F1 0.68–0.85, and models trained on that data reach roughly 0.68 F1 on CORD.

Pith tools