Pith. sign in

Paper Citation Record · LEDGER

LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2012.14740.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2012.14740 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:10:38.299321Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

59
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4e26d4db-73a1-449a-acec-44ccab6b8d48 · inbound

Nougat: Neural Optical Understanding for Academic Documents cites this paper.

Nougat: Neural Optical Understanding for Academic Documents LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:42:12.587348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:42:12.463309Z digest=sha256:73700b014406c6c4f9d07111e797c6bb22013c275808a9f946a20cb928f79983

Observation 85e4b954-d1b0-45fe-9935-4a5700f88a46 · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 271

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.321681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:edd0a1f1fef2698f80e0753e37ebf492448fb0a50beaf99bd51e8074f3e50f23

Observation dee30146-85d3-47fd-849a-dbc2a2bdfa95 · inbound

Performance Analysis of Traditional VQA Models Under Limited Computational Resources cites this paper.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.299321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.299321Z digest=sha256:29c843408ffec8691349c38c68b4816f743205e31afcd97405d44d07f4165ef0

Observation ccdb6357-b1f2-4e13-87ac-e85f2b070bae · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:29.988182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:29.988182Z digest=sha256:83445ed395a8953471bd19e1ba9cc5ce4990fbb1f8f6a53777aeb2984ee81086

Observation 3ceb7149-7d64-49fe-b295-13df1a98afe5 · inbound

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models cites this paper.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:12.315928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:12.315928Z digest=sha256:b7a5494c5d2bbdf3cf1535842a6d0d1bff779349650910689e0c551f90be1138

Observation 227c7055-8240-487f-ac35-2601503778e1 · inbound

Spatial ModernBERT: Spatial-Aware Transformer for Table and Key-Value Extraction in Financial Documents at Scale cites this paper.

Spatial ModernBERT: Spatial-Aware Transformer for Table and Key-Value Extraction in Financial Documents at Scale LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:49.169064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:49.169064Z digest=sha256:81013436408e0ced84aa309dd39ac6319f453ae9bb94cdcfccb89eae147c48bc

Observation 67f6bfd6-7632-4801-9fc8-fadf8f7056cd · inbound

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? cites this paper.

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T23:35:24.345477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:35:24.345477Z digest=sha256:0e4d3191e37648ce20de3f700ba3b1576a2ee232077eb797ee4bc378dfd81384

Observation ee736245-3fe3-4144-9d31-5926535a4c6d · inbound

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding cites this paper.

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.896872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.896872Z digest=sha256:de7991b2bd1be2260d29a41e3abe1687a1989b1f1f2838e9f5947123b415b622

Observation 6fc15024-0c61-4aa2-9ddf-8c353335db91 · inbound

OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models cites this paper.

OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:30:23.197949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T22:29:36.960961Z digest=sha256:8514a2e5179fd823e1c04250235d5e473980c0f6a471a8b66b0353564ff239c0

Observation 2fb0dedb-0a0e-4692-aab6-9d0c9d3a9044 · inbound

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA cites this paper.

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:44:02.473348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T04:41:55.030673Z digest=sha256:6dc72d87b48041a4a9fb57a0a9d62eccee0cb251e1fe7273d77093b9faa35625

Observation e2005fad-4789-424e-bc48-735e46b0e8f0 · inbound

DocAtlas: Multilingual Document Understanding Across 80+ Languages cites this paper.

DocAtlas: Multilingual Document Understanding Across 80+ Languages LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:02:58.730142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T21:02:45.148167Z digest=sha256:b425a63c49dad27a66e157b38232c153a65aac54d6462fd1391ff9b260de686b

Observation 3b42e5ef-c7b2-405a-af00-c73d368aa3f0 · inbound

DocAtlas: Multilingual Document Understanding Across 80+ Languages cites this paper.

DocAtlas: Multilingual Document Understanding Across 80+ Languages LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:54:47.176376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T09:51:40.160096Z digest=sha256:713976e12dede063b7a15f0756d3fad5764243ba308c0768d07e8c1b2d8cad7c

Observation 0d640502-46be-4f8b-9cc4-48bfaeed2348 · inbound

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding cites this paper.

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:05.090079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:48:34.771799Z digest=sha256:e494e0c47116f77eea18a36d51ee80155cec158c47b1dd1749726eeffa2ccb34

Observation 8a584c8d-0fbe-4019-964a-b06601adb99b · inbound

RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild cites this paper.

RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:45.176564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T09:19:38.285839Z digest=sha256:8a22ed0b7c51e6156b8be1984943ba2b428cc46fb484c6641a123a1386e831c7

Observation 9c50434a-5b04-44d3-bcf2-5f879b329869 · inbound

Structure-Preserving Document Translation via Multi-Stage LLM Pipeline: A Case Study in Marathi cites this paper.

Structure-Preserving Document Translation via Multi-Stage LLM Pipeline: A Case Study in Marathi LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:04:35.270473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T09:59:20.602618Z digest=sha256:bd955850f0c73689a5f940b93146937f68357dea9924e0a6b511ecf5b838511b

Observation 2e3eaee2-2b43-4485-8c4c-cf62fbf23850 · inbound

DocAnnot -- Accelerating the Creation of Key Information Extraction Datasets with GenAI-Powered Auto-annotation cites this paper.

DocAnnot -- Accelerating the Creation of Key Information Extraction Datasets with GenAI-Powered Auto-annotation LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T14:38:36.756737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:38:36.756737Z digest=sha256:a7f1763b551fe63a94548467b8a96ee4faa3c76b95c83f3392aab1e451985d20