Pith. sign in

Paper Citation Record · LEDGER

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding

As of 14 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2506.01388.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01388 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:49:10.777750Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T04:38:49.512293Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T04:42:04.436599Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 66858e4b-f5bc-4c16-86b1-5e9300c60366 · outbound

This paper cites Form-nlu: Dataset for the form natural language understanding.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Form-nlu: Dataset for the form natural language understanding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:13.861693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:49:08.059090Z digest=sha256:300239b72dfa5718a6e71ce9de8127f4441625808afd3c7c039a6de111f337ab

Observation ea21328d-8faa-4527-b7cd-6c93d3d69b14 · outbound

This paper cites 3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding 3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:49:11.108695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:49:08.445020Z digest=sha256:6ba92a72d7df1098e22911ac52881975608baddf476e3728c2fb588d4ed0caae

Observation 4ddf283f-3cd7-4135-9c16-1951ac6b46f3 · outbound

This paper cites Mask r-cnn.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Mask r-cnn

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:13.694463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:49:08.645933Z digest=sha256:4cb15d7b83bf191b3ce32335c5295ce5959084306471d8a37d3ee3553c0b008b

Observation d4a62393-152b-4421-88f4-980b42b1b12c · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36,.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Visual instruction tuning.Advances in neural information processing systems, 36,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:12.886130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:49:09.475403Z digest=sha256:f671fc5345c22d04a45abf2ff27bcaf9955471cfa6f1e8f1402fa880cecaa550

Observation dea4200a-6c0e-4ec7-b2fd-21b27c72125a · outbound

This paper cites Hello gpt-4o.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Hello gpt-4o

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:09.628609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:09.628609Z digest=sha256:90e3df6356e7f52907d613c145978e242581fcf522a3d4a62e5d788053027441

Observation cc5cc385-1a0a-4151-a9f2-c55e03706a25 · outbound

This paper cites Cord: a consolidated receipt dataset for post-ocr pars- ing.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Cord: a consolidated receipt dataset for post-ocr pars- ing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:12.630089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:49:09.707416Z digest=sha256:58e7be2f6b47f89415af83e643a8a66111b2ec1f2521ff943436f797bfe31217

Observation 53d1255f-7b73-44f8-a7b1-695661a9ab82 · outbound

This paper cites You only look once: Unified, real-time object detection.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding You only look once: Unified, real-time object detection

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:12.402963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:49:09.842186Z digest=sha256:95a1f39d876500eddf4bb9b13d2ffd1aa2f17a009207515d1eb5973dd3fbe179

Observation 543dd6a1-9e1e-436f-ab5a-916c46b4e8d3 · outbound

This paper cites Towards robust visual information extraction in real world: New dataset and novel solution.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Towards robust visual information extraction in real world: New dataset and novel solution

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:11.711186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:49:10.340936Z digest=sha256:e6260c207fd9fc1586d21ae9208d3eeef4f16847d7f08cffe0f2ecac91707319

Observation 8337f539-ca88-4a94-ac26-752cbbdbf8a1 · outbound

This paper cites LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:10.488825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:10.488825Z digest=sha256:399a3044cee6bdebdd5239df221e48d35dae06acbedd6bf39a82f07d485ccc34

Observation ed0ceba4-9e7c-4f19-95d3-5f1362107066 · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.arXiv preprint arXiv:2408.08872,.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding xgen-mm (blip-3): A family of open large multimodal models.arXiv preprint arXiv:2408.08872,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:10.632836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:10.632836Z digest=sha256:3f084e94db823db153f726a32c80d05dfcb757aab0971c8649f227fb71b571bf

Observation 66ac4e59-5f8b-464d-ad6b-1c3e5a998a2b · outbound

This paper cites Detrs beat yolos on real-time object de- tection.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Detrs beat yolos on real-time object de- tection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:11.441845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:49:10.777750Z digest=sha256:916f1c6efe0395bb4ee8929fa471ac3e89201385977f88d7073401f5daabdec7

Observation b44f0f5e-c3a8-4266-a7e9-e3cddaa59010 · outbound

This paper cites Kleister: key information extraction datasets involving long documents with complex layouts.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Kleister: key information extraction datasets involving long documents with complex layouts

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:12.141867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:49:10.104836Z digest=sha256:1ae740a715077625f3386005e92beeff862740dee1d083115b146cb8db9d5e3c

Observation 57ef1854-b6c7-403c-a274-757055004725 · outbound

This paper cites Faster r-cnn: Towards real-time ob- ject detection with region proposal networks.Advances in neural information processing systems, 28,.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Faster r-cnn: Towards real-time ob- ject detection with region proposal networks.Advances in neural information processing systems, 28,

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:09.961206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:09.961206Z digest=sha256:33d92bc0d8d125c63150091c9745b9d89c2bb55417b9f816379a884eece664bf

Observation a6a21c81-ceae-488b-adf1-bccb94a64354 · outbound

This paper cites Icdar2019 competition on scanned receipt ocr and information extraction.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Icdar2019 competition on scanned receipt ocr and information extraction

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:13.536986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:49:08.842306Z digest=sha256:f23fc764302ffe0e64d572c00dd5e5a3aa31f89247efa0172cb406cc08490e9c

Observation 27088a13-a74f-44e8-89d0-052d6c690446 · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Layoutlmv3: Pre-training for document ai with unified text and image masking

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:13.374992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:49:08.987749Z digest=sha256:e95ac696e8e15fdaa9a08cc13ef0c392671c23c5939e9d8b2383bdd30ccdba60

Observation f1697ce6-a1d4-4c7d-abb0-a4bcab3d657b · outbound

This paper cites Unifying vision, text, and layout for universal document processing.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Unifying vision, text, and layout for universal document processing

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:11.928334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:49:10.233981Z digest=sha256:ee910e072fa3b4e54abc937f1ded108a68a3fb82273a5beee671ab6577c3e390

Observation d13bf1e7-d242-4def-87fb-9122e04bcc48 · outbound

This paper cites Dit: Self-supervised pre- training for document image transformer.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Dit: Self-supervised pre- training for document image transformer

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:13.125103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:49:09.141603Z digest=sha256:56666a8f77a9d9ab163b74f0c3fbf4933b258aa24bda5d62d901751c6e46c8f5

Observation 29675e18-bc61-4d54-a1b8-bd3c025bcad1 · outbound

This paper cites David: Domain adap- tive visually-rich document understanding with synthetic insights.arXiv preprint arXiv:2410.01609,.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding David: Domain adap- tive visually-rich document understanding with synthetic insights.arXiv preprint arXiv:2410.01609,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:08.221446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:08.221446Z digest=sha256:ddfd9a2084c6c683385b4e8d0f7aa9feb707f7b86e6c67431b6f9d3315047a93

Observation 11227b6e-6f49-4fa6-9d1c-ed6ce3a5b060 · outbound

This paper cites PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:08.346682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:08.346682Z digest=sha256:8ee4cbaac6f8801c4bda745f6eb7c09471434814dff6d3c917d4d96b249ce2ef

Pith citing papers

Observation 6e2022bd-1519-4e72-9d7a-8af56266ec37 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.439042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:f3b9651122af6060185c24749546b69060d9b4dd1e058f64f2a65dd5e48a8c27