Pith. sign in

Paper Citation Record · LEDGER

DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2408.15045.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.15045 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:53:51.923682Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cea610ae-b9d6-4768-9db0-8826fbc32897 · inbound

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness cites this paper.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.303009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.303009Z digest=sha256:cd64b015a718c4530acd76892feac38ccd3ea7bbdf83f4f69e25a451a9a5dccf

Observation 840efc04-9dbd-4cac-bce6-e3a3808c3239 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.793729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:7cd44faab1381151902bb9c6d523cdc0d34b9d2b0451d32f0e4eac599fa0093f

Observation 195cf54e-b498-4da7-9066-a0ab28b3dd3f · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.802900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.802900Z digest=sha256:15b5480e23a97d17df685a2176c4e2d1a638015f35ae07b6d48a410fb7c1e0ac

Observation 55a9d781-b9c5-4bf2-b988-d2adb86f3101 · inbound

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence cites this paper.

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:00.818917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:06:00.818917Z digest=sha256:217fe6ef0797e9172b47cd9f42e81a2b2a2dfafbe963f78242e987e8b39e49da

Observation 028f4f98-e2e8-46ad-9a3c-18041556d76e · inbound

Document Image Rectification Bases on Self-Adaptive Multitask Fusion cites this paper.

Document Image Rectification Bases on Self-Adaptive Multitask Fusion DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T22:53:51.923682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:53:51.923682Z digest=sha256:dd24d6a3a44fb21c3bf5d071c1c9fbd74d3e02cd47da7cb1bd52ac606da27d99

Observation e2683014-4149-4068-9dce-8ebb7e5ef479 · inbound

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization cites this paper.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:04.645195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:04.645195Z digest=sha256:e7a57504f7460f11bc9c91deb3c3d3f1a013095bb333510945305290fbc2e9cb

Observation 08b6382f-8eb5-4395-bcee-e630f005b43c · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.050413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:005782782f0dc1dc3a9913056e48933d62df37cce384e92594da1fb373944105

Observation 34aaa92a-50c5-40f3-8837-6c4db8922b33 · inbound

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA cites this paper.

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:44:02.477273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T04:41:55.030673Z digest=sha256:7c880fc3970436a76e72a9c536b9cde22b758c370d20cd8079fafa6b8470d2a2

Observation c01bd59f-4028-4882-9156-c3badc769502 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.767230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:2a5872c5611d44d6eb2ad3da541d07c5340a479667823e25a6b11ebe8f2d4d73

Observation d3a90b89-c08f-4769-8f86-39e8221d455b · inbound

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth cites this paper.

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 182

Resolution
unresolved
no resolver link, observed 2026-08-02T14:55:39.244177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:55:39.244177Z digest=sha256:579c30b10c1337dac8e0099b0ee94165332f3d28a7098d79b82e877037487972