Pith. sign in

Paper Citation Record · LEDGER

DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2408.15045.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.15045 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:53:51.923682Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cea610ae-b9d6-4768-9db0-8826fbc32897 · inbound

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness cites this paper.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.303009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.303009Z digest=sha256:cd64b015a718c4530acd76892feac38ccd3ea7bbdf83f4f69e25a451a9a5dccf

Observation 840efc04-9dbd-4cac-bce6-e3a3808c3239 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.793729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:aae5825e85739e89239d9677f03cd3e7cefe5d23d14204f57b4c1a3456f99463

Observation 195cf54e-b498-4da7-9066-a0ab28b3dd3f · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.802900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.802900Z digest=sha256:15b5480e23a97d17df685a2176c4e2d1a638015f35ae07b6d48a410fb7c1e0ac

Observation 55a9d781-b9c5-4bf2-b988-d2adb86f3101 · inbound

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence cites this paper.

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:00.818917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:06:00.818917Z digest=sha256:217fe6ef0797e9172b47cd9f42e81a2b2a2dfafbe963f78242e987e8b39e49da

Observation 028f4f98-e2e8-46ad-9a3c-18041556d76e · inbound

Document Image Rectification Bases on Self-Adaptive Multitask Fusion cites this paper.

Document Image Rectification Bases on Self-Adaptive Multitask Fusion DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T22:53:51.923682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:53:51.923682Z digest=sha256:dd24d6a3a44fb21c3bf5d071c1c9fbd74d3e02cd47da7cb1bd52ac606da27d99

Observation e2683014-4149-4068-9dce-8ebb7e5ef479 · inbound

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization cites this paper.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:04.645195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:04.645195Z digest=sha256:e7a57504f7460f11bc9c91deb3c3d3f1a013095bb333510945305290fbc2e9cb

Observation 08b6382f-8eb5-4395-bcee-e630f005b43c · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.050413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:04362893cfdf0a1e690d6488784f13a83da583fa0734aa5288653f65fbfaf126

Observation 34aaa92a-50c5-40f3-8837-6c4db8922b33 · inbound

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA cites this paper.

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:44:02.477273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T04:41:55.030673Z digest=sha256:4dc36623a55f08fe0bd9e48c0506e0baa964d0c84dc0fd90f4247b6854f41af4

Observation c01bd59f-4028-4882-9156-c3badc769502 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.767230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:255b6071199a6c5112d4e04e09d112cb9748f3bb8e4cb6bda397fb1206f5e68a

Observation d3a90b89-c08f-4769-8f86-39e8221d455b · inbound

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth cites this paper.

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 182

Resolution
unresolved
no resolver link, observed 2026-08-02T14:55:39.244177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:55:39.244177Z digest=sha256:579c30b10c1337dac8e0099b0ee94165332f3d28a7098d79b82e877037487972