Pith. sign in

Paper Citation Record · LEDGER

DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2406.19101.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.19101 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:31:59.629310Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4cb094ab-17e0-442b-8e41-0cfaedb324df · inbound

FoPru: Focal Pruning for Efficient Large Vision-Language Models cites this paper.

FoPru: Focal Pruning for Efficient Large Vision-Language Models DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:31:59.629310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:31:59.629310Z digest=sha256:ef6398ae9c200a18bb78e64177f415a2b985aa0c51eee79253a40050c33dff63

Observation 5b1e0f2c-95f7-49dc-8cbd-9748e538a7cd · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.117286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.117286Z digest=sha256:2f5989666e49e8334f9adde8c3f8035c399622993dced78f14176cf9d58ba710

Observation ee43fcb4-06be-47be-9e95-8e88d068dc6f · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.332662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.332662Z digest=sha256:4a0f24af818529b2195e69c0d9ddf630fb21ba9396353112ab7475cd909bcf58

Observation e4ff13ee-acd1-48ef-b3e9-333580ea32ed · inbound

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs cites this paper.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.852342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.852342Z digest=sha256:43c4b8e472ea095a19b8976ea154643a837fa7e67375b0fcdf11340980394949

Observation 22404211-099a-45f2-9ca7-76d44f997139 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.092864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:930eb7ecd9825f3062fb92b70eab555dae8251b6486a6a13b253e784b68e1012

Observation a11f40bc-3d7c-4277-823a-7a34f8d3b502 · inbound

CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMs cites this paper.

CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMs DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T22:32:41.286530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:32:41.286530Z digest=sha256:b6bf0da4629ead9b0c3e0b68072a8e763e34238618d6f057c72ae61c3b0ea1ed

Observation bed1a920-6e8f-43a4-8415-845773168336 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.714114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:d45bce3c17c522021f19a29993723ad6a7af81668d5b483f0916db0697ab45b5

Observation c039b0e2-f46d-44e5-8c52-1added34df56 · inbound

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth cites this paper.

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming

Reference 181

Resolution
unresolved
no resolver link, observed 2026-08-02T14:55:39.161657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:55:39.161657Z digest=sha256:e8eabe706f7f4fa6bb2f5ed6af5046ac4e2b7d2158e6e857f5b382b3bd698c7e