Pith. sign in

Paper Citation Record · LEDGER

DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2311.11810.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.11810 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:00:01.242564Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f677e5cb-599f-4cbd-89dd-f60b879fa1ce · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.151987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:6a91d115cdb9faabc00aed4aac3dc1a140a1aa1bc29db303c3c21d247422beef

Observation 314c322e-2660-409d-b6a2-e2cf72007ac9 · inbound

Improving the Transferability of 3D Point Cloud Attack via Spectral-aware Admix and Optimization Designs cites this paper.

Improving the Transferability of 3D Point Cloud Attack via Spectral-aware Admix and Optimization Designs DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:00:01.242564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:00:01.242564Z digest=sha256:a280b4c829c69081b60fcdd338a831a7eee7a4b77040b632f2ef225b09404cb2

Observation 82d49e9b-a8ca-4c02-923a-26fdc7ad9a2f · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:33:26.782173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:79a987718d0929fcbe366afc86977af91f86cfb289c63699e20a109bec35c79b

Observation 6cfec1f6-86e8-4190-8d26-71614137f408 · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.570746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.570746Z digest=sha256:4949aff19646e3da941eb3bfa8f86bf455481e7f9e7b6aea9d18b3f684174898

Observation 7a6621d5-08e8-4837-bc1d-0883e5d4fb2b · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.300421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.300421Z digest=sha256:a696f00b0daf919994a5310ea046ba59912ea44d8fdff40b1176a10d5782577a

Observation 1b59d5e2-d19b-4111-b093-d28a829d7925 · inbound

EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition cites this paper.

EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T22:57:16.806336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:57:16.806336Z digest=sha256:10e3bf5b549865763ac3f75ae9f0d4f5a18e3da203e502016b37f808f43e65fa

Observation 9b46aa39-9f48-4bed-ad26-c4d3a64a6e59 · inbound

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence cites this paper.

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:00.719912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:06:00.719912Z digest=sha256:90460c4799a33daff1f8a9280ca0c44e15116e8332ad8216bfea5eb69debb8ac

Observation d67a0b45-f1c7-4a4b-b3e0-80abe85dfc15 · inbound

Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning cites this paper.

Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:17.350101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:17.350101Z digest=sha256:d84497ac23b3dc7d2483ad732c8097129f000c7a789a5de4804aa77167f88adc

Observation 435b0314-42ff-4404-a945-1bfbaa361a9c · inbound

MFH: Marrying Frequency Domain with Handwritten Mathematical Expression Recognition cites this paper.

MFH: Marrying Frequency Domain with Handwritten Mathematical Expression Recognition DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:37.883038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:37.883038Z digest=sha256:872f3c27350e28eafc8e8afa5e115fcd2474ed9e2abd89edd5ad65dd45f608ed

Observation a0c23f38-0caf-46c4-91e9-39e8418b7a07 · inbound

ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning cites this paper.

ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:36.874218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:36.874218Z digest=sha256:08485cbd96492cefd7e0db15c383bf957277aea1985691d4b242bf8de4c8fff3

Observation a9b00e8c-c3ad-4fb9-8ccc-e39fec10fd51 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.059565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:c520e112d4a72d2ba6f43c54cb93a3fa56c76bfb05a3d4c2d8612872a627c628

Observation 20a09f63-9701-414a-a48d-c2af8715b858 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:00.673034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:00.673034Z digest=sha256:6b5bc27d3284b4a3a80864a77faf317f6608f2e9bfcce834b52ce2fc521f3778

Observation 9cbaacce-12e0-474c-8aa0-f1d862ddba20 · inbound

Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models cites this paper.

Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:40:43.024304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T22:40:39.892802Z digest=sha256:a320011c264c9d733b68994a61e25af22235134de66886e23af7ae479e98b029

Observation 72be1701-116c-4c85-b92f-2d999396beb4 · inbound

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth cites this paper.

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:53:13.098663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T10:52:32.238241Z digest=sha256:97212961c37de253583fdb63f4e447ffbcb5db2c38567834ee9eb398a3fd553d

Observation 4ef64414-b84c-47bc-b9a1-4dafa3b11397 · inbound

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth cites this paper.

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T16:34:34.452430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T16:34:34.452430Z digest=sha256:f6e1e36d599f2ef0eb20749ff25a8851b7ca457c7d263470aab8c1984c0cdae7