Pith. sign in

Paper Citation Record · LEDGER

mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2409.03420.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.03420 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:43:17.502876Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T13:24:40.403192Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 851bce59-32ba-48a6-b376-3f99419b21cb · inbound

MinerU: An Open-Source Solution for Precise Document Content Extraction cites this paper.

MinerU: An Open-Source Solution for Precise Document Content Extraction mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:00:25.813195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T04:00:25.624430Z digest=sha256:24b7562cceac2f32fc35744a873ba633934a698ff3bf95542cfd89313cd4c32b

Observation a1bac248-8a5a-448f-9e6a-43dcbaa7a470 · inbound

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents cites this paper.

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:37:25.887847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T15:37:25.781240Z digest=sha256:9e8ca9e4db0c06375ce29429f2149a5a155f1ff34167fd0793b8583fb91442c0

Observation eb5d99ca-3ca3-4ed2-ac21-55e664e21c4e · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.227275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:431ea86691672c729ce0010d2af5d935e5460999e749334d41a8f2780dcfbebd

Observation 64294183-90b7-46e2-ba21-f46190dd377a · inbound

Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction cites this paper.

Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T17:33:45.080377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:33:45.080377Z digest=sha256:c3f9050a7f2a40bfe2726187327b32a5839ac9dbc1d3db375115f0b633e9dda8

Observation 00c8d2a0-970b-4ade-a554-71ce0b355193 · inbound

Chimera: Improving Generalist Model with Domain-Specific Experts cites this paper.

Chimera: Improving Generalist Model with Domain-Specific Experts mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:47.524267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:47.524267Z digest=sha256:57a30449e41ef69f68d5a5765abccd9d99dce10eb8ac612b462093573dce1fe2

Observation 85e3a444-7ea4-4776-9f00-05885719d524 · inbound

OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations cites this paper.

OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T18:41:26.470184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:41:26.470184Z digest=sha256:39e5d1898514dc8e8fd58eef1a2ea83531fd12921ee69d939807f1ed3ff9dd40

Observation ea4bc805-2138-49c8-b0d4-f514d986f7c1 · inbound

FILA: Fine-Grained Vision Language Models cites this paper.

FILA: Fine-Grained Vision Language Models mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.175229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.175229Z digest=sha256:850d55945782103322dba7ce13862bdf4f002473c16680375243ce1583341b4e

Observation 06682577-5388-4b29-b7fa-a806ecdc8dc3 · inbound

DocVLM: Make Your VLM an Efficient Reader cites this paper.

DocVLM: Make Your VLM an Efficient Reader mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.751688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.751688Z digest=sha256:e3c1ce687c642ae821339fbea29f59f01ec1df375ce1d7a511966cb0e8e21a76

Observation 639f9c5a-3974-4c55-a9ca-946c3a22b8e0 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.917214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:ee8cfe5ef255c5f5d5f9687cbf5f382c5c87815b9399b676cdfba1531bccaf64

Observation 2130ea05-3bd0-4f24-b5f2-f834062031b3 · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.645608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.645608Z digest=sha256:afbc7373203b01833abe7a1f4fa87c4687c5773c67dd1baee669af57aae89139

Observation c5de4ea7-eb2f-442a-bde9-424ba3531185 · inbound

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models cites this paper.

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:49.701778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:49.701778Z digest=sha256:3ea740d434e9f3b94863558981602f6a16c74afd4fd9ee8563a482499732f5cb

Observation 54713a2e-1baf-40b2-937a-1d7411c77327 · inbound

Ocean-OCR: Towards General OCR Application via a Vision-Language Model cites this paper.

Ocean-OCR: Towards General OCR Application via a Vision-Language Model mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:14:54.673534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:14:54.673534Z digest=sha256:047533c57557d0738a4d1f50ece8891c4a7cb05954d5e97082ce8462e008a5dc

Observation 4ee0b32f-9466-44de-912c-0d4252ff1857 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.674294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.674294Z digest=sha256:0fe72e45320ad43015b299d6a51d7d9d3fcd28ad698c82a9bb7f453da6c297c4

Observation e97737cd-0f35-4ae0-80d8-acfea68998ad · inbound

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding cites this paper.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.502876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.502876Z digest=sha256:924bb3f3feda10af821f46f00fb5cf5779387387acf0a98a766d95ac97c2b3e2

Observation c82a385b-598a-4c8f-87ef-8d3fcbdbcdda · inbound

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? cites this paper.

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T10:22:52.483040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:22:52.483040Z digest=sha256:ccf6e4f28f28c3e279154b8cd497afab01073929d37d56dd5c39bfaca678bdab

Observation 347854c3-2afa-4c36-812e-761718bb78d7 · inbound

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? cites this paper.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.491246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.491246Z digest=sha256:d0a7a12794c99e1a958e0b609386cdd0f99a27daba7e76f3fbfd10d106b4e186

Observation ff61d71a-9f7d-47a7-a885-c2a16f2f1f94 · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:25.692279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:25.692279Z digest=sha256:62247a40c0a9caa54b55843b2ae27ad9f87e0b4506df00f5d6819ef35e653dfd

Observation 434dc40d-4b84-4c9d-a5a0-3e4358cc30f1 · inbound

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning cites this paper.

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:17.619964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:17.619964Z digest=sha256:8467143d472ec2158f2424109ea7605f4267c19b3dd23a97a46188e63cf39394

Observation 7c7da621-157e-4f9a-9729-829ee48627e4 · inbound

CHAOS: Chart Analysis with Outlier Samples cites this paper.

CHAOS: Chart Analysis with Outlier Samples mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:53.854959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:53.854959Z digest=sha256:af8e2ac2a12e525a1d96d3f6b8746db40778e0f2ee1ab743013a1687f5e6c5c1

Observation d5c9bf1e-a4e8-4fb4-b9f3-a137212ed0c2 · inbound

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents cites this paper.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:37.142434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:37.142434Z digest=sha256:694028061163af4e30a0f80a101b258b85a862bef5f40a457fd09c69c016c2e4

Observation 1811b313-2de0-4609-84e8-33bc76d5fa5e · inbound

Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency cites this paper.

Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:28:11.394510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:28:11.394510Z digest=sha256:191c1a0569e3b5a996803c2783a1c71516a3292cb7a7cc675be4a01e52b925cc

Observation faf53257-4eb2-4980-a638-9bf324d05f69 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.341361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:ea70ee35b0bd0901442c2a247207f18b42a9ff111a82de438f55a952e7110eb1

Observation 1c5fe10e-e0f1-413f-b1ee-b83d1ece2d6d · inbound

ExpliCIT-QA: Explainable Code-Based Image Table Question Answering cites this paper.

ExpliCIT-QA: Explainable Code-Based Image Table Question Answering mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:15.333002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:15.333002Z digest=sha256:fbd5a85341fa6a4c4805ba656954defb27d09def4708f7386a75f434c3489ec7

Observation 1e4266db-48e4-4e6e-90ed-a8f540d82797 · inbound

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering cites this paper.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:52.666510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:52.666510Z digest=sha256:db2d879715b80acd23bb9ded0cbedd60e2301238d5e01cafb39ce6f93ffe31c8

Observation e5814c65-4bc7-47c9-818f-92d8ba83cbed · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:01.022352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:01.022352Z digest=sha256:06d5957e1f4a15c762bb127c394d5d89af91c6e0eddf72446d9e31d5bac2ffd1

Observation 2508b251-1ed1-4cf8-87d8-3812321e215f · inbound

Survey of Specialized Large Language Model cites this paper.

Survey of Specialized Large Language Model mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:37:49.232031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:37:49.232031Z digest=sha256:a61199e02f303fa50e610a727cf28a6d8cec9a0ee2270c2766feb4a3337934a2

Observation 8ffc5b73-2fb0-4973-9181-d92aa4eb664e · inbound

RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension cites this paper.

RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:48:00.804059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T14:45:20.535179Z digest=sha256:62d2b0dc46b561e45289a7108a421d1863faa1dde2b9bf64bcec40533730d154

Observation b95467a6-2bbb-4148-8252-87686a9f9a15 · inbound

MoDora: Tree-Based Semi-Structured Document Analysis System cites this paper.

MoDora: Tree-Based Semi-Structured Document Analysis System mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:10:15.582956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T19:09:41.482713Z digest=sha256:20e6822dab74ddb3e80cdafc5283f56cc2b94f35ae5c4be5dee15561efffd63e

Observation 13349e51-5bd2-4dc0-bb99-3977a5cabfc5 · inbound

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training cites this paper.

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:18:26.617773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T01:15:26.757215Z digest=sha256:858f8ac5160c1977749e185c9ac7699bfda4da808ea071038f7b1f58beb2b010

Observation b3fde4fd-881b-4d17-bee6-89b7105f20b6 · inbound

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval cites this paper.

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.404764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T13:17:04.441743Z digest=sha256:861cf9a9636294d27d16b603b942a1e396bf7dc5e59d41eedc5aba092c0a1056

Observation e796fad5-1ef1-43a3-ba26-0d6cd142a45b · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 190

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:00f9bbf6ad82a9283eda979b630a57f63a0cf45eba567c2c6f8b89053761859a