Pith. sign in

Paper Citation Record · LEDGER

mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2403.12895.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.12895 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:23:33.470903Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T17:07:25.743722Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9636f990-2707-4fc7-ac5b-fcd8b73293c8 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 158

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T02:56:41.782309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:32ab8d6ddf688f348517b8ee446664c5dc0532786370357ab7505b157190d2ba

Observation 87856638-81ac-452d-8d4d-5731764e1050 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T20:58:58.997584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:16bf4735109440cbf87ff946e0bef3556fc28d3b88fe6116495afa50066a133f

Observation 24372ced-89e2-4eb8-9ceb-ec8101d89e10 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T10:46:28.806610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:8960c2609f56f32d82ed0048249f8a045f3c4fc129a9cf4b9b39778debafd597

Observation 20d6da08-e96d-4de7-90c8-56bcb0da7189 · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:59:32.748963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:bc1825f6e6c5d0d8e9720796c3e485a051f7e62bf1fc9ee1d5e656dab1663250

Observation 29680334-4140-4af1-ae06-56f51675dc11 · inbound

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model cites this paper.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.864055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:0b177eff4082253930cefd36ab944e4b5e232f5af2d40d4de8eaa7131ff99e1b

Observation c2efe267-4983-441c-a96f-21bb893de652 · inbound

MinerU: An Open-Source Solution for Precise Document Content Extraction cites this paper.

MinerU: An Open-Source Solution for Precise Document Content Extraction mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:00:25.807671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T04:00:25.624430Z digest=sha256:fab83e3f4a9d0e6dd228ca6a711a21988741547f1d4be17f97776760a5e177f2

Observation 20b7ed55-3765-49f7-9512-ba4426200797 · inbound

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction cites this paper.

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:12:14.657006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T12:12:14.613620Z digest=sha256:4d7f60399547af7c376d55fe482cc098c5c14d5e5828618fea940b01a24679c2

Observation 0bb3dfa2-7fa4-495d-bb7d-4f3df067811d · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.181515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:9c452c4f28b3d999ad93c96ee0c3fa6f7e4482d33001569ae489b805f23dbdf4

Observation 36eb7878-7b97-440b-86f7-8d44812fc120 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.855416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:e3c79c2a60808a0e8334a29b84250ba898de71ca72cefaa44fba640534bcdfa3

Observation a94e4643-1970-4d16-8da9-39b3ed21684b · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:33:26.789867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:e66a6431f456f3b94c9c93e1f6ef87ae7ebb75c413db8eb3476e396ba440fa84

Observation 0f7ce9a3-e1cc-4166-8812-31f67a1a3b84 · inbound

OCSU: Optical Chemical Structure Understanding for Molecule-centric Scientific Discovery cites this paper.

OCSU: Optical Chemical Structure Understanding for Molecule-centric Scientific Discovery mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T14:23:33.470903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:23:33.470903Z digest=sha256:e5fef325f632e09a119ffab22ab78d362be82204368048b2a64cd1bfba5d0454

Observation 4d675fcd-16d9-4f0b-9ef9-a13f11b5db2e · inbound

Ocean-OCR: Towards General OCR Application via a Vision-Language Model cites this paper.

Ocean-OCR: Towards General OCR Application via a Vision-Language Model mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T14:14:54.670740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:14:54.670740Z digest=sha256:6a4f21e30584018360987d07cb16c05ac0104e2eadbf8d76e383353bd0e5d09e

Observation 87caebc6-9a09-42e7-85d4-f1a68256f4f8 · inbound

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs cites this paper.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.702360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.702360Z digest=sha256:0ec12bb82058c56f77548679c122f4ffa83fabe221f2ead16cb58792bfb1c680

Observation 04b58b47-9d1a-454e-a80d-43ac0c108524 · inbound

\'Eclair -- Extracting Content and Layout with Integrated Reading Order for Documents cites this paper.

\'Eclair -- Extracting Content and Layout with Integrated Reading Order for Documents mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T23:13:13.864169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:13:13.864169Z digest=sha256:2642f3f556632e9ff012ccdd0525d16af62dd90b51c43bb92a509b5cc165ec76

Observation a5d3e400-8976-49bc-9b90-6504d8b2b832 · inbound

EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition cites this paper.

EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T22:57:16.799161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:57:16.799161Z digest=sha256:355b1c642afba08fdbc423cb55796e74723f450dadb5bde67cd0d99085ba4225

Observation dd067796-af49-481d-a477-ddf1dea178ac · inbound

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence cites this paper.

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:00.747244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:06:00.747244Z digest=sha256:98aaf472f2b4616a3670f60d1d159d3b64a697f3aed6fbe7fd0714bf3a7f8cbe

Observation 196ae3f1-9380-4d33-b451-9513513ea221 · inbound

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM cites this paper.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:11.437146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:11.437146Z digest=sha256:31c3866e7b59287b7fd2fcfa21372b1d5da6176e7bace9ee88463721852cf81f

Observation 92ca35fa-89b2-4977-9350-01ceec4a6f3f · inbound

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning cites this paper.

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:17.552695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:17.552695Z digest=sha256:dbd51f93df0f10b2f97901dab7066acb2ff11f031ebb8144e52437c545b28515

Observation 7d1d87a7-911e-4369-ba43-a92325bf8a57 · inbound

CHAOS: Chart Analysis with Outlier Samples cites this paper.

CHAOS: Chart Analysis with Outlier Samples mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:53.742825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:53.742825Z digest=sha256:0184fcd77035d344c31a835e3fd862e48c02196cdb3f4b6b4baa2c4d81389843

Observation 913ffc16-e220-4270-96aa-c7f3c49aa2f9 · inbound

Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning cites this paper.

Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:40.157455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:40.157455Z digest=sha256:579fe81bb6255add7d4af8c0fce439971350bf4b9df3aaad0eaedae270b9f79e

Observation 25624c85-eccc-45a0-a88a-5a25ead4f616 · inbound

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement cites this paper.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:53.021225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:53.021225Z digest=sha256:6b0057011a7b0f50e69fc254f3d56a33effe529a9d3ff53623efd0ea7aab277c

Observation 13ce348c-6f4e-4ff0-a75c-36eed2ec4bee · inbound

CoMemo: LVLMs Need Image Context with Image Memory cites this paper.

CoMemo: LVLMs Need Image Context with Image Memory mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.234729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.234729Z digest=sha256:c0c38167789ddfe53f652155acaffbcce55bb6d72a98dbb3cd4e4e7c93b2f08d

Observation aa7ec550-3dc2-4674-a241-f86be58475fe · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:23.097635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:23.097635Z digest=sha256:ee9df8d4c88ab7529f745c16c69261f88769404cd483d66b4881d3fe203c2a0d

Observation 57b3f667-02d7-41d7-844d-fc169ac4be3f · inbound

Structured Attention Matters to Multimodal LLMs in Document Understanding cites this paper.

Structured Attention Matters to Multimodal LLMs in Document Understanding mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.065585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.065585Z digest=sha256:c4d2b1f7754e4a903fbfca3cd23bdb650f0120d844c280415ce2a837a5b60dda

Observation 29ef0fcd-2bc9-4672-916c-740cfec8368e · inbound

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models cites this paper.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:08.059391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:08.059391Z digest=sha256:408c2b50064729234fb6cde2f704d256aeebd2eb1bdecb1233e515edfc58e394

Observation 28ac8605-6fd0-46f2-bb60-72bcd742971a · inbound

ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning cites this paper.

ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:37.113140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:37.113140Z digest=sha256:49373faa196af7d2203643f37d06fcc166d90453333991bb368d907d14a415b0

Observation 310cdb34-bc96-4d48-ab61-5808c1d8a2b8 · inbound

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation cites this paper.

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:17.602697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:17.602697Z digest=sha256:d62a138623e99cb4880ed1018071e44cb1a218b840dc7beb5fb6f03d1600983f

Observation 3cb95d74-63eb-4d1c-bd59-1439d6c85896 · inbound

A document is worth a structured record: Principled inductive bias design for document recognition cites this paper.

A document is worth a structured record: Principled inductive bias design for document recognition mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 39

Resolution
malformed identifier
arxiv_id, observed 2026-05-19T05:02:05.215068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T04:57:35.441758Z digest=sha256:7e0be97e6ffc042c61d78fa719d6fad88cef69bd09ba53004cf8c375ae9cf227

Observation 2921b240-0064-4586-9758-4014f25367d4 · inbound

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation cites this paper.

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:49.641010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:49.641010Z digest=sha256:d1b56623205de1ee7e71b8712579faa8e4be62b33e7eb03e5c1abb7909152774

Observation 7dd154c6-d455-4066-8564-f221c43666c4 · inbound

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering cites this paper.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:14.796214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:14.796214Z digest=sha256:9bd76228947cebf4efa5c8119596a78477ac3904d61385a8f053496746795563

Observation acea7df4-fe20-4d75-9810-9ed67ee89d3c · inbound

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding cites this paper.

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:32:36.576223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:30:50.984873Z digest=sha256:86067663c141a9420015f3ebe969e5f800c675efe51b7b072e4b7a9fab532f64

Observation dfca068a-ce1c-4348-87f7-f1a55734cadc · inbound

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework cites this paper.

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T20:19:14.038805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:19:14.038805Z digest=sha256:a790c3c3eccbb7748d8f9f533be7104f3bfc7fde9336e8801230914127396b63

Observation 2ba7f8f2-2ea6-4e31-99fd-aa52ca51d8e0 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.692458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:9754d2dc6f1e28d7c6a1951a9e2c843cb25575873ec327942ffbeef1ad33b378

Observation 8ed5f5c4-6e72-483c-9ba1-909a80bff665 · inbound

SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models cites this paper.

SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:57:15.438914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T07:56:05.583390Z digest=sha256:a5247d9a7907c7395fd5892e4d482409f808dd433dc496ea520409a03e3babd3

Observation 9c6b683e-1544-444d-9dd6-40278a803d47 · inbound

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence cites this paper.

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:37:58.163050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:37:36.144960Z digest=sha256:69dcef49fe1b9d84bd647f7680a3900071fb1874b391b9711cad22497d2223cb

Observation 6492d638-4007-49ba-a24b-3196fe499c0d · inbound

Infinity-Parser2 Technical Report cites this paper.

Infinity-Parser2 Technical Report mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-10T17:07:25.745651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T17:02:28.089092Z digest=sha256:904947b8eb10396c74c38b2711ab784be3218d80bdc17b1997b2933fe2e971a7

Observation ccfee072-f5e6-40a4-be92-bdac0ad074c8 · inbound

Infinity-Parser2 Technical Report cites this paper.

Infinity-Parser2 Technical Report mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T08:03:51.775355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:03:51.775355Z digest=sha256:05ed70f649b50403c4b6e51c7681518e10d6d2208588e0d8347c89b84a8d3495