Pith. sign in

Paper Citation Record · LEDGER

UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2310.05126.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.05126 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:06:00.952910Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:18:37.543028Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7359fe39-2afa-448a-b8bd-ed7f0739d030 · inbound

DeepSeek-VL: Towards Real-World Vision-Language Understanding cites this paper.

DeepSeek-VL: Towards Real-World Vision-Language Understanding UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:58:54.738088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T17:58:54.177359Z digest=sha256:61884aa24687d7908a9d947cb10b426553ee7775ebf48e6ec8ef4198f7c818c6

Observation 4e34c82b-4a20-42d7-bde2-f41ceb3c4f19 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 127

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.271388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:df16d4560a076d199a443a3106cfd6de56b82b8bfa9c776944febd879a251d85

Observation 959f748c-3963-4c39-b326-c35b441f9251 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.796606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:645e0395e4f56fa384525d1d9f4efe7530084cfc367a04c1cf7919d1ae8bd510

Observation bf3a85f8-994d-4172-92fc-4ad13c8c5fd8 · inbound

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model cites this paper.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.903141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:578d3cdb2e2eb09e3eea91bd914ab8e51ec717b485d84cd1c31f962ef3cf1490

Observation dc054965-971e-4bc6-872d-ed373c5e9791 · inbound

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling cites this paper.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.831310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:79f9a421a4243379a072a72243cb50cf24217d3d0488b0f2814c89405d77a30e

Observation 3c622375-6f48-4b08-86a0-7efd163a6d25 · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 279

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.302097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:a9625caba3b59d9b2f111c6c1063023314b5826eff8f8f012f3afde56d762bdf

Observation 02db6aa8-8324-41a9-ae1c-46a9a356fac4 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.785865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:63322c99557882aa503c7eef15b558b401556a9287ee19345a7777e7f1d3433d

Observation bbae7ba5-2786-4931-9df5-154052e9c283 · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.711229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:a741c81f17b627dd65ce7610ca48030572235aa0d60f5a420efe6c2a86cd70c7

Observation a489b425-3008-48d3-9604-c616a7e11f3f · inbound

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence cites this paper.

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:00.952910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:06:00.952910Z digest=sha256:9632be06991728e0ecb25810f0c6e1f65736933da63a47a5f405c350fce835f6

Observation a14b27b0-8113-4d7f-8a65-b09e36415b5a · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:30.762169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:30.762169Z digest=sha256:6552ef6bc8bc85a71baf081a231cf6cd5b21564d2fe0f7ffcdc38df0a9e134a3

Observation 5a2db8e0-340b-4ab7-b439-f2c2c86854f8 · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.654963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.654963Z digest=sha256:91aa4774e2d4eff44fac401ce462d0b82858e856c6e3d5cb496f3cc9aa98c5a9

Observation 0190d4b3-deae-47a2-960d-9ef4e1052993 · inbound

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation cites this paper.

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:56.033608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:56.033608Z digest=sha256:557260e23c1d61bee359a9afcb3c242d2aa49fde094a6de5a1131325fb90717c

Observation 853d8cd9-fce2-459e-a3ef-e6e69e8209b9 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:04.980280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:04.980280Z digest=sha256:d28d098cf324e121270cd93ce81527041ae9f0755201f947a1ec8fe8fddc2c1e

Observation dc9c12fa-4573-4023-a931-b729e6d14542 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.666913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:a95b74f3a8610e402f3ac5c4d102bb621342881f2a924301d091803e567fc34c

Observation d962ee53-3e50-4bc5-864a-9d9354d1a2b4 · inbound

Evaluating Vision-Language Models as a Zero-Shot Learning Alternative to You Only Look Once and Optical Character Recognition for Nigerian License Plate Recognition cites this paper.

Evaluating Vision-Language Models as a Zero-Shot Learning Alternative to You Only Look Once and Optical Character Recognition for Nigerian License Plate Recognition UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:18:37.544571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T16:10:03.765621Z digest=sha256:2b1e4c930fb5bae39adf68610769c5039710251ea5ffc6e59113d62bda06f396

Observation 074c23a4-80a0-40e5-8eb7-63c922968231 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 147

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:89d463d280ba8bc783eda7a40bf7da55e7d1d0fa7476143aea94af916809c39d