Pith. sign in

Paper Citation Record · LEDGER

UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2310.05126.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.05126 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:06:00.952910Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:18:37.543028Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7359fe39-2afa-448a-b8bd-ed7f0739d030 · inbound

DeepSeek-VL: Towards Real-World Vision-Language Understanding cites this paper.

DeepSeek-VL: Towards Real-World Vision-Language Understanding UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:58:54.738088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T17:58:54.177359Z digest=sha256:5699469a604ffdae4fb1c7b42b448e7e9759dd14569d6880cd6b7e099ad38fa8

Observation 4e34c82b-4a20-42d7-bde2-f41ceb3c4f19 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 127

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.271388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:afaa735a9338337b605191024906b2535918fc71c828e27395dd726a9a527e7a

Observation 959f748c-3963-4c39-b326-c35b441f9251 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.796606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:82395cf9de20b56c19769a61146c0858c4f7900162f1b9614ef994f37d454114

Observation bf3a85f8-994d-4172-92fc-4ad13c8c5fd8 · inbound

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model cites this paper.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.903141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:64ccbaf4ec7dedb9e9b396548083dd651b117e896b2e3612d1cb65a566205028

Observation dc054965-971e-4bc6-872d-ed373c5e9791 · inbound

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling cites this paper.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.831310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:d0a8e6ed0a69ec760393dfc7f143182086feb9edf9c9493b28e6a67aa1e79a43

Observation 3c622375-6f48-4b08-86a0-7efd163a6d25 · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 279

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.302097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:6aa4c262932c8c854379470c30c9b55d403ecf4e32d7d05281aaf7a2965171b4

Observation 02db6aa8-8324-41a9-ae1c-46a9a356fac4 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.785865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:fe5bd56aa4cb4f4fcda1f748b395665fbdd084b879c4a0803b487e062d120768

Observation bbae7ba5-2786-4931-9df5-154052e9c283 · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.711229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:1a8adf5977e4d9d5ecf391ae98fbf4b78d9b602452c7764558c0665435620ac7

Observation a489b425-3008-48d3-9604-c616a7e11f3f · inbound

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence cites this paper.

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:00.952910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:06:00.952910Z digest=sha256:9632be06991728e0ecb25810f0c6e1f65736933da63a47a5f405c350fce835f6

Observation a14b27b0-8113-4d7f-8a65-b09e36415b5a · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:30.762169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:30.762169Z digest=sha256:edc4f2aeffc41678284725874262ba3fce01c66ae71ad76c9c34a62799632254

Observation 5a2db8e0-340b-4ab7-b439-f2c2c86854f8 · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.654963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.654963Z digest=sha256:6f2914824358ead0118ba5909ba887167a2fe0478b0f57b4e204d9a80461080f

Observation 0190d4b3-deae-47a2-960d-9ef4e1052993 · inbound

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation cites this paper.

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:56.033608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:56.033608Z digest=sha256:c4bcd42045e31990c20c9ebd507254e3d084ebead2f7c73c9e823e7abfc7f0a1

Observation 853d8cd9-fce2-459e-a3ef-e6e69e8209b9 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:04.980280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:04.980280Z digest=sha256:a5b214acc399d32c5a030f9cf6864242f7c2c730aa9aeb66f3cd976f4c47cc76

Observation dc9c12fa-4573-4023-a931-b729e6d14542 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.666913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:87cc84c120353f492c88069dcaacda3baa7bb619102abd4aaff53b180c43e10a

Observation d962ee53-3e50-4bc5-864a-9d9354d1a2b4 · inbound

Evaluating Vision-Language Models as a Zero-Shot Learning Alternative to You Only Look Once and Optical Character Recognition for Nigerian License Plate Recognition cites this paper.

Evaluating Vision-Language Models as a Zero-Shot Learning Alternative to You Only Look Once and Optical Character Recognition for Nigerian License Plate Recognition UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:18:37.544571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T16:10:03.765621Z digest=sha256:c7b1df745bd06cdd99999bd8fa72ee083bb47f64deec1cb8f64b9d30e35004a0

Observation 074c23a4-80a0-40e5-8eb7-63c922968231 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 147

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:89d463d280ba8bc783eda7a40bf7da55e7d1d0fa7476143aea94af916809c39d