Pith. sign in

Paper Citation Record · LEDGER

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?

As of 19 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2505.12766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12766 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:31:36.673705Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:11:51.138098Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T09:11:00.663569Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e898e83-0abc-4a92-97a5-49d6a7988c36 · outbound

This paper cites GPT-4 Technical Report.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.449612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.449612Z digest=sha256:af50527c77d0f5f8153a31ed88af88a15571709f6f33c1254d57593edfa3dd86

Observation f0310bbd-8149-442c-b476-7ecb8b67f4d5 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.484551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.454637Z digest=sha256:ba1df42f40cafbea7fefd91910bf511d8e37f2496955a2b1d68b25fd49f48ae0

Observation 47eba2d0-4ebc-4b2d-9214-692f035a5915 · outbound

This paper cites Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.470776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.459084Z digest=sha256:48a904093b5b14de761099cd67e52d54800e225eb9b6b39ae28d562c346b1b64

Observation a3b8a5c6-b2ab-4ff1-9338-5c64f034be4b · outbound

This paper cites Onechart: Purify the chart structural extraction via one auxiliary token.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Onechart: Purify the chart structural extraction via one auxiliary token

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.456499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.464033Z digest=sha256:430d00d89aa872bce67a1a346c24aa4424ef6bb7db376f4b6edd3ec3844db430

Observation 1edd40f8-57b2-448e-bdb9-ed62261577db · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.468772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.468772Z digest=sha256:25b65cc7ef84c1d6d5ab81e9c2f87074bfa2bde70c56e7546467045264405fb0

Observation d7bbec18-b5fd-41fa-97ec-b988064f9b9d · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.443532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.474441Z digest=sha256:86a9589aa03f7e87801a46fb5fdd619d3eb9f5649ffa92482733f95a12e007da

Observation e0b87790-3eff-4ca2-93cb-3059c250311e · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.479795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.479795Z digest=sha256:75099a5d2c82c946aeb35c25e076b7b84523b99af77a12b7157daabfd515dc5a

Observation fa295075-0556-4fdd-9c27-5e5fe8c4f06f · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.486216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.486216Z digest=sha256:de0fa80fa0e1243a410e4e0595657ad5952dbbcfbfd97160793ef2d04113b1d4

Observation 347854c3-2afa-4c36-812e-761718bb78d7 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.491246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.491246Z digest=sha256:fdd418444bb326ce7a6870a5842b5bdb5ebcd04c152bd231c48a50acd3044ca5

Observation 76aa7660-cc27-4fa7-8983-518bea5fd6e6 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? LLaVA-OneVision: Easy Visual Task Transfer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.500944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.500944Z digest=sha256:f0a544c611af59d5eb86fa8a23c9add814363b004decc6f7ef05c4b329826df8

Observation 0e6110f5-c63c-467d-9361-79217ce2c16e · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.506036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.506036Z digest=sha256:6edc29baed20a78b537ab49ce5304f6f2bed07a2ef891e138d052f44dc76b5fa

Observation 096892fe-06d3-441d-8eff-bc3abe7fd460 · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Monkey: Image resolution and text label are important things for large multi-modal models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.424399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.513878Z digest=sha256:31d43a52e370b638a58a31bace3cfbb0052ccf2bad8bbc62a130fff1e2422d94

Observation 74d3a0f9-77a2-4ba7-b166-6ca7cef8d207 · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.519704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.519704Z digest=sha256:cc447bf345b07b86c86ce5edd1549457d495c5cdefef3733b0dbc68b26544e26

Observation edf02ff3-4d44-46d6-8a90-5b240bf5f2c3 · outbound

This paper cites Mmc: Advancing multimodal chart understanding with large-scale instruction tuning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mmc: Advancing multimodal chart understanding with large-scale instruction tuning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.378581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.524813Z digest=sha256:09d75fe440cf581dcc261104c25eabfe5dd5c0b7e43ccc81fcd74bce359c57f0

Observation 028230fc-9476-4a55-a5db-4e189afedf76 · outbound

This paper cites Improved baselines with visual instruction tuning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Improved baselines with visual instruction tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.345907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.531578Z digest=sha256:8eefdff249624cf9149eac6d7ba1b8845c6e8e2594aaa991cf1c7e00f131ac46

Observation 01a93eae-8f7c-4358-b461-cb0b5dce8b8d · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.537271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.537271Z digest=sha256:9d120dfac92714dc8eb2f8569f805df856a742cdd1d8dd46e81ca8b4e915e38b

Observation 041c2c6e-6720-4b88-abc1-4d057104c833 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.260264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.541759Z digest=sha256:91e5f7742fc979506a41f4f9443fd0ac861ffe1d99eb8beaab9d054b870f31ce

Observation 38bb003f-b811-44fd-894f-1a7d08ef9f89 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.546422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.546422Z digest=sha256:bd91729945050500f72d1c479a134b78a2df6ca542a42a66c52c657809add85c

Observation f956029f-b1a5-4fab-b568-ccd13c9d57a9 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.196376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.550750Z digest=sha256:acb8cb4a8ed6d1a36b22b16aa25b129b5d4410e09c2b3aa6f57ad84c480a3b95

Observation 28531fba-c90d-422f-a967-cafcc9bee302 · outbound

This paper cites MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.555576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.555576Z digest=sha256:57e47d7e7c17cb8de59ddd1fd8df708bdf6822b79ab07c41ff14ae60f5d20261

Observation dbae85ca-7c13-41da-bc19-49eadac494a0 · outbound

This paper cites MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.562481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.562481Z digest=sha256:6e28bfe925da8a150cdd5fc274d4e33d926bb7db6815f325617b084c313e7f49

Observation 1540bb76-157e-493b-86b8-84f3145a3d91 · outbound

This paper cites Mmlongbench-doc: Benchmarking long-context document understanding with visualizations.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mmlongbench-doc: Benchmarking long-context document understanding with visualizations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.156732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.573654Z digest=sha256:d7f4564cbeb94e103b4b285299aff1420b5eae615d5e62f71471aecee63b6914

Observation 6dc2d162-4b3d-4398-8e12-eb5de8065b3d · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.143623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.579098Z digest=sha256:7c31dcc8ae4e5657d76788acff25c206629a0afd53a88e5f547ffa9d254c0763

Observation 786414d7-3b4b-4132-9100-d08dd49d2ef4 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Docvqa: A dataset for vqa on document images

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.130656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.584922Z digest=sha256:fad3ba2753984ebedf49ee182cdd7ed8a6374957b2a5fa288c7b3ec4c7fe64f3

Observation 05e07651-8e08-4917-849b-0d9a5d371654 · outbound

This paper cites Gpt-4o system card, 2024.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Gpt-4o system card, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.589136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.589136Z digest=sha256:265fb6c5f5dde1ca46c6fb71ddc1286dd108fc8e896d2c4e71f70f44b64bb353

Observation 154a248a-4893-4a57-a533-c88f53a68975 · outbound

This paper cites Towards vqa models that can read.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Towards vqa models that can read

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.106263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.594518Z digest=sha256:25eb80a3c73a4233e308dbeaaab40f38b914d2a28e9c26d9de22bf226d8cb9c7

Observation 784c6e98-c911-49f1-8058-a0620f5ce0b2 · outbound

This paper cites MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.598636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.598636Z digest=sha256:9ba6bafcf25dae3aa1ad41b79ce769baa68b6642bf3d36ba48918665aa98aa79

Observation e82d1835-91d0-4a76-96b7-a9fb8970244e · outbound

This paper cites Contextual: Evaluating context-sensitive text-rich visual reasoning in large multimodal models.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Contextual: Evaluating context-sensitive text-rich visual reasoning in large multimodal models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.090992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.603135Z digest=sha256:955e5a4e65b2133b4fd994791da001d8d658a552cd0537ce1fa184f727df65e7

Observation 104371b7-565b-4a82-a4f9-d4aaa8b3a135 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.609414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.609414Z digest=sha256:96c1c760641a6fc14b187f710d2f22cc435c45745a850a970f9fb993f495bfa5

Observation 2381fb54-50f9-4f49-9dc1-d820695f5bb7 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.615071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.615071Z digest=sha256:2b88b836d8c8b0ecc85d0493fa7fed43a94273ad2566fe516dea27754d7212cd

Observation 66f9c24f-d165-45a6-ab4e-ab777082572e · outbound

This paper cites Charxiv: Charting gaps in realistic chart understanding in multimodal llms.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Charxiv: Charting gaps in realistic chart understanding in multimodal llms

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.077729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.619496Z digest=sha256:5456532a30bf2fdb67a6c3d5edc7e5aac7f6b82a897b9f59dadcbef70a0b8929

Observation 6141f1d2-1314-4115-9085-0329bb18c23e · outbound

This paper cites General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.623999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.623999Z digest=sha256:00f23dab3b4b2e2e695c0be1494fea7040c620f1975373d067812a69368e06f0

Observation 5e8e4e41-ade1-42fe-9cb1-8d47ed0b8958 · outbound

This paper cites ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.630121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.630121Z digest=sha256:43a46ac2743868b95e57e27198866040a6ee8e4e2d62fae2dd2e041f9c78180b

Observation bc6e46c5-4558-4b8e-b9aa-5211e98293bf · outbound

This paper cites ChartBench: A Benchmark for Complex Visual Reasoning in Charts.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? ChartBench: A Benchmark for Complex Visual Reasoning in Charts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.635245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.635245Z digest=sha256:1cfc7c3632e6b2ad6fdaf43cf0d07f90f23cd1d9adb0c049ddf882ecca52c264

Observation 4f33518c-4841-4518-97d9-2af007e2bba8 · outbound

This paper cites If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.640042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.640042Z digest=sha256:139a1055d061921144b9b215563fae3e167cd80a6a8a3d143ef9adb208035a5a

Observation 1e8e66b1-1f2e-478f-b609-8a51f5b3f20c · outbound

This paper cites CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.645575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.645575Z digest=sha256:b52c9356877e2ff574a4ee62a5a5985dbf32ed665387402e06f91a3ba49bddfa

Observation ac2968f2-de27-417c-af0b-3a2764584527 · outbound

This paper cites Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.065030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.650427Z digest=sha256:817476c2fff3625eb708376607bc63e05abe24aed6c4822d38350b615ad93693

Observation d59324d4-2907-4833-a643-93d1a6f209fa · outbound

This paper cites Exploring the capabilities of large multimodal models on dense text.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Exploring the capabilities of large multimodal models on dense text

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.051841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.654661Z digest=sha256:eb5fd1034babe63bcf327c724d14fe958d28a7879802d77a82828c225cd9a86c

Observation c43f4e5d-04f6-473d-acc8-a7322714040c · outbound

This paper cites Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.659746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.659746Z digest=sha256:facad41ffd6a62718e06e56e61987513653ddcd5c35a372b32abdf783db04ddc

Observation f4c93a76-ec00-46f5-b44c-a038b50c953b · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In ECCV , pages 169--186, 2025.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In ECCV , pages 169--186, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.038092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.664683Z digest=sha256:a1668374b214cf01625ecd2c6877ce283aa992b967d2a1ab3161aa7c306d2418

Observation e740d921-a8e9-4ef2-bcff-3851a9807674 · outbound

This paper cites write newline.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? write newline

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.673705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.673705Z digest=sha256:cd118dafa3d3d0c06f3882b3647e826f74933ac1a06495346cf1a71adb32d3e5

Pith citing papers

Observation 5dade1ca-1000-4b86-8708-6d2536cc70fc · inbound

GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts cites this paper.

GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:11:00.667074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T16:11:51.138098Z digest=sha256:12636bce4be6e4ca760a6c04793a405c57d8979a3a95317cc1568c944b37aaab