Pith. sign in

Paper Citation Record · LEDGER

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?

As of 18 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2505.12766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12766 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:31:36.673705Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:11:51.138098Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T09:11:00.663569Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e898e83-0abc-4a92-97a5-49d6a7988c36 · outbound

This paper cites GPT-4 Technical Report.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.449612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.449612Z digest=sha256:a05697b51bc81dcac192a526d4aebe9c949b3ba1510cf5dc9b1037f7adb12ee2

Observation f0310bbd-8149-442c-b476-7ecb8b67f4d5 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.484551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.454637Z digest=sha256:464cd488dd3c75da8387f55d1fccf18e63caffd9e1da3d98d7a11df1b6965d9c

Observation 47eba2d0-4ebc-4b2d-9214-692f035a5915 · outbound

This paper cites Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.470776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.459084Z digest=sha256:13c2d6e33f572e7ea37da06514e05cead2cc5125c3ac48107e76f74d44431a71

Observation a3b8a5c6-b2ab-4ff1-9338-5c64f034be4b · outbound

This paper cites Onechart: Purify the chart structural extraction via one auxiliary token.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Onechart: Purify the chart structural extraction via one auxiliary token

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.456499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.464033Z digest=sha256:9c04fc53da87ffc258e1a405285eb11eb978158e547ade3c38f48e540c18b392

Observation 1edd40f8-57b2-448e-bdb9-ed62261577db · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.468772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.468772Z digest=sha256:3da36597a977a00288a78045e4c4f03efeb496986f5446221050fcd6c3dff161

Observation d7bbec18-b5fd-41fa-97ec-b988064f9b9d · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.443532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.474441Z digest=sha256:ea311b1a86935a5121f3706138b19f578ba69f36b352b14e20c9ba4dd88b2c4e

Observation e0b87790-3eff-4ca2-93cb-3059c250311e · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.479795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.479795Z digest=sha256:1214b9da0044cc33bad7149e7006a00244467c1a844b54b1bd75cb9d944e1975

Observation fa295075-0556-4fdd-9c27-5e5fe8c4f06f · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.486216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.486216Z digest=sha256:2782f86c09e052a322008f05772068f235dc66fa76da1267d03bad6df311d35d

Observation 347854c3-2afa-4c36-812e-761718bb78d7 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.491246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.491246Z digest=sha256:411bd4d00e3b857f853a0b268f75f27a10c733b387a2f24f2dec5028da00b998

Observation 76aa7660-cc27-4fa7-8983-518bea5fd6e6 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? LLaVA-OneVision: Easy Visual Task Transfer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.500944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.500944Z digest=sha256:e05db461650f40bd426637be7fbd482cbe6fff467ca5673606edb94245c322c3

Observation 0e6110f5-c63c-467d-9361-79217ce2c16e · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.506036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.506036Z digest=sha256:e4de22b8bb1801ef1a4b51c4eb63b41a1fe63140547a4cde4cf6beec6e16590c

Observation 096892fe-06d3-441d-8eff-bc3abe7fd460 · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Monkey: Image resolution and text label are important things for large multi-modal models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.424399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.513878Z digest=sha256:a1d672daf1ad0a48212511fd6a8f1b5f6ee2af8a145a2a850a4f97d44beb49f7

Observation 74d3a0f9-77a2-4ba7-b166-6ca7cef8d207 · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.519704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.519704Z digest=sha256:9cef6f4e696b269fc4b6e5bb49b609aae85168338527c832325c5aedd3e69709

Observation edf02ff3-4d44-46d6-8a90-5b240bf5f2c3 · outbound

This paper cites Mmc: Advancing multimodal chart understanding with large-scale instruction tuning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mmc: Advancing multimodal chart understanding with large-scale instruction tuning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.378581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.524813Z digest=sha256:7269fe28b5c93ba6d2b5836a330e8c97d678ee1ff69600ce9255ffbfc3765d95

Observation 028230fc-9476-4a55-a5db-4e189afedf76 · outbound

This paper cites Improved baselines with visual instruction tuning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Improved baselines with visual instruction tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.345907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.531578Z digest=sha256:0f811014f95dbd653a03d2929d1792d3cd082ab7248159c2cd2d69081d0d9554

Observation 01a93eae-8f7c-4358-b461-cb0b5dce8b8d · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.537271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.537271Z digest=sha256:e4dcf18354a92d213cd8c0142ea17f20318734f1c56d0f91d756666aa7b8c5d2

Observation 041c2c6e-6720-4b88-abc1-4d057104c833 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.260264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.541759Z digest=sha256:fd1cadfcf6d525125ece697243fd7974c27d1a4e2f16047d9d3239b6124839a7

Observation 38bb003f-b811-44fd-894f-1a7d08ef9f89 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.546422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.546422Z digest=sha256:4a2596b0f6c3ba0651ee9dc6be007eec3ae72da84e62cae4e6b1447f3c9d3f9a

Observation f956029f-b1a5-4fab-b568-ccd13c9d57a9 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.196376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.550750Z digest=sha256:fd51c6ac281a588a4d9c25bb9f9b3d8ac14c2047ad34255b9df840ae39ec341d

Observation 28531fba-c90d-422f-a967-cafcc9bee302 · outbound

This paper cites MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.555576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.555576Z digest=sha256:f61932b65ae356ad60212c7cdb335223accfacbc5bb07e0ebc1fee73d7474198

Observation dbae85ca-7c13-41da-bc19-49eadac494a0 · outbound

This paper cites MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.562481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.562481Z digest=sha256:1241fdf76c2eb47b2ceabb74d25882a8cd749d0cd537aeb8512c525e3c98d406

Observation 1540bb76-157e-493b-86b8-84f3145a3d91 · outbound

This paper cites Mmlongbench-doc: Benchmarking long-context document understanding with visualizations.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mmlongbench-doc: Benchmarking long-context document understanding with visualizations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.156732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.573654Z digest=sha256:48d360dd3bfc6e799cadc45e62e8c6ab8573d4291a69c6538dcab2f6c7634bc2

Observation 6dc2d162-4b3d-4398-8e12-eb5de8065b3d · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.143623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.579098Z digest=sha256:4ffa3ff3f5034b34645defe4a9ec8f479f7faec92ec2a999804365f8ddb24ac6

Observation 786414d7-3b4b-4132-9100-d08dd49d2ef4 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Docvqa: A dataset for vqa on document images

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.130656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.584922Z digest=sha256:b0e8fb3d49edacfc316a259eaba350a716eb9ae61017ee1ba22b28be8f593545

Observation 05e07651-8e08-4917-849b-0d9a5d371654 · outbound

This paper cites Gpt-4o system card, 2024.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Gpt-4o system card, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.589136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.589136Z digest=sha256:e17bd91824d601805fbd2a09d59a53ccb0f7ea410fbd3e41048cc475c7d12e5b

Observation 154a248a-4893-4a57-a533-c88f53a68975 · outbound

This paper cites Towards vqa models that can read.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Towards vqa models that can read

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.106263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.594518Z digest=sha256:9c69c2fba8eca4f5686646c07fd908ebc371e76f9279f083e1300c36e32f06e7

Observation 784c6e98-c911-49f1-8058-a0620f5ce0b2 · outbound

This paper cites MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.598636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.598636Z digest=sha256:99147a21e9f00e069e603272b203b69c825b0438967a01dfa9abff980c549805

Observation e82d1835-91d0-4a76-96b7-a9fb8970244e · outbound

This paper cites Contextual: Evaluating context-sensitive text-rich visual reasoning in large multimodal models.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Contextual: Evaluating context-sensitive text-rich visual reasoning in large multimodal models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.090992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.603135Z digest=sha256:56582a5bc756a76c43c591d36a59a063ad43203ed244866b55d651c18072dbf9

Observation 104371b7-565b-4a82-a4f9-d4aaa8b3a135 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.609414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.609414Z digest=sha256:d925630c7542b849d4d8de06a9d1e53d4b6fcf7f7e545c46599bf698dfcd8c5c

Observation 2381fb54-50f9-4f49-9dc1-d820695f5bb7 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.615071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.615071Z digest=sha256:c8f71576e9d7c2301789ce99560be4eac9600f508329a11b136556a04a49a26d

Observation 66f9c24f-d165-45a6-ab4e-ab777082572e · outbound

This paper cites Charxiv: Charting gaps in realistic chart understanding in multimodal llms.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Charxiv: Charting gaps in realistic chart understanding in multimodal llms

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.077729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.619496Z digest=sha256:539700d7ddf41ca4c08f1e5e35d747395a972ffad97a182aa23d8b7b6275e5df

Observation 6141f1d2-1314-4115-9085-0329bb18c23e · outbound

This paper cites General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.623999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.623999Z digest=sha256:5929383d2c4b96c57fd4e2cb992b0505474e34d9bbfce496237eec6d224a95d7

Observation 5e8e4e41-ade1-42fe-9cb1-8d47ed0b8958 · outbound

This paper cites ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.630121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.630121Z digest=sha256:cfc7b8b2333fb6f8da4bbd12f002e971e30571d83004b6ea825ba940d9cc30a4

Observation bc6e46c5-4558-4b8e-b9aa-5211e98293bf · outbound

This paper cites ChartBench: A Benchmark for Complex Visual Reasoning in Charts.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? ChartBench: A Benchmark for Complex Visual Reasoning in Charts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.635245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.635245Z digest=sha256:37ee38253672cf63480a83d85e2553d5da0bd46f59d092e06fb9643ed4e28fc5

Observation 4f33518c-4841-4518-97d9-2af007e2bba8 · outbound

This paper cites If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.640042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.640042Z digest=sha256:67111632fd21c3abd21c23715012d31698965441953f24a03927ee821293219a

Observation 1e8e66b1-1f2e-478f-b609-8a51f5b3f20c · outbound

This paper cites CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.645575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.645575Z digest=sha256:434adc8f6ad2fd041a75cbc8ab65aa311a62d79896c59099b9fb7f06317ba0b6

Observation ac2968f2-de27-417c-af0b-3a2764584527 · outbound

This paper cites Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.065030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.650427Z digest=sha256:9c71be4525f95b2565d500b4c635ec247f9312365ca4327da9de0b8d82833685

Observation d59324d4-2907-4833-a643-93d1a6f209fa · outbound

This paper cites Exploring the capabilities of large multimodal models on dense text.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Exploring the capabilities of large multimodal models on dense text

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.051841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.654661Z digest=sha256:3b7b1a24d889c39b4d635d20692e221c50259fb3355ee4012c1bd7a3ac4a9d4b

Observation c43f4e5d-04f6-473d-acc8-a7322714040c · outbound

This paper cites Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.659746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.659746Z digest=sha256:f0a7d762bdb9db74e7c372d5418a82905faeec042ab6db6677b26919e5499702

Observation f4c93a76-ec00-46f5-b44c-a038b50c953b · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In ECCV , pages 169--186, 2025.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In ECCV , pages 169--186, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.038092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.664683Z digest=sha256:a94e58985072570a4f2c99f6994f4fd0a39bc9808890d3177ca727869de26f50

Observation e740d921-a8e9-4ef2-bcff-3851a9807674 · outbound

This paper cites write newline.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? write newline

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.673705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.673705Z digest=sha256:bd00cf31e2c848bd1fa13d0bb03dbea5aca21bae489264ca241329bbbf512954

Pith citing papers

Observation 5dade1ca-1000-4b86-8708-6d2536cc70fc · inbound

GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts cites this paper.

GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:11:00.667074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T16:11:51.138098Z digest=sha256:3ec26491653820ae77ae7a92bdda30c8f72c50a5be1b0b79272d0ac6180b7b88