Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:31:36.673705Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2505.12766.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:31:36.673705Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T16:11:51.138098Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T09:11:00.663569Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4e898e83-0abc-4a92-97a5-49d6a7988c36 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0310bbd-8149-442c-b476-7ecb8b67f4d5 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 47eba2d0-4ebc-4b2d-9214-692f035a5915 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a3b8a5c6-b2ab-4ff1-9338-5c64f034be4b · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Onechart: Purify the chart structural extraction via one auxiliary token
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1edd40f8-57b2-448e-bdb9-ed62261577db · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7bbec18-b5fd-41fa-97ec-b988064f9b9d · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e0b87790-3eff-4ca2-93cb-3059c250311e · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa295075-0556-4fdd-9c27-5e5fe8c4f06f · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 347854c3-2afa-4c36-812e-761718bb78d7 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76aa7660-cc27-4fa7-8983-518bea5fd6e6 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? LLaVA-OneVision: Easy Visual Task Transfer
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e6110f5-c63c-467d-9361-79217ce2c16e · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 096892fe-06d3-441d-8eff-bc3abe7fd460 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Monkey: Image resolution and text label are important things for large multi-modal models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 74d3a0f9-77a2-4ba7-b166-6ca7cef8d207 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Focus Anywhere for Fine-grained Multi-page Document Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edf02ff3-4d44-46d6-8a90-5b240bf5f2c3 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mmc: Advancing multimodal chart understanding with large-scale instruction tuning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 028230fc-9476-4a55-a5db-4e189afedf76 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Improved baselines with visual instruction tuning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 01a93eae-8f7c-4358-b461-cb0b5dce8b8d · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Llava-next: Improved reasoning, ocr, and world knowledge, 2024
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 041c2c6e-6720-4b88-abc1-4d057104c833 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Ocrbench: on the hidden mystery of ocr in large multimodal models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 38bb003f-b811-44fd-894f-1a7d08ef9f89 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f956029f-b1a5-4fab-b568-ccd13c9d57a9 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 28531fba-c90d-422f-a967-cafcc9bee302 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbae85ca-7c13-41da-bc19-49eadac494a0 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1540bb76-157e-493b-86b8-84f3145a3d91 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mmlongbench-doc: Benchmarking long-context document understanding with visualizations
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6dc2d162-4b3d-4398-8e12-eb5de8065b3d · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 786414d7-3b4b-4132-9100-d08dd49d2ef4 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Docvqa: A dataset for vqa on document images
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 05e07651-8e08-4917-849b-0d9a5d371654 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Gpt-4o system card, 2024
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 154a248a-4893-4a57-a533-c88f53a68975 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Towards vqa models that can read
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 784c6e98-c911-49f1-8058-a0620f5ce0b2 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e82d1835-91d0-4a76-96b7-a9fb8970244e · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Contextual: Evaluating context-sensitive text-rich visual reasoning in large multimodal models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 104371b7-565b-4a82-a4f9-d4aaa8b3a135 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2381fb54-50f9-4f49-9dc1-d820695f5bb7 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66f9c24f-d165-45a6-ab4e-ab777082572e · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Charxiv: Charting gaps in realistic chart understanding in multimodal llms
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6141f1d2-1314-4115-9085-0329bb18c23e · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e8e4e41-ade1-42fe-9cb1-8d47ed0b8958 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc6e46c5-4558-4b8e-b9aa-5211e98293bf · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? ChartBench: A Benchmark for Complex Visual Reasoning in Charts
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f33518c-4841-4518-97d9-2af007e2bba8 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e8e66b1-1f2e-478f-b609-8a51f5b3f20c · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac2968f2-de27-417c-af0b-3a2764584527 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d59324d4-2907-4833-a643-93d1a6f209fa · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Exploring the capabilities of large multimodal models on dense text
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c43f4e5d-04f6-473d-acc8-a7322714040c · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4c93a76-ec00-46f5-b44c-a038b50c953b · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In ECCV , pages 169--186, 2025
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e740d921-a8e9-4ef2-bcff-3851a9807674 · outbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? write newline
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dade1ca-1000-4b86-8708-6d2536cc70fc · inbound
GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.