Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2404.09797.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T06:06:07.136455Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T15:09:55.324968Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation a3550c13-5430-4d16-beb1-4f025fe774c8 · inbound
CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7428d43e-8f64-49c9-ac41-3279c7a5462a · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 14d2e6a5-43c5-4b83-9c53-e5550f74b49f · inbound
Training-Free Reasoning and Reflection in MLLMs TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2480a51-1e1a-4a8a-a84c-8fb6af04d3b8 · inbound
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 768f4ee9-819d-409a-82a3-6511bcca86bf · inbound
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c86508c6-2314-49f8-aa28-9a37f1f5abea · inbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c58f9d50-3dee-4194-a524-81aad9ff36cf · inbound
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf724eff-60da-4cd5-9d17-31b711af7b41 · inbound
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 968f33d1-246c-4337-bb3f-7db71a08302a · inbound
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 956a5d75-d46e-45d8-8fac-b1b44f4ea0a1 · inbound
M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a53a9d3-7f3c-4e1b-9c85-eef5ff213a55 · inbound
HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc652840-7b6c-4a37-b668-10437765a6d4 · inbound
Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9a6ef3ee-0bca-49a3-98e9-2309dbe33b32 · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.