Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2403.12966.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:27:03.174749Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T15:59:57.253374Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation a0251e79-613a-493e-b49f-b8077e2b6f91 · inbound
Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c529c18e-e9da-48dc-b009-f7231aa7fe71 · inbound
VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a157d044-e05c-4ff9-afad-74503f0d4c9a · inbound
Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12ae62c6-8899-485a-86ff-e0ce6deaca68 · inbound
Ola: Pushing the Frontiers of Omni-Modal Language Model Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd69af3a-c658-4de7-b693-ad706aefd14b · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f80a9b77-b39d-4501-92d5-8c31e8374029 · inbound
ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3194040f-a74e-4226-818d-3234125cc836 · inbound
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0fdd9d3-d2de-4c1f-b2c4-56f6c581b613 · inbound
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb5ea1f2-c357-474d-8ffb-11d85425ac2d · inbound
AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b64d36b5-fdc8-43d7-9c26-6476c5b90630 · inbound
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c23b8a61-b5cb-4a6c-8acc-8171becc0267 · inbound
CoT-Segmenter: Enhancing OOD Detection in Dense Road Scenes via Chain-of-Thought Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd23a595-bf5b-4d0f-8550-15b82f3f5b94 · inbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2e4a3485-1b56-4e99-860b-5402c90fa88f · inbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c23e29bd-49d4-4ee2-a256-716b3e5727d9 · inbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0d61b9f-e1be-4f8d-9c4f-fbb619126a18 · inbound
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc6636ce-071f-46a3-91a6-23c9f979f512 · inbound
M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 684c86e5-0d91-499a-85b5-0a00f9524f72 · inbound
HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f45d257b-0d11-45a7-813b-85ed3811a840 · inbound
Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23e646b7-5992-4713-afa8-5cd784aa1b6c · inbound
Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a961c15c-59e3-4194-9a80-6740ef13ffcc · inbound
Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 12dd8a4c-9c84-456d-94f0-cd1823df109c · inbound
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bbef1c4a-05ba-41c9-84bd-d1f8555e4ab0 · inbound
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 02bf47d2-71cf-4189-acf2-9dc74a589aea · inbound
MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6fec5a22-b305-436b-9315-e96dc9f8b8d2 · inbound
An LMM for Precisely Grounding Elements in Documents Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b6d50c90-17fe-415b-8956-ed10cf4791f2 · inbound
ActiveScope: Actively Seeking and Correcting Perception for MLLMs Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b4d01059-1aac-4bd7-882c-e1b6a0ef5325 · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 130
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4cf2e4d6-aaca-4331-a20f-c6aba144b36b · inbound
Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4a13e359-7a5d-4ebf-ba38-dea51b07c98f · inbound
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 210
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 125a6848-5382-4caa-adf9-477492677fed · inbound
Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43170df6-d614-4c33-94f2-bfce8ce52198 · inbound
Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9741293f-4a3e-4b3d-8c62-acc4f2a1ba01 · inbound
ToolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fed3ace4-7f7a-4068-a766-df90f1678d4d · inbound
Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.