Pith. sign in

Paper Citation Record · LEDGER

Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2403.12966.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.12966 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:27:03.174749Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:59:57.253374Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a0251e79-613a-493e-b49f-b8077e2b6f91 · inbound

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models cites this paper.

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:06.857653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:06.857653Z digest=sha256:c206607a45027dff2c359f50e712bd8876584f673d5d4ec2e33ecd720cdb83bf

Observation c529c18e-e9da-48dc-b009-f7231aa7fe71 · inbound

VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning cites this paper.

VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T23:50:01.313217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:50:01.313217Z digest=sha256:cc92dd8a3cb784c8ed061cb0e7ab0b683269fc469aa4a1185f1d4014b52c037d

Observation a157d044-e05c-4ff9-afad-74503f0d4c9a · inbound

Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives cites this paper.

Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T21:44:55.875468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:44:55.875468Z digest=sha256:1d5b447fe6bb555089ff756d1a2dbd761b0b949b212c137e303bd0b1bcf50cbd

Observation 12ae62c6-8899-485a-86ff-e0ce6deaca68 · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.212186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.212186Z digest=sha256:48bcf886accdebf0841e9c0400f46943ed0aadf70470255836bbf3db883563da

Observation cd69af3a-c658-4de7-b693-ad706aefd14b · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.710895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:050800808976d923aa60b555169bdc513c77fda05e29ddbb19594852d25d7c31

Observation f80a9b77-b39d-4501-92d5-8c31e8374029 · inbound

ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding cites this paper.

ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:55.262895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:55.262895Z digest=sha256:07bc26f4d96d4ac4094c419a26d06b30673481cdfa0434e40a29bc388c42ae53

Observation 3194040f-a74e-4226-818d-3234125cc836 · inbound

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking cites this paper.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.978031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.978031Z digest=sha256:0e23a4f1c236b97bdce69ec939cfd1232d8427c4e732a8914f144f0411e31964

Observation b0fdd9d3-d2de-4c1f-b2c4-56f6c581b613 · inbound

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning cites this paper.

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:50.881392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:50.881392Z digest=sha256:121c1c05780265f9b732846fc50717e85d8a42ebf2ec32d06b8c5b25e4f147bb

Observation eb5ea1f2-c357-474d-8ffb-11d85425ac2d · inbound

AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions cites this paper.

AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:28.301957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:28.301957Z digest=sha256:1ea9d407e368b7f40a1f676f038b5e9a9e4063783904a477f9b328127a4b551b

Observation b64d36b5-fdc8-43d7-9c26-6476c5b90630 · inbound

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning cites this paper.

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:32.299228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:32.299228Z digest=sha256:b8430ceca88933d0a08d8d87545bb0324af8a69a048c27646e454bb4c2de10d7

Observation c23b8a61-b5cb-4a6c-8acc-8171becc0267 · inbound

CoT-Segmenter: Enhancing OOD Detection in Dense Road Scenes via Chain-of-Thought Reasoning cites this paper.

CoT-Segmenter: Enhancing OOD Detection in Dense Road Scenes via Chain-of-Thought Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:21.639221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:00:21.639221Z digest=sha256:10d566930a5b92b8f696ce4193c496f34c70f6404b7794573bb8a9e0d791a2d1

Observation cd23a595-bf5b-4d0f-8550-15b82f3f5b94 · inbound

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning cites this paper.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.081795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:c97fcf80aa4789d150979b218f4e917a5ea972a4f310523affbdb74750de4cc9

Observation 2e4a3485-1b56-4e99-860b-5402c90fa88f · inbound

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance cites this paper.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.779348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.779348Z digest=sha256:73148a76300991d24dc0a4aa9d9510f749c390b0f6da428c7bcb9122b5a23cb5

Observation c23e29bd-49d4-4ee2-a256-716b3e5727d9 · inbound

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding cites this paper.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.294475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.294475Z digest=sha256:65e710ccbf0d6ff95756a9940ddf4cb71c819fc9dc1508a70766b483b66ab6bf

Observation a0d61b9f-e1be-4f8d-9c4f-fbb619126a18 · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:30.540242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:30.540242Z digest=sha256:f189784129daf2bfdfc6e6a074dba8caa57fbd8419e9af01c3196c98e7a95745

Observation fc6636ce-071f-46a3-91a6-23c9f979f512 · inbound

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding cites this paper.

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T10:49:58.803374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:49:58.803374Z digest=sha256:9499f1720689b5531b8b1f08fb619a40acb3a8ac7bd362cbde7b7f927d5db296

Observation 684c86e5-0d91-499a-85b5-0a00f9524f72 · inbound

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework cites this paper.

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T20:19:15.127415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:19:15.127415Z digest=sha256:fda0d1e661687cf0630e1843b062666dbad4e95b7c5ab9ed7756889637a991d9

Observation f45d257b-0d11-45a7-813b-85ed3811a840 · inbound

Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking cites this paper.

Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T16:55:20.099628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:55:20.099628Z digest=sha256:f464d2f78a8977778f5e33ed7d2ba801aa20647c4f7cb37cf91677a07be47199

Observation 23e646b7-5992-4713-afa8-5cd784aa1b6c · inbound

Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs cites this paper.

Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:00.943697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T21:35:12.859669Z digest=sha256:adac62c3a356f13c42d10ae3620f61d5eb060c08f43766d3636049afdb39294e

Observation a961c15c-59e3-4194-9a80-6740ef13ffcc · inbound

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search cites this paper.

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.293323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-12T02:58:46.728868Z digest=sha256:e7a9b82659af0456c95fa8b6e6bc5ff3c7b5b12553a75f9349ef397f69465ed7

Observation 12dd8a4c-9c84-456d-94f0-cd1823df109c · inbound

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation cites this paper.

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:13.393013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T10:58:01.621488Z digest=sha256:2420b2af948785cf5bd12614022b456958693c75301e4334e969322969c7e429

Observation bbef1c4a-05ba-41c9-84bd-d1f8555e4ab0 · inbound

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation cites this paper.

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.704886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T18:28:21.605646Z digest=sha256:94c9b4873b0620123a519627e3279ffab12b8b1061d2a19ed8814f6a1fb22055

Observation 02bf47d2-71cf-4189-acf2-9dc74a589aea · inbound

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning cites this paper.

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:57.800982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T01:23:40.564561Z digest=sha256:7f7b34fd82377580f1a98337dff002e79f1dd7c94c7e792cd7e188863d060299

Observation 6fec5a22-b305-436b-9315-e96dc9f8b8d2 · inbound

An LMM for Precisely Grounding Elements in Documents cites this paper.

An LMM for Precisely Grounding Elements in Documents Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.732696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T01:51:50.043216Z digest=sha256:a1c87505f80e6175e4298ee3b5dc24959d19284a75a92d0477ad2dd1443a61c0

Observation b6d50c90-17fe-415b-8956-ed10cf4791f2 · inbound

ActiveScope: Actively Seeking and Correcting Perception for MLLMs cites this paper.

ActiveScope: Actively Seeking and Correcting Perception for MLLMs Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:59:57.255666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T01:06:40.481169Z digest=sha256:b048bf9e8311e5cd846069f71d0335fff0477a4fce769a1081c2ef5a697f1c6a

Observation b4d01059-1aac-4bd7-882c-e1b6a0ef5325 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.117287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:8ef8b779d19c645be5948d4027a485428c8faf4c90f580669ceb32ef7d7fe561

Observation 4cf2e4d6-aaca-4331-a20f-c6aba144b36b · inbound

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning cites this paper.

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:26:58.461700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-02T13:24:17.538850Z digest=sha256:9b9b92474539b75f93fbc45427f0189c425530f79ace68ade8c535d76b57af74

Observation 4a13e359-7a5d-4ebf-ba38-dea51b07c98f · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 210

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:c0f7378516765a2fc7c4246df210cf25fd18aa990f53a5cc81e2625e8eadaba6

Observation 125a6848-5382-4caa-adf9-477492677fed · inbound

Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs cites this paper.

Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T06:18:15.882601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:18:15.882601Z digest=sha256:8ac3edd09473a9f94e5bde60b7cd27118df5d44272c88d259c8ca05362fcbe32

Observation 43170df6-d614-4c33-94f2-bfce8ce52198 · inbound

Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models cites this paper.

Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T11:57:40.838682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:57:40.838682Z digest=sha256:aeef7750dd81891b5866d6117862f1bbf9e32baf91214fff651b3fd4977c6d70

Observation 9741293f-4a3e-4b3d-8c62-acc4f2a1ba01 · inbound

ToolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision cites this paper.

ToolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:27:03.174749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:27:03.174749Z digest=sha256:ee573c6684ec6f1f5fffd0c874105796c158b60c6784c05d9e64c286d49863fc

Observation fed3ace4-7f7a-4068-a766-df90f1678d4d · inbound

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning cites this paper.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.107674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.107674Z digest=sha256:5f1aa258e92700cdcbf3a56cc1ed6b23c91c9d965f2b20aa66782b592a07344f