Pith. sign in

Paper Citation Record · LEDGER

Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2410.03321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.03321 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:10:10.150431Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:19:57.722772Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5a63e436-a5bf-45da-9eda-54aaa23bbb3f · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.742553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:aafb706e169dd8b14a8dd634971df4bad2814540f7f41b1dea8998e6298775d8

Observation 0562d9e9-61e4-4f1f-9d53-cc899a91a916 · inbound

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought cites this paper.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:10.150431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:10.150431Z digest=sha256:986711947f20c4e67eb23bb2977cba4979102530db6f68459a475da4c13d4794

Observation 319f97d0-691f-45aa-bcb7-0271ddc1840b · inbound

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought cites this paper.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:24.327499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:24.327499Z digest=sha256:1dcb4b0a9c807c9c1d75077bac23ce158d3bbe6c60d22764379b1b6b546a10fe

Observation c88c960c-7624-418f-870a-1bde9f3e909d · inbound

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning cites this paper.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.075376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.075376Z digest=sha256:2f1062c0f148f4a84bf196374df7b74907574e3e225e04ac7b1bb6f483bb06c4

Observation 70c288f9-a8eb-4fdc-b88e-b5b704dabdd5 · inbound

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? cites this paper.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.318843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.318843Z digest=sha256:96e89d23bf44e4c55a4f0879b3224dd6a4ad3cf721dcda71aa2a48ff69b760f1

Observation f4b37619-8d58-40f3-b06f-d0b65eff3cfa · inbound

Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework cites this paper.

Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:32:36.364468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T12:31:25.257879Z digest=sha256:ccd413b746c5a584ac27c2ff3ae8258c323dec67f132aa0bc0a5386cd20ca93c

Observation 09dd9144-ac73-4db2-a9f2-b0be7b0f6c04 · inbound

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought cites this paper.

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:19:57.724720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T00:46:17.094339Z digest=sha256:05971617a90ed0dc60c5b233f6bf36cb41c90676e6fb8b090251e0f97b37f00f