Pith. sign in

Paper Citation Record · LEDGER

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning

As of 18 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2505.16579.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16579 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:27.579774Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-09T21:32:08.503584Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T14:36:05.322735Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d392d12d-df48-447d-b2c8-55afe1872500 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:24.783223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:24.783223Z digest=sha256:49b2e2cae653eaffc8f7ab94f805f2cecefaf249b4d481b3cd6b8dd4dbfcbee8

Observation d369d101-b043-43a5-ac8e-35b385e3d5bc · outbound

This paper cites Qwen2.5-VL Technical Report.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:24.843868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:24.843868Z digest=sha256:b06a732fa4e5b2affe1aa1cabe2d3acc978514341a018a1fc9f32fbf301c468f

Observation a768fc52-f4a4-45e4-8ead-41ade8a9bbb9 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:24.955656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:24.955656Z digest=sha256:b6df4f74d296a4ca01a2e1d6f41264ee2207dcf1cefa5c99d5fc97c0d64a0f65

Observation 724ccaf1-5152-4020-b225-fbaa6e536ca7 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:25.132191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:25.132191Z digest=sha256:3dbe8c6673a64ce204e2dcf595b06e95df33e1d923be2e294c655c6f3b5dc890

Observation 4e7f2367-3101-4041-9abb-96fe64d90401 · outbound

This paper cites Interleaved-Modal Chain-of-Thought.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning Interleaved-Modal Chain-of-Thought

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:25.200640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:25.200640Z digest=sha256:0e033d2b17b1eec9e159b92094dd6c2f048595cb15d9b770ac14d984d2f40a0a

Observation 9f366451-7cae-4ceb-860e-2eb4ee718d4e · outbound

This paper cites The Abduction of Sherlock Holmes: A Dataset for Visual Abductive Reasoning.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning The Abduction of Sherlock Holmes: A Dataset for Visual Abductive Reasoning

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T15:02:28.069695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T15:02:25.246792Z digest=sha256:9ff3fab4a12a183fbfca20693ac4853dc205d7cc5603dfc17b9a69b351f68cf3

Observation e7c8bc35-93ae-49e7-8dd3-60f36a5cfcb2 · outbound

This paper cites MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:25.409128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:25.409128Z digest=sha256:9a6e947d94bb6b7bfa53f406850cca99c1334b49d6d4c8936709c3988564061e

Observation 904ec99a-5f78-4d84-8995-bd187d72ae45 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning LLaVA-OneVision: Easy Visual Task Transfer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:25.544884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:25.544884Z digest=sha256:305db63ad19a1480e54b1cd58a89fab5bdd9c0b750d855679bdd86bf7f0d92fd

Observation 2d134f9e-5c17-4b86-9d2c-5ff156aa1530 · outbound

This paper cites Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:25.601130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:25.601130Z digest=sha256:1a9e474dc519c928dbea564dee9a52df479b84378153f6250080c6c81a297ea7

Observation 760a5f88-7882-4a49-ba54-e8a8f01cc2d8 · outbound

This paper cites an unresolved cited work.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:25.691311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:25.691311Z digest=sha256:9a874c22dd41735dd0125ee606f5260157339e421bc80edde2897cdbb73b0c7d

Observation c39c6697-9f26-4516-900b-2db79dd05e48 · outbound

This paper cites Med-PMC: Medical Personalized Multi-modal Consultation with a Proactive Ask-First-Observe-Next Paradigm.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning Med-PMC: Medical Personalized Multi-modal Consultation with a Proactive Ask-First-Observe-Next Paradigm

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:25.738836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:25.738836Z digest=sha256:7b4f46104bf5d6705fd939da6dbb84b847d374abcacf064c8987a1890b24384b

Observation 65a724fb-6abd-4279-8576-c22f418abf0d · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning NVILA: Efficient Frontier Visual Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:25.841387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:25.841387Z digest=sha256:74199dd5c9d5de91631cd340cbce2715d25338e5f5e6eeb3102d86013843f3de

Observation 20fc7cfd-c3b1-4448-9b4a-6f89d7d9f9c1 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:25.944057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:25.944057Z digest=sha256:90821c659909d6d234aa634b2e8c26dde1e655bf68febbdea861b4a187366d37

Observation 122c7166-4311-4d08-a4a3-f36094e61cd0 · outbound

This paper cites Compositional Chain-of-Thought Prompting for Large Multimodal Models.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning Compositional Chain-of-Thought Prompting for Large Multimodal Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:26.083572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:26.083572Z digest=sha256:ad73c13ae0dc6fb372b6b746e52fa04bba7acc30d55cc25c8fe1e4344d0bd5a0

Observation 2f609678-7fec-42bd-9efb-9440d1536144 · outbound

This paper cites an unresolved cited work.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:02:28.503460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T15:02:26.167566Z digest=sha256:cf3e1c262df36a3af44325b96c0c17c3ca1c354ed930fd70c3802f78c2cc1e61

Observation 5c8956ee-8a8d-4793-8f62-05ccc2f48825 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:26.254694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:26.254694Z digest=sha256:c0fbf2c83bcb2438439ff15bd0108e86680dfedb5b9c6f4291e3e7d5f80d42c4

Observation 3a676439-2414-4e49-95cd-ea8f163b8edd · outbound

This paper cites an unresolved cited work.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:02:28.305760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T15:02:26.330414Z digest=sha256:226a0e8ffffb1bddf0540bca46b236dd16618c6d1b5a3885e129c8ba59f7243b

Observation b695ae8d-aa12-48c9-a326-f4abd40500b4 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:26.398066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:26.398066Z digest=sha256:ca2f9b411c7ffd15b8da24b8148919190d7b039141c69eced118de3780b180bc

Observation ce0a4132-05fa-4f6b-81e4-15ea8a5f2294 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning Qwen2.5-Omni Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:26.475930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:26.475930Z digest=sha256:1112487467ccda6b989106f2644f45f08c9f599bcbb1c1fd4473781dacd048e3

Observation 259f02f7-c6b6-437b-b6b3-bbfaca65695f · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:26.564151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:26.564151Z digest=sha256:d5a59341e253759df13abf48cce3bade92b0fb1639a67ee17aa85ca7319882e2

Observation 18c2e927-af90-4a3f-8cb7-9d6565d0d76e · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:26.636292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:26.636292Z digest=sha256:ac6066b1b975d8565fd0e1d76ba9b74b692922da618c5debfa1b3ebbc27794c3

Observation 35bdbbd9-123a-4424-8122-d3144942bcae · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:26.697184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:26.697184Z digest=sha256:8c2c87398fc6416f940c6acd4d1a3c55c35b8e807345cdcbb457d8965b07ff62

Observation 565bbe37-6430-40e2-af36-f08aad36afa9 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:26.814318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:26.814318Z digest=sha256:f3c4a0359505970a25880b2f15f3114378b2ec53f7f14d2d5f0cb3aa495c5e5e

Observation 78815e43-b6a4-4974-be74-be751308dd65 · outbound

This paper cites an unresolved cited work.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:26.931367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:26.931367Z digest=sha256:f10bd9fd65ef3ed83d5a109e5ed0ba951aba5dd9eb20b83c288d4067df246f71

Observation 215337b5-4fd6-4475-95f7-3aba6a64c905 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:27.056855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:27.056855Z digest=sha256:a23713c7a81cc39b3ea8f5a90e21b64038426c6c94f15dc7442c417a5ac3259e

Observation 4887c9c4-572d-4dee-8cd3-c88ea7972825 · outbound

This paper cites DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:27.176868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:27.176868Z digest=sha256:30eaa9a8286df4d58a3ab4dd51e75d24e724a47efc8d23e04c327242cf91e7cb

Observation 75822fd1-1674-4244-a18f-b958e6bb4628 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:27.349532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:27.349532Z digest=sha256:22baa4df47eae58bbf865c9caff38770d26e5b4ed99d9e6d75c05b89d3fce9c7

Observation f22aed27-8280-444e-becf-9813a3e5a6f0 · outbound

This paper cites online" 'onlinestring :=.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning online" 'onlinestring :=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:27.475806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:27.475806Z digest=sha256:c6f2bad1776be3eeb29ec19b5281fa577c29b4893f624a0c88d363338c5b5319

Observation 548b94c8-3f8f-4a2e-8a38-ceb2116087ed · outbound

This paper cites write newline.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning write newline

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:27.579774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:27.579774Z digest=sha256:371f81315107ae799ba5cd19c397c28731975932c03abc85363377e08276a8bd

Pith citing papers

Observation a3be3925-d3a2-4896-8e84-2d466f961232 · inbound

SketchVLM: Vision language models can annotate images to explain thoughts and guide users cites this paper.

SketchVLM: Vision language models can annotate images to explain thoughts and guide users Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:36:05.326174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T21:32:08.503584Z digest=sha256:f4b5c9943b73cb7b1808a2cda99f451c01daf4cad227e6d47ee7a556eaca678b