Pith. sign in

Paper Citation Record · LEDGER

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

As of 14 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 4 inbound Pith citation observations for arXiv:2507.13362.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13362 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:54:48.162766Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-25T05:51:18.233043Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:55:25.118073Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 026491a9-7bf9-4eb7-93f4-0d6ff127268f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.058714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.058714Z digest=sha256:a92d0858ddf0057c8c10d9475d3203cd095968942edf1ed0d86563cb80b08ef0

Observation 660692e1-7f6f-4f16-a9a7-57b0b0c008ba · outbound

This paper cites Visual Spatial Reasoning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Visual Spatial Reasoning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.062289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.062289Z digest=sha256:d3b4e00bb907ea561680e801ce261142827fef141db8b417404b858b5c81a7a8

Observation 7237e9ee-d699-4820-a06a-8563f308b19f · outbound

This paper cites Sat: Dynamic spatial aptitude training for multimodal language models,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Sat: Dynamic spatial aptitude training for multimodal language models,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.065213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.065213Z digest=sha256:43ae6d7ac4bcfc44f62cddc322ac6503347fb371ae590fe5b14a7e896d6d6687

Observation 2cede708-7895-4775-8ce0-f9e6f9ce8811 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.068292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.068292Z digest=sha256:5c047127e197d12f3bf2b6ea114e4d4facf49449f684b867f577a1f5841f8554

Observation 8b8cc7d2-446f-4491-bc6b-7dfa4da6239e · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.613029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:54:48.071617Z digest=sha256:6c83b9bc10dcd3d636b4a9afce89a9cae280ff2a437cb36b3ddc02cf965b2922

Observation 344944d5-44eb-4806-a839-9e0064ee6aed · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.074302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.074302Z digest=sha256:4f660a373d6f504691dc8eff34ffc775c1195bb3851ba9f9b4dba7e513da5bc5

Observation 31ce173d-4c71-403f-8b5b-69c00dc94c58 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.077400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.077400Z digest=sha256:cfda60e4f917a7f55a8e2702db90ac24a4e58879ade4af3f5bc7fa5ecf6c3bdd

Observation f0f284c0-8e9b-41f3-93a9-c12961eda8de · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than $3,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning R1-v: Reinforcing super generalization ability in vision-language models with less than $3,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.604078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:54:48.079868Z digest=sha256:22c6bc221806d603ac26ad9aaeb19223806a9733060568a15226554bf760b277

Observation b5752c06-a821-4f2b-832b-c9732a089370 · outbound

This paper cites Rlvr in vision language models: Findings, questions and directions,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Rlvr in vision language models: Findings, questions and directions,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.595750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:54:48.082191Z digest=sha256:354ba109dcdabf4125f9640576fe6d74513fa0feaf22a0b23ecb03b0f21ed362

Observation 93e83563-3570-4ed1-b0e7-64ee2a7f3db7 · outbound

This paper cites R1-zero’s.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning R1-zero’s

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.587733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:54:48.084730Z digest=sha256:53164a2c7260af0f13b8553a4a7c517ec1825adc64639b775d8770c56690cd88

Observation c9714968-8465-4d2e-960e-7030c0c7d297 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.090566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.090566Z digest=sha256:dd8603abd63665f37f31a873db226b8b7f6b55924ebbedabe1e5676b27e02ead

Observation 1aeb14d0-f97e-4d0c-b62a-c21146162283 · outbound

This paper cites Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.093206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.093206Z digest=sha256:cd9f87005cfc4f2d49081a85a6ee4e720fa1bff5b2dc7733ca2ad879797056e9

Observation 459ba278-8570-4a00-a6cf-778ca1f6a78e · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.095753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.095753Z digest=sha256:683ce1fbf76b1691bfa95b0e429ef724ac9c7470ad7c0e2517ce62d090c75b82

Observation 403e28e2-fd40-4f71-9f1f-50b839a7a5a1 · outbound

This paper cites Spatial-r1: Enhancing mllms in video spatial reasoning,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Spatial-r1: Enhancing mllms in video spatial reasoning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.579259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:54:48.098417Z digest=sha256:73e97a74da8050b274958d3f3fe8771038619514c8d4a281bb4cb87941da16b2

Observation d1911141-ba6d-45a0-9592-05121f0c56cf · outbound

This paper cites Improved Visual-Spatial Reasoning via R1-Zero-Like Training.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.103936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.103936Z digest=sha256:e16d1ccde4238128a0eaccc3461d6081f5d3593d3d8ee5718a91c71451081728

Observation 08e55a4f-c9fb-4548-899a-300f02ff56b2 · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.101125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.101125Z digest=sha256:35a7f807cc674570f8516e53a1a481e65dda3dcb0e18c5ac512f33405d7744b6

Observation ef361559-baec-48aa-b250-31ffba03bdd6 · outbound

This paper cites Compositional Chain-of-Thought Prompting for Large Multimodal Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Compositional Chain-of-Thought Prompting for Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.110228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.110228Z digest=sha256:675df6d4d630b973f9cbd9f640bacacd16ccc34004b90a7f4c23f6598428c0e9

Observation 8bcc650d-7dcb-4299-bb70-9ff14e903ad3 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.107118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.107118Z digest=sha256:93c8c0df740cae5b44354b0dd86d8164d67ad654ceebc601a8dc32ce87de91d2

Observation 5666202e-5beb-4160-8ff1-4c59f3288bbb · outbound

This paper cites Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.116134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.116134Z digest=sha256:ff247d5b5e44f0c55b4cbbc66e039aa5fd96805e2ead09f897d4db70dc88445e

Observation 474139ab-99b5-46cc-801e-99b140d37b0b · outbound

This paper cites SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.113277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.113277Z digest=sha256:40cb370f0b2f26a5ba7da2e516feaa3f572d0246ae4dc42618a5dd279241281b

Observation b4a23978-495e-42ab-a85a-5013e648f521 · outbound

This paper cites Clevr cogent valb,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Clevr cogent valb,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.571061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:54:48.122022Z digest=sha256:42a9b8c137b58b308457e74dd286d22f32e0249b2142ec533ceaeb87e51c866d

Observation a241cba3-ce08-431e-90ed-5525cdbb95cc · outbound

This paper cites Chain-of-Symbol Prompting Elicits Planning in Large Langauge Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Chain-of-Symbol Prompting Elicits Planning in Large Langauge Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.118975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.118975Z digest=sha256:bb64d2f8f551c09739d835d3a6f695b4a79f1a5614666941ef488b119a608493

Observation 7f53ab88-b6a3-4746-9d32-22d246d838f3 · outbound

This paper cites Depth Anything V2.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Depth Anything V2

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.127358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.127358Z digest=sha256:a30ed2960b160b594b34e83892ca83d7d2bf69fd1b24eef43d80cf33a69cf422

Observation e63589c4-b6e1-4889-b392-73cd413e2956 · outbound

This paper cites Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.124661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.124661Z digest=sha256:a5ae4cb65bcdd5f6f85f3188f6d8cceec876824f887830e68db43814160bcbcf

Observation 6fe160ae-13f8-4435-baf3-d27e065c7463 · outbound

This paper cites ClipSAM: CLIP and SAM Collaboration for Zero-Shot Anomaly Segmentation.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning ClipSAM: CLIP and SAM Collaboration for Zero-Shot Anomaly Segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.132799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.132799Z digest=sha256:b5a821a626bfb7e0458e4dfc1978c18a88423663a38ec0ae3276b3ae74442f0c

Observation 426118d5-58bd-4f7f-a067-ed12597e2dc4 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.130017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.130017Z digest=sha256:2523e0d4e079ebb1cb4ccd9ae6d259d25306b98e7cec2d09bc2b4fcace9b1bd2

Observation a27f093c-9609-49fb-8719-68d3b7ce78f9 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.139134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.139134Z digest=sha256:e10f310e6afeb81ca3098d78b7913c6b36633a77c0d1249ffba74c8cf00de5e2

Observation 5b1b369b-9af6-4c77-bc28-85882706c23c · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.136009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.136009Z digest=sha256:d26042091bee6f6ed48966e590b182122b6f8c9c09a1dd73db26aa6ae3d7ac76

Observation 50ae8837-e520-4134-81e2-36491a27c56e · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.144642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.144642Z digest=sha256:247fab15dc4e4583f263e665c1ccbdb85fd49cdbd104bd9775f0b04d0aa5e5fd

Observation fb0e9815-b80d-47f6-b11a-6cf7316cedc1 · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.142033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.142033Z digest=sha256:5f65e764283cfeadba3407e07494f26940d34ab6b2255f66d579070c606501b3

Observation 248d13b4-d54a-449a-91f8-49431fa02145 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Training Large Language Models to Reason in a Continuous Latent Space

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.150801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.150801Z digest=sha256:019d4390ac3160b66634fbf67923b87927273b62fc2bdb4866acb8f262afa380

Observation 25ecb92a-5ce5-4dd3-9db3-cc0770f60934 · outbound

This paper cites EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.147566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.147566Z digest=sha256:aacb32704e66d5675af9f533999d7236c97c60b2c5dc6c1cf1a93482b4867a76

Observation 46a2075a-2a04-4750-af4a-544bd439c12c · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.156627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.156627Z digest=sha256:146c474ba4312d3d3a9550430ebe99af580c28b124261c8aa342cafcbd594ba9

Observation 9faaa57f-7cfc-4f53-a10a-cf3e13efeb9a · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.153646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.153646Z digest=sha256:422341c3d81d43ac9a4f75fa4ff1e6d360411e257c90eb6fe720c5797ccddf5d

Observation e7f85b5b-2a86-40c6-b069-8646390f3e0c · outbound

This paper cites SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.162766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.162766Z digest=sha256:c1f96ab174143a9438a46cf954168038dc64d2db1aecd182bff458545234cdeb

Observation bd595542-d108-48c3-8ae1-71279158c07e · outbound

This paper cites SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:54:48.204284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T19:54:48.159895Z digest=sha256:bde3b92b0b4a1ed69a72441ced70b8bb23832882b79d25b571606df3e1f8e3b4

Observation d72bffca-7bf1-4666-bb7f-eb04a5452b88 · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.087387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.087387Z digest=sha256:fd0019d2d7c201fdac57800462388f665e2886513a7366f5fa3b096c7f09b260

Pith citing papers

Observation 205bf49e-2912-4f14-97fd-2c0ad26eddea · inbound

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models cites this paper.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:00:34.244407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:caec191245932a44934e2904ba499037843e978cae6c1c09bcd4d844e5db26fa

Observation 8864a7cd-4258-41a7-8c93-e9db60e40cd3 · inbound

SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning cites this paper.

SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:08:13.509580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T11:06:29.022994Z digest=sha256:81a0026570e6066e3cd57c6fc18680c548b5a5a425e59854fb4c38d88957197a

Observation fafb401a-850a-41f4-a92c-12801b397889 · inbound

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models cites this paper.

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:58:04.874738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T05:57:45.964145Z digest=sha256:1263e1659fe370d37f5edd1c744afedeb784de62eac156133012d123204ab2cf

Observation 12a1fb11-346b-4f01-b440-4f0280f224b8 · inbound

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models cites this paper.

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:55:25.121803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T05:51:18.233043Z digest=sha256:f0f8fc72acb7d97481c33895510c91ea3a4b864bfb26b01027f1584d6e803e7e