Pith. sign in

Paper Citation Record · LEDGER

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 4 inbound Pith citation observations for arXiv:2507.13362.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13362 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:54:48.162766Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-25T05:51:18.233043Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:55:25.118073Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 026491a9-7bf9-4eb7-93f4-0d6ff127268f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.058714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.058714Z digest=sha256:7fd392d9cfa419790e9eef0ce3070a24ee6f81cde6323ac727ac06ee71bc0d44

Observation 660692e1-7f6f-4f16-a9a7-57b0b0c008ba · outbound

This paper cites Visual Spatial Reasoning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Visual Spatial Reasoning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.062289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.062289Z digest=sha256:eab0bcbf4e8b9363638950221b8a87bbb31f4878b6b454f4a65f0d801d340085

Observation 7237e9ee-d699-4820-a06a-8563f308b19f · outbound

This paper cites Sat: Dynamic spatial aptitude training for multimodal language models,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Sat: Dynamic spatial aptitude training for multimodal language models,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.065213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.065213Z digest=sha256:a538649506c6998b8c62a5433a46c3a1e08f25f83374e49c191e0b1ebd1634c3

Observation 2cede708-7895-4775-8ce0-f9e6f9ce8811 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.068292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.068292Z digest=sha256:30343d5f7a3ef025e42134b5cb4a8ab8b3f553abdc774c771ec5a634e4c844ad

Observation 8b8cc7d2-446f-4491-bc6b-7dfa4da6239e · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.613029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:48.071617Z digest=sha256:1a9f9f6d36c5df3fa4367c2db1866316411008e00b47d71f2b6e3db4e78c75f1

Observation 344944d5-44eb-4806-a839-9e0064ee6aed · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.074302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.074302Z digest=sha256:e3d3577cad4b63d7d49bded00b16e64ee2834759f5007ef3c61f05aadd4d0b3f

Observation 31ce173d-4c71-403f-8b5b-69c00dc94c58 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.077400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.077400Z digest=sha256:68a6a905078c33305be99be044966b3a511e8f554c1f7080f2fb1aaa93935e39

Observation f0f284c0-8e9b-41f3-93a9-c12961eda8de · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than $3,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning R1-v: Reinforcing super generalization ability in vision-language models with less than $3,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.604078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:48.079868Z digest=sha256:63af2843a2b1b9b500dea47e2af1086135428b2e858a2bb58d42bfaa1798b2c5

Observation b5752c06-a821-4f2b-832b-c9732a089370 · outbound

This paper cites Rlvr in vision language models: Findings, questions and directions,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Rlvr in vision language models: Findings, questions and directions,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.595750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:48.082191Z digest=sha256:c85f9addd12149618b32cfeaf06f971b063de172b415549b2e9e774152d43fec

Observation 93e83563-3570-4ed1-b0e7-64ee2a7f3db7 · outbound

This paper cites R1-zero’s.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning R1-zero’s

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.587733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:48.084730Z digest=sha256:f7e65e79e8e7de2fc834e5acb5ba09122ab6bf01b577366320693c92afd123e6

Observation c9714968-8465-4d2e-960e-7030c0c7d297 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.090566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.090566Z digest=sha256:eb60dea7bba6c8a04aa4ff66b6dc95a424eee50862df874323b011b5a5444662

Observation 1aeb14d0-f97e-4d0c-b62a-c21146162283 · outbound

This paper cites Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.093206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.093206Z digest=sha256:060cf4a585607fc3a4eac36b99a1880b69d5a91938ec508ab7c9fc269dd8bee6

Observation 459ba278-8570-4a00-a6cf-778ca1f6a78e · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.095753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.095753Z digest=sha256:93f00bdfd7b35ab3d89aa3d6d0b88559bab1bb0caeccb20b83a5a9a53295f925

Observation 403e28e2-fd40-4f71-9f1f-50b839a7a5a1 · outbound

This paper cites Spatial-r1: Enhancing mllms in video spatial reasoning,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Spatial-r1: Enhancing mllms in video spatial reasoning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.579259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:48.098417Z digest=sha256:a83107c1f3b0f6ce7ed5a28595547c1eb2124964bb5a5b6b4a73c02881e46fb1

Observation d1911141-ba6d-45a0-9592-05121f0c56cf · outbound

This paper cites Improved Visual-Spatial Reasoning via R1-Zero-Like Training.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.103936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.103936Z digest=sha256:dda8e71c91b12f337ef4235b13109c0bb8c26bc34a70db756219dbaecdaef131

Observation 08e55a4f-c9fb-4548-899a-300f02ff56b2 · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.101125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.101125Z digest=sha256:ae8ecfffb664656f88c1d7b1c713540a27f3e92eef146130455abb3d667d31d2

Observation ef361559-baec-48aa-b250-31ffba03bdd6 · outbound

This paper cites Compositional Chain-of-Thought Prompting for Large Multimodal Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Compositional Chain-of-Thought Prompting for Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.110228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.110228Z digest=sha256:0e882afcbd4df204f969ec96f8b804afb44bf2985963a79ae5a4b5866ce9d6fb

Observation 8bcc650d-7dcb-4299-bb70-9ff14e903ad3 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.107118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.107118Z digest=sha256:28b1616b73bf2f03b539c7f9e329bb803c3f015635f75ff211edc1d4938c690f

Observation 5666202e-5beb-4160-8ff1-4c59f3288bbb · outbound

This paper cites Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.116134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.116134Z digest=sha256:84a692202c5c9359acdcda1730772c2996b8cf8adbbf48b3975bbfe1e2dc5789

Observation 474139ab-99b5-46cc-801e-99b140d37b0b · outbound

This paper cites SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.113277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.113277Z digest=sha256:c78a74a54892f80d9700ace9c66967c38d6db4206a55c46dc584aabedb009858

Observation b4a23978-495e-42ab-a85a-5013e648f521 · outbound

This paper cites Clevr cogent valb,.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Clevr cogent valb,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:48.571061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:48.122022Z digest=sha256:67debcf6a016c5a21499146cc7f0859a70e7379394aa8d7a8a3eeadecf1acf1e

Observation a241cba3-ce08-431e-90ed-5525cdbb95cc · outbound

This paper cites Chain-of-Symbol Prompting Elicits Planning in Large Langauge Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Chain-of-Symbol Prompting Elicits Planning in Large Langauge Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.118975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.118975Z digest=sha256:b0d7063e957fae933871787674ab8ebf437cd61b7defb7dada0b8855c75dba20

Observation 7f53ab88-b6a3-4746-9d32-22d246d838f3 · outbound

This paper cites Depth Anything V2.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Depth Anything V2

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.127358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.127358Z digest=sha256:b4631163c2e14d2168a66f3d24adc2ec83363d36f4403ba4a5d27804ae04f26b

Observation e63589c4-b6e1-4889-b392-73cd413e2956 · outbound

This paper cites Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.124661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.124661Z digest=sha256:70d3d6787eab1e414ce6323050bc1c0a32b176a7c305279714e4bd5f2cc72181

Observation 6fe160ae-13f8-4435-baf3-d27e065c7463 · outbound

This paper cites ClipSAM: CLIP and SAM Collaboration for Zero-Shot Anomaly Segmentation.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning ClipSAM: CLIP and SAM Collaboration for Zero-Shot Anomaly Segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.132799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.132799Z digest=sha256:4cea6cc185b775540459b7f1befe6b9844b5749641b57441dc41eb83ea9dc3b6

Observation 426118d5-58bd-4f7f-a067-ed12597e2dc4 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.130017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.130017Z digest=sha256:368180facde59ce8a082f17d8b9fedba4877829b85b87b62106c79b33eef014a

Observation a27f093c-9609-49fb-8719-68d3b7ce78f9 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.139134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.139134Z digest=sha256:2f1a795b94b2c1fe945b82cda17bffee6a8527b35f2ed7e9d67148e3571eef79

Observation 5b1b369b-9af6-4c77-bc28-85882706c23c · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.136009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.136009Z digest=sha256:f7acb6e4076c3d858a5651ff2f42fcf52239f1455ffd21d00fcdfecfff56ea0a

Observation 50ae8837-e520-4134-81e2-36491a27c56e · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.144642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.144642Z digest=sha256:1dd7fc59e5d53e201adc0654fd54d926961543f613581e27d3066f3b4f74c4a3

Observation fb0e9815-b80d-47f6-b11a-6cf7316cedc1 · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.142033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.142033Z digest=sha256:d01019b9684a38c01c72745f3eb889b4e8ce9027f1ddae5f439e1dba76d09555

Observation 248d13b4-d54a-449a-91f8-49431fa02145 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Training Large Language Models to Reason in a Continuous Latent Space

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.150801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.150801Z digest=sha256:41cec78026e926a486138baac2c3ee07ec79a35a3c2b073153a327cfb75addc9

Observation 25ecb92a-5ce5-4dd3-9db3-cc0770f60934 · outbound

This paper cites EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.147566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.147566Z digest=sha256:3763ca2d4bd85727f2434616c137d143b7f1d0efe73a267772953f41ca13bb57

Observation 46a2075a-2a04-4750-af4a-544bd439c12c · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.156627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.156627Z digest=sha256:f1cd6f8bb3ef006750b4ee3fc82baccbc73605e993e4397f90e1691517443b1b

Observation 9faaa57f-7cfc-4f53-a10a-cf3e13efeb9a · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.153646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.153646Z digest=sha256:cc03763c8a58335d8d716b811d48e457ecd718a9a38b07edd139e54d1c089fe5

Observation e7f85b5b-2a86-40c6-b069-8646390f3e0c · outbound

This paper cites SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.162766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.162766Z digest=sha256:1fa8e190c08c0bb8b911d85389fb367da12e1d3117dbeec463c56aff9a6569e9

Observation bd595542-d108-48c3-8ae1-71279158c07e · outbound

This paper cites SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:54:48.204284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:48.159895Z digest=sha256:fca19b22d0e0c3099833b2dc07e1b581c718cc9560ef6ef20d734ae82fc54eec

Observation d72bffca-7bf1-4666-bb7f-eb04a5452b88 · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.087387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.087387Z digest=sha256:b25c2b009208e8f00f564ce1122d571fd3dcd6ad9582b0c237a16e5f4aa35438

Pith citing papers

Observation 205bf49e-2912-4f14-97fd-2c0ad26eddea · inbound

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models cites this paper.

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:00:34.244407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T00:59:29.584299Z digest=sha256:d123321e41662b6575a61ac83fec24abafe3b80a73df1a6496a618b8bd9c7652

Observation 8864a7cd-4258-41a7-8c93-e9db60e40cd3 · inbound

SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning cites this paper.

SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:08:13.509580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T11:06:29.022994Z digest=sha256:dee9997babff2f707ebb7c142a78c235a81f51a6ee8782afa9b7a358c07312bb

Observation fafb401a-850a-41f4-a92c-12801b397889 · inbound

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models cites this paper.

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:58:04.874738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:57:45.964145Z digest=sha256:57c15b43e486f1fdc234825d43bf7452c5d229c24d67c2430ccc554531436ee1

Observation 12a1fb11-346b-4f01-b440-4f0280f224b8 · inbound

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models cites this paper.

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:55:25.121803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T05:51:18.233043Z digest=sha256:708c3f2c3babdfad2f861991585186b8a9e32289dacb3bdc884d6acbae0e745f