Pith. sign in

Paper Citation Record · LEDGER

Improved Visual-Spatial Reasoning via R1-Zero-Like Training

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2504.00883.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.00883 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:18.424642Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:18:37.298889Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cfa03b87-9845-4c98-a800-14a7f564b40f · inbound

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding cites this paper.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.550790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:bf6b02cf4654f7ae5cd3ee865ef4f89ddb6967609f92bc4e69fabae6eb38d947

Observation f7b0dbec-41e6-4426-b8ee-c8dd450953f1 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:18.424642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:18.424642Z digest=sha256:d800902cffd26656cf04fe7825e499c5b429daf5c8b297ec0538d883144a0f19

Observation ff15f0c3-de02-41ae-b59b-5e0d8646e711 · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:40:56.029917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:62783f65d1ddd0ca89a8fa0e545cf4bf49851ac6e27fef25cafd7c6ed526b4d7

Observation 292307aa-50aa-4b8e-ad88-d44a0372db00 · inbound

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames cites this paper.

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:17.956196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:17.956196Z digest=sha256:550cdc0c3cf1b165a2f1114a336e78e90d4da01a8de91409bf488394be6b7db8

Observation d04d06b1-776c-47a2-be9f-0e489c873a9d · inbound

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning cites this paper.

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:14.722576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:14.722576Z digest=sha256:e01dbdee03f9fc5b8a16cfddc46c8c0f464f5af649d5c35f4944cf136f26b124

Observation 900592f3-cb75-485e-8f6d-57cc9329113b · inbound

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning cites this paper.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.114508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.114508Z digest=sha256:8ea592ef6841910e6e07574606679e1a80005e3eb2911d5a0bbf8b6383caf848

Observation 1d5a5c88-f935-41fe-ae91-41cd19f4e1c3 · inbound

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning cites this paper.

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:52:08.116778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T06:50:02.607136Z digest=sha256:335ec4e2e8a800d24fe636435af4028b23c3a221e71c6c73797b7a32bd9fbba1

Observation d1911141-ba6d-45a0-9592-05121f0c56cf · inbound

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning cites this paper.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.103936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.103936Z digest=sha256:58d394d6700ff109bae8bbcdeca5438d9e60958fee2aaa57f27463287c02d8cc

Observation c3a40141-58cc-4308-b02d-db22c41b6f97 · inbound

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture cites this paper.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:22.714354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:22.714354Z digest=sha256:1348f556a781990262287add5c50f801e7ae0ba6bca17141d4a2bcace6c6131d

Observation 2a7ed2cc-0db1-417c-8b68-aae1f1a6ccf7 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.995648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.995648Z digest=sha256:9023e5af7546934abca481136d9ab734c1c63131ad34b3f193133444b59967da

Observation 85ebea19-bc89-4b8f-a73c-c97ba66a00b4 · inbound

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding cites this paper.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:05.079698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:05.079698Z digest=sha256:9e254926f238ed2280da0c4d53b91faf55f789a8eec8234540b0528a3850759c

Observation b4d6e639-687f-48bd-b459-b9d41f707997 · inbound

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs cites this paper.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.926420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.926420Z digest=sha256:78007d1e521acc6ae55e56961ca5ee3625bc4940b2506724e20276ce89e34759

Observation 87bc1d6f-4537-455f-bca5-9b40c5226ec0 · inbound

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models cites this paper.

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T03:16:49.999803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:16:49.999803Z digest=sha256:a98c19481951b7b5249c4ece49a39c991f169fc1675019802592b5f6fb80fa01

Observation 873458e3-a56d-4bd4-a4d4-ed0ab4552919 · inbound

Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models cites this paper.

Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:00:04.278492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T11:56:29.743281Z digest=sha256:cde61e3477caa2f26b72b992d9d76d732823296234c8454333589311de07056b

Observation 28551100-a395-4bf3-8a26-bb319e01047b · inbound

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding cites this paper.

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T17:54:05.713076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:54:05.713076Z digest=sha256:fb7686e6e301b0e509df70351a4ffe58ecb1782ffea66853bc0c74d13cf24bbd

Observation a1dff7ec-a172-4dc1-b672-2ab691354125 · inbound

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs cites this paper.

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:43:22.842223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T22:41:09.840792Z digest=sha256:27e6a31ded0d7d556cf79fec545f627f07a45eb7310cbfdaa30cbe38c8fe3e80

Observation d8ac2c06-1c86-44a5-9652-b424b447625c · inbound

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs cites this paper.

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T14:39:55.177552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T14:39:55.177552Z digest=sha256:43e2ae4b19a8bea2d5ca3fec79b97b152c46413cc8a464caa71c5b569a62bbb2

Observation 3548575d-a02f-41b1-b7c1-3d3883a46ef6 · inbound

Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning cites this paper.

Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:15:47.770329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:16:46.753641Z digest=sha256:b2630a3baeee8657363141826d6361d8a8f9c3cdb215f27b0b88d2f059bbb6d0

Observation 68324a95-075f-4510-afde-f682474e04b7 · inbound

Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models cites this paper.

Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:46:01.362635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:20:47.418874Z digest=sha256:4378cec828e459c9d9557fee75516d7e51330d1680357b71514f4e1e64ded329

Observation 75226c71-ce97-427c-87db-a0944021055f · inbound

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models cites this paper.

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.748348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T01:23:32.849326Z digest=sha256:625805c944d2d8b3d16c0fd75c6efd9563b834ce0cb5518af78c53134a527a19

Observation 1ec4db68-218b-4ec9-968b-c9b9a5e89eb6 · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:28:04.514201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:8dc45e51fa2f333d056cb8141924445763cfceb784d930108f64db42f1c78dc2

Observation ff9b5c72-0230-4538-ab8f-ae71625a189f · inbound

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration? cites this paper.

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration? Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:12:24.798806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T17:08:59.111306Z digest=sha256:606476c3a5a8c6b074e291b605ea98bc8a6ca9c2410ae2af57ed9dd1206176c0

Observation b5b5edfe-f574-464e-9f75-e5a3662b22e0 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 219

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.583697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:ed9c20e2aa2a699539e55d903e2a5be82dd910336e59e1ac765fa13802c15f67

Observation 83a601e4-40bf-4523-8d9a-22308548fa91 · inbound

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning cites this paper.

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:47:50.499205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T10:41:41.697216Z digest=sha256:bdf2f4a59ed897f5e4adc4902f1146a8f3d0d3ae27b197d4bdc03272c82074c8

Observation 11d33447-b8ce-4e57-bc5e-7f3061ada16f · inbound

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video cites this paper.

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:18:37.300491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T16:16:41.412451Z digest=sha256:686db84b31def2c1fd0fc4cd1bda1a9d0d0c5b5e911bfcbb5ad2f1320bb08b1a

Observation 8c4a1fa3-70d6-499e-b42b-d96f56b54eb2 · inbound

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding cites this paper.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.979540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.979540Z digest=sha256:edf1988747e8f3e4b431356b576cc4d82f3de0eb703d132e42d1ff2e7af6a72a