Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2406.01584.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:24:04.777438Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation db059a5a-634b-4ba9-aff2-7ec41000d777 · inbound
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6ee4ee5c-9a39-4345-90d7-00518300fbf6 · inbound
Can Multimodal Large Language Models Understand Spatial Relations? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c1db7ac-7239-46d4-a0be-41b8f3b2a0d3 · inbound
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaf3a9d2-5463-4376-a6bc-5326c43abcfe · inbound
GenSpace: Benchmarking Spatially-Aware Image Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61d2f55d-3e29-4091-92b3-383a5c5938ae · inbound
BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e85307f9-bf8d-42ae-aaf8-e44fdc65d69a · inbound
AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f89895c-f4bf-432a-8c13-6e42b8eabd64 · inbound
Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e0ac93-8634-4490-81b6-3162bc71aa63 · inbound
RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cb6f7bd-1dfe-4715-881b-fbe525c78b9a · inbound
Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca1bd479-40fe-484f-a892-1beea35e5999 · inbound
UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b71a8db-4c94-48c2-b200-c1b04c17a942 · inbound
Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eff6343f-b388-4618-810d-b99b1296daa9 · inbound
AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa4cd893-691c-4236-96ff-78c5eb4d968e · inbound
3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a214032c-b48a-448e-b668-586e250128f4 · inbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cabf21b-16b5-4f30-849a-44cecbdfbe34 · inbound
Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e5a4d9-4f94-4b70-b6f5-839031e9bc6e · inbound
Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fbbdb503-c116-4f08-9fb4-c62944dd3a0b · inbound
Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9cb7737-9013-4111-b31c-9d0ec45dee76 · inbound
DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 185a0f71-71f6-47b5-b1c0-81a63a0152c0 · inbound
DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c11b4357-5a09-4755-be09-f9b025ac2240 · inbound
Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58ef9cc4-cba5-4ace-bfdc-b6888ef921e0 · inbound
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a763a5cf-2318-4778-b72d-9a76454f1ad9 · inbound
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aac82ce6-25c0-4e8d-b671-7897e406c611 · inbound
TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9a45f726-1e72-4ffd-bf4c-fc25f0c75fab · inbound
TableVision: A Large-Scale Benchmark for Spatially Grounded Reasoning over Complex Hierarchical Tables SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f57c5c60-8de1-43ac-a63d-0d3b49227210 · inbound
Spatio-Temporal Grounding of Large Language Models from Perception Streams SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 886c959b-7470-4a40-af9b-df541b56054b · inbound
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ead79631-e448-4237-9af1-29b412302b17 · inbound
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6f538434-ae76-48b7-8472-071c30261078 · inbound
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a3a9c864-aa3e-4fbc-9b2e-78d242c02496 · inbound
Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7e319e28-7db0-4b32-a131-154b894762ed · inbound
Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 72c8c96f-3697-4353-bab3-fce7fbdce1e8 · inbound
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6c6e2dfe-20c2-49cd-a532-764928834525 · inbound
Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 80dead48-e9dc-418f-9f8d-49be12347821 · inbound
LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d2bd64cf-6304-4103-8a82-22c262e82591 · inbound
The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b9b16f0b-1593-4c89-b43d-c9450caed3c8 · inbound
Reinforcing Dual-Path Reasoning in Spatial Vision Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b8566166-c3ac-4e6a-b48d-b1de1d6526c3 · inbound
Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 93afacff-302a-46a6-a0f7-3e0e5f9b8d17 · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d5f799bb-513d-43f3-aa90-2c83bb44d230 · inbound
VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.