Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2401.12168.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:33:58.479385Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T03:06:43.700595Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 48438ee1-5f7e-4dd8-b156-f1dc65cc05a8 · inbound
LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98641f75-569f-4748-b8a6-73654f86fa24 · inbound
Explainability for Vision Foundation Models: A Survey SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 131
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a45179ff-a548-46c8-bfcc-2f6311b025a4 · inbound
Learning the RoPEs: Better 2D and 3D Position Encodings with STRING SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecee111c-99a8-4f8e-82a8-7493c4edd222 · inbound
Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a82e9a24-ed16-4b9a-a427-25cd479fe798 · inbound
Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1d998d1-f294-4f63-905f-d4b06ff9f1f6 · inbound
Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f85b5b-2a86-40c6-b069-8646390f3e0c · inbound
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d56ca6d2-3da4-48fc-83ca-1c00617ec8ad · inbound
Sustainability assessment using multimodal AI agents SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 033b7dc6-1001-45df-ae01-844f0722ad04 · inbound
Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e0d4a50-0c88-4a17-9bce-c82ec548426a · inbound
TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation dc63545e-e5f0-43f7-a984-90ee3e098ed3 · inbound
SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f806e86-4c61-4cc9-b2f2-229ee754e391 · inbound
Multimodal Language Models Cannot Spot Spatial Inconsistencies SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4042bd62-411d-4827-a6a5-d300abfe0de5 · inbound
Exploring Spatial Intelligence from a Generative Perspective SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8115d2dd-7059-40e7-ae26-f6c89343a181 · inbound
Fast Core Identification SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 33c744aa-d26f-4a17-a6a6-16932b24b563 · inbound
Latent State Design for World Models under Sufficiency Constraints SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 69da90aa-aaa9-48f4-b247-5247b855e41d · inbound
PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 23a5b44b-64c5-49ad-be51-9ce0c3695052 · inbound
PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b696d70e-7b11-4944-861a-72ea37bed505 · inbound
Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8b054447-c513-4c7b-96e8-94c20819320c · inbound
GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c1ccb55f-fade-44c5-a941-4f5298ab4061 · inbound
Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 19c5ae1c-b4f1-42ca-830b-a59f7df338e2 · inbound
LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 63d4cc74-765a-4972-8a29-1c166fdb8eb9 · inbound
Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e2c309cb-eb6c-44c3-bca1-bf0e98e3f9da · inbound
Reinforcing Dual-Path Reasoning in Spatial Vision Language Models SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 71a61b51-b882-40bd-952c-b3fcc22ebe5f · inbound
HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation abacc9ee-aef9-4d82-8fd3-78487bc291c6 · inbound
Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6861a3c7-0b02-429d-b5b7-34abebf4b381 · inbound
Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e5a2891b-07d2-407c-b3a2-24b431e394d9 · inbound
ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5f835b9-a579-4434-bb92-bb0e2de903a8 · inbound
Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 703e5350-68a9-4d22-a1c6-6143e6ef89bf · inbound
PhysX-CoT: Structured Physical Reasoning from a Single Image to Simulation-Ready 3D Assets SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.