Pith. sign in

Paper Citation Record · LEDGER

SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2406.01584.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.01584 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:24:04.777438Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation db059a5a-634b-4ba9-aff2-7ec41000d777 · inbound

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought cites this paper.

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:09:34.859572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T23:09:34.805552Z digest=sha256:6bcc5b3d9dced9734a557f89121024f66c935213f35b14f5c08ea157febbf4a8

Observation 6ee4ee5c-9a39-4345-90d7-00518300fbf6 · inbound

Can Multimodal Large Language Models Understand Spatial Relations? cites this paper.

Can Multimodal Large Language Models Understand Spatial Relations? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:04.777438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:24:04.777438Z digest=sha256:953084cbe4a52f633959061a2c513121e0dfd3795de6a85ad7d7c31c467975a2

Observation 6c1db7ac-7239-46d4-a0be-41b8f3b2a0d3 · inbound

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation cites this paper.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.933165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.933165Z digest=sha256:62d72f3da366a27db4753bab4f1458ef37e647c52eea02bbe9799c9bbd08775e

Observation eaf3a9d2-5463-4376-a6bc-5326c43abcfe · inbound

GenSpace: Benchmarking Spatially-Aware Image Generation cites this paper.

GenSpace: Benchmarking Spatially-Aware Image Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:21:20.914495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:21:20.914495Z digest=sha256:b8f01bd0ee007beef02904ded701b7d40b03981002827059bac64872e7e2f14b

Observation 61d2f55d-3e29-4091-92b3-383a5c5938ae · inbound

BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models cites this paper.

BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:44.249542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:44.249542Z digest=sha256:3e1665826b54e082d627c4d4181cc0ef5839f57b09bd62e745fdbd23d284dede

Observation e85307f9-bf8d-42ae-aaf8-e44fdc65d69a · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.127004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.127004Z digest=sha256:c19f959a9046a215105cdafc26530194f9b5e13c171eead3471ec023525b845f

Observation 9f89895c-f4bf-432a-8c13-6e42b8eabd64 · inbound

Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments cites this paper.

Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:25:05.011646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:25:05.011646Z digest=sha256:ab548829fc346d9d146fc91520bb095a8420c011673936a880a00c9161daf0bf

Observation 91e0ac93-8634-4490-81b6-3162bc71aa63 · inbound

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills cites this paper.

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:29.384793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:29.384793Z digest=sha256:9ca5989527dc139c0523a8a99081748f295166ca1a5b9c4cc8b6cccbca3ec972

Observation 1cb6f7bd-1dfe-4715-881b-fbe525c78b9a · inbound

Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning cites this paper.

Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:12.750753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:12.750753Z digest=sha256:6bd90b05062ec73a8ba76543092316587077ddc4ca7228fad7bca39cd76a3dd0

Observation ca1bd479-40fe-484f-a892-1beea35e5999 · inbound

UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding cites this paper.

UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:20.579394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:20.579394Z digest=sha256:14566d4e210fbdd4661ebd137d45b588a36e88984f0fe3c684334bd091d6a74c

Observation 9b71a8db-4c94-48c2-b200-c1b04c17a942 · inbound

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models cites this paper.

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:27.097127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:27.097127Z digest=sha256:baea0d784db316fd1c6420de6f1bd643ba2868a80b0de53d325d5f1a300a1177

Observation eff6343f-b388-4618-810d-b99b1296daa9 · inbound

AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning cites this paper.

AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:56:14.629466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:56:14.629466Z digest=sha256:8eb6be3d7587267526c071c611f6f5279f32d64d2212a06e343c5d78b5f8b4a6

Observation aa4cd893-691c-4236-96ff-78c5eb4d968e · inbound

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds cites this paper.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:25.638480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:25.638480Z digest=sha256:58acad5e4a880cb8ecbacf841333bc5200597d767079a6b0fdbbc8aa6b90ed47

Observation a214032c-b48a-448e-b668-586e250128f4 · inbound

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation cites this paper.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.450996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.450996Z digest=sha256:6120a97b4ba51a1240466b012a87b43952a7c380ff80b855e0f090d428f5b42d

Observation 6cabf21b-16b5-4f30-849a-44cecbdfbe34 · inbound

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models cites this paper.

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:16.913232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:16.913232Z digest=sha256:0a0f50a3b5636f287f286a6b8b0196dff799a15bb6bd030e3643be714e9ae108

Observation 49e5a4d9-4f94-4b70-b6f5-839031e9bc6e · inbound

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation cites this paper.

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:06:51.724391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T22:04:34.235731Z digest=sha256:0aedc8357880aedd3dc872b6552aaaa4cc92ca9a0126714c889bf756671355e0

Observation fbbdb503-c116-4f08-9fb4-c62944dd3a0b · inbound

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks cites this paper.

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:08.100667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:08.100667Z digest=sha256:abe69294024a9527c9ed152b8862c82bc1cae4dfcc4cecfd50fae4551858cd9f

Observation d9cb7737-9013-4111-b31c-9d0ec45dee76 · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:20:34.446572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T01:18:44.523602Z digest=sha256:76b35a6ff2d96ae8761dac398c4e78eac068889aae27bb78ba3b4148a48da2ac

Observation 185a0f71-71f6-47b5-b1c0-81a63a0152c0 · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T00:11:50.465239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:11:50.465239Z digest=sha256:83a64c3740fde8a29ef378e4659f3fd1009e1dd0ff47bb6eb121d0f10d855f50

Observation c11b4357-5a09-4755-be09-f9b025ac2240 · inbound

Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation cites this paper.

Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T21:41:28.734178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:41:28.734178Z digest=sha256:3dfad09941205a12e0c16b686fe11659c2c602c425ebe3c79a931edd589af917

Observation 58ef9cc4-cba5-4ace-bfdc-b6888ef921e0 · inbound

Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation cites this paper.

Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:32:41.261611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:31:39.548562Z digest=sha256:b4c65420e198ad9d5d1b85372e552776c43f102e736751c5570583ffc8690aec

Observation a763a5cf-2318-4778-b72d-9a76454f1ad9 · inbound

When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning cites this paper.

When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:25:12.110307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:25:12.110307Z digest=sha256:9d9147afbcf32a6c518f22a870ea09f76151d39ce927f1a3d73c4507c0c3a7b4

Observation aac82ce6-25c0-4e8d-b671-7897e406c611 · inbound

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization cites this paper.

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:16:09.802485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T15:12:23.459579Z digest=sha256:d8ee8a28a93df062e16b39c20f6707f8f13dd266bd18a81b61b4b8dc053b85c8

Observation 9a45f726-1e72-4ffd-bf4c-fc25f0c75fab · inbound

TableVision: A Large-Scale Benchmark for Spatially Grounded Reasoning over Complex Hierarchical Tables cites this paper.

TableVision: A Large-Scale Benchmark for Spatially Grounded Reasoning over Complex Hierarchical Tables SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:23:02.419126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T17:20:28.036531Z digest=sha256:ed6c1457f0e30ffc6628309303e003a95c91374d7e82d5d4ab07ba1865894e27

Observation f57c5c60-8de1-43ac-a63d-0d3b49227210 · inbound

Spatio-Temporal Grounding of Large Language Models from Perception Streams cites this paper.

Spatio-Temporal Grounding of Large Language Models from Perception Streams SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:26:02.689307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:10:45.837684Z digest=sha256:8b2397cfe08c77229c24b8db8ae6cf100f8a7cbb86eb3816ea40cbdfd4c7d21f

Observation 886c959b-7470-4a40-af9b-df541b56054b · inbound

CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark cites this paper.

CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:38:12.293094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T10:37:27.926364Z digest=sha256:3c4fa2d4d4aef002bcc190fc2783c2a5edf8997df151cbe87fd1369d4e75ef7e

Observation ead79631-e448-4237-9af1-29b412302b17 · inbound

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop cites this paper.

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:53:13.215752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T10:52:22.778489Z digest=sha256:85464a6e93f8094a187c3549a16514ece6204ed325a270092261ad7cfbae158d

Observation 6f538434-ae76-48b7-8472-071c30261078 · inbound

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop cites this paper.

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:05:47.188163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T18:25:17.831116Z digest=sha256:8f2e60eff8ab6394975281ae934daa5c3ed6d6f3c9752f8d0e7c0bac3865b2ee

Observation a3a9c864-aa3e-4fbc-9b2e-78d242c02496 · inbound

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? cites this paper.

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:59:45.735753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T06:55:04.657347Z digest=sha256:cc636f9d18d9feedeb390c79fa5cee7286addefbbd2d19f83dd768050050d212

Observation 7e319e28-7db0-4b32-a131-154b894762ed · inbound

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? cites this paper.

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:54:57.801081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T17:52:26.785086Z digest=sha256:dc0c38d3d2b963e07f9073eae5fa3b2a9e4e74960ee4d085e7a099e3c0d605d5

Observation 72c8c96f-3697-4353-bab3-fce7fbdce1e8 · inbound

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)? cites this paper.

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.548789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T07:48:19.295578Z digest=sha256:5bc20b872cb4c3b7f709248bd317d22b319ecd5cd8645692764031b43a22f401

Observation 6c6e2dfe-20c2-49cd-a532-764928834525 · inbound

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning cites this paper.

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:15.673197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T15:40:05.730181Z digest=sha256:3e487e1bc538a9dccfba2f369e9cfeb54ce62dcc358ff8e7067d06f5a9372241

Observation 80dead48-e9dc-418f-9f8d-49be12347821 · inbound

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video cites this paper.

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:57.256251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T02:16:25.730555Z digest=sha256:33fd81d01c47b84ad5bff8b4c11b82102249fa5ed9422833c336a8fe9a177470

Observation d2bd64cf-6304-4103-8a82-22c262e82591 · inbound

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models cites this paper.

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.468135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T21:58:53.702009Z digest=sha256:f0aa92bb51a1ac82a580aa897c029bd2eb5d919730d4075ddf9dddc58b22596a

Observation b9b16f0b-1593-4c89-b43d-c9450caed3c8 · inbound

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models cites this paper.

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:08:55.590121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T01:42:30.005911Z digest=sha256:f0f6d150b760e85205bcc07cfa8d09c320e1a4b779f75f8a18dc46ee1db256ef

Observation b8566166-c3ac-4e6a-b48d-b1de1d6526c3 · inbound

Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation cites this paper.

Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-26T10:39:18.313055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T10:37:24.946718Z digest=sha256:f2853dd78156551624ac3dffdd2044f641d1ac5ba3f9b4d7eb106b17bc85a4dc

Observation 93afacff-302a-46a6-a0f7-3e0e5f9b8d17 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.087107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:5762f12bb2cb5cafb60b8bec2b198190206557a72e84105e9f58a7b46dec6dfd

Observation d5f799bb-513d-43f3-aa90-2c83bb44d230 · inbound

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation cites this paper.

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T06:36:27.565496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:36:27.565496Z digest=sha256:bbba65cac81f696b36816d6a41b035183e7309b9231946ec053ffcb981fb0b60