Pith. sign in

Paper Citation Record · LEDGER

SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2406.01584.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.01584 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:24:04.777438Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation db059a5a-634b-4ba9-aff2-7ec41000d777 · inbound

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought cites this paper.

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:09:34.859572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T23:09:34.805552Z digest=sha256:0844c810bf6ee9648ebcfe94a968e18f4123a303dbc589f748ca26efbf822096

Observation 6ee4ee5c-9a39-4345-90d7-00518300fbf6 · inbound

Can Multimodal Large Language Models Understand Spatial Relations? cites this paper.

Can Multimodal Large Language Models Understand Spatial Relations? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:04.777438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:24:04.777438Z digest=sha256:1ed5754615c4ef4d2a6c9add1bfe8c0d7f26673c3284d879f321ee560e6c5c86

Observation 6c1db7ac-7239-46d4-a0be-41b8f3b2a0d3 · inbound

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation cites this paper.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:39.933165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:39.933165Z digest=sha256:f9298c1b59a79c78cee806cbfc5032195cb30cf546bc13bfc232f50d6bf59500

Observation eaf3a9d2-5463-4376-a6bc-5326c43abcfe · inbound

GenSpace: Benchmarking Spatially-Aware Image Generation cites this paper.

GenSpace: Benchmarking Spatially-Aware Image Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:21:20.914495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:21:20.914495Z digest=sha256:457b5b1df8aaefd0f9859892d70bec7e1e0c66249e6d12e67347702189a33362

Observation 61d2f55d-3e29-4091-92b3-383a5c5938ae · inbound

BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models cites this paper.

BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:44.249542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:44.249542Z digest=sha256:a8ef715afaf5c3b83f60f919958cecfaa815d16ef36ac757e26b7e27dee506c3

Observation e85307f9-bf8d-42ae-aaf8-e44fdc65d69a · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.127004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.127004Z digest=sha256:dd39ff62393b21d9da8582818055554038b80e6992fa7625ba2cc34695639d13

Observation 9f89895c-f4bf-432a-8c13-6e42b8eabd64 · inbound

Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments cites this paper.

Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:25:05.011646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:25:05.011646Z digest=sha256:5bdeb81f35228df50dc64cb5e3f3a9c9c66e0558e6bec14dee852a32a84930be

Observation 91e0ac93-8634-4490-81b6-3162bc71aa63 · inbound

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills cites this paper.

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:29.384793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:29.384793Z digest=sha256:9c016bf3818bfcb209dad69ea8192aece8dc79c39e82455fe09f1b2a898eab05

Observation 1cb6f7bd-1dfe-4715-881b-fbe525c78b9a · inbound

Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning cites this paper.

Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:12.750753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:12.750753Z digest=sha256:7d3ea6104981221e7135cf26c20674e54f237d24d93365cf0407a74617b75f15

Observation ca1bd479-40fe-484f-a892-1beea35e5999 · inbound

UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding cites this paper.

UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:20.579394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:20.579394Z digest=sha256:0c6506a24a26a3327bb65dae724f405d066691821995772494ed993eb29cdcac

Observation 9b71a8db-4c94-48c2-b200-c1b04c17a942 · inbound

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models cites this paper.

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:27.097127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:27.097127Z digest=sha256:f5130bae298816e0a7178f16a0b2b55ea7f0cfc8bc7c49b53fd683faa55d77d5

Observation eff6343f-b388-4618-810d-b99b1296daa9 · inbound

AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning cites this paper.

AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:56:14.629466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:56:14.629466Z digest=sha256:f67e30ca1d9b74424f8466f76f8ee9306bd9a4a23f525292112b489633641fa0

Observation aa4cd893-691c-4236-96ff-78c5eb4d968e · inbound

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds cites this paper.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:25.638480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:25.638480Z digest=sha256:44c6e6611e556ca2d73d864e6d1628886cb40133f71f3f25182e3a9069e82b4c

Observation a214032c-b48a-448e-b668-586e250128f4 · inbound

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation cites this paper.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.450996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.450996Z digest=sha256:1ecee96fc481301ab9867bd0afc879e4993ee758e06b2af283f4639a0538a23d

Observation 6cabf21b-16b5-4f30-849a-44cecbdfbe34 · inbound

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models cites this paper.

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:16.913232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:16.913232Z digest=sha256:aba7f0bb318b73b6a2ab8d7c946c23784a311d6bac01147ffb3ff122e0ceaabb

Observation 49e5a4d9-4f94-4b70-b6f5-839031e9bc6e · inbound

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation cites this paper.

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:06:51.724391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T22:04:34.235731Z digest=sha256:61991781304b98a2a6929d5917bf42bcab898e6397b23d492750b654203fa669

Observation fbbdb503-c116-4f08-9fb4-c62944dd3a0b · inbound

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks cites this paper.

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:08.100667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:08.100667Z digest=sha256:66d704eb26bac33caab8478f41aad988d025d5ca773d7d33e92841b163be4c31

Observation d9cb7737-9013-4111-b31c-9d0ec45dee76 · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:20:34.446572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T01:18:44.523602Z digest=sha256:853ce67aa81a61a81951de1b76861beae3a2f979f1f7e8902617a814a5b61c77

Observation 185a0f71-71f6-47b5-b1c0-81a63a0152c0 · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T00:11:50.465239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:11:50.465239Z digest=sha256:86094a4e3eb18c7e245a0cb70d1098ff75ed01e266c45ee7fd08ef7274dc1b1f

Observation c11b4357-5a09-4755-be09-f9b025ac2240 · inbound

Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation cites this paper.

Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T21:41:28.734178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:41:28.734178Z digest=sha256:81934f5d57fa732e8c2ac59f69291ef12f76ae4f40a5422b3c10a8e659e20841

Observation 58ef9cc4-cba5-4ace-bfdc-b6888ef921e0 · inbound

Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation cites this paper.

Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:32:41.261611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T09:31:39.548562Z digest=sha256:e6d9333738f66231bcee60114bc6f526a594b2dd9f4d12e7c32d05093785e241

Observation a763a5cf-2318-4778-b72d-9a76454f1ad9 · inbound

When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning cites this paper.

When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:25:12.110307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:25:12.110307Z digest=sha256:0f6432c2e010368a2e0c8eb5651e027cb6b71963d25419509f4281bddae896c7

Observation aac82ce6-25c0-4e8d-b671-7897e406c611 · inbound

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization cites this paper.

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:16:09.802485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T15:12:23.459579Z digest=sha256:37da6fb0d23af277f05ada7e961221cd9dc7f57aba06933bd66d23b091d83082

Observation 9a45f726-1e72-4ffd-bf4c-fc25f0c75fab · inbound

TableVision: A Large-Scale Benchmark for Spatially Grounded Reasoning over Complex Hierarchical Tables cites this paper.

TableVision: A Large-Scale Benchmark for Spatially Grounded Reasoning over Complex Hierarchical Tables SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:23:02.419126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T17:20:28.036531Z digest=sha256:98c93dad00d90616138592d8178cdb7c033745c34a88f557757d95a582357494

Observation f57c5c60-8de1-43ac-a63d-0d3b49227210 · inbound

Spatio-Temporal Grounding of Large Language Models from Perception Streams cites this paper.

Spatio-Temporal Grounding of Large Language Models from Perception Streams SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:26:02.689307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:10:45.837684Z digest=sha256:688b923d907259e4f9c295244cb0908fc2451dda9e137fbbd9f7959503d84905

Observation 886c959b-7470-4a40-af9b-df541b56054b · inbound

CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark cites this paper.

CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:38:12.293094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T10:37:27.926364Z digest=sha256:a0aa456ec05f68321f4526a35b5a8f50aab540d877753ffa0195f15b96b4d887

Observation ead79631-e448-4237-9af1-29b412302b17 · inbound

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop cites this paper.

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:53:13.215752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T10:52:22.778489Z digest=sha256:9bbe2b45706a7cac35fe536ac7e6972c7968fba7fd968e93b9c516ce0cb46c69

Observation 6f538434-ae76-48b7-8472-071c30261078 · inbound

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop cites this paper.

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:05:47.188163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T18:25:17.831116Z digest=sha256:5af14b3436b721ddcfde38d4897208aaa89299dc81e361670159c5af838210a2

Observation a3a9c864-aa3e-4fbc-9b2e-78d242c02496 · inbound

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? cites this paper.

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:59:45.735753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T06:55:04.657347Z digest=sha256:9e35adef600c175261d4df0a842c4cfb5b530b1b10b21cde40707a437d7e9af9

Observation 7e319e28-7db0-4b32-a131-154b894762ed · inbound

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? cites this paper.

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:54:57.801081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:52:26.785086Z digest=sha256:1f2aa305ebd19914deb32a5d85c178b31191eabe213f7ddcda4b642cda303f34

Observation 72c8c96f-3697-4353-bab3-fce7fbdce1e8 · inbound

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)? cites this paper.

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)? SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.548789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:48:19.295578Z digest=sha256:051b1e5ba04281171566c7e0b1561029284ca7e8fb024c7089e5dc34f734e348

Observation 6c6e2dfe-20c2-49cd-a532-764928834525 · inbound

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning cites this paper.

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:15.673197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T15:40:05.730181Z digest=sha256:8cd9cf9c548f8a4da356c23668e90da30120a1465dbd601f2db9fae65fdb5e01

Observation 80dead48-e9dc-418f-9f8d-49be12347821 · inbound

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video cites this paper.

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:57.256251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T02:16:25.730555Z digest=sha256:2ba8472f07224b47011ea4a9c7de461bcc067e69329b7afad877b2c41f1150d4

Observation d2bd64cf-6304-4103-8a82-22c262e82591 · inbound

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models cites this paper.

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.468135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T21:58:53.702009Z digest=sha256:4eb9e37292146453497689aebe27d8cdbe5230bee181f0a99dafc5ec498cb47b

Observation b9b16f0b-1593-4c89-b43d-c9450caed3c8 · inbound

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models cites this paper.

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:08:55.590121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T01:42:30.005911Z digest=sha256:2bd44e099c3eaab16a6202361a342bd83afb28470ff2d9b72490fe1dadc7583d

Observation b8566166-c3ac-4e6a-b48d-b1de1d6526c3 · inbound

Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation cites this paper.

Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-26T10:39:18.313055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T10:37:24.946718Z digest=sha256:f47c76f90655e91edb1854215ccb6e57e9477758d8af1c0d5718e74a7f80d3e7

Observation 93afacff-302a-46a6-a0f7-3e0e5f9b8d17 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.087107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:26bc4cd807d9de308778cfc323e49baeda4ea669506987b8a2b879d879db5287

Observation d5f799bb-513d-43f3-aa90-2c83bb44d230 · inbound

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation cites this paper.

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T06:36:27.565496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:36:27.565496Z digest=sha256:e178b8a6ea004841154ef96de957ca3e7a5f0164b4dd891329e3b0020111dc19