Pith. sign in

Paper Citation Record · LEDGER

SpatialBot: Precise Spatial Understanding with Vision Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2406.13642.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.13642 v7

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:38:02.242912Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:18:37.306944Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e251e3d7-f2cf-4e7b-a04c-63d90c83f091 · inbound

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces cites this paper.

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:27:44.172737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T09:27:43.919941Z digest=sha256:c702789421c27a5729b9fd25c08d8199d7d174cc4f6232f752788ed252c39c00

Observation bbc29c12-0a68-4762-a83f-ceb18684cf1e · inbound

OscNet: Machine Learning on CMOS Oscillator Networks cites this paper.

OscNet: Machine Learning on CMOS Oscillator Networks SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T13:38:02.242912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:38:02.242912Z digest=sha256:55acfffac6f6dfda9e46ca9a4d8371b5f23edd446a8b35659975a832793e9aac

Observation 7b406e20-d445-4681-a415-e8c4673aa066 · inbound

Can Multimodal Large Language Models Understand Spatial Relations? cites this paper.

Can Multimodal Large Language Models Understand Spatial Relations? SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:04.556539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:24:04.556539Z digest=sha256:b03b8ddcb6cd309f7bb3fdebe13573840e45c45ac708aa6de0565bf9ea9b4056

Observation a02f28ca-6b43-4393-b617-ad0ba9419286 · inbound

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames cites this paper.

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:16.202407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:16.202407Z digest=sha256:3fabbaa106c994a82851eb7c7e91caf636a5795c3035cac150e0604c2d664026

Observation e726471f-2f61-4b87-ae35-606d52eb03d2 · inbound

A Spatial Relationship Aware Dataset for Robotics cites this paper.

A Spatial Relationship Aware Dataset for Robotics SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:51:24.318444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:51:24.318444Z digest=sha256:f2a0d4c59db84b8d513ff7926eb13e4519b30488b3b1ea1099b53cfc678119b2

Observation 64ec50b4-7244-4a6f-b574-b2061b2a6d40 · inbound

OscNet v1.5: Energy Efficient Hopfield Network on CMOS Oscillators for Image Classification cites this paper.

OscNet v1.5: Energy Efficient Hopfield Network on CMOS Oscillators for Image Classification SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:51:48.968184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:51:48.968184Z digest=sha256:7afff4bd721eaa132150c84c465e1bf28fedc26b2d0e9a4384ab1f75b6376f69

Observation ae73add2-f905-443d-b7aa-abbab6e06216 · inbound

ToSA: Token Merging with Spatial Awareness cites this paper.

ToSA: Token Merging with Spatial Awareness SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.887022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.887022Z digest=sha256:f87a921fe136e0371b359176e185a87dc577d0198c676365e9eb213f348d7399

Observation 27c51cb1-594a-4ea0-9033-28489f4484f4 · inbound

Depth Anything at Any Condition cites this paper.

Depth Anything at Any Condition SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:52:02.423508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:52:02.423508Z digest=sha256:4758edd52cc3682053000e9add2013e705f0aac1cd65bd22b98cdc261307bf4d

Observation 5404da64-bf32-4d31-9561-f7a31b58e557 · inbound

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models cites this paper.

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:27.079986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:27.079986Z digest=sha256:dc864052017b973103fcee3c67df9cd6dc7ee59c0180dfac6a55138350007f4d

Observation c569df7c-8152-4629-8373-70c74568eca3 · inbound

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset cites this paper.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.733114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.733114Z digest=sha256:8167c0603944e68df17fed09d26d0d8e3f4a784d01e35887d5ba560399587de6

Observation dbb6d299-169c-4016-a284-3201659f70f7 · inbound

RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot cites this paper.

RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:38.185326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:38.185326Z digest=sha256:a29ee551efaed00b891055364118120515422bfbd1bef57bd28b17f1dead1f7f

Observation 79ba07b9-c5b4-4f00-afe7-84d27c3f0500 · inbound

PRISM: Pointcloud Reintegrated Inference via Segmentation and Cross-attention for Manipulation cites this paper.

PRISM: Pointcloud Reintegrated Inference via Segmentation and Cross-attention for Manipulation SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:48:07.793262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:48:07.793262Z digest=sha256:754dce927998e985327d6d4a7ed88c1baf687afddebf5f36c2d7c14bc63e368b

Observation e2ed1f8c-2a8d-4720-84e0-5316589660b5 · inbound

Warehouse Spatial Question Answering with LLM Agent cites this paper.

Warehouse Spatial Question Answering with LLM Agent SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:29:38.536084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:29:38.536084Z digest=sha256:869adc75f88daeca29a8e71702b8629feb35063101ae5b0f7d3efc66ece4ec5a

Observation 46a2075a-2a04-4750-af4a-544bd439c12c · inbound

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning cites this paper.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.156627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.156627Z digest=sha256:4b53a9a25627397554c8903fa957ec4d3e72195fef21ef738b7c4ced63967434

Observation 5957f5e9-d878-47cd-9323-840f7158ce10 · inbound

BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models? cites this paper.

BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models? SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:38:59.749275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:38:59.749275Z digest=sha256:f9f17faddf7aac2c11eff250f9d7f2f4e1153539c93ebd3d99262eea163cc724

Observation 1cfeaa5f-2de8-42bf-b014-002bbc930982 · inbound

Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas cites this paper.

Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T22:23:09.481651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:23:09.481651Z digest=sha256:8d06ccc08f014669519fd739ca7398e3f3531372e34d05b7209f7e42ce963b56

Observation 452ba9fc-3031-4cb0-aa31-3c394b634645 · inbound

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation cites this paper.

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:06:51.720215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T22:04:34.235731Z digest=sha256:4fff732917eab0b390d4a6276395fe59fe01ee6e9c79af417c8c4d45d11d6505

Observation 69047047-b5e5-45cd-8174-14a6659374bd · inbound

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks cites this paper.

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:07.891723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:07.891723Z digest=sha256:9f0c8b805d82da621c536390bc80d3b8f9558c1f3ff40080819684dfc5cdd355

Observation 82cd167f-a1ad-452d-a6d8-9646c13e431c · inbound

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert cites this paper.

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:11.546062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:38:11.546062Z digest=sha256:4db3c8a43fdee25a557384b8b07a717bf2a960da943e894e80832cc967e625cc

Observation a73e99a7-7552-4f1a-9e21-59d69faf6d4b · inbound

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation cites this paper.

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:25:32.994831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T00:22:41.611893Z digest=sha256:a995ed13091138140d75748260417295b58ce9eb976708ee40c7cdd86e56021f

Observation c89204c0-a66f-433b-9146-90382ced24ee · inbound

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards cites this paper.

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:49.001243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:08:49.001243Z digest=sha256:8e3b5fd840988f192615aae0dcb6625a7396663c9fa5d42296848e1a48d57c5d

Observation f4991661-264f-47e3-b6ef-96f837d1bf93 · inbound

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding cites this paper.

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:04.620503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:21:12.375936Z digest=sha256:926a100527a38d5a271b9d2203c68d0d4cbe72f8f0139c5a9bfa485976d9e56f

Observation 89fc6ba7-bae3-419e-9de8-09a4738a8d4a · inbound

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images cites this paper.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:54.925300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:54.925300Z digest=sha256:5460d6340ff9be99541ded319e1ab34b12018b9854b3b32f1db1c62349ee9be3

Observation 0a578d09-6ef1-47c0-acc8-fb7471255bc4 · inbound

Spatio-Temporal Grounding of Large Language Models from Perception Streams cites this paper.

Spatio-Temporal Grounding of Large Language Models from Perception Streams SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:26:02.678079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:10:45.837684Z digest=sha256:59acb855084023bb374f7bcc362dcb7a68075056585a1c632e79cc74a2e22d50

Observation 64bbe8e8-686b-4ab3-8c2a-8a9ecec42f57 · inbound

EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training cites this paper.

EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:11:03.815073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:16:08.687340Z digest=sha256:46f2db2e9540a9f4e94f8612b9b23a725ff082214e18f150d75d70c4c37bb282

Observation 54209770-f05e-4e37-ba5e-3fc297cfd849 · inbound

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding cites this paper.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:11.692955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T14:20:08.404090Z digest=sha256:f86ed02f85456cdce4465b7b17d962a608933eb10f10ff12a165bdc9023fc89b

Observation fa29e92c-af28-48a0-a4ed-1caaaf27aa2f · inbound

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding cites this paper.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:16:39.992535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:a4f0e484f1217da8fe889004931421f8e8811d1909fad3ea23a4cb4a7a806fc8

Observation 9aa7b8c5-3f88-49ab-bdd7-ffe4c9f2a117 · inbound

Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence cites this paper.

Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.382386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T03:24:41.877312Z digest=sha256:edd7a3c882d3ce09446dc184ed53ee226599630aaa3f36c70eeebdc55bfd3481

Observation 88753f51-bdb3-4f4a-892d-3971c0915305 · inbound

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop cites this paper.

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.222326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T10:52:22.778489Z digest=sha256:1fa4eb3db3140af1721c6f281b67e60fe54b4b099aaec5be6d425eda1c04aba0

Observation b0018adf-4f46-4670-9535-fda813036c2d · inbound

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop cites this paper.

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:05:47.182481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T18:25:17.831116Z digest=sha256:869d61dfb9d1b1271e2dab7d5d5567e9523d3e84c2c6c6ebfa283f7b1cda9429

Observation 016b2ce0-c946-4469-add7-ec1b9b10e22c · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.560592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:23369f4585f0123c519a2ee4a1428db1f17505acbba61f368c333658ac5eb6c4

Observation effb4d79-bfe3-4ed5-b356-8d1378a6ecba · inbound

VLM3: Vision Language Models Are Native 3D Learners cites this paper.

VLM3: Vision Language Models Are Native 3D Learners SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:14.019914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:45:31.978215Z digest=sha256:f08c3f02a5b59f2980716d7f8d12628e1a7279862222e541d2583ad9704cc004

Observation 6f97f28b-55ff-41db-b7ee-b719312b5659 · inbound

Brick-Composer: Using MLLMs for Assembly with Diverse Bricks cites this paper.

Brick-Composer: Using MLLMs for Assembly with Diverse Bricks SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:26:48.424339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T05:59:29.302038Z digest=sha256:3e84ed4a707e66ba8d541d03a9f301c60f9e751ea4a72a081889a8b18494a11d

Observation 443ec8bd-ee16-464b-b4fb-1deb9e64acf5 · inbound

Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning cites this paper.

Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:41.110274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:23:00.372251Z digest=sha256:db22f283ed0888a43d48331e0edc4d9ffba156f0b5cbd26b57dc4758ab67f52d

Observation bb376c85-f6f5-470b-b01f-51d3b0834949 · inbound

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video cites this paper.

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:18:37.308596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T16:16:41.412451Z digest=sha256:28c7399afbac8398427f941d461badfa200ce9d6243a0293e8744be818aa683e

Observation 83309848-96fa-4915-a231-5899935daab2 · inbound

When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents cites this paper.

When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:20.587662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:49:20.587662Z digest=sha256:670fd0dc81e502a8cc8690de3e1074d7400f1e79b7e78dc3ed338c617b367cde