REVIEW 19 cited by
SpatialBot: Precise Spatial Understanding with Vision Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Vision Language Models (VLMs) have achieved impressive performance in 2D image understanding, however they are still struggling with spatial understanding which is the foundation of Embodied AI. In this paper, we propose SpatialBot for better spatial understanding by feeding both RGB and depth images. Additionally, we have constructed the SpatialQA dataset, which involves multi-level depth-related questions to train VLMs for depth understanding. Finally, we present SpatialBench to comprehensively evaluate VLMs' capabilities in spatial understanding at different levels. Extensive experiments on our spatial-understanding benchmark, general VLM benchmarks and Embodied AI tasks, demonstrate the remarkable improvements of SpatialBot trained on SpatialQA. The model, code and data are available at https://github.com/BAAI-DCAI/SpatialBot.
Forward citations
Cited by 19 Pith papers
-
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images
MonoSR is a 1M-question benchmark for spatial reasoning from single photos across indoor, outdoor, and object-centric scenes; current VLMs score roughly 30-40%, and giving models 3D box coordinates lifts them near perfect.
-
GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert
A frozen, large-scale pretrained diffusion policy converts sparse 3D waypoints from a VLM into dense robot actions, enabling zero-shot reuse of the action expert on new tasks and environments.
-
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
A new egocentric benchmark shows vision-language models fail at spatial reasoning across disjoint frames, falling 28 points behind humans and only improving sharply when handed ground-truth 3D coordinates.
-
When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents
Spatial-memory staleness is a measurable safety failure for VLM agents: stale memory increases deaths, and visual auditing of stale entries is highly model-dependent.
-
SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
Dense scene-graph-grounded rewards let a 7B multimodal LLM trained on 7K synthetic questions beat SFT and sparse-RL baselines and outscore GPT-4o on average across 12 spatial/real-world benchmarks.
-
Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks
A new contrastive real-image benchmark shows most vision-language models fail spatial relation tasks, while chain-of-thought reasoning models approach human-level accuracy.
-
Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas
Canvas3D lets users arrange objects in a 3D canvas generated from a text prompt, then feeds depth, skeleton, and lighting constraints to diffusion models to produce images that match the layout.
-
BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?
BenchDepth evaluates eight depth foundation models by their performance on five downstream tasks, finding Depth Anything V2's relative version to be the most practically useful.
-
RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot
A generative model and wrist camera turn human hand videos into robot gripper demonstrations that train manipulation policies at success rates close to those trained on real gripper data.
-
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
BMMR provides a 110k-question bilingual, multimodal, college-level dataset across 300 subjects where state-of-the-art models score at most about 50%.
-
Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models
Vision-language models can handle some 2D shape puzzles but nearly all fail at multi-step 3D spatial deformation reasoning.
-
ToSA: Token Merging with Spatial Awareness
A training-free token merging method that adds depth-derived spatial similarity to ToMe's bipartite soft matching, improving VQA accuracy at high token reduction rates.
-
Can Multimodal Large Language Models Understand Spatial Relations?
SpatialMQA, a new spatial-relation benchmark, shows the top MLLM reaches 48.14% accuracy versus 98.40% for humans.
-
Depth Anything at Any Condition
A fine-tuned Depth Anything V2 model using perturbation consistency and spatial distance constraints improves monocular depth estimation under adverse conditions without any labeled data.
-
Warehouse Spatial Question Answering with LLM Agent
An LLM agent equipped with lightweight distance and inclusion perception models achieved 95.86% accuracy on the 2025 AI City Challenge warehouse spatial QA benchmark, ranking first.
-
PRISM: Pointcloud Reintegrated Inference via Segmentation and Cross-attention for Manipulation
PRISM trains a diffusion policy on segmented point-cloud object tokens fused with joint states via cross-attention, reporting 82.0 percent average success across six RoboTwin tasks versus 58.4 percent for DP3 and 22.3...
-
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
Scene-graph-based chain-of-thought prompting and GRPO training improve spatial reasoning accuracy in vision-language models, and GRPO degrades less than supervised fine-tuning when question wording is flipped.
-
OscNet v1.5: Energy Efficient Hopfield Network on CMOS Oscillators for Image Classification
A forward-only Hebbian Hopfield classifier with per-class weight matrices and learned sparsity reaches 75.3% on binary MNIST under simulated oscillator inference, and up to 97.4% when its features are fed to a finetun...
-
A Spatial Relationship Aware Dataset for Robotics
A new robot-acquired, spatial-relationship-labelled dataset is released and benchmarked, with qualitative evidence that explicit spatial cues improve ChatGPT 4o robot planning.
Discussion (0). Sign in to comment.