Pith. sign in

REVIEW 19 cited by

SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.01584 v3 pith:E6MOC4LL submitted 2024-06-03 cs.CV

classification cs.CV
keywords spatialspatialrgptvlmsreasoninglanguageregiontasksvision
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision Language Models (VLMs) have demonstrated remarkable performance in 2D vision and language tasks. However, their ability to reason about spatial arrangements remains limited. In this work, we introduce Spatial Region GPT (SpatialRGPT) to enhance VLMs' spatial perception and reasoning capabilities. SpatialRGPT advances VLMs' spatial understanding through two key innovations: (1) a data curation pipeline that enables effective learning of regional representation from 3D scene graphs, and (2) a flexible plugin module for integrating depth information into the visual encoder of existing VLMs. During inference, when provided with user-specified region proposals, SpatialRGPT can accurately perceive their relative directions and distances. Additionally, we propose SpatialRGBT-Bench, a benchmark with ground-truth 3D annotations encompassing indoor, outdoor, and simulated environments, for evaluating 3D spatial cognition in VLMs. Our results demonstrate that SpatialRGPT significantly enhances performance in spatial reasoning tasks, both with and without local region prompts. The model also exhibits strong generalization capabilities, effectively reasoning about complex spatial relations and functioning as a region-aware dense reward annotator for robotic tasks. Code, dataset, and benchmark are released at https://www.anjiecheng.me/SpatialRGPT

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning

    cs.AI 2025-11 unverdicted novelty 7.0 of 10

    DecompSR is a large, symbolically verified benchmark dataset and generation framework that independently varies productivity, substitutivity, overgeneralisation, and systematicity to probe compositional multihop spati...

  2. GenSpace: Benchmarking Spatially-Aware Image Generation

    cs.CV 2025-05 conditional novelty 7.0 of 10

    GenSpace benchmarks spatial awareness in image generation with a 3D reconstruction-based evaluator, showing models struggle with allocentric relations and metric measurements.

  3. S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation

    cs.CV 2025-05 conditional novelty 7.0 of 10

    S4-Driver uses a multimodal LLM with a sparse 3D spatio-temporal volume representation to achieve self-supervised motion planning that rivals supervised methods on nuScenes and WOMD.

  4. When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning

    cs.CV 2026-02 conditional novelty 6.0 of 10

    Adaptively gating and scaling world-model imagination at test time matches or outperforms always-on imagination on spatial reasoning benchmarks while using substantially fewer world-model calls and tokens.

  5. Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation

    cs.CV 2025-11 conditional novelty 6.0 of 10

    A frozen VLM's dual-query Yes/No log-odds act as a differentiable semantic-and-spatial critic, improving alignment and geometry in both SDS-based and feed-forward text-to-3D pipelines.

  6. Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A new contrastive real-image benchmark shows most vision-language models fail spatial relation tasks, while chain-of-thought reasoning models approach human-level accuracy.

  7. Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation

    cs.GR 2025-07 conditional novelty 6.0 of 10

    Fine-tuning open vision-language models on 240K synthetic question-answer pairs with exact camera-object labels improves camera-object recognition by 33.4% on average over GPT-4o and Claude-3-Sonnet on the paper's benchmark.

  8. 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

    cs.GR 2025-07 conditional novelty 6.0 of 10

    A self-improving vision-language-model policy iteratively crafts 3D environments from text, and renderings of those environments serve as effective synthetic pretraining data for vision models.

  9. AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning

    cs.RO 2025-07 conditional novelty 6.0 of 10

    AutoLayout combines slow reasoning with fast evolutionary placement and a self-correcting loop of LLM-generated relation checks to produce physically plausible, semantically matched tabletop layouts.

  10. Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Vision-language models can handle some 2D shape puzzles but nearly all fail at multi-step 3D spatial deformation reasoning.

  11. UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A fine-tuned small multimodal LLM outperforms much larger general models on urban tasks in a new benchmark, with caveats about benchmark overlap with training data.

  12. Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Surgery-R1 uses supervised fine-tuning and reinforcement fine-tuning to give a surgical visual question answering model chain-of-thought reasoning, improving accuracy and localization on two EndoVis benchmarks.

  13. RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills

    cs.RO 2025-06 conditional novelty 6.0 of 10

    RobotSmith autonomously designs, 3D-prints, and uses task-specific tools for robotic manipulation, raising task success from 2.8% (no tool) to 50% in simulation.

  14. Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments

    cs.RO 2025-06 conditional novelty 6.0 of 10

    Narrate2Nav uses Barlow Twins alignment to distill language-based reasoning from a large teacher into a small RGB-only navigation model, reporting lower trajectory error and higher goal-reaching success than four baselines.

  15. AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making

    cs.RO 2025-06 conditional novelty 6.0 of 10

    AntiGrounding lifts candidate robot trajectories into the VLM's visual space via multi-view rendering and structured VQA, and reports 57.5% average success across eight manipulation tasks, beating three intermediate-r...

  16. BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A Blender-based diagnostic toolkit that tests VLMs on fine-grained visual skills by varying one visual attribute at a time, exposing failure modes that coarse benchmarks miss.

  17. Can Multimodal Large Language Models Understand Spatial Relations?

    cs.CV 2025-05 conditional novelty 6.0 of 10

    SpatialMQA, a new spatial-relation benchmark, shows the top MLLM reaches 48.14% accuracy versus 98.40% for humans.

  18. VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation

    cs.RO 2026-07 conditional novelty 5.5 of 10

    VistaVLA lifts multi-view vision-language features into 3D Gaussians, compresses them 99% via Merge-then-Query, and improves real-robot manipulation success by ~23% over baselines.

  19. Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Spatial 3D-LLM adds a progressive spatial awareness scheme to a 3D vision-language model, improving several 3D understanding and grounding metrics and introducing new distance and layout-editing tasks.

Pith tools