Pith. sign in

REVIEW 4 cited by

EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.05756 v1 pith:AKNZIVKZ submitted 2024-06-09 cs.AI cs.CLcs.CVcs.MM

classification cs.AIcs.CLcs.CVcs.MM
keywords embodiedlvlmsspatialunderstandingbenchmarkcurrentembspatial-benchlarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent rapid development of Large Vision-Language Models (LVLMs) has indicated their potential for embodied tasks.However, the critical skill of spatial understanding in embodied environments has not been thoroughly evaluated, leaving the gap between current LVLMs and qualified embodied intelligence unknown. Therefore, we construct EmbSpatial-Bench, a benchmark for evaluating embodied spatial understanding of LVLMs.The benchmark is automatically derived from embodied scenes and covers 6 spatial relationships from an egocentric perspective.Experiments expose the insufficient capacity of current LVLMs (even GPT-4V). We further present EmbSpatial-SFT, an instruction-tuning dataset designed to improve LVLMs' embodied spatial understanding.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models

    cs.CV 2026-01 conditional novelty 7.0 of 10

    Using a simple action-token adapter, nine VLMs are compared as robot policy backbones, showing general VLM ability transfers poorly to control and the vision encoder is the key bottleneck.

  2. RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A unified transformer model generates language-and-image planning sequences for embodied tasks, and shows real-robot manipulation without large-scale action pretraining.

  3. Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A new contrastive real-image benchmark shows most vision-language models fail spatial relation tasks, while chain-of-thought reasoning models approach human-level accuracy.

  4. Can Multimodal Large Language Models Understand Spatial Relations?

    cs.CV 2025-05 conditional novelty 6.0 of 10

    SpatialMQA, a new spatial-relation benchmark, shows the top MLLM reaches 48.14% accuracy versus 98.40% for humans.

Pith tools