Pith. sign in

REVIEW 7 cited by

LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.19990 v3 pith:RA6IRTQF submitted 2025-03-25 cs.AI

classification cs.AI
keywords reasoningspatialmllmslego-puzzlesunderstandingmulti-stepsequentialtasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-step spatial reasoning entails understanding and reasoning about spatial relationships across multiple sequential steps, which is crucial for tackling complex real-world applications, such as robotic manipulation, autonomous navigation, and automated assembly. To assess how well current Multimodal Large Language Models (MLLMs) have acquired this fundamental capability, we introduce LEGO-Puzzles, a scalable benchmark designed to evaluate both spatial understanding and sequential reasoning in MLLMs through LEGO-based tasks. LEGO-Puzzles consists of 1,100 carefully curated visual question-answering (VQA) samples spanning 11 distinct tasks, ranging from basic spatial understanding to complex multi-step reasoning. Based on LEGO-Puzzles, we conduct a comprehensive evaluation of 20 state-of-the-art MLLMs and uncover significant limitations in their spatial reasoning capabilities: even the most powerful MLLMs can answer only about half of the test cases, whereas human participants achieve over 90% accuracy. Furthermore, based on LEGO-Puzzles, we design generation tasks to investigate whether MLLMs can transfer their spatial understanding and reasoning abilities to image generation. Our experiments show that only GPT-4o and Gemini-2.0-Flash exhibit a limited ability to follow these instructions, while other MLLMs either replicate the input image or generate completely irrelevant outputs. Overall, LEGO-Puzzles exposes critical deficiencies in existing MLLMs' spatial understanding and sequential reasoning capabilities, and underscores the need for further advancements in multimodal spatial reasoning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SpaceVista: All-Scale Visual Spatial Reasoning from mm to km

    cs.CV 2025-10 conditional novelty 7.0 of 10

    SpaceVista contributes a 1M-QA, 38K-video all-scale spatial reasoning dataset spanning mm to km, a manually verified benchmark, and a fine-tuned 7B MLLM with scale experts and progressive reward training.

  2. Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Mild visual corruption can boost a vision-language model's top-1 accuracy while its confidence–correctness alignment (measured by the new RAS score) degrades.

  3. SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Dense scene-graph-grounded rewards let a 7B multimodal LLM trained on 7K synthetic questions beat SFT and sparse-RL baselines and outscore GPT-4o on average across 12 spatial/real-world benchmarks.

  4. CharacterShot: Controllable and Consistent 4D Character Animation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A new pipeline generates pose-controlled, view-consistent 4D character animations from one reference image and a 2D pose sequence, backed by a new 13,115-character dataset and benchmark.

  5. Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Vision-language models can handle some 2D shape puzzles but nearly all fail at multi-step 3D spatial deformation reasoning.

  6. Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models

    cs.AI 2025-05 conditional novelty 6.0 of 10

    Current vision-language models fall far short of humans on spatial reasoning, especially when they must generate answers directly instead of choosing from options.

  7. STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

    cs.CV 2025-05 conditional novelty 6.0 of 10

    STAR-R1 uses single-stage reinforcement learning with fine-grained rewards to improve spatial transformation reasoning in multimodal LLMs, outperforming supervised fine-tuning on cross-view TVR tasks.

Pith tools