Pith. sign in

REVIEW 8 cited by

Does Spatial Cognition Emerge in Frontier Models?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.06468 v2 pith:YO3YJBVE submitted 2024-10-09 cs.AI cs.CVcs.LG

classification cs.AIcs.CVcs.LG
keywords modelsspatialbenchmarkcognitionfrontiercognitiveevaluateslarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Not yet. We present SPACE, a benchmark that systematically evaluates spatial cognition in frontier models. Our benchmark builds on decades of research in cognitive science. It evaluates large-scale mapping abilities that are brought to bear when an organism traverses physical environments, smaller-scale reasoning about object shapes and layouts, and cognitive infrastructure such as spatial attention and memory. For many tasks, we instantiate parallel presentations via text and images, allowing us to benchmark both large language models and large multimodal models. Results suggest that contemporary frontier models fall short of the spatial intelligence of animals, performing near chance level on a number of classic tests of animal cognition. Code and data are available: https://github.com/apple/ml-space-benchmark

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GenSpace: Benchmarking Spatially-Aware Image Generation

    cs.CV 2025-05 conditional novelty 7.0 of 10

    GenSpace benchmarks spatial awareness in image generation with a 3D reconstruction-based evaluator, showing models struggle with allocentric relations and metric measurements.

  2. VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations

    cs.CV 2026-03 conditional novelty 6.5 of 10

    Sparse multi-view reasoning is largely unsolved for VLMs; grounded CoT with visual evidence improves moderate cases and transfers to real data, but deep multi-hop reasoning still scales poorly.

  3. MentisOculi: Revealing the Limits of Reasoning with Mental Imagery

    cs.AI 2026-02 conditional novelty 6.0 of 10

    Visual thoughts — latent tokens, interleaved images, or video rollouts — do not currently improve multi-step reasoning over text-only baselines in frontier models.

  4. CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates

    cs.CV 2025-12 conditional novelty 6.0 of 10

    VLMs perform near chance at detecting and correcting an erroneous step in visual sequential planning, and incrementally updating scene graphs step-by-step (SGI) yields small, partially inconsistent gains.

  5. Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations

    cs.CV 2025-06 conditional novelty 6.0 of 10

    STARE is a 4K-task benchmark showing multimodal LLMs perform near random chance on multi-step spatial simulation tasks such as cube net folding and tangrams, despite strong 2D transformation results.

  6. Can LLMs Learn to Map the World from Local Descriptions?

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A 0.5B LLM trained on templated local descriptions from a synthetic grid city infers unseen distances and directions, encodes coordinates in its hidden states, and plans shortest paths, but fails under navigation pert...

  7. ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    ZeroBench is a hand-built 100-question visual reasoning benchmark, adversarially filtered so every evaluated frontier LMM scored 0% at release.

  8. RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A new 9,121-case benchmark of road-marking tasks shows most multimodal LLMs perform near or below simple rule-based baselines in fine-grained urban spatial reasoning.

Pith tools