REVIEW 8 cited by
Does Spatial Cognition Emerge in Frontier Models?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Not yet. We present SPACE, a benchmark that systematically evaluates spatial cognition in frontier models. Our benchmark builds on decades of research in cognitive science. It evaluates large-scale mapping abilities that are brought to bear when an organism traverses physical environments, smaller-scale reasoning about object shapes and layouts, and cognitive infrastructure such as spatial attention and memory. For many tasks, we instantiate parallel presentations via text and images, allowing us to benchmark both large language models and large multimodal models. Results suggest that contemporary frontier models fall short of the spatial intelligence of animals, performing near chance level on a number of classic tests of animal cognition. Code and data are available: https://github.com/apple/ml-space-benchmark
Forward citations
Cited by 8 Pith papers
-
GenSpace: Benchmarking Spatially-Aware Image Generation
GenSpace benchmarks spatial awareness in image generation with a 3D reconstruction-based evaluator, showing models struggle with allocentric relations and metric measurements.
-
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations
Sparse multi-view reasoning is largely unsolved for VLMs; grounded CoT with visual evidence improves moderate cases and transfers to real data, but deep multi-hop reasoning still scales poorly.
-
MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
Visual thoughts — latent tokens, interleaved images, or video rollouts — do not currently improve multi-step reasoning over text-only baselines in frontier models.
-
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
VLMs perform near chance at detecting and correcting an erroneous step in visual sequential planning, and incrementally updating scene graphs step-by-step (SGI) yields small, partially inconsistent gains.
-
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
STARE is a 4K-task benchmark showing multimodal LLMs perform near random chance on multi-step spatial simulation tasks such as cube net folding and tangrams, despite strong 2D transformation results.
-
Can LLMs Learn to Map the World from Local Descriptions?
A 0.5B LLM trained on templated local descriptions from a synthetic grid city infers unseen distances and directions, encodes coordinates in its hidden states, and plans shortest paths, but fails under navigation pert...
-
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
ZeroBench is a hand-built 100-question visual reasoning benchmark, adversarially filtered so every evaluated frontier LMM scored 0% at release.
-
RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
A new 9,121-case benchmark of road-marking tasks shows most multimodal LLMs perform near or below simple rule-based baselines in fine-grained urban spatial reasoning.
Discussion (0). Continue with ORCID to comment.