Pith. sign in

Paper Citation Record · LEDGER

Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2503.19707.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.19707 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:24:47.666879Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T09:37:00.841150Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e770b902-a529-42f7-91e2-4410b4695509 · inbound

SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence cites this paper.

SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:11:35.751741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T13:07:11.548885Z digest=sha256:029a6185f28cdc216ed25012019b8bef4bcb387b15f232ccfebab4478ebc4911

Observation fc5d705f-b0e8-4de9-a506-d73312103c29 · inbound

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models cites this paper.

Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:33.010530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:33.010530Z digest=sha256:939a7b8e70fc7426dcd655115d628e6b6d31ad7cd8ab74894a00610ca73b5210

Observation 0cd3df1a-b09a-4cb4-b3a5-a6eaa402bac7 · inbound

BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models cites this paper.

BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:43.691323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:43.691323Z digest=sha256:1ba82b36724834f1078a097d240e938950efa596755a69ac9950e2f38876a75d

Observation 6e20543d-3d2d-4381-8c77-b14dd9b70932 · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.247850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.247850Z digest=sha256:921f1c49aa2d5c9bbb057cba9b931aebc1b93fa81fb439f38043719a4ccf649b

Observation f1ff9a6f-f2be-416f-a1c5-4836384f0f1a · inbound

11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis cites this paper.

11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:54.990372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:54.990372Z digest=sha256:d5a7a22d9dc1bc88514e20e2681f28c718f9a1cf652f496c44c0205b00e7fc57

Observation e2578177-12ce-4060-8b52-e60f60d175a6 · inbound

The N-Body Problem: Parallel Execution from Single-Person Egocentric Video cites this paper.

The N-Body Problem: Parallel Execution from Single-Person Egocentric Video Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T16:55:22.331240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:55:22.331240Z digest=sha256:65dd7e32aec7641a5eaedcade3c30dfb6d280879818dfe1624f61c2b1f1d707c

Observation 7b3245e6-35e9-49b4-9611-369e9c34a280 · inbound

A Deep Learning Model of Mental Rotation Informed by Interactive VR Experiments cites this paper.

A Deep Learning Model of Mental Rotation Informed by Interactive VR Experiments Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:36.756028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:36.756028Z digest=sha256:713aced1309a265489e5d55b6bec076e8655f8f8a5d039da215ab31724e7fe63

Observation f20761b5-2bc6-43be-86e5-777dedd3d6af · inbound

Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation cites this paper.

Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:40:33.119630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T03:39:28.364183Z digest=sha256:4f306bdb6c4eff0ca094acaa9253ac2c2eb13710de0eed7b9891152dae9933d2

Observation 35b1f545-9ce4-4c8f-8068-6e79edf83bff · inbound

It's Time to Get It Right: Improving Analog Clock Reading and Clock-Hand Spatial Reasoning in Vision-Language Models cites this paper.

It's Time to Get It Right: Improving Analog Clock Reading and Clock-Hand Spatial Reasoning in Vision-Language Models Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-15T13:03:40.459461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:03:40.459461Z digest=sha256:e0a19664114cc70f8ab3065af659c40f621599bf82b652a89e9afa3760207914

Observation 684a2e35-f9b4-4673-823d-6c853d42bafb · inbound

Agent-Aided Design for Dynamic CAD Models cites this paper.

Agent-Aided Design for Dynamic CAD Models Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:15:10.441699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T11:13:01.515933Z digest=sha256:d46440f9182d35391b51a6378417f2fbcc3db9f8ef1a911e1ab7f9f1647aeada

Observation 276ae1ac-3d73-47c1-8eeb-35a9d0672089 · inbound

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems cites this paper.

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:48:01.663869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T08:45:54.303143Z digest=sha256:2aebba988eda41aa297c98f88e2ca6f3e27c58f590cd166020da63ed8b5e02a4

Observation 958a3f9d-f2a8-4c50-9831-c97aa925b0de · inbound

State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading cites this paper.

State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:16:26.452568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T11:45:57.291112Z digest=sha256:61ad2231d13dc46cd2610ebe521b0d4e8713833aadf9ac10e04a9d9285b3373c

Observation 2d20f7e4-2d5f-43a6-9df0-955f61c6ee48 · inbound

SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs cites this paper.

SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:56:45.651024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:43:51.900721Z digest=sha256:ca729610869d024a6be05182caec70401331fb646590d06edd05c37745a7cc7d

Observation 8263d4c3-7d96-4686-bf6a-5eedab9bdb88 · inbound

Do multimodal models imagine electric sheep? cites this paper.

Do multimodal models imagine electric sheep? Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.535821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T03:24:01.933339Z digest=sha256:5def794c611242ccc65c4c9f0a3722842758e720d0740d982d685e5362c21a24

Observation 82189804-e2b1-4d7a-9bb5-05fbf4032280 · inbound

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World cites this paper.

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:42:57.648243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:40:59.877854Z digest=sha256:9a93699c391d2d8a430396e1ec44756e846f860b0a83d46a17c399bb0ff4c6ca

Observation 4008e5b3-43d7-4822-adda-ec54db1b1db0 · inbound

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World cites this paper.

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.176444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T16:57:03.172340Z digest=sha256:28f05d44d6c38e8b09f07e29f527a2b22cebf2cbde3bd670da2eae55eb4bbd8e

Observation ebe4fb80-1f1c-49cd-835c-d0a57767b29e · inbound

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? cites this paper.

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:59:45.742535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T06:55:04.657347Z digest=sha256:ef923abcb4055ecff1effdb5d88d989aa8684d7980b8173cfe7c18dd601fbb8e

Observation 64dbe084-5d50-408c-acab-981e84e68547 · inbound

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? cites this paper.

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:54:57.791672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T17:52:26.785086Z digest=sha256:38c57529cc028539aa6bf6eff5ef609f1de599094aa3cb1095188e6e5596a691

Observation 2730b754-ea6a-4b0e-8042-3c4808aaee64 · inbound

VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images cites this paper.

VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:20:24.981533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:17:45.034348Z digest=sha256:e98f31de2d785f14bc03252e637d6048b36be08e9a3bf1dd8f43ac9c27024896

Observation 7577280a-8cca-43d9-a91f-8c3dab8ac721 · inbound

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving cites this paper.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:15:22.790290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:10:32.522453Z digest=sha256:eddd67f26410cb1f81212f85f0e16a26d5085f38426c5d1c3a66775548de0677

Observation d18c83e4-783f-4d0f-898b-e86c4781003e · inbound

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving cites this paper.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:44:56.045464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:40:22.441025Z digest=sha256:ba1ba58ef613bc00128d5ea28ab7698708ebe3f75146b8b4c1a91a7d54744f0e

Observation c0bcbcf9-3c93-4616-9b86-1034cf286b57 · inbound

PhotoFlow: Agentic 3D Virtual Photography Missions cites this paper.

PhotoFlow: Agentic 3D Virtual Photography Missions Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:35:21.022926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:33:24.355622Z digest=sha256:a71c2463a71d9f9d4ae81f8624f8c0f2f3afb09cb15f77626c3f90a6f095ed69

Observation 27886b22-3ddf-456c-9d55-27e77cd0e545 · inbound

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)? cites this paper.

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)? Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.577911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T07:48:19.295578Z digest=sha256:459563c076b2b71ae65adae178bab0fc35491148fa041798de41f792dd3332cc

Observation 63fd505f-af17-4c8f-b2e4-3dc7eb652d1b · inbound

Readable Yet Unpredictable: Rotated-Outcome Prediction in Vision-Language Models cites this paper.

Readable Yet Unpredictable: Rotated-Outcome Prediction in Vision-Language Models Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:26:18.105066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T15:21:16.342973Z digest=sha256:de434ae4cc8b431ac635a04716750d484516a93429b2b12db6a3c4f55fd4a4b1

Observation 68fdc119-0b54-4194-b323-bcbc6f4668fc · inbound

The Art of Interrogation: Consistency Amplifies Factuality in Spatial Reasoning cites this paper.

The Art of Interrogation: Consistency Amplifies Factuality in Spatial Reasoning Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:03.046297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T09:46:16.088490Z digest=sha256:911a2002b46c8cf1d4b8f02148b9ce2197380cb4f21aceab7af44f1bc0a66aef

Observation 75c0083b-0814-4336-b217-bedf8ab6c357 · inbound

OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning cites this paper.

OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:14:19.130332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T06:11:09.693576Z digest=sha256:bf9bd1321da3807cbae3dc63c2da84d72996840bf95369e7d39f47e3d443f7e2

Observation 706e41e4-1ee3-4f9c-85a8-d07ae0816756 · inbound

AeroVerse-SatAgent: UAV-Satellite Collaborative Spatial Reasoning Inspired by the Dual Visual Pathway Theory of Cognitive Neuroscience cites this paper.

AeroVerse-SatAgent: UAV-Satellite Collaborative Spatial Reasoning Inspired by the Dual Visual Pathway Theory of Cognitive Neuroscience Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:45:39.862470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T06:18:35.494240Z digest=sha256:201edd4a252e20db870cc77366e79b1a944d136c7788cef94f883ef248f30e95

Observation a6d702a7-d72f-4393-af86-8a1340ef24b8 · inbound

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models cites this paper.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.842655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:c762cc5e822a143981609909ca0d3c29f8769b491f5b68635b936590d795fe90

Observation a2f61674-1d1c-4f5f-9868-4bde23c20bfe · inbound

OrientSAM: Mitigating Camera-Centric Shortcut in Multimodal Spatial Reasoning via Orientation-Aware Spatial Alignment cites this paper.

OrientSAM: Mitigating Camera-Centric Shortcut in Multimodal Spatial Reasoning via Orientation-Aware Spatial Alignment Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T17:25:10.666535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:25:10.666535Z digest=sha256:ca6d6fbe489558971033511f4f10fcf2ac7b42517a9da052230309c07bf4f338

Observation ce3e6dab-3e1f-4d8f-a056-01001d16d180 · inbound

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text cites this paper.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.509832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.509832Z digest=sha256:9299b5c2430aa3e23583cf23a1d9adb336134ffbbe7956da9daeaf2eb9d05bbb

Observation 591a34d0-2c22-4a3f-8869-c2f4ae3232e4 · inbound

Foveated Probes Recover Localized Binding Information in Vision Foundation Models cites this paper.

Foveated Probes Recover Localized Binding Information in Vision Foundation Models Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:24:47.666879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:24:47.666879Z digest=sha256:cf5d79261ed9c31ac198974a467ecb41e11f66bde0331592abdc5d529443938d