Pith. sign in

Paper Citation Record · LEDGER

Vision language models are unreliable at trivial spatial cognition

As of 16 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2504.16061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16061 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:17:06.512196Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T16:18:15.917659Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T16:26:20.752416Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a357382b-e330-4dd0-9785-b4ffd30cc74d · outbound

This paper cites Amant, J Gregory Trafton, et al.

Vision language models are unreliable at trivial spatial cognition Amant, J Gregory Trafton, et al

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:07.007823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.408369Z digest=sha256:ac5a3f840d234a2bc32a4c0b624e92a05f0f2428cc21316296feb0ab29dae033

Observation f3212152-2377-407c-8f96-e4782da0eaae · outbound

This paper cites Spatial reasoning.

Vision language models are unreliable at trivial spatial cognition Spatial reasoning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.995868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.412890Z digest=sha256:c04546cd996bb7abb5ebd8b528f4d1a207f45a164f5a7f2a39edc6e585e52d3c

Observation c13173e1-e7ea-4349-b5ab-7b82556d935e · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

Vision language models are unreliable at trivial spatial cognition SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.417789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.417789Z digest=sha256:49a9d9014e8301568c7cc1d4a13c27fafc740a0a671493673c8ee49d76834665

Observation 86d68f78-f7db-4f96-bc6c-fa2f127ea92b · outbound

This paper cites Spa- tialvlm: Endowing vision-language models with spa- tial reasoning capabilities.

Vision language models are unreliable at trivial spatial cognition Spa- tialvlm: Endowing vision-language models with spa- tial reasoning capabilities

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.985035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.422254Z digest=sha256:675f3d4b106d40735a18174b334023776f48f012ffe26af2649af9f3c64fb192

Observation 3c787524-bae6-48bb-a611-bfe85bd05b99 · outbound

This paper cites Large language models are visual reasoning coor- dinators.

Vision language models are unreliable at trivial spatial cognition Large language models are visual reasoning coor- dinators

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.973691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.425782Z digest=sha256:aacf972baf9e1b70fc43d3ffeb3134ece9da3abf6082572c68987f61e0035b1e

Observation 279db72f-026a-4e35-a91b-dd96ba8541ec · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

Vision language models are unreliable at trivial spatial cognition SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.429836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.429836Z digest=sha256:9b12ec4229cc807473e69378198e0163b4773595a43bb09c9697df211d79f7b4

Observation b77c39ae-db78-44ce-b606-f7e7a0d9daa2 · outbound

This paper cites What makes mental modeling difficult? normative data for the mul- tidimensional relational reasoning task.

Vision language models are unreliable at trivial spatial cognition What makes mental modeling difficult? normative data for the mul- tidimensional relational reasoning task

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.962533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.433824Z digest=sha256:cfc0f1a21714f7920357301f07a115f49122fb3eba510b241be6f0f8a6eb6bc6

Observation 45e848a0-3ba5-4b6a-b9f9-4984f3f8d969 · outbound

This paper cites Spatial com- munication systems across languages reflect universal action constraints.

Vision language models are unreliable at trivial spatial cognition Spatial com- munication systems across languages reflect universal action constraints

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.952223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.437677Z digest=sha256:460901ea7e5b30653407b37725c093c3a5213f25ea6a433751ebc633ffea27db

Observation 365d2330-92f4-4e46-944b-402a5fa80f1f · outbound

This paper cites Objaverse: A universe of annotated 3d ob- jects.

Vision language models are unreliable at trivial spatial cognition Objaverse: A universe of annotated 3d ob- jects

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.941305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.441113Z digest=sha256:352472596ffd2bb63644cabe241c0aa2a43f1844b8dd2e25ec0cf9d8621e43b4

Observation 59f2b213-f65a-4d5a-8dfe-bb83c2465e25 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Vision language models are unreliable at trivial spatial cognition Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.445081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.445081Z digest=sha256:ad8b58b05b124f2cd3a20cadf2908cde8a78c95d7b0a5e789e18e6c7ebb0b721

Observation d78ec83f-5de6-4d97-b4f2-ed4ad02f7744 · outbound

This paper cites Exploring the fron- tier of vision-language models: A survey of current methodologies and future directions.

Vision language models are unreliable at trivial spatial cognition Exploring the fron- tier of vision-language models: A survey of current methodologies and future directions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.448839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.448839Z digest=sha256:f266f7b0cc2e408d38ba1f87052670e939dd95c60d6ac68e6923490b4011b120

Observation 1dc2dd4f-ed37-4fcf-9e9a-fd5fedb816d8 · outbound

This paper cites Spatial lan- guage and spatial representation.Cognition, 55(1):39– 84, 1995.

Vision language models are unreliable at trivial spatial cognition Spatial lan- guage and spatial representation.Cognition, 55(1):39– 84, 1995

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.929509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.452384Z digest=sha256:8bafc97789381f24c14643346580b52b01faf50d7c039a9702f10a1dcc43d846

Observation 9bdba29e-97ae-425e-9e6c-b11b6cc2df26 · outbound

This paper cites Correctness comparison of chatgpt-4, gemini, claude-3, and copilot for spatial tasks.

Vision language models are unreliable at trivial spatial cognition Correctness comparison of chatgpt-4, gemini, claude-3, and copilot for spatial tasks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.917474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.455942Z digest=sha256:98f17c780882ea0d6348d043a21195abfb52c9e501d3e411a2ada5bbd91def35

Observation 6a3bac8e-e1a5-4b5a-b989-233a96832039 · outbound

This paper cites What's "up" with vision-language models? Investigating their struggle with spatial reasoning.

Vision language models are unreliable at trivial spatial cognition What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.459331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.459331Z digest=sha256:0f1f8d93a1749ffb2981f5b8ec1b9f9f24b9410d2ad4363f1d6d5403f58eadff

Observation fb4dba2f-a4ca-4658-a531-6c77f33fa41b · outbound

This paper cites Space to reason: A spatial theory of human thought.

Vision language models are unreliable at trivial spatial cognition Space to reason: A spatial theory of human thought

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.905447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.462976Z digest=sha256:22eb0304cfc01a79e127698ed55f461bf6c8de26e60877716b8aa5179f650b1b

Observation 43d752fc-0c63-4975-a338-8b112bdd0a31 · outbound

This paper cites Whence and whither in spatial language and spatial cognition? Be- havioral and brain sciences, 16(2):255–265, 1993.

Vision language models are unreliable at trivial spatial cognition Whence and whither in spatial language and spatial cognition? Be- havioral and brain sciences, 16(2):255–265, 1993

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.894674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.466324Z digest=sha256:a63b4fd10e9c1db8dbba569524d144d24ca2b42f57e14901a35280629f42156b

Observation a02cbd10-c82a-47ab-a211-84f641d0a0b4 · outbound

This paper cites What matters when building vision-language models?.

Vision language models are unreliable at trivial spatial cognition What matters when building vision-language models?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.469788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.469788Z digest=sha256:8eb6e97cc811f4891e99c71be2445e238bde32dbe4284fccfb01c0b8b13468f6

Observation fe77a344-09ca-4285-b9bf-69d7366a100f · outbound

This paper cites BLIP: Bootstrapping language-image pre- training for unified vision-language understanding and generation.

Vision language models are unreliable at trivial spatial cognition BLIP: Bootstrapping language-image pre- training for unified vision-language understanding and generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.881789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.473387Z digest=sha256:8f784384ff87cbe1a382d3f498b60ebef9769d861b0242bbfc0bf6cd27200fd2

Observation ef0e08f4-99df-4b9b-afe0-1d96547d1f23 · outbound

This paper cites A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends.

Vision language models are unreliable at trivial spatial cognition A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.477196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.477196Z digest=sha256:5688d481c5e8536b8264c8428d295f3266621f97e6c1f77265978ac84beeedfc

Observation c035cd08-07a2-450e-b9aa-ff4bc77294d7 · outbound

This paper cites Visual spatial reasoning.

Vision language models are unreliable at trivial spatial cognition Visual spatial reasoning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.869517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.481010Z digest=sha256:d0a0f25f0b04bbad30855fcac3cfef5308633e90c2761ab0783e703a14bd725e

Observation a825e976-011b-4371-983f-452a91303b21 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Vision language models are unreliable at trivial spatial cognition A Survey on Hallucination in Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.484550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.484550Z digest=sha256:d34a8d30d0527cf834f30798f85b043dcb5465b3b593dc5fd976210f8e5ae619

Observation 2e236229-42f0-4454-9cca-8f285cd87247 · outbound

This paper cites Zero-shot visual reasoning by vision-language mod- els: Benchmarking and analysis.

Vision language models are unreliable at trivial spatial cognition Zero-shot visual reasoning by vision-language mod- els: Benchmarking and analysis

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.857728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.488056Z digest=sha256:c23bf678fb81c2d12b3ef1b2f16509f8c74d0355fc5028415286695b2ed75165

Observation 1ecb8195-7411-402b-9ec5-44089044ae8a · outbound

This paper cites A theory and a computational model of spatial reasoning with pre- ferred mental models.

Vision language models are unreliable at trivial spatial cognition A theory and a computational model of spatial reasoning with pre- ferred mental models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.846222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.491502Z digest=sha256:4a8937bfdaa62118f275813d63c227d27fc8de128c053ba636c7a4fbebe06707

Observation da906f70-248b-4cbf-8bd5-825a4d8b1813 · outbound

This paper cites Sparkle: Master- ing basic spatial capabilities in vision language mod- els elicits generalization to composite spatial reason- ing.

Vision language models are unreliable at trivial spatial cognition Sparkle: Master- ing basic spatial capabilities in vision language mod- els elicits generalization to composite spatial reason- ing

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.494988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.494988Z digest=sha256:3f7b5ed43e79c08649db105372a88a1253e2269943e232dace70fe798fa5a372

Observation bfe9250d-697b-4c28-bdd9-53055b172d26 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Vision language models are unreliable at trivial spatial cognition Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.498562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.498562Z digest=sha256:58f7d1094f1b26a0d8d3fdbacea02bc853420735e4cd3f2bc96f9da54a39558b

Observation 73f7707a-ccb3-4450-af0a-1d2dc1b946ca · outbound

This paper cites Learning phys- ical parameters from dynamic scenes.

Vision language models are unreliable at trivial spatial cognition Learning phys- ical parameters from dynamic scenes

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.834817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.501992Z digest=sha256:2237974c6dc3bb6565c6c42cfbb49b915e91e17d3cbb013e1a57e2492a823e9d

Observation c69e8c64-5800-4136-bf54-54f998527666 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?, 2023.

Vision language models are unreliable at trivial spatial cognition When and why vision-language models behave like bags-of-words, and what to do about it?, 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.823137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.505609Z digest=sha256:ad9e11d73cd00c767f819f42b36bca8e1bb14356d58d0b837ccc6a5107319b19

Observation ca3535d1-a0d7-43fe-a784-cfca6ba4d0dd · outbound

This paper cites Vision-language models for vision tasks: A sur- vey.

Vision language models are unreliable at trivial spatial cognition Vision-language models for vision tasks: A sur- vey

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.809897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.508700Z digest=sha256:7b905d02bf4519be182cd77995db6638b56da6f3df7154480b86d301303a4b7c

Observation ec4b1981-974c-4b0c-896d-d2fb7111b640 · outbound

This paper cites Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities.

Vision language models are unreliable at trivial spatial cognition Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.512196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.512196Z digest=sha256:26a7b0afcaae86fcba117bc63a49e5050e99ab4eef666e1bc144261458507a98

Pith citing papers

Observation 400f398f-56c9-4e8a-b2d4-a96bd4f7b03c · inbound

Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models cites this paper.

Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models Vision language models are unreliable at trivial spatial cognition

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-09T16:26:20.757332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-09T16:18:15.917659Z digest=sha256:707e65b2d0dcaade90acbe825b2de02485f4305e0193c9880f4a4607af4de30d