Pith. sign in

Paper Citation Record · LEDGER

Vision language models are unreliable at trivial spatial cognition

As of 16 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2504.16061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16061 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:17:06.512196Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T16:18:15.917659Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T16:26:20.752416Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a357382b-e330-4dd0-9785-b4ffd30cc74d · outbound

This paper cites Amant, J Gregory Trafton, et al.

Vision language models are unreliable at trivial spatial cognition Amant, J Gregory Trafton, et al

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:07.007823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.408369Z digest=sha256:a95c83e2bfd441b3231699411b0db5e3e6296757ff11502f1b92824703bd41b5

Observation f3212152-2377-407c-8f96-e4782da0eaae · outbound

This paper cites Spatial reasoning.

Vision language models are unreliable at trivial spatial cognition Spatial reasoning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.995868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.412890Z digest=sha256:42d00fa2ad62a2a29e22d547a0673a40c327d567a0d4663f5843b44104396b58

Observation c13173e1-e7ea-4349-b5ab-7b82556d935e · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

Vision language models are unreliable at trivial spatial cognition SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.417789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.417789Z digest=sha256:1819841573325660b90c480923e02133adfce0fdbcc007a359bdfc40ca9b017f

Observation 86d68f78-f7db-4f96-bc6c-fa2f127ea92b · outbound

This paper cites Spa- tialvlm: Endowing vision-language models with spa- tial reasoning capabilities.

Vision language models are unreliable at trivial spatial cognition Spa- tialvlm: Endowing vision-language models with spa- tial reasoning capabilities

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.985035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.422254Z digest=sha256:edaa5652ec2fd7fd17a14dc005086190e34ea99628ad2848563fcf5efd615c79

Observation 3c787524-bae6-48bb-a611-bfe85bd05b99 · outbound

This paper cites Large language models are visual reasoning coor- dinators.

Vision language models are unreliable at trivial spatial cognition Large language models are visual reasoning coor- dinators

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.973691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.425782Z digest=sha256:5599cad115540ebe86abe9cc95992ee3a65511cd3833f2d12f38c304fbe70554

Observation 279db72f-026a-4e35-a91b-dd96ba8541ec · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

Vision language models are unreliable at trivial spatial cognition SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.429836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.429836Z digest=sha256:c3f00ea9e8bfef911bbb7b016bf2ce482932914b447944ae24f6d0b8d6f125d0

Observation b77c39ae-db78-44ce-b606-f7e7a0d9daa2 · outbound

This paper cites What makes mental modeling difficult? normative data for the mul- tidimensional relational reasoning task.

Vision language models are unreliable at trivial spatial cognition What makes mental modeling difficult? normative data for the mul- tidimensional relational reasoning task

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.962533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.433824Z digest=sha256:bee72bc29615a5133195ff1d12f55dc567bf65297cc4a47181e6790a52ef76c1

Observation 45e848a0-3ba5-4b6a-b9f9-4984f3f8d969 · outbound

This paper cites Spatial com- munication systems across languages reflect universal action constraints.

Vision language models are unreliable at trivial spatial cognition Spatial com- munication systems across languages reflect universal action constraints

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.952223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.437677Z digest=sha256:69c41cf87cfdce72868e2b1b40af2af50bc249edf21de8264c7a7f521f9d41d8

Observation 365d2330-92f4-4e46-944b-402a5fa80f1f · outbound

This paper cites Objaverse: A universe of annotated 3d ob- jects.

Vision language models are unreliable at trivial spatial cognition Objaverse: A universe of annotated 3d ob- jects

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.941305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.441113Z digest=sha256:79f3e90f21de9be290b7d26fe3523579c14f0f2467d3638dac7420778c386b19

Observation 59f2b213-f65a-4d5a-8dfe-bb83c2465e25 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Vision language models are unreliable at trivial spatial cognition Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.445081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.445081Z digest=sha256:e17f1841ebd06c3ca896b29cc687d010a7a58e103309fa9ccd37102cd83fb846

Observation d78ec83f-5de6-4d97-b4f2-ed4ad02f7744 · outbound

This paper cites Exploring the fron- tier of vision-language models: A survey of current methodologies and future directions.

Vision language models are unreliable at trivial spatial cognition Exploring the fron- tier of vision-language models: A survey of current methodologies and future directions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.448839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.448839Z digest=sha256:e9b463cf1d9bfa3cec47228a3152edb1deea6107433d95e75d52166fdc0323cd

Observation 1dc2dd4f-ed37-4fcf-9e9a-fd5fedb816d8 · outbound

This paper cites Spatial lan- guage and spatial representation.Cognition, 55(1):39– 84, 1995.

Vision language models are unreliable at trivial spatial cognition Spatial lan- guage and spatial representation.Cognition, 55(1):39– 84, 1995

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.929509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.452384Z digest=sha256:97805cde3608555fe55904ff71eaeffcb7e3aedad7b3fe38ae8f0f0f2c77d2e6

Observation 9bdba29e-97ae-425e-9e6c-b11b6cc2df26 · outbound

This paper cites Correctness comparison of chatgpt-4, gemini, claude-3, and copilot for spatial tasks.

Vision language models are unreliable at trivial spatial cognition Correctness comparison of chatgpt-4, gemini, claude-3, and copilot for spatial tasks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.917474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.455942Z digest=sha256:6ca41af09ceebeebbd7d037737695562ed2db0857f92423a500ba9c80d6559ba

Observation 6a3bac8e-e1a5-4b5a-b989-233a96832039 · outbound

This paper cites What's "up" with vision-language models? Investigating their struggle with spatial reasoning.

Vision language models are unreliable at trivial spatial cognition What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.459331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.459331Z digest=sha256:b6bb499a54160da47610c1ef5765bbbce4a48e36d3d08b311a49469ca93f43c3

Observation fb4dba2f-a4ca-4658-a531-6c77f33fa41b · outbound

This paper cites Space to reason: A spatial theory of human thought.

Vision language models are unreliable at trivial spatial cognition Space to reason: A spatial theory of human thought

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.905447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.462976Z digest=sha256:416e4022260972dfc98d82dc30570f29be0ad7f5c0654b70747b00047f7a4ad3

Observation 43d752fc-0c63-4975-a338-8b112bdd0a31 · outbound

This paper cites Whence and whither in spatial language and spatial cognition? Be- havioral and brain sciences, 16(2):255–265, 1993.

Vision language models are unreliable at trivial spatial cognition Whence and whither in spatial language and spatial cognition? Be- havioral and brain sciences, 16(2):255–265, 1993

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.894674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.466324Z digest=sha256:1d9c418cc4fca6a8b172922ad7793fc5bd02ecba3389d166e387d9ac1f66a99a

Observation a02cbd10-c82a-47ab-a211-84f641d0a0b4 · outbound

This paper cites What matters when building vision-language models?.

Vision language models are unreliable at trivial spatial cognition What matters when building vision-language models?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.469788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.469788Z digest=sha256:55b4ed3bd7781f6b0d0b5ccf8c3fcf549d055a1dc3c30696e9c50a9976e52f7e

Observation fe77a344-09ca-4285-b9bf-69d7366a100f · outbound

This paper cites BLIP: Bootstrapping language-image pre- training for unified vision-language understanding and generation.

Vision language models are unreliable at trivial spatial cognition BLIP: Bootstrapping language-image pre- training for unified vision-language understanding and generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.881789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.473387Z digest=sha256:1b0637e112b1d0c53fea6c1b2d2c4ed032a2ae1077bdd7c446b03893feca5b80

Observation ef0e08f4-99df-4b9b-afe0-1d96547d1f23 · outbound

This paper cites A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends.

Vision language models are unreliable at trivial spatial cognition A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.477196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.477196Z digest=sha256:a86d09d37e4ebfaaf28dbcce60cabb325261a343de6ed3cb482043eb69cff9f4

Observation c035cd08-07a2-450e-b9aa-ff4bc77294d7 · outbound

This paper cites Visual spatial reasoning.

Vision language models are unreliable at trivial spatial cognition Visual spatial reasoning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.869517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.481010Z digest=sha256:af2b217a1684cd2615f3082c0053661d51cd87764b32f67883b0c12ed214133e

Observation a825e976-011b-4371-983f-452a91303b21 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Vision language models are unreliable at trivial spatial cognition A Survey on Hallucination in Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.484550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.484550Z digest=sha256:e45e116d912464db423aa881c6c7a9eaec824873b7d5e96986ede59f109f5b05

Observation 2e236229-42f0-4454-9cca-8f285cd87247 · outbound

This paper cites Zero-shot visual reasoning by vision-language mod- els: Benchmarking and analysis.

Vision language models are unreliable at trivial spatial cognition Zero-shot visual reasoning by vision-language mod- els: Benchmarking and analysis

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.857728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.488056Z digest=sha256:64ef628172df9cbc3e7b0b45180868fe10e538b095ae47f624759cd9f1d148a5

Observation 1ecb8195-7411-402b-9ec5-44089044ae8a · outbound

This paper cites A theory and a computational model of spatial reasoning with pre- ferred mental models.

Vision language models are unreliable at trivial spatial cognition A theory and a computational model of spatial reasoning with pre- ferred mental models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.846222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.491502Z digest=sha256:39305a3e1765a3c534157c38bc2d0287b36a49b1762146e7e2669df38d363771

Observation da906f70-248b-4cbf-8bd5-825a4d8b1813 · outbound

This paper cites Sparkle: Master- ing basic spatial capabilities in vision language mod- els elicits generalization to composite spatial reason- ing.

Vision language models are unreliable at trivial spatial cognition Sparkle: Master- ing basic spatial capabilities in vision language mod- els elicits generalization to composite spatial reason- ing

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.494988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.494988Z digest=sha256:e8faec663f3061c92dede020e55a7997be1940c3969eef1d914ebd0f44ae54a4

Observation bfe9250d-697b-4c28-bdd9-53055b172d26 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Vision language models are unreliable at trivial spatial cognition Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.498562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.498562Z digest=sha256:0abdff740cea4f0d3a78d8091d9f7df0798cbe791faf63aed51f73c571a3961f

Observation 73f7707a-ccb3-4450-af0a-1d2dc1b946ca · outbound

This paper cites Learning phys- ical parameters from dynamic scenes.

Vision language models are unreliable at trivial spatial cognition Learning phys- ical parameters from dynamic scenes

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.834817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.501992Z digest=sha256:b5e3c2fb0b666fd3f419fa6af460cacf77207dda9f1a59bb6fa4988907291a38

Observation c69e8c64-5800-4136-bf54-54f998527666 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?, 2023.

Vision language models are unreliable at trivial spatial cognition When and why vision-language models behave like bags-of-words, and what to do about it?, 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.823137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.505609Z digest=sha256:ed6f5910f195f55305e1e748f2e8c1df0c58389e7ec93ae9219f243bb2657f31

Observation ca3535d1-a0d7-43fe-a784-cfca6ba4d0dd · outbound

This paper cites Vision-language models for vision tasks: A sur- vey.

Vision language models are unreliable at trivial spatial cognition Vision-language models for vision tasks: A sur- vey

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:17:06.809897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:17:06.508700Z digest=sha256:0bb820a291e0862ecc9d73f157b7cebfb85911de2afaaccef79a85e25fba7d88

Observation ec4b1981-974c-4b0c-896d-d2fb7111b640 · outbound

This paper cites Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities.

Vision language models are unreliable at trivial spatial cognition Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.512196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.512196Z digest=sha256:db128653c3359c46590c7f2ed26c553932ad19024e53b04bd7e590056ff32477

Pith citing papers

Observation 400f398f-56c9-4e8a-b2d4-a96bd4f7b03c · inbound

Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models cites this paper.

Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models Vision language models are unreliable at trivial spatial cognition

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-09T16:26:20.757332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-09T16:18:15.917659Z digest=sha256:97f506c94ee8ef6f79b06ebaa70bcecd275aaf5a77685b1a79e7c99f8d58b15a