Pith. sign in

Paper Citation Record · LEDGER

Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2410.17385.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.17385 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:37:03.205845Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T21:07:24.216562Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e58a5a0d-5842-4afd-bd71-a9b4930b7537 · inbound

AutoSpatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning cites this paper.

AutoSpatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:27:17.808374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T00:26:58.273861Z digest=sha256:54c7cad1c41d75abcb12af3c8694cc5869ed988c65653be3c59d4d070c00e705

Observation 1697cb18-7585-47e4-ae01-aecf876f8eac · inbound

Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes cites this paper.

Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:37:03.205845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:37:03.205845Z digest=sha256:876e22ef13b76b51d50a355ccd05372079f3e93ee8deaa6ebfd959fe1fe23e5b

Observation ec4b1981-974c-4b0c-896d-d2fb7111b640 · inbound

Vision language models are unreliable at trivial spatial cognition cites this paper.

Vision language models are unreliable at trivial spatial cognition Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:06.512196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:06.512196Z digest=sha256:26a7b0afcaae86fcba117bc63a49e5050e99ab4eef666e1bc144261458507a98

Observation c6ecd8fb-e5d4-45e4-a9ae-ae54aa22204e · inbound

SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data cites this paper.

SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T05:29:57.046039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:29:57.046039Z digest=sha256:1174a6d9d225ea6728c100ae0a0f1e1c3dd613caa5c0ff083fb9618a295bae13

Observation c809b1ae-52e3-4d17-9928-ba1826c2861f · inbound

GenSpace: Benchmarking Spatially-Aware Image Generation cites this paper.

GenSpace: Benchmarking Spatially-Aware Image Generation Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:21:26.903958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:21:26.903958Z digest=sha256:9d9573520e613b7ea1e987cc7ea8888aac9f5480dbc951fc9f86fcca24cf4d6d

Observation 5d8d4ce8-1ead-46aa-9a7a-4320a09cc668 · inbound

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills cites this paper.

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:37.513161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:37.513161Z digest=sha256:7a81d61dca0a96d5a6690a0140e9b713dbcaafe6e8d9420bd89407ef14054630

Observation 6cf7047d-dbdd-4062-b0e6-73561138ce58 · inbound

LLM-Based Social Simulations Require a Boundary cites this paper.

LLM-Based Social Simulations Require a Boundary Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T18:28:06.277031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:28:06.277031Z digest=sha256:8dcbba1c8147e1ff994f20e962cc178001d3e849ca6a140b7e7ccbe427b5c4fb

Observation 5d45d58b-8271-4058-967f-89ce26013993 · inbound

Augmented Vision-Language Models: A Systematic Review cites this paper.

Augmented Vision-Language Models: A Systematic Review Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:44.971203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:44.971203Z digest=sha256:e30780cb809887b1a0dc5c49393d794479b0714dca30a75350d59e6b422e6cd6

Observation c9abb6af-23de-4bc7-9c52-8f506d4a8604 · inbound

When Do Diffusion Models learn to Generate Multiple Objects? cites this paper.

When Do Diffusion Models learn to Generate Multiple Objects? Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:15:31.946430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-01T08:07:10.345273Z digest=sha256:f6d4e9e271819a5c0cfdb3e01ecac1e5b2b115d438b52d0570de50daa8eff1a4

Observation 64b0d813-e0cc-449e-95eb-d68b53d55d79 · inbound

ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs cites this paper.

ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:23:59.865695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T22:23:38.178195Z digest=sha256:4d73f40c57a6e6df25df370340d571d462cd0739baa9fcf711abde6262ad481f

Observation 91c09a56-d392-490c-a84f-47e490910923 · inbound

IntentNav: Learning Spatial-Visual Object Navigation from Human Demonstrations cites this paper.

IntentNav: Learning Spatial-Visual Object Navigation from Human Demonstrations Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:07:24.218032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T19:54:58.135809Z digest=sha256:7f7f43e09cb4d804698d1f4ad9823b9fe03183b954bb60d5ff06d82bea64e673

Observation 1e17cee6-cb6a-4eac-bc91-2b58da4e2a1a · inbound

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping cites this paper.

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:17:02.402742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-02T14:16:39.649823Z digest=sha256:545c816790e060a4455b3e47a9b6f0a32e17ad550bfa65d3265716de8251b99b