Pith. sign in

Paper Citation Record · LEDGER

SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2401.12168.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.12168 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:33:58.479385Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T03:06:43.700595Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 48438ee1-5f7e-4dd8-b156-f1dc65cc05a8 · inbound

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations cites this paper.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.385412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.385412Z digest=sha256:0776afd71cb8bec176f995fe3ce7b7c8835862150869db707d309147a6565c4a

Observation 98641f75-569f-4748-b8a6-73654f86fa24 · inbound

Explainability for Vision Foundation Models: A Survey cites this paper.

Explainability for Vision Foundation Models: A Survey SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 131

Resolution
unresolved
no resolver link, observed 2026-08-10T17:26:35.470855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:26:35.470855Z digest=sha256:cda037968e011b86ef8cd1c5a7fac0d17eb7c026af9411becf8880c61862b46f

Observation a45179ff-a548-46c8-bfcc-2f6311b025a4 · inbound

Learning the RoPEs: Better 2D and 3D Position Encodings with STRING cites this paper.

Learning the RoPEs: Better 2D and 3D Position Encodings with STRING SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T11:50:03.189919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:50:03.189919Z digest=sha256:72102e5d344a684cb160375a01048434bac315cf27d12f712d53e4ae44bbf9f5

Observation ecee111c-99a8-4f8e-82a8-7493c4edd222 · inbound

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT cites this paper.

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:36:20.778339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:36:20.778339Z digest=sha256:1fd890b78abab79f24ca967c066474685cdf2e61f33a20ab71416049284871a8

Observation a82e9a24-ed16-4b9a-a427-25cd479fe798 · inbound

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization cites this paper.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:11.687427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:11.687427Z digest=sha256:c155c32538b76ecdc1cacc5c5304843698e629990dbdfbb45d9200fb4c7c7407

Observation e1d998d1-f294-4f63-905f-d4b06ff9f1f6 · inbound

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models cites this paper.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.657856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.657856Z digest=sha256:dceb53ef5ec94ca21658f7e02d015b214bc566610244ea6e5d00d58dd2687a87

Observation e7f85b5b-2a86-40c6-b069-8646390f3e0c · inbound

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning cites this paper.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.162766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.162766Z digest=sha256:df76f9cf8696f2dff3f7dbd2ce891e543a6db6a2706bc056b7664f80019d36c8

Observation d56ca6d2-3da4-48fc-83ca-1c00617ec8ad · inbound

Sustainability assessment using multimodal AI agents cites this paper.

Sustainability assessment using multimodal AI agents SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:03:27.082032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:03:27.082032Z digest=sha256:c7c03c1112737ecf3511871a0b86b6af8316e0bf87fad9ce3b0fb5fd0943cc54

Observation 033b7dc6-1001-45df-ae01-844f0722ad04 · inbound

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks cites this paper.

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:07.947697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:07.947697Z digest=sha256:3141dc5116923bd693cbe1b0bacd2a1536ff07e4bb41e108f42c418359bf7dc5

Observation 6e0d4a50-0c88-4a17-9bce-c82ec548426a · inbound

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization cites this paper.

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:16:09.680303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T15:12:23.459579Z digest=sha256:3239a43e0dc52a1f77caf049d2abf7e44d343f3d15d346ec912ce002adead608

Observation dc63545e-e5f0-43f7-a984-90ee3e098ed3 · inbound

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning cites this paper.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:d522ff5f029daef391b168438f3f1bd50a405bfd758ff74f291e439e67baf6ac

Observation 1f806e86-4c61-4cc9-b2f2-229ee754e391 · inbound

Multimodal Language Models Cannot Spot Spatial Inconsistencies cites this paper.

Multimodal Language Models Cannot Spot Spatial Inconsistencies SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:13:24.784848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-13T23:12:27.333405Z digest=sha256:3064c0a80ca32b40712a962a5b2c89d960477a9a90db533b42a8fbcbf498ba93

Observation 4042bd62-411d-4827-a6a5-d300abfe0de5 · inbound

Exploring Spatial Intelligence from a Generative Perspective cites this paper.

Exploring Spatial Intelligence from a Generative Perspective SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:49:48.849225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T00:45:46.261005Z digest=sha256:7dc7d0c5caf37307479addb41fddcf8b9fd70b4f2e98b72ce4269bb3ad8d32dd

Observation 8115d2dd-7059-40e7-ae26-f6c89343a181 · inbound

Fast Core Identification cites this paper.

Fast Core Identification SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:01:14.738363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T07:11:30.553696Z digest=sha256:43bd0e85827401d805b231e384bb4bed7f937f1bcd90391399a9ce1c14643975

Observation 33c744aa-d26f-4a17-a6a6-16932b24b563 · inbound

Latent State Design for World Models under Sufficiency Constraints cites this paper.

Latent State Design for World Models under Sufficiency Constraints SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:04.461629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:55:31.825583Z digest=sha256:3488345344422990474667d847582db93ec38148b367670a18f8b357b5aa7cb7

Observation 69da90aa-aaa9-48f4-b247-5247b855e41d · inbound

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World cites this paper.

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:42:57.776205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T20:40:59.877854Z digest=sha256:676e66a406d7fcf928b427fd1f296a047ba24733c5d1a5e6446448e8b5b551bf

Observation 23a5b44b-64c5-49ad-be51-9ce0c3695052 · inbound

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World cites this paper.

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.149820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T16:57:03.172340Z digest=sha256:6fde6b292a685f707535f3bbd5b565aa6673f6178eb9864fdc9a20abbb3761a1

Observation b696d70e-7b11-4944-861a-72ea37bed505 · inbound

Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning cites this paper.

Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:13.715412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T17:46:21.821764Z digest=sha256:622b47b17a9e8b37eabf4fa6cab9dc2c2383399192897b2fbd9807c68e1c16c5

Observation 8b054447-c513-4c7b-96e8-94c20819320c · inbound

GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models cites this paper.

GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.202939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T10:04:10.627420Z digest=sha256:61ddc7be5e4d0f609b8fadb416273010c1d939cec2ba845f270b2a6ceaf99f59

Observation c1ccb55f-fade-44c5-a941-4f5298ab4061 · inbound

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching cites this paper.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.141917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:648495504fd71b9a9553e4e6f65c393e0f06c3d154fe474d2de87eeca4c3d297

Observation 19c5ae1c-b4f1-42ca-830b-a59f7df338e2 · inbound

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video cites this paper.

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:57.214345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T02:16:25.730555Z digest=sha256:5a6112af01cf7309dc3fa4a7ee4ac87630bd0b41b95ec60f4012fc76252eb736

Observation 63d4cc74-765a-4972-8a29-1c166fdb8eb9 · inbound

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models cites this paper.

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:57.509150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:13:53.219095Z digest=sha256:d3f7ac82a33fbe796017a01d42a1db566972d25226f2909bf1e46f715a121e34

Observation e2c309cb-eb6c-44c3-bca1-bf0e98e3f9da · inbound

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models cites this paper.

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:08:55.661477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T01:42:30.005911Z digest=sha256:1a94e093735b7cef8165e42978dcbbf826d8efa109922cfed4ce5864981063ff

Observation 71a61b51-b882-40bd-952c-b3fcc22ebe5f · inbound

HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory cites this paper.

HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:49:46.492839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T08:27:15.889885Z digest=sha256:b08aa69683bdb6b7956dcab520aa1d3bbb9151363460602c4c69893e9ab62603

Observation abacc9ee-aef9-4d82-8fd3-78487bc291c6 · inbound

Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping cites this paper.

Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:55:41.219848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-01T06:01:37.880943Z digest=sha256:789719a91780cf9fef25d82c93d8c722d04e8959addcc6bdd1d32248042ddd73

Observation 6861a3c7-0b02-429d-b5b7-34abebf4b381 · inbound

Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference cites this paper.

Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-10T03:06:43.701881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T02:59:21.353359Z digest=sha256:48d2ee10cedb7d50c74d275f12c83cbaf96dc7cd46c2a29bf9fb554a2cc80975

Observation e5a2891b-07d2-407c-b3a2-24b431e394d9 · inbound

ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams cites this paper.

ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T07:19:14.164642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:19:14.164642Z digest=sha256:13e63976f154865f118d127cc8b23f8c3293b7c9ef7eec943a4eac7b675bf916

Observation c5f835b9-a579-4434-bb92-bb0e2de903a8 · inbound

Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO cites this paper.

Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T18:44:50.379910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:44:50.379910Z digest=sha256:bc2a4ad614120aa643c4b9aad01c9554f060bdb29f9672ed49e72e46b034f7f5

Observation 703e5350-68a9-4d22-a1c6-6143e6ef89bf · inbound

PhysX-CoT: Structured Physical Reasoning from a Single Image to Simulation-Ready 3D Assets cites this paper.

PhysX-CoT: Structured Physical Reasoning from a Single Image to Simulation-Ready 3D Assets SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:58.479385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:58.479385Z digest=sha256:38760416d65e8a7e625b662a3fa03e691c13d014718d14cd194623274ccb62b1