Pith. sign in

Paper Citation Record · LEDGER

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance

As of 14 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2608.00502.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00502 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:54:48.841652Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact3
  • verified fuzzy3
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4c5d8d37-3172-49eb-a030-76bbe9e0695d · outbound

This paper cites Qwen3-VL Technical Report.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.786424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.786424Z digest=sha256:b7807325b6ea40f2faf29c592c07cbc669ef430b48ddeba2a165d7b9431dc0f8

Observation fe9811cd-c017-4fd8-bc46-6ae4eec4e4b8 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.790316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.790316Z digest=sha256:d7ea92be53cbfc2bd3cdf783c904119e475938828a18b801420ef118a7e8cc96

Observation 1d74063a-18b6-467c-9346-ec035d883fc5 · outbound

This paper cites One-Shot Affordance Detection.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance One-Shot Affordance Detection

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-05T00:54:49.129302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T00:54:48.811578Z digest=sha256:86a7dc4eb174eb47c197fe9c357f01c1bd407ed0f00cc0217121fb6e32df67b0

Observation aa4483fb-ef56-478a-89eb-4ce8a5f5083c · outbound

This paper cites Qwen2.5-VL Technical Report.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Qwen2.5-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.817378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.817378Z digest=sha256:2d836dbce78b2eabe6ae6bad3f0c9cde6879d703a1764e52ec5c2b34cbd37e6b

Observation 8ceeb62a-a73c-4ed5-9df3-906602311fc8 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.823565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.823565Z digest=sha256:3cb5ff503ce53f6cf0673a764c3e2a8f02c65a1e9b48716d6c401f2aea9c595e

Observation b4127c37-5561-4372-bae1-2f28034ef8cf · outbound

This paper cites Aligning large multimodal models with factually augmented rlhf.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Aligning large multimodal models with factually augmented rlhf

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:54:49.198559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T00:54:48.826852Z digest=sha256:0b0f07acb3284533fd7fc231af89261efb96dbae67d2c79d60010c4e203e3d2d

Observation a08943d3-b8cf-46ad-87ba-8b807a0a6bdb · outbound

This paper cites RoboBrain 2.5: Depth in sight, time in mind.arXiv preprint arXiv:2601.14352,.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance RoboBrain 2.5: Depth in sight, time in mind.arXiv preprint arXiv:2601.14352,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.829656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.829656Z digest=sha256:f909836e601a8346654c962e9c1f7b9a5c9c4f4fe4244318518b29698c893b91

Observation ce239547-b0fe-49b5-8120-2637e6fedde0 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.835533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.835533Z digest=sha256:f313e572d3ce3eae8fa33f525e290ecb11e4e7a096d1a3666e607588ed5a90b8

Observation 47da41f5-3b1f-4c2d-bce4-0c8231d80e48 · outbound

This paper cites PartAfford: Part-level Affordance Discovery from 3D Objects.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance PartAfford: Part-level Affordance Discovery from 3D Objects

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.838732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.838732Z digest=sha256:6646d7600f5790009e3cf3cdfdab3b95b93cc6cebf5baf8615cf9e88acf17d2a

Observation 603f4ed2-4932-46af-8a9f-68303b365f65 · outbound

This paper cites Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.841652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.841652Z digest=sha256:0d0d2c475c65965667797a6b6d3f459040ba0776a8c5e132a193ab98ddca79bf

Observation c4d9ff90-75a3-45c6-b2d8-b90ff19a7653 · outbound

This paper cites Mgpo: Thinking with images via multi- turn grounding-based reinforcement learning.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Mgpo: Thinking with images via multi- turn grounding-based reinforcement learning

Reference 1979

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:54:49.218206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T00:54:48.797363Z digest=sha256:d6aa8658c104963656101de36a448a9b22c1d6d5bc19c08622d26de4a13a7747

Observation 34304f19-c3d3-4887-9ef7-3aadb0dfa4fd · outbound

This paper cites Do MLLMs really see it: Reinforcing visual attention in multimodal LLMs.arXiv preprint arXiv:2602.08241,.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Do MLLMs really see it: Reinforcing visual attention in multimodal LLMs.arXiv preprint arXiv:2602.08241,

Reference 2015

Resolution
verified exact
raw_fallback, observed 2026-08-05T00:54:49.115128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T00:54:48.814582Z digest=sha256:6f6a6614b9873a6f07878363eb8748f27518d24c6024a042ae3f369be8d801bc

Observation 9d870beb-a6c7-4cc2-947e-3c64e26ccd69 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.820591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.820591Z digest=sha256:4aa7c2f4929f9568fc5ac2296e79b207e991d7ca435b7bad34c620d48550acee

Observation bfc3d2c6-72be-4554-a5b9-3b0d19f72fe9 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance PaLM-E: An Embodied Multimodal Language Model

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.793928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.793928Z digest=sha256:9ad81537b907956b02cea7ff2bc82bff72947f7223b76a00b2f2b74cd2bdd12f

Observation 97d9db3f-cfeb-4faf-b572-984c4dfa35f4 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.800997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.800997Z digest=sha256:eae72c0266e478b08b497b136c74a3631c09a6b932528abf8ecace1fa1ebf655

Observation 08652df0-1ecb-4470-b233-cf5c238f4dd9 · outbound

This paper cites Improved baselines with visual instruction tuning.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Improved baselines with visual instruction tuning

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:54:49.208446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T00:54:48.808322Z digest=sha256:0635ad46670e84b4717ceea60312a73fc4685d87da2d3143edc77e093e4f65de

Observation 7f743f60-859f-4db7-a84e-11729fbe09f1 · outbound

This paper cites Token-Based Affordance Grounding with Large Vision-Language Models.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Token-Based Affordance Grounding with Large Vision-Language Models

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-05T00:54:49.142536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T00:54:48.804739Z digest=sha256:6fe2e3e62329a9ac824919b0d287e14184f170d13937a539a17a8a5cc5be7703

Observation 29606749-d338-4124-8d6e-841411590730 · outbound

This paper cites RoboBrain 2.0 Technical Report.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance RoboBrain 2.0 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.781637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.781637Z digest=sha256:d1b87c949687b30c0ddc58166e470ba175701f8b99dc2ac28e46328e3b07d4f6

Observation 9323f021-c714-4397-a258-cbc5f52e805d · outbound

This paper cites Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models.

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:48.832527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:54:48.832527Z digest=sha256:a3cc1de9ada2353685441d4512593f554e4279bd15704ae84e3267dae8251859

Pith citing papers

No inbound Pith citation observations are available.