Pith. sign in

Paper Citation Record · LEDGER

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning

As of 13 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2606.25360.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.25360 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-25T21:23:44.253416Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T15:45:43.483298Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa56e012-6ce6-40a4-b48b-d5e80e006be1 · outbound

This paper cites π 0: A vision-language-action flow model for general robot control,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning π 0: A vision-language-action flow model for general robot control,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:799cd6a8a20a950fbb7e4ae42071b26e1b1801dfab3f05f6037ae7a66766d27e

Observation 034c7b5b-f9de-4e31-82d0-fc7538c38318 · outbound

This paper cites OpenVLA: An open-source vision-language-action model,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning OpenVLA: An open-source vision-language-action model,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:1f9e2b8b1546e7d7ee88afb376456b4b6f816c591aed6b0542886c44a2373dfd

Observation c6e1ab2f-2136-4386-9e4c-5fea8e64ece3 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning SAM 3: Segment Anything with Concepts

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:20:07.157528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:d9f9dc4af558d8bfcbccf0171305ae2152f6fc4ab10f05677fadcb86f64da243

Observation 74b78b5e-a78e-4084-aa7a-c55801a491c4 · outbound

This paper cites ALVINN: An autonomous land vehicle in a neural network,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning ALVINN: An autonomous land vehicle in a neural network,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:f41af362cbf3d881432ece22985533fb95fab67f2a790cd2b49f7e88f098f1df

Observation 3239d6f8-2b3f-4200-a385-b3431935b6c8 · outbound

This paper cites Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:07.152003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:161f45dc1a08cdcb51501d92884ddb3bc2078566499692e503295f9665388f9a

Observation bb255f11-f354-47d4-84f7-7d2e69fc3472 · outbound

This paper cites Learning fine-grained bimanual manipulation with low-cost hardware,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Learning fine-grained bimanual manipulation with low-cost hardware,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:54e5d3c980f3611c85a9c863999a8e3f490a94a776645685c0136b91d3449f3b

Observation 6a2444a1-2ec6-403e-b4a2-ba672a804eca · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Diffusion policy: Visuomotor policy learning via action diffusion,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:063de4b2d06b3e9e43dd5e9fa559546761156ad95fc8a2c9b824acdc2f62aee6

Observation 0d921627-8eff-4cbe-a854-7f02fd358531 · outbound

This paper cites FiLM: Visual reasoning with a general conditioning layer,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning FiLM: Visual reasoning with a general conditioning layer,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:416184ed037319995d33823866117709d8116178c7dba83a7b27db5f247bddc9

Observation 988e34b6-fd3d-4425-9b4a-a98309529043 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Learning transferable visual models from natural language supervision,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:fd41ca887710666fc71ab0f06b82d186a6086068973d909584f709f753a4c013

Observation fb07ed91-7e93-4a20-a6f1-c0e0ac4f51c6 · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:20:07.149183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:42854155e29a209ae0ff83a09a444eb58a30d76bbac2f6f55aeb2fc7d4f1d592

Observation 2c10252d-d4a1-421f-a607-2e71311693ec · outbound

This paper cites 3D diffusion policy: Generalizable visuomotor policy learning via simple 3D representations,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning 3D diffusion policy: Generalizable visuomotor policy learning via simple 3D representations,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:cf225a07cedefa97809e69574934a11f4c228baa27b2ca5869af1dbf608f5bc9

Observation 3a72c63c-cce9-42f8-88fb-9f599b8d6e4c · outbound

This paper cites RISE: 3D perception makes real-world robot imitation simple and effective,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning RISE: 3D perception makes real-world robot imitation simple and effective,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:3cc5b20fbe6aa80fcd3cff6d56c50a124fc01228dadf37eb6585cc236b886a5b

Observation ed1ddf3b-764b-42c5-9e64-83c472a43a57 · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning RT-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:e659a111a88df5ed003e8e628c2ad16690d810dc7f1d2751298305de54fd24af

Observation e85e71e3-8618-49c5-a24a-8e2cc70719ec · outbound

This paper cites Dense object nets: Learn- ing dense visual object descriptors by and for robotic manipulation,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Dense object nets: Learn- ing dense visual object descriptors by and for robotic manipulation,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:978590d48828b2d11e7f794033d829dd847be31e35543fbc0b1c7f62e1e148d3

Observation f67ef300-7f55-4893-b806-b931d961480b · outbound

This paper cites VIMA: General robot manipu- lation with multimodal prompts,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning VIMA: General robot manipu- lation with multimodal prompts,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:65a3232f219394839887327c934e65f79d3d683b35f13badd6538f36a3f9ed57

Observation 68320909-8301-4fd9-8cb8-0ffbc1d4656d · outbound

This paper cites MOKA: Open-vocabulary robotic manipulation through mark-based visual prompting,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning MOKA: Open-vocabulary robotic manipulation through mark-based visual prompting,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:ad96a189b085ce0557fc23e1ef5515cdd976841ac928c7eacd58751635848891

Observation 5ad94462-be3c-472f-a3b8-350a8bda9438 · outbound

This paper cites ReKep: Spatio- temporal relational keypoint constraints for generalizable robotic ma- nipulation,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning ReKep: Spatio- temporal relational keypoint constraints for generalizable robotic ma- nipulation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:1faddd1d63705f558e4510c341fe9ab2885cddf7dd1690d3996fc4d746c5a0ed

Observation 596185ce-c585-4a7d-952d-58411ee7f0e5 · outbound

This paper cites ProtCLIP: Function-Informed Protein Multi-Modal Learning.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning ProtCLIP: Function-Informed Protein Multi-Modal Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:07.154799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:5a81ad743a12d8eb4a5c10f4a0f3c3cd314d01ea9daac45493b95c1d7e7c964d

Observation 1001b5fe-7477-4d1d-80f9-c7768c154381 · outbound

This paper cites Spatial forcing: Implicit spatial representation alignment for vision- language-action model.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Spatial forcing: Implicit spatial representation alignment for vision- language-action model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:07.143749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:cd62a58189ad815c8df3267fcdc1e307e3001a0da62bee0676e9c902a30bb275

Observation 26a75245-f23b-4850-902a-b68bd98c7382 · outbound

This paper cites More than a point: Capturing uncertainty with adaptive affordance heatmaps for spatial grounding in robotic tasks,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning More than a point: Capturing uncertainty with adaptive affordance heatmaps for spatial grounding in robotic tasks,

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:07.146577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:56ec7107da2bd25b79932e45528f3d8966816185d19f3f277334d5f0c330c8a0

Observation 5d95502a-351d-48bb-abf4-29b917a635b1 · outbound

This paper cites SAM2Act: Integrating visual foundation model with a memory architecture for robotic manipulation,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning SAM2Act: Integrating visual foundation model with a memory architecture for robotic manipulation,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:123e1df1961d6b7ee01aa751043744424b3b2ae4987b0d26d8ca6ea4cc3d2f1b

Observation 31ecf627-bde9-4b78-adf1-12745b35a77c · outbound

This paper cites Deep residual learning for image recognition,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Deep residual learning for image recognition,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:a124ad9d4eaa61982076432de5d0aff200056b260b7e504b379a4db4a516cd0e

Pith citing papers

Observation eb82373d-f568-47a7-b9b6-aecb5a80be73 · inbound

KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation cites this paper.

KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-11T15:45:43.483298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T15:45:43.483298Z digest=sha256:72d0535c8adeab887faa31d78e53c5da6268e0b650cf9cfb40fc95dac13a8499