Pith. sign in

Paper Citation Record · LEDGER

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning

As of 13 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2606.25360.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.25360 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-25T21:23:44.253416Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T15:45:43.483298Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa56e012-6ce6-40a4-b48b-d5e80e006be1 · outbound

This paper cites π 0: A vision-language-action flow model for general robot control,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning π 0: A vision-language-action flow model for general robot control,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:f35293f0f068fbf1e95da439f32214c9745decf19b01841b0133d56949fae123

Observation 034c7b5b-f9de-4e31-82d0-fc7538c38318 · outbound

This paper cites OpenVLA: An open-source vision-language-action model,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning OpenVLA: An open-source vision-language-action model,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:a3b10fe2fddb9c89d1835eb6161d5e43edf40ef7b98e06cecd2ef918e65161e0

Observation c6e1ab2f-2136-4386-9e4c-5fea8e64ece3 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning SAM 3: Segment Anything with Concepts

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:20:07.157528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:a5de5f047e53afe01b871a95df0382e0bd6555ced206c917e22d68a1a7ccd077

Observation 74b78b5e-a78e-4084-aa7a-c55801a491c4 · outbound

This paper cites ALVINN: An autonomous land vehicle in a neural network,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning ALVINN: An autonomous land vehicle in a neural network,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:066e63a0238d60a21f5a10dd9db888031e43a06c79ae5247ae1cb75ec3b11f8e

Observation 3239d6f8-2b3f-4200-a385-b3431935b6c8 · outbound

This paper cites Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:07.152003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:02db38b501dd59b7cff6585331be6b042765edfce35e54a90e7df4f8b0a34c46

Observation bb255f11-f354-47d4-84f7-7d2e69fc3472 · outbound

This paper cites Learning fine-grained bimanual manipulation with low-cost hardware,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Learning fine-grained bimanual manipulation with low-cost hardware,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:7733261c9508c3791267365c3620e91a436f566c9deb9f48d8e606c74d7a5f2f

Observation 6a2444a1-2ec6-403e-b4a2-ba672a804eca · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Diffusion policy: Visuomotor policy learning via action diffusion,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:067525beb377ca17adfa040be99ebd57a8593a18fce843bba73bbc92ef8b694b

Observation 0d921627-8eff-4cbe-a854-7f02fd358531 · outbound

This paper cites FiLM: Visual reasoning with a general conditioning layer,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning FiLM: Visual reasoning with a general conditioning layer,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:999234434a2c9af48afe19b078fd2dd6047137424e31b190c7ec5ba6cd042968

Observation 988e34b6-fd3d-4425-9b4a-a98309529043 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Learning transferable visual models from natural language supervision,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:900da27e273c8df692ccbc94f1066853f8c73e0d487f096f12f3827cbe6e5a13

Observation fb07ed91-7e93-4a20-a6f1-c0e0ac4f51c6 · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:20:07.149183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:76b8ba92ee485f9a98db91cc7b019c025dd58e65ac715146b6211a2a2a546a22

Observation 2c10252d-d4a1-421f-a607-2e71311693ec · outbound

This paper cites 3D diffusion policy: Generalizable visuomotor policy learning via simple 3D representations,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning 3D diffusion policy: Generalizable visuomotor policy learning via simple 3D representations,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:bc5a203fb04c9aebc2c3a5485f2985471714506340d88605224da54d141f52ff

Observation 3a72c63c-cce9-42f8-88fb-9f599b8d6e4c · outbound

This paper cites RISE: 3D perception makes real-world robot imitation simple and effective,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning RISE: 3D perception makes real-world robot imitation simple and effective,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:fb757376b26a6ea2d02f1371fae0886245c655fc6704144f6a3a9d3e35ce8582

Observation ed1ddf3b-764b-42c5-9e64-83c472a43a57 · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning RT-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:9ea581a48e7aab946852231326ba5e2e2ac954b889b606b505c483de492e541f

Observation e85e71e3-8618-49c5-a24a-8e2cc70719ec · outbound

This paper cites Dense object nets: Learn- ing dense visual object descriptors by and for robotic manipulation,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Dense object nets: Learn- ing dense visual object descriptors by and for robotic manipulation,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:76582dff486ca3ae951b7c8acd81d5d2f3b19b3bd02be9fbfa92ce7b7523bf05

Observation f67ef300-7f55-4893-b806-b931d961480b · outbound

This paper cites VIMA: General robot manipu- lation with multimodal prompts,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning VIMA: General robot manipu- lation with multimodal prompts,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:2d163a9b1aa85c1b2808bc2240dcfe8392ed9f15631cd9af0abd817f022ebd15

Observation 68320909-8301-4fd9-8cb8-0ffbc1d4656d · outbound

This paper cites MOKA: Open-vocabulary robotic manipulation through mark-based visual prompting,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning MOKA: Open-vocabulary robotic manipulation through mark-based visual prompting,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:c98a568e9d385864f1cacb19ca7cc6865b36e52101ab7c4db330c6906215e54f

Observation 5ad94462-be3c-472f-a3b8-350a8bda9438 · outbound

This paper cites ReKep: Spatio- temporal relational keypoint constraints for generalizable robotic ma- nipulation,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning ReKep: Spatio- temporal relational keypoint constraints for generalizable robotic ma- nipulation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:b877855caf7b952d11324fb3c977657ab9d3cce3f822b41406022630d4a2635f

Observation 596185ce-c585-4a7d-952d-58411ee7f0e5 · outbound

This paper cites ProtCLIP: Function-Informed Protein Multi-Modal Learning.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning ProtCLIP: Function-Informed Protein Multi-Modal Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:07.154799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:4fb05bea5d22e86e1288763eee5f6f2dca1cef833055fcd7405c5aa872899a33

Observation 1001b5fe-7477-4d1d-80f9-c7768c154381 · outbound

This paper cites Spatial forcing: Implicit spatial representation alignment for vision- language-action model.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Spatial forcing: Implicit spatial representation alignment for vision- language-action model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:07.143749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:bf03854b493f2967aba9f5407f4fd8aadcda586187ddea61a1c496542bd3471c

Observation 26a75245-f23b-4850-902a-b68bd98c7382 · outbound

This paper cites More than a point: Capturing uncertainty with adaptive affordance heatmaps for spatial grounding in robotic tasks,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning More than a point: Capturing uncertainty with adaptive affordance heatmaps for spatial grounding in robotic tasks,

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:07.146577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:6e7f6c31f8fdc7e18d956ca7f19c4785fa1160947ae611da83008c2b1f139d22

Observation 5d95502a-351d-48bb-abf4-29b917a635b1 · outbound

This paper cites SAM2Act: Integrating visual foundation model with a memory architecture for robotic manipulation,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning SAM2Act: Integrating visual foundation model with a memory architecture for robotic manipulation,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:0830af01a58f064bfd25a867715c0b2f76746979b86e66cc8c8a7a95f83d786a

Observation 31ecf627-bde9-4b78-adf1-12745b35a77c · outbound

This paper cites Deep residual learning for image recognition,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Deep residual learning for image recognition,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:cdccf5c1ec85ef6f932578151001e847a6f91c5aa5067e345de286c359441ec9

Pith citing papers

Observation eb82373d-f568-47a7-b9b6-aecb5a80be73 · inbound

KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation cites this paper.

KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-11T15:45:43.483298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T15:45:43.483298Z digest=sha256:d4fc304acc61f1bb6df050a9371c95fe969752af959ce6932308cdb65d0eb7a9