Pith. sign in

Paper Citation Record · LEDGER

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy

As of 19 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 4 inbound Pith citation observations for arXiv:2508.13103.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.13103 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:20:58.997506Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:35:40.081807Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T21:20:17.200356Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d85a76ce-9693-48c6-bfb0-b800b98b7e45 · outbound

This paper cites 4" FUNCTION default.is.dash.repeated.names #1 FUNCTION default.name.format.string.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy 4" FUNCTION default.is.dash.repeated.names #1 FUNCTION default.name.format.string

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:20:59.786970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.738794Z digest=sha256:2b9eb65e9d48ccc8bb436d7177eb2f28f94e7122e7c8ff9365c554660fb506e2

Observation 9f2509d9-f7a6-49bb-aa1c-665aac7f6649 · outbound

This paper cites write newline.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.744008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.744008Z digest=sha256:e177593951fffadaae12bbf90ed7133e6407119f97d89e77e2cb7d4e6431f26d

Observation 48b5803c-ac39-4333-918c-d40671f9e2ac · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy RT-1: Robotics Transformer for Real-World Control at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.748350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.748350Z digest=sha256:7bba0b81b045f19c3c4c54a02a000f213d3d81685ac0b53d6fffa497f9ee501f

Observation 734b01f0-f168-4917-8cba-58bf6431d86f · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy OpenVLA: An Open-Source Vision-Language-Action Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.752671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.752671Z digest=sha256:ee8ac0618d8f634bd36955f99fe409948b1b0e469a661ececfd3884afa0deba4

Observation 23354231-5499-4c66-8595-2e6a3e19a374 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Octo: An Open-Source Generalist Robot Policy

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.756736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.756736Z digest=sha256:d5c685222537568eaf89138737d4fe8ae76a7a8ede12f37f2c58ababa88bc08a

Observation dddb3e91-544a-4ed8-9d89-f076557ec2e4 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy RT-H: Action Hierarchies Using Language

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.760786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.760786Z digest=sha256:726c560dda28af619df4e8c8a64b3cec75bb1033b290b1467df27cb54202a08f

Observation fa9071f4-453a-4c30-88f2-a3b46583fa9c · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.764768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.764768Z digest=sha256:e9f4aadf2c2401be7bbb39085787e986d70f47c1b9e6e8ae73841032d927b02d

Observation 3834d317-7c77-4a80-85f6-af2a9466ee0e · outbound

This paper cites Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.769098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.769098Z digest=sha256:4cf0219357c41a6e7dff73a417b4711498de19762bff1bb04f28b0f7db1029f6

Observation 9acb5f3b-5285-4121-8b04-d1df902c316f · outbound

This paper cites O’Neill, A.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy O’Neill, A

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:20:59.768740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.772747Z digest=sha256:6b61261573a681c396d2bec971072abff41c88a6da7ff1aaf753c077aa59eafc

Observation 3ca030ed-ce2f-45fc-82dd-4373e3ece7bb · outbound

This paper cites an unresolved cited work.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:20:59.758405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.776588Z digest=sha256:14c0eaaeb05fc7ce2b08ae913435cbe82d4824cb6785846c245f872462487ea7

Observation 9a8114fe-0a7b-4afc-af96-7d5ed672a3c5 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.780368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.780368Z digest=sha256:a29972ff1648c0662938e7456559ef21e6a7546a624c635ea1021cf663836d77

Observation 7ef80c0a-27a0-42cd-88e4-c054994dd4c5 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.784060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.784060Z digest=sha256:e301764540a44e60001420087e386c680189440f332d933a018336e9ea7a2790

Observation cfd8da52-0ebb-4a16-b9ba-d1fb25567db9 · outbound

This paper cites Kroemer, S.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Kroemer, S

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:20:59.747977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.787946Z digest=sha256:28ecef5ebe9af61b27cba50f0a8588690262d382f54df06c58d6c2d2958708b5

Observation 10549dd3-a620-47d7-a082-e86f3f2ab065 · outbound

This paper cites Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.791408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.791408Z digest=sha256:59666bb8bf2487285d72160b70bf927531209363eb0666b577f63a65d343e972

Observation 20019985-bdd6-48c0-878a-9bfa32ad0a36 · outbound

This paper cites Yamada, Y.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Yamada, Y

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:20:59.736360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.795123Z digest=sha256:1c16e45e24288d2bb9c98f222a2eb10888904a1245920d1c3d6e2c1b480fc2d6

Observation 8d954c0f-f57f-4d8a-99e4-f84ed3273d49 · outbound

This paper cites ReLMoGen: Leveraging Motion Generation in Reinforcement Learning for Mobile Manipulation.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy ReLMoGen: Leveraging Motion Generation in Reinforcement Learning for Mobile Manipulation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.798696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.798696Z digest=sha256:bad1f2a2c2afb761fe2606244dde37392afe19248ca14f7e6df4d41e0a80c115

Observation 30061129-5f04-4f11-b1f4-6df388d5f1c6 · outbound

This paper cites Shridhar, L.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Shridhar, L

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:20:59.725438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.802552Z digest=sha256:0f673b7d6a8520fe16eb965a48239e6e520b50d2e2dd874569d16b0da8a57a49

Observation 67326880-6e9d-4d4a-aab0-9819958d8cb5 · outbound

This paper cites 1em plus 0.5em minus 0.4em PMLR, 2023, pp.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy 1em plus 0.5em minus 0.4em PMLR, 2023, pp

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:20:59.714489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.806237Z digest=sha256:f726af9a025131e9be53a178557083c7c1134288ed6f87fbf03e3371868b3bcd

Observation ee1037bd-3e72-422f-8a97-ed3b136b9992 · outbound

This paper cites an unresolved cited work.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:20:59.701937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.810306Z digest=sha256:a66a692ed7a8a6373430f71a6ec2489bb0f571951fca5cf126f40db5e566c025

Observation 00f2c73a-a693-49c1-a132-dbdac4a44cb0 · outbound

This paper cites an unresolved cited work.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:20:59.690616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.813744Z digest=sha256:ebcbe191fb5c5acf9b7777e24ab3a48926c066d9feef113ce183bed414b7f523

Observation d40f5942-a82f-4e1a-8bac-3f7af25819b2 · outbound

This paper cites Singh, V.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Singh, V

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:20:59.679196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.817071Z digest=sha256:66652faa92413226669b1ea220022f15167843e78fa30f8b43e086f6a5038f9f

Observation e2c743aa-aa58-461a-ac72-d5b9b86cea18 · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Vision-Language Foundation Models as Effective Robot Imitators

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.820651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.820651Z digest=sha256:545cead2acf2a96cb5c73b3341c3caee91d9915fb2f5ab6a813eb49a9fcb49a9

Observation 8bc92a44-f2b0-4e57-97a0-8e62412361e2 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.824487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.824487Z digest=sha256:d74576421574b5a73a794904ab05ceb73190c068ca78571ca73912321a27aed7

Observation 2887d720-1e8e-40c6-958b-4226c0c30297 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.828433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.828433Z digest=sha256:d11e66f3c61d79dc779b8b6c8fe6acb27af55ee2016f2008720e5436f03e16e0

Observation 7c897f87-319f-476d-b7e9-9ef03a48c61d · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.832177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.832177Z digest=sha256:3cdba0442cb34c6713d0fbe7a0be2b428c154d334462a99e47429f3e8b2ca4eb

Observation f72dd4de-f631-4155-93ad-cffd7ff8bcbe · outbound

This paper cites an unresolved cited work.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:20:59.667004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.837248Z digest=sha256:075fa6925a396e721f93266054e95208ca15146648259f7f4207a0feded2de23

Observation 4593b79c-12ac-4420-89e8-c588ddb08110 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy PaLM-E: An Embodied Multimodal Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.844841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.844841Z digest=sha256:d83513aa13f836bbfabf891211fd2f6c1d5b11ef376eac82bd3139b6bce7d8e4

Observation db764539-b10d-45bb-9ed1-8fe0d46e2c6c · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.849618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.849618Z digest=sha256:09c052e35bb9f77bbcc09da67be7bfd66b9040d796b957ce007f293f34874f19

Observation 56226417-d62c-4142-9a24-1b0e1cce870c · outbound

This paper cites an unresolved cited work.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:20:59.655677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.853573Z digest=sha256:cbc5a10287f2f9cf0639bfff6064507a0271ea85b0e42b9b2a940d6698c3d8ab

Observation 6354200b-4ebb-4e15-9720-81670b801a45 · outbound

This paper cites Huang, L.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Huang, L

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.857277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.857277Z digest=sha256:bf899edefe7b231b80d9f3e47a6511a505ffbd75ce95e395fb441de8e1687adf

Observation f5f69a06-e112-4a58-9cbc-3f0bc725937e · outbound

This paper cites Language Conditioned Imitation Learning over Unstructured Data.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Language Conditioned Imitation Learning over Unstructured Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.861001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.861001Z digest=sha256:a28d3ad1f1f24a83ddf9a17eefd9781060cd10fc492a59a53ae3c55697a6e36f

Observation d5b97641-f8b7-4a65-9c33-9f18b39b064e · outbound

This paper cites Goal-Conditioned Imitation Learning using Score-based Diffusion Policies.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Goal-Conditioned Imitation Learning using Score-based Diffusion Policies

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.864808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.864808Z digest=sha256:4611d7621af2a7686ec81ed586d04df0f7709e89fc0e2a3f49a335039409e61f

Observation e6538e44-2fc5-428f-aa93-9eddacf40e08 · outbound

This paper cites an unresolved cited work.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:20:59.644458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.868761Z digest=sha256:278afcb5616fd0df63c5bbd5ffba1f143a882e9ee01c326f6539fecf18a6fdbe

Observation 45d9f000-910d-42f2-95f5-6a09a6cc90f4 · outbound

This paper cites Myers, A.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Myers, A

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:20:59.633513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.872556Z digest=sha256:be66f4f1c0c6175a19355598a90c39aa7486fa35fe0ee4d79ccd8c3b03b1d03a

Observation 05ed046d-9846-4fde-bff5-8672c7a80afe · outbound

This paper cites Language Control Diffusion: Efficiently Scaling through Space, Time, and Tasks.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Language Control Diffusion: Efficiently Scaling through Space, Time, and Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.876860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.876860Z digest=sha256:aba3435135ca87f21ec751d3f106c310c9c7a3d958c8fcef5da3441beded53bf

Observation 81f3e618-f8ef-4199-81bd-e9cac7c679f8 · outbound

This paper cites an unresolved cited work.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:20:59.622997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.880951Z digest=sha256:e15986674ee71fa23f725d6cbb821c035f9225c0657ee96ddc8ba39229408292

Observation 87bcbfa7-3e06-4941-90f1-ad98c0e0b519 · outbound

This paper cites Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.884621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.884621Z digest=sha256:37b0888c904e3b827d49aa66a33371a6f620344987e357d24dfe9b2395ad144f

Observation b6fb59da-9b13-45f1-952a-264a854db01c · outbound

This paper cites an unresolved cited work.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.888405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.888405Z digest=sha256:341013ac692051dbff0dbde69261065cb83c79477f03ccae3ebd9bc78df2819c

Observation cc01c332-dc6a-4477-ace3-3d6b8e71f465 · outbound

This paper cites Rombach, A.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Rombach, A

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.892684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.892684Z digest=sha256:60b989d606db6cc7ca17db04f33f9d8e28bd2da226d78f08e42a255203f8e681

Observation f60f9ea7-1ebc-4279-9f4b-480e2ef007e3 · outbound

This paper cites Dhariwal and A.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Dhariwal and A

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:20:59.597879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.896306Z digest=sha256:644102af80e22e0f41f1cece5976b8c1e05f8a1bcb5e4befbe97287ceab6ba94

Observation 24c92f8a-4c56-4013-9ffa-682c3a042552 · outbound

This paper cites Peebles and S.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Peebles and S

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:20:59.586622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.900017Z digest=sha256:6dd83d75158303031f8aa3c09a6f235532c97bdaecdf40ece1ae843232893dac

Observation 66209e9b-2b3e-4dda-86ca-3cf3b0f87c7b · outbound

This paper cites Brooks, B.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Brooks, B

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:20:59.575609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.903808Z digest=sha256:964d8be48ff5b552fccb5997391ed7c431467640ff21966cbe29f6ec0b1deb70

Observation bb9ce109-b65c-42f6-be5a-cc8f382df2ee · outbound

This paper cites Liang, Y.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Liang, Y

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:20:59.563925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.907413Z digest=sha256:da653da3de5b6b3912bf38c9d2ec3ebbe9b52970eb631c39edc7548c25e23ec9

Observation 888903d9-95b4-46ba-b323-8d42d91a1ecc · outbound

This paper cites One-Step Diffusion Policy: Fast Visuomotor Policies via Diffusion Distillation.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy One-Step Diffusion Policy: Fast Visuomotor Policies via Diffusion Distillation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.911181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.911181Z digest=sha256:c94eb9161b279fc8a86df749bad6502f7885751e9eb209f0cd36c2397313b678

Observation a0ef240d-295e-49be-9294-3e23d8b39bc2 · outbound

This paper cites Mamba Policy: Towards Efficient 3D Diffusion Policy with Hybrid Selective State Models.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Mamba Policy: Towards Efficient 3D Diffusion Policy with Hybrid Selective State Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.915138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.915138Z digest=sha256:cfa9932b3fa01110de68497ab0d535180023974075a4b2a07f506722e72b2841

Observation 9e73b499-37f0-4d09-a4a7-09fca6792b35 · outbound

This paper cites Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.919019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.919019Z digest=sha256:ae1bc5b8b931a61ee8deb09e81df13c0c6c01dfa53a325643437f64fef417cb2

Observation ba85b343-affd-450c-9826-d7b59b566817 · outbound

This paper cites Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.923194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.923194Z digest=sha256:29d39f330ab89f8bd69ef5cdbed28ce68ad45d10da6c5617eb5bd2889f055546

Observation ca6eb620-3589-4a0f-a1ff-0c258c238e53 · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.927160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.927160Z digest=sha256:421ae497604c4e14e09608e5dc0c7a1f38d5394899933629620a1ad1cb22e0b6

Observation e8baf7c1-f9ca-40c1-a7a8-d2d287d826a0 · outbound

This paper cites an unresolved cited work.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:20:59.553395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.930890Z digest=sha256:ea2f2ad6114d6c972eb97095ff69c6096c291fbb50cd67da24fd2dee0339cd11

Observation acde0b73-d592-4848-9326-e332913009ad · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.934359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.934359Z digest=sha256:bcb79de2c5edd11f1959a5f7b57fea1c076687dcdee30c30126e200de312f706

Observation 1186a65c-d6f8-494a-85e2-81a28c269ab8 · outbound

This paper cites Reuss, \"O.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Reuss, \"O

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:20:59.543357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.938694Z digest=sha256:5e1b1a6dbc87bd9933a5f4d054aa53ea44184ed5cb0083bc2e465ff242d33052

Observation 42f61a0d-d56f-4aee-b9b8-f7f21055eb35 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.943602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.943602Z digest=sha256:fcf9a1f9de6453711af12f273975fc3545c0a00f899476e9aac35eb1e2b2d0b0

Observation 82910864-0c88-4399-8365-67430e518d7a · outbound

This paper cites DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.947794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.947794Z digest=sha256:0d06ac86a3834516a2d639b623a1ae9c20dccb3ed3f42e989974e70ab20dcaac

Observation 57d86c72-065c-4732-9470-a42dcad4b8be · outbound

This paper cites Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.951601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.951601Z digest=sha256:cf6774dd46c07f735f85aa5a631a3be015409f5d17ff70e65cebd8cccd91264f

Observation 2e1dd8f3-45b6-41fb-8d37-bb1178ed4d31 · outbound

This paper cites Vaswani, N.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Vaswani, N

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:20:59.532414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.955632Z digest=sha256:f2bfa0a5d9b8199bc5cba63311ad1ca6b34ce04b9440e22fdcd1316789a58aef

Observation 895ac6f9-c772-4c00-aa44-191b40606214 · outbound

This paper cites The Ingredients for Robotic Diffusion Transformers.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy The Ingredients for Robotic Diffusion Transformers

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.959249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.959249Z digest=sha256:1aa6e4c2c53e335bb7bc4b50eda8fb60d19265b6f12ea094feb3350368e7087a

Observation ebddbd0a-01ce-4ee5-994e-f9d86936d032 · outbound

This paper cites Radford, J.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Radford, J

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:20:59.520716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.963630Z digest=sha256:f6e2c2eabab014951ef65c56e486dc17bbaab80b4cdeb6f936ac858ac5beaaab

Observation 9dfd9416-ada3-408e-bb7e-f4dbae5d7cdf · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy DINOv2: Learning Robust Visual Features without Supervision

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.967132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.967132Z digest=sha256:1e8d5540f27ede2c06b21d1f7518d4fd7cf901e9aec4ccd7c7d8c4a2bc37425b

Observation 2768103f-b7fb-4ad0-8c7c-54a971dc66cf · outbound

This paper cites an unresolved cited work.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:20:59.510218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.971114Z digest=sha256:9c530a3a9a82ac00d5fbaf4e770759164e7ff489672107d0719b69c2cd121f86

Observation e29ae840-db82-411f-b256-fe27388aaa09 · outbound

This paper cites Perez, F.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Perez, F

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:20:59.498900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.975523Z digest=sha256:19405eb539448b83c051ef329324533fd1b6eae92ee91358d29a6965caa2d50d

Observation 59dc8ea7-0ee5-4382-a913-a1d709ccea14 · outbound

This paper cites an unresolved cited work.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:20:59.488043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.979238Z digest=sha256:60c7114094e09481258b781379bf00db7fb997f10f7d078a69df6474165e6b2d

Observation eeaf38b9-a8a8-43d8-9311-8193bf6c3abd · outbound

This paper cites an unresolved cited work.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:20:59.475700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:20:58.982651Z digest=sha256:9cb0011f01035938e38acc650135c0563a4613087117f9d2e62343de1c2a7948

Observation 53eab859-fe6c-480c-988b-cdfe15eb9514 · outbound

This paper cites Loshchilov and F.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Loshchilov and F

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.986164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.986164Z digest=sha256:4746d0ca2e65f898891782d5d0c858e1ff3b033faa853cafd7416950473ce041

Observation b8767a5f-23d6-49b2-a52f-27926a768585 · outbound

This paper cites ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.989671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.989671Z digest=sha256:b6f74e718f1f32c79935da8311ce293370f7edb4585adf7db9547e4c1842efe7

Observation 81ac2842-ba41-4f9a-96bd-2cfbca5b1466 · outbound

This paper cites ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.993573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.993573Z digest=sha256:c351c475ab100bdaf7325589ab84eb37186e045e4bd788b4890b00e7f6634a0b

Observation 94a6902e-67c1-4359-b61c-1b3d483f2407 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy LLaMA: Open and Efficient Foundation Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.997506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.997506Z digest=sha256:6970001f00ee55fad9cc24af685e526fb6433377b5ca59e89eb1dfa071fdd139

Pith citing papers

Observation cb4fe4bc-0492-4874-81b6-d0d012b26025 · inbound

Unify Robot Actions in Camera Frame cites this paper.

Unify Robot Actions in Camera Frame Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:20:17.204124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:45a19b1c71a36ba8947320bcfd11af1a22b126d78ede4e12c83420afdae6b699

Observation 10dc769e-4614-447e-be7e-30a273905e7a · inbound

CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors cites this paper.

CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.882356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T22:07:24.208555Z digest=sha256:9004f33ac478374c394077118c2adff20b610001677756cf6f93c18b3cb3d048

Observation 700ee7b0-2398-431f-ac70-b38bf1fa5d0c · inbound

CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors cites this paper.

CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T18:39:49.173828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:39:49.173828Z digest=sha256:98a0c6f1a1ce5b7f330337b07deb43bbf1eed8f4990c040467eaa3e9c06cedb4

Observation 3c5b2d5d-7484-41fd-b0d1-f98a5ad20c98 · inbound

Cross-View Action Consistency for Camera-Robust Vision-Language-Action Policies cites this paper.

Cross-View Action Consistency for Camera-Robust Vision-Language-Action Policies Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T17:35:40.081807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:35:40.081807Z digest=sha256:15dcf05292c2fb09d3d96d0f796750364626426587880440738a39b27b6a71e9