Pith. sign in

Paper Citation Record · LEDGER

How Should Vision-Language-Action Models Use Proprioceptive State?

As of 10 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2608.03052.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03052 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T01:01:44.151987Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact7
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation afc253b5-1f6b-43c1-bb74-6511dc44c14d · outbound

This paper cites HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation.

How Should Vision-Language-Action Models Use Proprioceptive State? HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.071866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.071866Z digest=sha256:5e6ba598a2ebf10e83d70e8c8d2e5f8880bba8a73489c588e620a1620fa0b23d

Observation 8a3b5760-2bc7-46e0-a334-855dcc767c0e · outbound

This paper cites RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies.

How Should Vision-Language-Action Models Use Proprioceptive State? RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.084075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.084075Z digest=sha256:121c514ee09d9b8880917bb6d8fae5a0e3ea773f5ca661272404151eac5e0578

Observation cc10f7a7-1860-46b7-8577-790329b85c56 · outbound

This paper cites Davies, Y.

How Should Vision-Language-Action Models Use Proprioceptive State? Davies, Y

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-08T01:01:44.989531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T01:01:44.087910Z digest=sha256:f4dd4f525ec2dc062460cff01eef2b9c7a4130978e40b5b28c14ee4e6a0a0e72

Observation af6f9f4e-56ad-43f8-817a-cb06c02ca0cd · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

How Should Vision-Language-Action Models Use Proprioceptive State? $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.094721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.094721Z digest=sha256:17a0c71d769fd7bc31f790ad1038429659253e6235a7a790fcdc14b5eadf5df0

Observation 203c76f1-4b45-4a9f-ac77-aeb8ac7954d1 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

How Should Vision-Language-Action Models Use Proprioceptive State? OpenVLA: An Open-Source Vision-Language-Action Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.098166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.098166Z digest=sha256:b9ec96131591cfb87e77647c605505bf65ab2bba958ddd34383cc2f3c2a6f399

Observation c83c44b3-fcb4-4aba-8afe-dc0c84019d49 · outbound

This paper cites 2602.12032.

How Should Vision-Language-Action Models Use Proprioceptive State? 2602.12032

Reference 11

Resolution
verified exact
doi, observed 2026-08-08T01:01:44.528761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T01:01:44.107978Z digest=sha256:a2d890db315eb4cfc8157babb31ceb389d00f301e1ef56f7e85d80b4079d349a

Observation 0d5b8cb8-3eeb-498b-9b0a-d553c6637214 · outbound

This paper cites an unresolved cited work.

How Should Vision-Language-Action Models Use Proprioceptive State? Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T01:01:45.035967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T01:01:44.111421Z digest=sha256:f6f724a02470217a7afd31e6d1c9115dd59f6ac8de3c2d0412bdde476e7e24dc

Observation d9ce61a8-bbd2-4c14-8313-63116db4ecd5 · outbound

This paper cites MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.

How Should Vision-Language-Action Models Use Proprioceptive State? MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.115129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.115129Z digest=sha256:d55724c62ccc91032a2d1e79dbe6c238905359a82253a90d3371975410e49d50

Observation e54940ea-f858-4887-9295-6bd8d2cf92cc · outbound

This paper cites Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies.

How Should Vision-Language-Action Models Use Proprioceptive State? Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.118690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.118690Z digest=sha256:8f19c51d02270ed3fb2cdacc8db66a7b542e1e2fa3068b728c4f0bd4f41dfa5f

Observation 8c5ad290-f9e0-4386-bc10-c1d1fc7b6ee2 · outbound

This paper cites Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies.

How Should Vision-Language-Action Models Use Proprioceptive State? Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.122151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.122151Z digest=sha256:fe7bf1f7dfe1f2a30a0f35a1ee954939e1dd0d77fcec6f27a2b0f4faa52b418b

Observation c3e3f441-ea11-4b9d-870b-847031eba6c7 · outbound

This paper cites an unresolved cited work.

How Should Vision-Language-Action Models Use Proprioceptive State? Unresolved cited work

Reference 16

Resolution
verified exact
doi, observed 2026-08-08T01:01:44.433068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T01:01:44.125411Z digest=sha256:f2ac1c9e8b4f7147014dec2d8e44fe9b1d885baf85dd89ca0125d99e2104f06e

Observation 35f2b077-8ae6-4598-8b4e-67b33e60a9da · outbound

This paper cites Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models.

How Should Vision-Language-Action Models Use Proprioceptive State? Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-08T01:01:44.721229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T01:01:44.132304Z digest=sha256:4b6b07dc8ba9ad214de994d5ede62849cf0b186b58d8f1dc03e6661b2c549822

Observation 30f870f2-dca5-419c-9ab7-5c38ad0279b5 · outbound

This paper cites an unresolved cited work.

How Should Vision-Language-Action Models Use Proprioceptive State? Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.135594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.135594Z digest=sha256:39c8bfa7ab8ad703bb031e5faea07b56e132fcd07332c37e815b3f7def26d5f2

Observation 463d4e7e-96de-4c03-b6ea-a74878bc71d5 · outbound

This paper cites 2601.14133.

How Should Vision-Language-Action Models Use Proprioceptive State? 2601.14133

Reference 20

Resolution
verified exact
doi, observed 2026-08-08T01:01:44.358840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T01:01:44.138844Z digest=sha256:62b6ec19da0ba6511b6ab7c0e3105980a31ff8db27053378db9b54e1ebf26655

Observation 66820ca3-7f3c-4dce-9585-fba23ab4730d · outbound

This paper cites Zhang, J.

How Should Vision-Language-Action Models Use Proprioceptive State? Zhang, J

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.142251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.142251Z digest=sha256:ae60541016418a055b670cfead92a6f9bce94cafbf7e84dea4d87d18424cb799

Observation 6d897209-cf3b-4e60-a0cb-24de3d6d59ec · outbound

This paper cites URL https://doi.org/10.

How Should Vision-Language-Action Models Use Proprioceptive State? URL https://doi.org/10

Reference 22

Resolution
verified exact
doi, observed 2026-08-08T01:01:44.284230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T01:01:44.145383Z digest=sha256:539c238da1d4dd7adc344aab761bb73500a7f6ce98e9ef0387e209390f659978

Observation 5d2a6ede-04cb-4506-b734-b0bf1f0fd6a4 · outbound

This paper cites an unresolved cited work.

How Should Vision-Language-Action Models Use Proprioceptive State? Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.148650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.148650Z digest=sha256:a9cfa90f3bc49040193d7efaf3ac3b718f3bd87709e1e3b90222c76c4d94838e

Observation 85208fef-0207-490f-8633-3851c3c81abe · outbound

This paper cites an unresolved cited work.

How Should Vision-Language-Action Models Use Proprioceptive State? Unresolved cited work

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.091348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.091348Z digest=sha256:381142af0f1c7c0322d098a30d5d8b9db8ea4ded17cf360e7ed182d2547e5a8a

Observation 87fdd124-616f-4096-9bdc-3ba1d7801856 · outbound

This paper cites ROSA: Harnessing Robot States for Vision-Language and Action Alignment.

How Should Vision-Language-Action Models Use Proprioceptive State? ROSA: Harnessing Robot States for Vision-Language and Action Alignment

Reference 2020

Resolution
verified exact
local_arxiv, observed 2026-08-08T01:01:44.735748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T01:01:44.128884Z digest=sha256:18fe845efaef568ec508a40ae888469ddc2928282176b30044f61f5ade4ca740

Observation b93173d7-4de2-45eb-9873-d8c5239c2712 · outbound

This paper cites Whenwouldvision- proprioception policies fail in robotic manipulation? CoRR, abs/2602.12032,.

How Should Vision-Language-Action Models Use Proprioceptive State? Whenwouldvision- proprioception policies fail in robotic manipulation? CoRR, abs/2602.12032,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.104786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.104786Z digest=sha256:b9bc2a838ba80824bf5e97e42e7d7c88ddc419d1852d0652d622ea34f9d2e6af

Observation c42ebbaa-e9d2-4ae7-919b-8806c2da7a10 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

How Should Vision-Language-Action Models Use Proprioceptive State? Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.101602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.101602Z digest=sha256:babdc276265cc5d6bc4a46a29a3a264d141c8eb8a98602b3e31dd4e323246cc9

Observation 854015ee-14eb-4e0b-8f34-60d21ed41dfd · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

How Should Vision-Language-Action Models Use Proprioceptive State? $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.080292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.080292Z digest=sha256:050c80b982bda4cfa81ac4a5073cf770ca89384f1e1bd73983a113878297610f

Observation e72c9459-5697-4a3c-a778-220795f9843f · outbound

This paper cites HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation.

How Should Vision-Language-Action Models Use Proprioceptive State? HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:44.076699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:01:44.076699Z digest=sha256:87b0e358ce0480d8d49e524fe9584f3a79baf2f5128df8d1d12170a4bf06b8bf

Observation 808705d3-d049-4661-9671-42e1387f8e10 · outbound

This paper cites an unresolved cited work.

How Should Vision-Language-Action Models Use Proprioceptive State? Unresolved cited work

Reference 4096

Resolution
unresolved
raw_fallback, observed 2026-08-08T01:01:45.026043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T01:01:44.151987Z digest=sha256:8385038c969eeb21c9148479db86389582cfff6e7be8ca97558b03d8a0925c2b

Pith citing papers

No inbound Pith citation observations are available.