Pith. sign in

Paper Citation Record · LEDGER

PointVLA: Injecting the 3D World into Vision-Language-Action Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2503.07511.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.07511 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:04.124888Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T15:18:33.003700Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dcd4bf35-08f6-4584-8d82-ba1130d069e3 · inbound

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge cites this paper.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:04.124888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:04.124888Z digest=sha256:ba4e9ab7f84f46cf82de2974fc9d86d71fbfb54e2d355074489cfb2b0704c005

Observation ff242380-cf6f-43b4-8272-bc9b797fab99 · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.913716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.913716Z digest=sha256:361b14d8d6cd8ad9299ce4261f31eb99c1ef39ba66af37613c6e3e45a97e8ee2

Observation c0b9c9b2-fb93-4682-8288-15a48054d178 · inbound

Block-wise Adaptive Caching for Accelerating Diffusion Policy cites this paper.

Block-wise Adaptive Caching for Accelerating Diffusion Policy PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:37:14.012841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T09:36:09.790248Z digest=sha256:e5e1c487eb601c01a64dd10ef7c925a54710e25211710b5633d002bb6aeae503

Observation 5a081d2e-86fc-45f3-801d-ec30236c942d · inbound

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models cites this paper.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.478000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.478000Z digest=sha256:6b0bbb81848d965e2433739fa48baa2604eb521bb95f22dcb083994ae761b23c

Observation 4acb9048-e194-4083-bc7f-d33873d302b6 · inbound

A Survey on Vision-Language-Action Models for Autonomous Driving cites this paper.

A Survey on Vision-Language-Action Models for Autonomous Driving PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T21:31:04.553726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:31:04.553726Z digest=sha256:79a91652d6721e5faf0427e9af1054d100f10a12e7f2f2b3f42a9562e88c0be6

Observation a35ca375-9e95-4252-9aff-04c085ca2630 · inbound

AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies cites this paper.

AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:44:57.369165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:44:57.369165Z digest=sha256:fdcf130833feedbceb8f851c058d591e2f9ceb2aae8d4c587eb1d3042dfc1c02

Observation f9c6fa5d-81b5-4426-873a-e0eae9c2790e · inbound

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models cites this paper.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.635037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.635037Z digest=sha256:b0c42a444c5ea68e873d8ea5c07aa454551646f26842fbb5e30bd2dde9ab34e1

Observation aa7accc5-6d43-4a9f-aa6e-b4bec3e39f33 · inbound

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning cites this paper.

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:44.697712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:44.697712Z digest=sha256:95035b25e7942aaa83417ab6c9b830c47108f638214b3197aa3cc8bc64bcfa58

Observation 3de2c2ee-50b8-430b-aaa8-d90b0f28cf0e · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 161

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:16.020169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:0af3f85a432d7d50e676807958f4b2d094f8bedaa4518f434dad58b76b622317

Observation e0d84149-dcb9-4c5f-995f-a824256616a8 · inbound

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models cites this paper.

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:43.166142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:34:43.166142Z digest=sha256:6733674d8525a0a8635ef57fa71783a120b62daa1be44eaf7b6860a5edec04a5

Observation 68888964-eab6-4441-904e-4ebba7a9d12e · inbound

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision cites this paper.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.825243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.825243Z digest=sha256:89714e1f62b67bff3187104a12af856961776577ee3cb2d00dbce4091c478fa1

Observation d6e2d2d9-a70f-45c9-80e7-c8fc88203900 · inbound

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation cites this paper.

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:41:10.984675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T18:39:59.449746Z digest=sha256:6957693ed5a8cb96cae1a90be5dbbfcd36e3197f053f8e0a2ee01851c0357108

Observation 1e3dd686-6eab-4812-8a1b-70c84eb8d3bb · inbound

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation cites this paper.

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T13:35:54.124177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:35:54.124177Z digest=sha256:9a8af86d82c9a6858f4279eff38fb12ecd6209e2f23c4da13f04745064f638d6

Observation 849bca38-52c1-4ca2-9d1e-388de3769f06 · inbound

VLA Knows Its Limits: Adaptive Execution Horizons for Robot Policies cites this paper.

VLA Knows Its Limits: Adaptive Execution Horizons for Robot Policies PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T21:06:09.471941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:06:09.471941Z digest=sha256:62e1838acef44f85a06c3452875716c40f2231486246191912165770c134ee88

Observation aadd17e3-827e-458d-8cd3-43ee312c2195 · inbound

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes cites this paper.

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:48.402416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T20:21:45.365156Z digest=sha256:fe79197d03a173c66d4fdf9f7905f3c67162cbc08a9c5fd379f6ab0aff0e9863

Observation fc36ecab-8d5d-474f-b84b-f282c404de6f · inbound

R3D: Revisiting 3D Policy Learning cites this paper.

R3D: Revisiting 3D Policy Learning PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:00:04.190007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T10:57:54.273960Z digest=sha256:b3d269c64a642a4a17357326c91ff789f0972e122b4fa1304fcb9925ea4b4d14

Observation 1d7d34ec-203c-4022-af06-631cd113fe33 · inbound

ST-$\pi$: Structured SpatioTemporal VLA for Robotic Manipulation cites this paper.

ST-$\pi$: Structured SpatioTemporal VLA for Robotic Manipulation PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:40:19.869296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T04:46:21.784769Z digest=sha256:f4f9977a04461cb448ae0bee6098fef21131546741dd041e54d4fa2631d11b04

Observation 74418ad3-5c5e-444d-82f1-3f654360937b · inbound

PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance cites this paper.

PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:29:47.562283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:04:16.073230Z digest=sha256:4a18df20366fb8911bdec5fec0bc23636f25bc5f83196b1f6177466accc63c04

Observation e5357876-7653-4d70-aa16-56aa03c37e24 · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:06:06.268981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:452f1142da93810496a7626f3540d0451c8763b0aa2494f9ef056dd460d1c260

Observation 34217eea-514c-41a9-bdcc-01e7036fedf2 · inbound

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction cites this paper.

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:52:17.020022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T04:50:07.792757Z digest=sha256:7dd3c94d37509031242c47e8d124c8376280b8748f73e3c402f9a42b9350bc8a

Observation 63376b24-316f-4f5f-afcc-dc4a37b175d1 · inbound

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model cites this paper.

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:06:08.703355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T06:05:52.461348Z digest=sha256:248f44a26220c4c209635a9b73c2768c7a4441b618ed888dce791cf21f13aa3a

Observation 134aa164-34ca-4cba-aeb5-4e13ff9a65d3 · inbound

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model cites this paper.

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:54:58.317685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T16:52:05.565650Z digest=sha256:d0432c83eb2a46f588951b36a49a0d356cc4c7e4974c12cbad08b6bbcf4a4a27

Observation 79850471-bffb-49f2-9a5e-c01b70001734 · inbound

Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models cites this paper.

Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:44:40.743959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T13:40:59.330788Z digest=sha256:aea1ef7b25745d33b020c319a24fb9d0a8460c4329b3e9c1567d4228b8c5273c

Observation 74cb6a2f-f619-431b-a572-ae0c900b4345 · inbound

OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation cites this paper.

OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:59.337303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T21:27:26.682689Z digest=sha256:e7b32bdeed4733fac9192ac4a47aa62cdde9b6246a7178c1dc976f6bd095928d

Observation c1da69b1-9aa2-4bac-ab1b-fcabffb90382 · inbound

World Pilot: Steering Vision-Language-Action Models with World-Action Priors cites this paper.

World Pilot: Steering Vision-Language-Action Models with World-Action Priors PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:18:03.254393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T09:40:02.137152Z digest=sha256:31db2d9175c5928e39fb5a4d91e34bbcc7f9cfbfcaa499d86da527953069c9ad

Observation cd6b747c-1cdd-4abe-8ffd-4e9a3261f804 · inbound

GeoHAT: Geometry-Adaptive Hybrid Action Transformer for Mobile Manipulation cites this paper.

GeoHAT: Geometry-Adaptive Hybrid Action Transformer for Mobile Manipulation PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:33.005241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T06:37:21.516021Z digest=sha256:441040a2225a2109d6667f416735ffe7800121584eeafd0900ae4d0acb7a02b7

Observation 4e0972ee-631b-49a8-b983-2cc8cdd6ba3f · inbound

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation cites this paper.

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T06:36:24.361775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:36:24.361775Z digest=sha256:32bba8b77585f49b4cb12e0ca1244a5cb8a3b4bbd7dc9f26b127c2520fdd8f36

Observation 866f949d-1f4b-448d-9bf3-77f4d8a5c3ea · inbound

Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models cites this paper.

Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T11:57:40.822184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:57:40.822184Z digest=sha256:2c0b7cc7c9344115e3565555892b5ef97d424e8a7aa4f9d29a63d9e37999d3b9