Pith. sign in

Paper Citation Record · LEDGER

PointVLA: Injecting the 3D World into Vision-Language-Action Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2503.07511.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.07511 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:04.124888Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T15:18:33.003700Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dcd4bf35-08f6-4584-8d82-ba1130d069e3 · inbound

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge cites this paper.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:04.124888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:04.124888Z digest=sha256:ba4e9ab7f84f46cf82de2974fc9d86d71fbfb54e2d355074489cfb2b0704c005

Observation ff242380-cf6f-43b4-8272-bc9b797fab99 · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.913716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.913716Z digest=sha256:361b14d8d6cd8ad9299ce4261f31eb99c1ef39ba66af37613c6e3e45a97e8ee2

Observation c0b9c9b2-fb93-4682-8288-15a48054d178 · inbound

Block-wise Adaptive Caching for Accelerating Diffusion Policy cites this paper.

Block-wise Adaptive Caching for Accelerating Diffusion Policy PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:37:14.012841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T09:36:09.790248Z digest=sha256:769147403b568abbbc2d4aa37631e40b62633d33d1841f66081a4c780a22856b

Observation 5a081d2e-86fc-45f3-801d-ec30236c942d · inbound

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models cites this paper.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.478000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.478000Z digest=sha256:6b0bbb81848d965e2433739fa48baa2604eb521bb95f22dcb083994ae761b23c

Observation 4acb9048-e194-4083-bc7f-d33873d302b6 · inbound

A Survey on Vision-Language-Action Models for Autonomous Driving cites this paper.

A Survey on Vision-Language-Action Models for Autonomous Driving PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T21:31:04.553726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:31:04.553726Z digest=sha256:79a91652d6721e5faf0427e9af1054d100f10a12e7f2f2b3f42a9562e88c0be6

Observation a35ca375-9e95-4252-9aff-04c085ca2630 · inbound

AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies cites this paper.

AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:44:57.369165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:44:57.369165Z digest=sha256:d8d132d8dfbabefe04d217b84dce8ab9bc1b155934feb16372d07f0626b41cbb

Observation f9c6fa5d-81b5-4426-873a-e0eae9c2790e · inbound

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models cites this paper.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.635037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.635037Z digest=sha256:b0c42a444c5ea68e873d8ea5c07aa454551646f26842fbb5e30bd2dde9ab34e1

Observation aa7accc5-6d43-4a9f-aa6e-b4bec3e39f33 · inbound

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning cites this paper.

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:44.697712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:44.697712Z digest=sha256:d247f219d27d2377920229dbe46a85578f055c44a83ee0b4548a471979a5973e

Observation 3de2c2ee-50b8-430b-aaa8-d90b0f28cf0e · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 161

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:16.020169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:148ab827328d3c99cd00810abdbec070529e4e526eb64c4cbdf14593e34ab075

Observation e0d84149-dcb9-4c5f-995f-a824256616a8 · inbound

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models cites this paper.

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:43.166142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:34:43.166142Z digest=sha256:6733674d8525a0a8635ef57fa71783a120b62daa1be44eaf7b6860a5edec04a5

Observation 68888964-eab6-4441-904e-4ebba7a9d12e · inbound

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision cites this paper.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.825243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.825243Z digest=sha256:2c6b6267c1ff2d5e97408b4f5244378c27d833cc91187554fecd0809823a7b7f

Observation d6e2d2d9-a70f-45c9-80e7-c8fc88203900 · inbound

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation cites this paper.

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:41:10.984675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T18:39:59.449746Z digest=sha256:829b438bf697d1f815f7ea6453848c766357de1d1d55bc01b2f3502c9919e374

Observation 1e3dd686-6eab-4812-8a1b-70c84eb8d3bb · inbound

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation cites this paper.

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T13:35:54.124177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:35:54.124177Z digest=sha256:9a8af86d82c9a6858f4279eff38fb12ecd6209e2f23c4da13f04745064f638d6

Observation 849bca38-52c1-4ca2-9d1e-388de3769f06 · inbound

VLA Knows Its Limits: Adaptive Execution Horizons for Robot Policies cites this paper.

VLA Knows Its Limits: Adaptive Execution Horizons for Robot Policies PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T21:06:09.471941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:06:09.471941Z digest=sha256:62e1838acef44f85a06c3452875716c40f2231486246191912165770c134ee88

Observation aadd17e3-827e-458d-8cd3-43ee312c2195 · inbound

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes cites this paper.

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:48.402416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T20:21:45.365156Z digest=sha256:947c98e73656f656e338cd64f049530c2c49d1610476dc1f5e8ef93eb0fc5897

Observation fc36ecab-8d5d-474f-b84b-f282c404de6f · inbound

R3D: Revisiting 3D Policy Learning cites this paper.

R3D: Revisiting 3D Policy Learning PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:00:04.190007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T10:57:54.273960Z digest=sha256:f5ea055c53ff0acadedd2be5a4bf9e0ddcf88393d9d859952a536b157028ae0a

Observation 1d7d34ec-203c-4022-af06-631cd113fe33 · inbound

ST-$\pi$: Structured SpatioTemporal VLA for Robotic Manipulation cites this paper.

ST-$\pi$: Structured SpatioTemporal VLA for Robotic Manipulation PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:40:19.869296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T04:46:21.784769Z digest=sha256:9260c9915aa91815a97dec5e361c2d5bf6d58f7a814e43a1150a486c7123e30b

Observation 74418ad3-5c5e-444d-82f1-3f654360937b · inbound

PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance cites this paper.

PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:29:47.562283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T00:04:16.073230Z digest=sha256:b9d25407504eac7739d2fc52af6c173f5895addee8cc9d6ea27e5ae0422754ee

Observation e5357876-7653-4d70-aa16-56aa03c37e24 · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:06:06.268981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:a269f1bbe23a5b3da6aec4b4a3150fb7397b96a312fa47675c6cef92a4d157a0

Observation 34217eea-514c-41a9-bdcc-01e7036fedf2 · inbound

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction cites this paper.

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:52:17.020022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T04:50:07.792757Z digest=sha256:f78dc94a309a9addc5535469018b5deec3847030ac8c001f003e6e1ac2a28aad

Observation 63376b24-316f-4f5f-afcc-dc4a37b175d1 · inbound

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model cites this paper.

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:06:08.703355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T06:05:52.461348Z digest=sha256:9589f6b4955a3b20f36094565bc0a68c5a19b207fe97756526aeedc3fce3c88b

Observation 134aa164-34ca-4cba-aeb5-4e13ff9a65d3 · inbound

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model cites this paper.

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:54:58.317685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T16:52:05.565650Z digest=sha256:17f5f30c9e4065a5a6788284f6be5143e1dae009510778f91f70d04621f95b3e

Observation 79850471-bffb-49f2-9a5e-c01b70001734 · inbound

Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models cites this paper.

Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:44:40.743959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T13:40:59.330788Z digest=sha256:15ce338f26172ae096d1cc77cc76bed018571e6514663b1c18eee65c107a00c1

Observation 74cb6a2f-f619-431b-a572-ae0c900b4345 · inbound

OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation cites this paper.

OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:59.337303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T21:27:26.682689Z digest=sha256:5f55e7f82b0f7297e0af8f5b41c1c899549b1008ce923e747b28ab07f264d51e

Observation c1da69b1-9aa2-4bac-ab1b-fcabffb90382 · inbound

World Pilot: Steering Vision-Language-Action Models with World-Action Priors cites this paper.

World Pilot: Steering Vision-Language-Action Models with World-Action Priors PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:18:03.254393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T09:40:02.137152Z digest=sha256:a551cea73508ef9d9f052b392bf771c6e3cf4b8457889209905c742ddc704003

Observation cd6b747c-1cdd-4abe-8ffd-4e9a3261f804 · inbound

GeoHAT: Geometry-Adaptive Hybrid Action Transformer for Mobile Manipulation cites this paper.

GeoHAT: Geometry-Adaptive Hybrid Action Transformer for Mobile Manipulation PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:33.005241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T06:37:21.516021Z digest=sha256:781aa58ce71a69068cee4a439c051d9d674ee8c1ced18f4a24caaae3b5eceed9

Observation 4e0972ee-631b-49a8-b983-2cc8cdd6ba3f · inbound

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation cites this paper.

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T06:36:24.361775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:36:24.361775Z digest=sha256:32bba8b77585f49b4cb12e0ca1244a5cb8a3b4bbd7dc9f26b127c2520fdd8f36

Observation 866f949d-1f4b-448d-9bf3-77f4d8a5c3ea · inbound

Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models cites this paper.

Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T11:57:40.822184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:57:40.822184Z digest=sha256:2c0b7cc7c9344115e3565555892b5ef97d424e8a7aa4f9d29a63d9e37999d3b9