Pith. sign in

Paper Citation Record · LEDGER

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

As of 15 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 38 inbound Pith citation observations for arXiv:2508.09071.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09071 v2

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:16:50.651526Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:08:25.768932Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:38:43.940925Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 044b3559-17b0-4d81-81da-b5245e869a54 · outbound

This paper cites Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.631795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.631795Z digest=sha256:ce9e32d8e8e65f5b7c0e944b1dfdd2c5bdc5b643fb247c5d3b4f8b90b2000442

Observation 3f75680b-eb5e-4213-93c0-4d7397703433 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.637420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.637420Z digest=sha256:7202c66d5d9fa2aef64b3a750f68e7e4bec65ff6eadf42e5d274008f175d5835

Observation 19dc6192-36bc-4fbf-995e-a2ea50decd90 · outbound

This paper cites Flow Matching for Generative Modeling.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models Flow Matching for Generative Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.640176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.640176Z digest=sha256:906432caf7ca1ad383668f87d1b1b1d50d68a2c78699c47579ce0e0ebbd92c96

Observation 0953712d-9267-4135-882f-c3bfcbf75c5f · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.642960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.642960Z digest=sha256:fb29ed9eff91eab8a1392dd3329976843e82c17ecb297f426d5d948e4a0a5e1d

Observation 4a3bf52f-d29f-42f7-a82d-f45e792260a3 · outbound

This paper cites an unresolved cited work.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T21:16:50.739395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:16:50.651526Z digest=sha256:93b4960aa575a5a8b29c44e84d9fe0154645769e36c85177a913b9567b94b61a

Observation cc1f00d7-649d-4957-920c-df764718e823 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.647280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.647280Z digest=sha256:1c8161854c5bdb036670c56e7a0608b3384876020804342312153dfdbe86d430

Observation 8fb6f571-f63e-4466-985d-9a52c58cb6d2 · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.649376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.649376Z digest=sha256:c091364b27e876088e505f523f3d93f614f641c3f7783289d3e5e97cbdf31eba

Observation 8779decc-64d5-427b-a144-f8ce6795cb08 · outbound

This paper cites Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.629382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.629382Z digest=sha256:b0fc838139d4af4746e1fb0cfab0a0060ebad6dbdff423d593edecae01966d4b

Observation 845a691d-df31-4f89-9d08-7449379b87df · outbound

This paper cites Qwen Technical Report.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models Qwen Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.623102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.623102Z digest=sha256:3dc8b60eec03393a092b63402f49994f0ae16ccb07a33675818992b513aedf77

Observation 75752dd1-39ad-493c-888b-7bad0658224e · outbound

This paper cites DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.645267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.645267Z digest=sha256:1b00eaac4790fb9db04408f634c4e8b8c4696bb129f7f6504ffda7b346e519b1

Observation f9c6fa5d-81b5-4426-873a-e0eae9c2790e · outbound

This paper cites PointVLA: Injecting the 3D World into Vision-Language-Action Models.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.635037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.635037Z digest=sha256:b0c42a444c5ea68e873d8ea5c07aa454551646f26842fbb5e30bd2dde9ab34e1

Observation 48171647-2b2b-4126-b3c1-14667dfd248c · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.626399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.626399Z digest=sha256:f872cb0588b73db52588f9952869337b66ad96a7f8e66c18dac70864c177f45e

Pith citing papers

Observation 62c69a0d-4cd7-467b-b9d5-3b8d2080edb3 · inbound

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation cites this paper.

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:43:24.471673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T20:43:24.417901Z digest=sha256:5f29f4b68fde3639d6a4ed10371799fd848c35e216a43684a8bb9a96f4018a6b

Observation a014a169-5699-4e65-b9f4-7d9f9f2c16aa · inbound

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization cites this paper.

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:20:02.012837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T06:20:01.885711Z digest=sha256:d3443e671a78f7c560b971c374894713daa1a150eedf9024906ccd4b7f90b929

Observation 1e1e29c0-0145-4860-8940-ad6e07cf9bed · inbound

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization cites this paper.

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:55.445345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:38:55.445345Z digest=sha256:7cc20eae10c9c6deb742dbfb98e591c05fe44b10cc8389a1e2a3da0535f674e9

Observation 8d85dd0d-8aee-4bdf-875b-303f0b8d2231 · inbound

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models cites this paper.

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:44.272979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:34:44.272979Z digest=sha256:766a57b9e9716f17ce3ce1ab90d0eed24cf4bd78989316b8c9331e4d8853f123

Observation 71d97c6e-4e7d-4b10-b9ff-f83eedfed65a · inbound

A Pragmatic VLA Foundation Model cites this paper.

A Pragmatic VLA Foundation Model GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:16:59.751177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T21:16:59.712007Z digest=sha256:3d3c662e1dd47985c54ab64d53c2fcaddf3b33d75f30392c082ecfbe3ac46c38

Observation 83ff0ae9-ccf5-42c4-9dee-05e79c17ddda · inbound

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes cites this paper.

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:00:48.382348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T20:21:45.365156Z digest=sha256:500db66f44b8cb34e213cce3a2dcbb263a4ee7d905eb8e8bd29b70a8b667c895

Observation 7d9bab16-e82e-41f8-93dc-9dcfc1d2a85d · inbound

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment cites this paper.

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:55:52.326722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T19:27:12.286456Z digest=sha256:c2f69635b48a022506e81069f12b11b1f4f1ca41a784e4bf8fbdab1a9a070ee2

Observation 48bd2dfc-665b-498f-8670-62cde3646a6f · inbound

Robotic Manipulation is Vision-to-Geometry Mapping: Vision-Geometry Backbones over Language and Video Models cites this paper.

Robotic Manipulation is Vision-to-Geometry Mapping: Vision-Geometry Backbones over Language and Video Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:46:01.810013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:20:47.418874Z digest=sha256:5da8f4b4741f2fa545c5c0df736a3fe735fcd80afab0cceaa0dd2afd3393ffa9

Observation 3c4e72ad-2694-47f3-8c08-22e5abc0a37c · inbound

R3D: Revisiting 3D Policy Learning cites this paper.

R3D: Revisiting 3D Policy Learning GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:00:04.200200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T10:57:54.273960Z digest=sha256:d3204968bb6d8f7c6118e324acd0a9efbc8b9c12f796ee83ce34f799a41a382b

Observation e8d9e6ef-2437-443e-abcf-a7a7f98caced · inbound

PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance cites this paper.

PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:29:47.564969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T00:04:16.073230Z digest=sha256:01b0dc512805f78ca7efa3efab252d48b74a2adcd94288b8e3f24fbc1cd1be6c

Observation 19e63f52-6734-438b-a66a-4cbef4048519 · inbound

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising cites this paper.

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:26.281607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-07T10:40:04.767657Z digest=sha256:1b10ce96b676c2af517bd47abbc13d7a09d24e27f9453038f82c721211910a5b

Observation d57c6f33-91f7-46fa-97f4-0e8527a07f2e · inbound

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising cites this paper.

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:12.310771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T03:20:04.435904Z digest=sha256:c6e4c7975fbf0f257bdbef84c69ea0430e0f8c91602fd5f9a9f925ca95b0cc5c

Observation fa371804-0d82-4488-a3c2-42452d784ade · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:06:06.398109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:b2680b2f7d8e5d00350f50533e2e1c6aeb37d45fab6e5f30331dd5aee9e099a5

Observation 27f97e89-9f07-455e-b991-4f3b9dbc9460 · inbound

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models cites this paper.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:26.179579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:5090461f5c99789e590ef7371b67119aa655cb9d54cc662d44fff78930b61445

Observation 4fcfcc17-bca0-4983-90e1-8e69f46bb3db · inbound

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation cites this paper.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.495377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:e43490c5831dc21b28a0889b719aa17e292d47a1191e4be613f9868a5c37ae78

Observation 230cea23-6c8d-4fe1-8421-db1277ed9a20 · inbound

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction cites this paper.

X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:52:17.039700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T04:50:07.792757Z digest=sha256:9b9afa20abab205a33cb652d080480109a222a95fec3b122289984006326c4d8

Observation af347835-7fd8-4fa8-b65a-10715cf232bb · inbound

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization cites this paper.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:57:12.743980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T03:57:10.338617Z digest=sha256:ffb5b174c44ce77326711110bf86905570edbf6025589d50b307352b39a722c1

Observation 02fd77d4-1ee8-499e-9500-af46074635e6 · inbound

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization cites this paper.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:46.741680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:99bf0203279e80481afb627b78c8d4f6b5218012290abe58bc9a7ff8ff326cb8

Observation ff736a69-2844-4053-9402-17fddd5d9379 · inbound

PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction cites this paper.

PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-21T03:39:29.563473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T03:36:23.117616Z digest=sha256:c51604a2444f60ad8274d1059ff03c45335cee4149e5b096e7bd6f369b78a93d

Observation cdff4942-d152-4ee1-bdc1-672ff08d89e0 · inbound

ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation cites this paper.

ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:13:15.294796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T06:57:41.245418Z digest=sha256:ec6ed6b642dc79813b5a2245cfd9f99caeb42518eacf6a955225597b6c2cff80

Observation 774ff56d-cb86-4531-ac0a-e37f7093b0b8 · inbound

Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning cites this paper.

Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:20.364540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T14:41:50.084254Z digest=sha256:a6b686bf21828e00d8f3d3b7a8202f69f913fb15f99cec406c09553ac3211ba8

Observation f163582c-b0d7-43c9-905a-83044c62c510 · inbound

GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models cites this paper.

GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.193553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T10:04:10.627420Z digest=sha256:8a9e12c05098ac7263578b57ab3fef2244482c3d901e34b8119943d8ec6f6bfd

Observation 3fcff667-0757-4581-994a-7d30ee88679a · inbound

3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training cites this paper.

3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:36:44.079773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T07:20:40.998337Z digest=sha256:9fb5359cc83b33c22f2536b7129379b04c778b4beffefdf4e2018ba93857b367

Observation 3921b61e-cb8c-4e63-85b0-ca7ec9126dd7 · inbound

Robots Need More than VLA and World Models cites this paper.

Robots Need More than VLA and World Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:46:59.303478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T01:01:33.530167Z digest=sha256:34d22ea8b62c798f9f11da73af68249365d74553333ccff312698e7e75e90392

Observation 8e4205c0-1ddb-4ee1-8255-a97cdf6b1ef5 · inbound

MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model cites this paper.

MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:25.707491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T19:24:59.678266Z digest=sha256:35818dbb20674fa68543706eb32be83215243451534cd93ac6ed4f9e1f5cbbc2

Observation 9143f34f-a725-4658-85f3-38421e58cb5e · inbound

MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models cites this paper.

MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.294172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T16:13:19.238535Z digest=sha256:254fd6a9ab600622cdea77124856b689df0092dfabab5d0525d9c9e56cd2843a

Observation deb79e72-96fd-4fb6-a5db-a938de7c19f2 · inbound

VeriSpace: Spatially Grounded Action Verification for Vision-Language-Action Models cites this paper.

VeriSpace: Spatially Grounded Action Verification for Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:17:41.839593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T12:47:35.314486Z digest=sha256:6eb0e04c67484b49c21e53c73cdd75b7fa58e152935b18f3d7bb03ccef677d1f

Observation a0f12d5f-9ade-463e-ac63-4ee602e7259a · inbound

GeoHAT: Geometry-Adaptive Hybrid Action Transformer for Mobile Manipulation cites this paper.

GeoHAT: Geometry-Adaptive Hybrid Action Transformer for Mobile Manipulation GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:33.009592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T06:37:21.516021Z digest=sha256:a0bcbb35056eaacb51036dce2185034b269c94107d74a7433a5359a53e545600

Observation 51bb3c66-f216-4640-a45b-a17b4c375d58 · inbound

Geometric Action Model for Robot Policy Learning cites this paper.

Geometric Action Model for Robot Policy Learning GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:38:43.942419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T03:59:17.983035Z digest=sha256:a6985cf0516303ad960978c32b212b3542c854f17a15e89acd1ef7332ee47011

Observation d36d26c3-0f0d-41e6-8210-ae9cac506e7e · inbound

Learning 4D Geometric Priors for Inference-Efficient World Action Models cites this paper.

Learning 4D Geometric Priors for Inference-Efficient World Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-11T14:23:57.266710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T14:23:57.266710Z digest=sha256:bc2e882208529866a90bba6a655d52695ba313d0a3999f6755fc2b684c999472

Observation f7dcc86c-a1b0-4b2b-a6fb-29884bb9807e · inbound

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models cites this paper.

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T05:11:57.092685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:11:57.092685Z digest=sha256:adc5e03d2725e105b7ac4ffff8551fd95a456ae06a4d1bf2fd9eac8b05d0d510

Observation 856e1487-93bc-4916-9ee1-2bf9a0884165 · inbound

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation cites this paper.

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T06:36:26.629789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:36:26.629789Z digest=sha256:34e278a5550af3e3cc28a5cf66d8e870cd194bee989d9c32caa6942d861d7f8c

Observation 049d2ea7-780b-4e41-95b4-8a945fbce450 · inbound

SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models cites this paper.

SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T01:10:40.118599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:10:40.118599Z digest=sha256:7be12f106aecbde949e98a5a2a00222027d0cf45ee3ea4bebdb803a1c472ce5a

Observation 68d6609a-eb1b-4dcc-a7a6-803c180eb262 · inbound

Multi-View Unified Camera Fields: Geometry-Shaped Action-Facing Representations for RGB-Only Multi-Camera VLA Policies cites this paper.

Multi-View Unified Camera Fields: Geometry-Shaped Action-Facing Representations for RGB-Only Multi-Camera VLA Policies GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T15:08:25.768932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:08:25.768932Z digest=sha256:9a092ebfb3f090cbf562271a0d13189d3be1575672e0346c3a5aedf36e3cf01c

Observation 7dd3cd96-9da4-4c72-b559-74a2723c3c96 · inbound

VR3D: View-Robust 3D Representation Learning for Aerial-Ground Person Re-Identification cites this paper.

VR3D: View-Robust 3D Representation Learning for Aerial-Ground Person Re-Identification GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T03:21:35.555813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T03:21:35.555813Z digest=sha256:3b69f2dcb6659bd3a7f7f29fa1b94bc7b72d31f1c08d4398d3f38be78780f330

Observation db459a88-3610-46fd-9aaa-6187dbaf2234 · inbound

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models cites this paper.

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:19:46.803171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:19:46.803171Z digest=sha256:8bd2050f13cb0a11119d346746a43f01ec5abebeef922bd1cd28f24e1226e195

Observation b7df52d0-2bea-422c-a38b-49f06c97be77 · inbound

DreamWAM: Beyond RGB Future Prediction for World Action Models cites this paper.

DreamWAM: Beyond RGB Future Prediction for World Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:27.996941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:59:27.996941Z digest=sha256:a6f54778c76af72fdf79b91dc879d767b141e22a35af57a880848ac2288cb422

Observation 684f71ff-c72e-4af1-9211-ad0792094b98 · inbound

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models cites this paper.

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:33:38.696655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:33:38.696655Z digest=sha256:aa7b88c30df209a6ffdf8f20cd7b2aebd195ef3e868ce15a8ba00d672e4cdc6f