Pith. sign in

Paper Citation Record · LEDGER

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models

As of 10 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2605.10485.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.10485 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T05:09:21.028373Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:34:31.166730Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact19
  • verified fuzzy30
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 288c0cb6-07f0-4c68-9fa7-dce3b86da573 · outbound

This paper cites 3d cavla: Leveraging depth and 3d context to generalize vision language action models for unseen tasks.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models 3d cavla: Leveraging depth and 3d context to generalize vision language action models for unseen tasks

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:36:24.923919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:8ce252122ff05532d3e2869fe64e47244d0349ca2cf2ede14673b2d9dae0bd3c

Observation bf4bf8f1-22f4-4bcd-b2cd-4fd870ece669 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:36:24.912557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:61f390fddf2461e5cb999c2ccbee2c2d0d55bd47b5d07fe03ed459b0d324c404

Observation 86a7a6e7-2c61-4028-be60-8a39c949787a · outbound

This paper cites π0.5: A vision- language-action model with open-world generalization.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models π0.5: A vision- language-action model with open-world generalization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.065834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:6477c9c93b434093175076f52807c3c30bf18b3b502ec36e76f55102bce31d47

Observation f48eba0b-4134-4598-8b20-2b2756793505 · outbound

This paper cites π0: A vision-language-action flow model for general robot control.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models π0: A vision-language-action flow model for general robot control

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.069312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:29d011c819602790441225416b0f891f68e068980dd86c00f618de2285d27447

Observation 9fa21a8b-2f14-49c7-be3f-87dac521f47d · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:36:24.888491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:bdb60619f2d99b13f1adb62461c81b6e6abfce997bce7af15abcf79d5bf53c82

Observation 77d23117-d408-43d4-9d1a-9757cd1f4a01 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.061909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:32f26152976d33884f4a8001991e4471c28b12b6a468932dcd4de4adefc17b2a

Observation d3f023ae-bd48-4772-a659-21a03d8fb49e · outbound

This paper cites Knowledge distillation with the reused teacher classifier.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Knowledge distillation with the reused teacher classifier

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.071961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:882a226bc08f755d025b867ffb82680628d5110d4ed24e3dd31e10ed228c3e47

Observation a70c30b5-eb5f-4749-a389-89f8d8610c2f · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:36:24.899722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:47d34442429921f10532c526331f2f6241f3f897a996fd898b75eeee46fbdf83

Observation 52b97257-e3af-4ce2-b985-68a1af823548 · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:36:24.871267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:ae338827c8ab34de8c72ec09dedf200e518d0945ae0541b764738090b54927ef

Observation 6bc36c34-3ce5-459c-8525-608760fe5bb3 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Diffusion policy: Visuomotor policy learning via action diffusion

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.074517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:f87662786a99a95bae4796ec84bb1e897f0e4cd1b9728ad87a4151e60acc3fdc

Observation 454005b6-74b7-4281-a547-e1eeef9b7a20 · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Objaverse: A universe of annotated 3d objects

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.148408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:1748f45520c2189e129d297b1205e4054e5246ec71bacff7e1ad8b76b36a74cc

Observation 0c08331c-e246-4fa6-9091-a023d0744d34 · outbound

This paper cites Rvt: Robotic view transformer for 3d object manipulation.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Rvt: Robotic view transformer for 3d object manipulation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.164911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:5a22904d92986c5f2018a21eeaa79613af6e612b5cba2d6de82d65dfca5418cb

Observation d886c025-cf7e-4a93-8481-fb7bd91d3c31 · outbound

This paper cites arXiv preprint arXiv:2512.09619 (2025).

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models arXiv preprint arXiv:2512.09619 (2025)

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:26.193107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:86fdfcf8c1d1014d20604ed58aadab7596530b26e35a625aca210b074e701fdd

Observation 715b4cce-4370-4acc-9fc9-1b6d8a1b2c7a · outbound

This paper cites Lora: Low-rank adaptation of large language models.Iclr, 1(2):3.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Lora: Low-rank adaptation of large language models.Iclr, 1(2):3

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.153632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:eac26b191f62b2a64ac59c3cc890bdecf0c566a0c04354bc1b3a4884c2714f54

Observation 940efacc-92b9-400c-bc45-90cea4912d07 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models An Embodied Generalist Agent in 3D World

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:22:18.773673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:4d1a5daad1f956ac336c08aceec8f42c6031375d42b12656a2f85d2712413349

Observation ba798ace-4dc1-4890-a2af-bf2dcc5632aa · outbound

This paper cites Mllms need 3d-aware representation supervision for scene understanding.arXiv e-prints, pages arXiv–2506.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Mllms need 3d-aware representation supervision for scene understanding.arXiv e-prints, pages arXiv–2506

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.137843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:c2342a355d2fa2408ebed9d6af89661b53cc71c3bd49fae15befefc10d270270

Observation eab724c8-8dae-4cc4-9b6d-888b28e2c9f6 · outbound

This paper cites What’s “up” with vision-language models? investigating their struggle with spatial reasoning.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models What’s “up” with vision-language models? investigating their struggle with spatial reasoning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.101141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:06216dd6529dc195f7519f98cec3c1d89c2cae1c53d3bf93a42115d74437ca02

Observation 7ab7566c-37b2-44ca-a372-912b59a37155 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.104776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:b6f189562aaa755a3f3313bd2b047d7fdf73345e9d24deb6ed37b94a5ec559fe

Observation b45b0076-b0c6-4a4e-a783-2dddc83475aa · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.ACM Trans.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models 3d gaussian splatting for real-time radiance field rendering.ACM Trans

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.108503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:77fa57d3644f85f65bd692a0db98195cfb4d6966616f49f4b74f22a28e99bc02

Observation df3b2a92-fde4-4ee7-8998-fa1d832ede1a · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:31:26.205311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:bab0e2d52599d046440191405883dd626993df2e8adbc05277abfb72440c51b1

Observation 3bea8396-6ecf-4c27-922f-6191b3b69979 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:31:26.139493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:be46bb975dae633c028160c4b6d928906e526699f3c845dd6d2f3efecaf94782

Observation 2506cb1c-e9dd-46af-bc32-65ff85fbb04c · outbound

This paper cites A review of robot learning for manip- ulation: Challenges, representations, and algorithms.Journal of machine learning research, 22(30):1–82.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models A review of robot learning for manip- ulation: Challenges, representations, and algorithms.Journal of machine learning research, 22(30):1–82

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.119351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:93d8b00a33c5a2cf261b290c78dd37cd0594fb3a9bf31dff0ae5b8d268bf5308

Observation 8175d114-beaa-4b83-bc1a-5511238af15d · outbound

This paper cites A review of spatial reasoning and interaction for real-world robotics.Advanced Robotics, 31(5):222–242.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models A review of spatial reasoning and interaction for real-world robotics.Advanced Robotics, 31(5):222–242

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.123395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:e49401a3694fa33620ae79c2b5b187b8e7ce61e39fa6844348911223dc9fc6c9

Observation e5835fe9-69b1-4615-a6eb-99a27f7d540b · outbound

This paper cites Pointvla: Injecting the 3d world into vision-language-action models.IEEE Robotics and Automation Letters, 11(3):2506–2513.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Pointvla: Injecting the 3d world into vision-language-action models.IEEE Robotics and Automation Letters, 11(3):2506–2513

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.141966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:c820744e48b38e816c77a322c0311a835a90e1ad6bb5c80b2aba84dae6958e68

Observation 2a8b4649-1302-4a36-8aca-c3dedbe5e7d3 · outbound

This paper cites Spatial forcing: Implicit spatial representation alignment for vision- language-action model.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Spatial forcing: Implicit spatial representation alignment for vision- language-action model

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:26.188928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:eae26e6704a10dfc3a295ccd103177109369829cd3f61bae34c9e3df18010181

Observation 1f7d639f-3222-424c-9e96-3549df67ffae · outbound

This paper cites Evo-0: Vision-language-action model with implicit spatial understanding.arXiv preprint arXiv:2507.00416.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Evo-0: Vision-language-action model with implicit spatial understanding.arXiv preprint arXiv:2507.00416

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:26.174886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:00a0eaef05949577cacbb581b9e2567bf014995f5c30a3da02f5c15e6eec4dd5

Observation aeb4f825-3c86-4c40-a6c6-ced04b25825f · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.093663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:c34b9154eb858b499dbad4769c6f242e555a8e7a5574f03d02d2660ff1603114

Observation dbe476cc-2772-48f6-b10c-9808912ccfa6 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:31:26.219409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:369bf95354b7ac2a8bc1588cfb24bb17727e204381961620c932771d9ba5dde3

Observation 5f0b34c4-c348-44be-8e12-221c47f87d3d · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models DINOv2: Learning Robust Visual Features without Supervision

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:36:24.928992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:8fe538b7b22ad55b37a322f6fe1e8f67904ebabca8d7296219ed992eae00ea1c

Observation 734766e6-bd05-4d34-bfbe-5cc6ed4bc686 · outbound

This paper cites Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.145253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:ff40abd4613ee19f40c6984beb7e059b9c84c7b505091765b6f6428811632ff6

Observation b1798aea-963e-4d10-9c5e-9539590e036e · outbound

This paper cites Film: Visual reasoning with a general conditioning layer.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Film: Visual reasoning with a general conditioning layer

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.157223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:93b324bd16e49e6f6098845d3b95ea175ad5193ffd4d9256862d7fe900aa7bb0

Observation d5e96b2a-3d85-4d7c-a0e7-fc27ade84f9d · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:12:22.898827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:9167f514d2e0022a7334515654c378219fd8652a6a760416a30d705e1faab729

Observation 7304605e-475b-4f2f-ae88-4d10a8a2145c · outbound

This paper cites Perceiver-actor: A multi-task transformer for robotic manipulation.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Perceiver-actor: A multi-task transformer for robotic manipulation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.126906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:ebccaa97e25399c460f994a57bb786bd277862307ddc94b888dcc69b4f6ad2be

Observation df22501b-42a1-44d2-b133-cf43852717d7 · outbound

This paper cites Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.150931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:507502a76dbcb36f46697948d281c6ea516e44ec39c47b0caed88258b161bca3

Observation f5979ade-eb3a-4b41-93fe-c6b66ec3e213 · outbound

This paper cites Rocket: Residual-oriented multi-layer alignment for spatially- aware vision-language-action models.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Rocket: Residual-oriented multi-layer alignment for spatially- aware vision-language-action models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:26.168690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:fd2365970a32c6d81d75748744008047c641c66bbe3c7cfd32c6009a6e8e6cba

Observation 27f97e89-9f07-455e-b991-4f3b9dbc9460 · outbound

This paper cites GeoVLA: Empowering 3D Representations in Vision-Language-Action Models.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:26.179579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:a8ec19012d90f32ee630088f80d88ca896233d75754fca64cb5a66cb29868d23

Observation 553f3816-b02f-45ec-b513-2f8cf2bab091 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Octo: An Open-Source Generalist Robot Policy

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:31:26.154335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:24ea1a056f07aa7cbf253d04ce008c0afc3ac82c4a614ee16f11b81e698221f8

Observation 57f4e4bb-af80-4afe-b382-23f7831ed237 · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Bridgedata v2: A dataset for robot learning at scale

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.130223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:62094b3da220e5191a74f88bb1305f16edb67775bf335a41be982af4cd9c0211

Observation 689ab1b1-4cb2-410b-b924-eac27ae39780 · outbound

This paper cites Vggt: Visual geometry grounded transformer.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Vggt: Visual geometry grounded transformer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.090295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:acc93c5686ebebe7887b1c04b1269a185cf109fab71788bc667f476df87aa1ae

Observation 482c5687-db62-43d7-8405-4cb4c7ae365d · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Depth anything: Unleashing the power of large-scale unlabeled data

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.096839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:78833b9aed777a739ab23eb3da09edb1f14894f03f5d872a5333610ae0b1b948

Observation 903892b5-6838-4477-a974-7a5381729265 · outbound

This paper cites Depth anything v2.Advances in Neural Information Processing Systems, 37:21875– 21911.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Depth anything v2.Advances in Neural Information Processing Systems, 37:21875– 21911

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.112063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:48e2cb541b73af753e4f8b24f3041994b783323bc5dd0c85a8bbb42e2a129662

Observation bbe0a506-6e8f-4564-9f9f-dce213bb1c31 · outbound

This paper cites Scannet++: A high- fidelity dataset of 3d indoor scenes.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Scannet++: A high- fidelity dataset of 3d indoor scenes

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.133373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:f2d7122e07cd992629b7ba52ac86f9aba48f783c68b2ef2cae0d85c686a81e9f

Observation 97c48c6d-4df8-4c1a-86fd-39b946f8bd03 · outbound

This paper cites RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:26.143646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:ec911996be0907eb1d5f4d9c208b0b135cf603892f9065fb4b00887247763607

Observation 18578f66-026b-4a86-9dd7-05bb52f171a7 · outbound

This paper cites Improving 2d feature representations by 3d-aware fine-tuning.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Improving 2d feature representations by 3d-aware fine-tuning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.115308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:f96feccc2b5c8737f7c17463a614154097a528fd4c5de344c47ee8083b16b4c2

Observation 9b5d40de-7725-4c42-a780-9c37402bdccf · outbound

This paper cites Sigmoid loss for language image pre-training.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Sigmoid loss for language image pre-training

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.159779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:c4c3c51b99c36d8189ec62cbf4f9b1273689fc082254fb206e3a244c9f325ec5

Observation d462be3f-9b37-4448-b91f-275895efa924 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:6cccfb0b294cf8f757421e7c7b686ec712d499269cdfdd941b0a191d37392502

Observation a8584f01-c829-465d-920b-0c1b98325f7b · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.162351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:9f2f7bdab1aede4255d0458a573ebb8092cc2b75145cc712fa262aa770224fbf

Observation ae111cc0-b8df-4cbf-b2d2-735e49b36e8f · outbound

This paper cites This task requires precise spatial perception to locate the screen and hinge, as well as smooth and controlled motion to avoid damaging the articulated structure during contact.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models This task requires precise spatial perception to locate the screen and hinge, as well as smooth and controlled motion to avoid damaging the articulated structure during contact

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.083821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:fbece83b0fb8b30bdf06cf45a96c76e2f6273252b8e2ffe2997cadfdba2ee40d

Observation f9c5205d-ad76-4eef-adde-3a561af8b304 · outbound

This paper cites This task requires accurate object localization and a smooth transfer trajectory to ensure stable grasping and precise placement without dropping the object.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models This task requires accurate object localization and a smooth transfer trajectory to ensure stable grasping and precise placement without dropping the object

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T12:11:33.087208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:10f52ca9b8015610681a1efc1d47ac135ed687a5c104c6ae84f225ac7146164b

Observation c17223a3-08d7-41d7-842e-a7adb38e5216 · outbound

This paper cites an unresolved cited work.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-12T12:11:33.080104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:a2c65458cc89bc48a69aac00532013b2b88d901911d4cbd1a23aa0d1c7cb12c4

Observation 556c2877-d691-4ec2-a8da-7b4e56801933 · outbound

This paper cites an unresolved cited work.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-12T12:11:33.077238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:39584b9b3842234af3a6b1ff707bf2942a12e1c281d4a5cf9ca52d828bd0e337

Pith citing papers

Observation 9637b991-567d-4564-8175-23d732540a7d · inbound

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models cites this paper.

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:31.166730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:31.166730Z digest=sha256:9cb94c894f3bc13e03f8b7104dfa76a11c3f9831ac2f40eef7d627c1750f10df