Pith. sign in

Paper Citation Record · LEDGER

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models

As of 14 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 4 inbound Pith citation observations for arXiv:2605.19678.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19678 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T04:48:39.069675Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:27:58.462125Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact28
  • verified fuzzy18
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation e57640bf-08e9-4183-ad62-debaad09c2fb · outbound

This paper cites Robot manip- ulation based on embodied visual perception: A survey.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Robot manip- ulation based on embodied visual perception: A survey

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.505050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:374c9108430661ce8df4f2792b98abee2ab46c7fdfd7edc79e697113158eb130

Observation d1acaf5d-ebe7-49fa-b3e8-ebc8a2b25beb · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.469762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:f3043abf1d2ec19697a2c6c0d330b7890dc26312d1951a1371e129fe57190bfc

Observation 18cd10a6-cb70-42fb-b6f3-6442a03c0f72 · outbound

This paper cites Pointnet++: Deep hierarchical feature learn- ing on point sets in a metric space.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Pointnet++: Deep hierarchical feature learn- ing on point sets in a metric space

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.500887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:df86555020f5cc887d56ad3f3abeae6b75279af6e169cceb6d595fd38f1620c0

Observation af129ecb-b1ec-4f44-b6d0-53f56b6f953c · outbound

This paper cites Dspnet: Dual-vision scene perception for robust 3d question answering.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Dspnet: Dual-vision scene perception for robust 3d question answering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.483514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:df4cc6c8f752acfbc8d6326b38ae2221bf41fdadf5f2850651bed30dd1cad4d7

Observation 99f3fa20-5457-4054-a13a-3bc8c71d2708 · outbound

This paper cites A Survey on Large Language Models for Automated Planning.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models A Survey on Large Language Models for Automated Planning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:53:05.028904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:2c1eeb626829ca47086f466ded13b4a1cf1c5bd534911a72ff8d3b10c394da44

Observation b41d9bf2-095e-4f6a-9512-fe5ff9f4ad38 · outbound

This paper cites Qwen3 Technical Report.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Qwen3 Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.098240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:841d838b53e306fcd56374375652a72f1b467effad6ba92cb922b4e1892214da

Observation 3db5227c-06f3-44f0-942b-b0599a6a8346 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.955065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:4af7da67711cd12e48573cc436fe0ab9d77933e0630f259a12082fa048f7fb9e

Observation b4914052-5b62-4881-b68c-41af3e22c01a · outbound

This paper cites Scalable diffusion models with transformers.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Scalable diffusion models with transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.484971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:0540e1cab1f047c597785cc7e5d57cbadacd13aa407a9cdc5b93d5f30b021a8c

Observation f2708e3c-5574-48c3-94c3-0094d5ab062b · outbound

This paper cites Flow Matching for Generative Modeling.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Flow Matching for Generative Modeling

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.961336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:b4061a945e8f424746e3fea459ce386e60d62039eb8ee1732ef431aaac929f27

Observation f875bbee-ee14-4c57-9460-60ec2d321e7d · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Diffusion policy: Visuomotor policy learning via action diffusion

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.491411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:baf13d610ad345aadef51e1387bb1bd9b97d6a98e65e0ce8926dbe97550aa4c0

Observation 9999ba4d-16d6-4794-83c8-b0537da292a3 · outbound

This paper cites A Survey on Vision-Language-Action Models: An Action Tokenization Perspective.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models A Survey on Vision-Language-Action Models: An Action Tokenization Perspective

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.059092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:9814e35a0cec66d226ec361a59f4497405ae3bf2ff80935931551fc11c1e2d3b

Observation 4a2abc5f-5600-481a-8793-8fb32cf11e7c · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.034940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:2aaf6a64de4d06563684eedfb1eef74047b5f4e952dc0567e0eb4da9c76efcb2

Observation 8ae1af9a-4cc1-4cd8-b608-32cbbeb489f5 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.980530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:5530f03bec95c8da353c3fbd2d7d367d069504ab271df516b4e5407972832c41

Observation 8d9e74a5-ccef-4af3-9c75-182bc6dd3547 · outbound

This paper cites InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.948828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:86fb21c5195231cf7b63ca40eb8454be728046aebb230b0a8eefe4e91aa0983c

Observation 3804e7f5-f1ed-46b1-b6a0-4fca197f46a0 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.084084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:71ebd705996cee1f0458560a2ec6984d8a312d80dbc8a8970882ec85adcf95af

Observation 38d43cec-182e-4832-bbab-38dc98379317 · outbound

This paper cites 𝜋0.5: a vision-language-action model with open-world generalization.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models 𝜋0.5: a vision-language-action model with open-world generalization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.473722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:efd8e0acabcc0f66bdee21f9965ef4572541b7c30aa5f41f8fe166dfa170e820

Observation 2ea9bde2-168c-42eb-8016-4e8055baa41d · outbound

This paper cites GR00T N1.6: An Improved Open Foundation Model for Generalist Humanoid Robots.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models GR00T N1.6: An Improved Open Foundation Model for Generalist Humanoid Robots

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.486709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:e5e06aec8d5899eb703579c7037aa54f464cd78d6962bd529d27e901fb79f483

Observation 7f251330-dae6-420f-8dba-36fecf44ff78 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.065003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:041c327c61c68a0c7159ddb7385592ba45161d3930c32f6d8d5485beed17c0d9

Observation a1cfc27f-2e73-4acd-8106-53cf64716e85 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.465903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:4e822e4d10090556ee178b505927af3a135eb2190fcdbda9ce0f7a81a1591a2e

Observation c20cca0b-4517-466a-94cb-990df1101f8a · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.016329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:2c4a843f645b8abfb8c02cd7c7962e34662caf88eddb25d478b419cc5032dabf

Observation 4eb40f09-a60f-4907-b16f-d8f321b61cea · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.942661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:a42abd0bbf29a873fd1d1f9ce1b736bf45851cd1c8ab7337928ba02393d76756

Observation d31d1d5e-4cfa-4f65-99fa-ecff09788d4d · outbound

This paper cites Exploring the adversarial vulnerabilities of vision-language-action models in robotics.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Exploring the adversarial vulnerabilities of vision-language-action models in robotics

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.493662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:4c034777dd764cf2851c0661337eee685279d1a9bc6a19322d9d352ed07480c3

Observation 7fac4a71-de89-4aaf-bb46-0c121708859f · outbound

This paper cites Instructvla: Vision-language-action instruction tuning from understanding to manipulation.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Instructvla: Vision-language-action instruction tuning from understanding to manipulation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:53:04.968738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:65f40c61121b05dd56dbcb431956baf471707a4e9dddda0890db925c8f2cf8e5

Observation 36728258-acf3-42fc-b438-0cc387b4ead3 · outbound

This paper cites Interactive Post-Training for Vision-Language-Action Models.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Interactive Post-Training for Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:25:47.344448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:dac18c1e43732550c299cc4f0640b4c01f5752a2a07386891b7934122d86c50e

Observation bdfc9829-895f-48e9-9865-735dd8eb8ebd · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models WorldVLA: Towards Autoregressive Action World Model

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.936624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:3b5baa2b49a71e7fd47b82acc048425e30022d507d7cbb22f047fbb15606df15

Observation 8d9aed89-8314-4974-815b-6ecdb9d97484 · outbound

This paper cites Unified Vision-Language-Action Model.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Unified Vision-Language-Action Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:53:04.974302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:7d89fa1b7306cdadf795c7d0f84f2d4c7821c24c28c5a61df317efffd7637e7c

Observation d9def460-4881-4c81-a809-3ed73587fa4b · outbound

This paper cites Aligning cyber space with physical world: A comprehensive survey on embodied ai.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Aligning cyber space with physical world: A comprehensive survey on embodied ai

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.444517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:1b11ec217837a72ca6f57ead0bc444b4a2f7886a53948313b8346e8d21cf1984

Observation 5adfb2df-49cb-4ed2-8800-96e320924b09 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.466587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:4b0b3ca0d291971b83abd24450ee7e5abc9ef26049f71f3a09b998873cf8f23b

Observation 100cdfe6-13e6-4f98-a106-dacec8f07232 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Octo: An Open-Source Generalist Robot Policy

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.921971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:1d6b618af41db528ce79df266e8ce56e0300f9a691100a6d6ca7c37f48b4ef4f

Observation de6ad2ba-2729-45f1-b6b7-23bbae718b82 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.994309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:7e8fcace1ef6a7c0be3db1e854d1d9814e9bdb31285534d937c08e74797fa586

Observation 1bbed8e8-9328-4cbb-9cdf-cc0c8222db22 · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.916497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:72433216ef9e82b32d14a25eed7cdecb8e089079a6ae27da6897c68b131bf636

Observation aaae344f-c28c-4e9c-adf7-157f92a7a0d1 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.988204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:47bdf112a7280f5f19c0e77e9964a658afe68e1aa6645a801bb0941a656f6879

Observation 63d9bbee-b054-4230-8624-9a32f21dd066 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.053263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:26291bce12c9c9ca4b7caee89f504288fdc2bb1df54afa113c42125b854debd2

Observation 2973567a-d2b0-4726-94cb-f1b680de1706 · outbound

This paper cites Vlatest: Testing and evaluating vision-language-action models for robotic manipulation.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Vlatest: Testing and evaluating vision-language-action models for robotic manipulation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.502640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:d38b3f128f9e6cab06c66091dc0c2922899c4919ac6d530bcbb3281fa2745abd

Observation 430ae5d8-3844-4f86-887b-010754764c0b · outbound

This paper cites RynnVLA-002: A Unified Vision-Language-Action and World Model.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models RynnVLA-002: A Unified Vision-Language-Action and World Model

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:36.254972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:19ee8d8f89ed088d448765a1ceee7646e0a8f530340f9a1ff08c0cf730517788

Observation e962c671-c9d7-4baa-b19e-ac5311ca4fd9 · outbound

This paper cites Motus: A Unified Latent Action World Model.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Motus: A Unified Latent Action World Model

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.909357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:957e29d2db87324289775b23a4a8b66dc9958a125f8b7f52fa14d4bee9611458

Observation 56dc571f-f0fb-4aed-8b64-3999e63c1f19 · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.074152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:a5880ff783cb322f93f37ab56b7e78ea52a30dc9ee830e9962857bb910d44c60

Observation 7fd2a87a-9997-4dc3-9544-66ab4c0c7f21 · outbound

This paper cites SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.009602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:6dc3165791e0a0173409629b98fc9ddd5bddbf3c41e6de570dc7d51d37eef5c2

Observation 1b34c62f-a977-4c36-b31c-df5317b9ebfd · outbound

This paper cites Robustvla: Robustness- aware reinforcement post-training for vision-language-action models.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Robustvla: Robustness- aware reinforcement post-training for vision-language-action models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:53:04.928533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:ef2266de11fe8a923a573f62f1431b4d11b199cac086bfa67d8feaaafb4e28c8

Observation 967c2efc-9665-496c-bb1a-d3a15c81b749 · outbound

This paper cites Mean teachers are better role models: Weight- averaged consistency targets improve semi-supervised deep learning results.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Mean teachers are better role models: Weight- averaged consistency targets improve semi-supervised deep learning results

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.490131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:29dbdeca8ca3cdfeec40acfa45f71be11f9c2161ae399b9110d2865381b76023

Observation 9f6a548f-18fb-4bcb-8ec1-ea1abdd79833 · outbound

This paper cites Virtual adversarial training: a regularization method for supervised and semi-supervised learning.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Virtual adversarial training: a regularization method for supervised and semi-supervised learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.481359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:194d08beade06552ba9e5aa7bc2b0bbd52482ee3a62f8b587e9bcef1732d6f88

Observation 044070dc-4a57-4f35-84b8-249db13d47d0 · outbound

This paper cites Fixmatch: Simplifying semi-supervised learning with consistency and confidence.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Fixmatch: Simplifying semi-supervised learning with consistency and confidence

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.497406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:46e189551f926c02bf26fa59ac95310f5d73d0e3c18f3692fe441ac853ea18a6

Observation 62bfdd38-53b8-4dfe-b9e2-150272b6debf · outbound

This paper cites Image augmentation is all you need: Reg- ularizing deep reinforcement learning from pixels.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Image augmentation is all you need: Reg- ularizing deep reinforcement learning from pixels

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.488231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:6ce1ab91ec3518626802793f523be10b07d3343c3b4cf20037f8d5694926f9e9

Observation 6be07c4e-cdab-407c-a83f-ea04d3444c97 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.090887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:88ca59eded8d195fc48dcbcb534c37d5472dddf095badd3b64db3bbb6c59474f

Observation 8ea2901a-c186-40e0-bcb5-984fc0e36e62 · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Explaining and Harnessing Adversarial Examples

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.040820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:b034ee4684e90b5d427b96b867e48ace52348afdf834af40940f6cb88610f39e

Observation 5e8f34f8-3642-4613-a558-7c3dacab4437 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Libero: Benchmarking knowledge transfer for lifelong robot learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.498790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:b7abb44ed3fd8015e81bcc14c994448f8a194c7a4d69df626f5455225ea9ac1d

Observation bc487d13-178a-4d5a-b86a-e50d35ff3eed · outbound

This paper cites Decoupled Weight Decay Regularization.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Decoupled Weight Decay Regularization

Reference 47

Resolution
malformed identifier
local_arxiv, observed 2026-05-20T04:53:05.021961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:a7e319159b781c810184afa3d4d677b8944ca908a39ef0249a6e3ef200d1e188

Pith citing papers

Observation 5e01aef6-0534-4d92-b406-61603fcd4d26 · inbound

PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution cites this paper.

PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-08-01T20:33:34.710891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-01T20:28:18.919704Z digest=sha256:f699a2c9041a8a0a3309e3f774e0163ecfa8a97a04ef6b5f789b1e08525cfb6e

Observation 70fd44b0-e1af-4694-89c6-37006110f8a1 · inbound

Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models cites this paper.

Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T06:10:33.302184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:10:33.302184Z digest=sha256:304b8c37f74e1c0e5516cd3c0bbf3adafda4f7728413c94a442f436b1a83d94f

Observation 75fb5270-f809-4809-ac2d-dcdb8efcabcd · inbound

Probabilistic Reachable-Action Verification of Visuomotor Policies via Set-Based Training cites this paper.

Probabilistic Reachable-Action Verification of Visuomotor Policies via Set-Based Training RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T05:07:02.883441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T05:07:02.883441Z digest=sha256:e6916cbe9c12a5879cf9808cfad819b0be63777b514533193dd913e81f8122c2

Observation 926b2cab-789c-4ed3-8301-018a6cbbb54f · inbound

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling cites this paper.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.462125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.462125Z digest=sha256:22c02063e90d3da48e87b3e011d220754183421c1230bcf78d720d50daebc639