Pith. sign in

Paper Citation Record · LEDGER

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

As of 13 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 8 inbound Pith citation observations for arXiv:2510.00037.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.00037 v6

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:58:38.554805Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T11:18:08.362562Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:19:48.867283Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 93d059bf-220b-455a-8068-86b1ca67a759 · outbound

This paper cites Robustness to Multi-Modal Environment Uncertainty in MARL using Curriculum Learning.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Robustness to Multi-Modal Environment Uncertainty in MARL using Curriculum Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:36.376470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:36.376470Z digest=sha256:76a1508cf0e2c00f6a96b2357037ede50234eaf9847b522a2f6ccb717dbecde2

Observation 13b5c165-cb7a-43c6-9a7c-9599299a8a2e · outbound

This paper cites Flow Matching for Generative Modeling.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Flow Matching for Generative Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.024872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.024872Z digest=sha256:09cb44059dc31e106a7e40f117fb1da650f5f42fb70ef111ea4b27999b8ffe75

Observation aed448bd-e510-4354-be6d-baa4945462da · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.250038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.250038Z digest=sha256:3f1deeb2f650e8b7e5f1ddba08c33e8b6c4946b6527847c656502a6e3b55d80a

Observation 0de6fb67-aa8b-45f4-8da8-a62bb926f755 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.522955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.522955Z digest=sha256:e3d8d3b8fe1134c139dfc1b572ffa75bd6c5e282fc7192a16bca606a50dfe193

Observation 2d5cdbe7-f040-4892-af65-3a7cf7e363ea · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.704747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.704747Z digest=sha256:4e4cacfd10baac1e200ffa1f6687461426cf9807ffd44b1ee74e28c398d2cba2

Observation 7672aea4-10b7-4194-b65f-2e08c9250624 · outbound

This paper cites Vision-language-action models: Concepts, progress, applications and challenges.arXiv preprint arXiv:2505.04769,.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Vision-language-action models: Concepts, progress, applications and challenges.arXiv preprint arXiv:2505.04769,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.834746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.834746Z digest=sha256:412b0f46af9d686cb810026fbd2312422ebe782f9fbf009c68160be8d8fc7817

Observation 528b7f4b-f742-4a42-be34-deed34c3db36 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Octo: An Open-Source Generalist Robot Policy

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.974007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.974007Z digest=sha256:2d87d69d061af78d6bfa8e0ef46f57c76fbf7637fea496e085f243053116f37b

Observation 4a1e4bdd-335c-4b95-8e03-7a2c428acd87 · outbound

This paper cites Robust Reinforcement Learning on State Observations with Learned Optimal Adversary.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Robust Reinforcement Learning on State Observations with Learned Optimal Adversary

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:38.134753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:38.134753Z digest=sha256:5223c7712f5ef368aa26505914cd0fc955b6940c8c732dd5ac19d264511fa205

Observation 07df484b-414d-4dc7-ab4c-71b1074eeebf · outbound

This paper cites In robustness against VLA input, we additionally use UCB algorithm to select the best perturbation.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations In robustness against VLA input, we additionally use UCB algorithm to select the best perturbation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:38.554805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:38.554805Z digest=sha256:00fd9ac66d0865f399634ed3176b52cb2698504fdfb497b7c49ecf2124d48b2f

Observation 623bb6f6-37c1-47b5-8afd-102391bce1c6 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:36.452731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:36.452731Z digest=sha256:2abf9c8d237fc9d8d51d38159dbaf9ea34fbbc0b7d51fbe496ac7a97acab0e0b

Observation a1ca3ec4-4b91-4633-a0fa-9f62958c0589 · outbound

This paper cites Robust Reinforcement Learning for Continuous Control with Model Misspecification.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Robust Reinforcement Learning for Continuous Control with Model Misspecification

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.384743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.384743Z digest=sha256:6dabc3e3f6472a149202669e628249630b812fa7fc972faa9830859466025527

Observation 4bf2ccd7-86ba-40a2-9672-9a0f50fc6633 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:36.912617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:36.912617Z digest=sha256:c30de479964359ad277820ea0ae669e340198ab0056bbf69f903da0718963b0a

Observation a8290e40-00db-4c03-a400-30c712074d45 · outbound

This paper cites A Survey on Vision-Language-Action Models: An Action Tokenization Perspective.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations A Survey on Vision-Language-Action Models: An Action Tokenization Perspective

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:38.253146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:38.253146Z digest=sha256:d35e60ec535980e4950a66bb9ec34525bddf52c33defe1fc1e8ae46a5f0f6055

Observation 4a3fec9a-378b-4b3f-a59a-966e6d35fca0 · outbound

This paper cites Maximum Entropy RL (Provably) Solves Some Robust RL Problems.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Maximum Entropy RL (Provably) Solves Some Robust RL Problems

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:36.682378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:36.682378Z digest=sha256:dc72a0ebf8818e855aed546172ee28581cd4f860ea61eec45307aa7e9e319616

Observation 43db3ebd-a31d-4ac1-b4c1-fd96d6a97837 · outbound

This paper cites Rectified Flow: A Marginal Preserving Approach to Optimal Transport.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Rectified Flow: A Marginal Preserving Approach to Optimal Transport

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.144747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.144747Z digest=sha256:e40d56e2036de6d6f55c9cace82c70d546164d5c4c91589035745e6413327869

Observation 368931d7-7fb9-48d5-a52f-578c8297282d · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations RT-1: Robotics Transformer for Real-World Control at Scale

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:36.544334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:36.544334Z digest=sha256:10d2faa06769587ef12a6d2688e4fb9329bf066e9c932a706142eba96b62de98

Observation 54cacf35-c3c2-465d-a4e0-e4b79061e6bc · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:36.786424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:36.786424Z digest=sha256:b75930242c957fc4be51253a5c5fdc8061c08bc491d66b74af7757ac8ec6d83d

Pith citing papers

Observation 41a02e7a-a149-45f3-9f0c-e1abf2527fb8 · inbound

STRONG-VLA: Decoupled Robustness Learning for Vision-Language-Action Models under Multimodal Perturbations cites this paper.

STRONG-VLA: Decoupled Robustness Learning for Vision-Language-Action Models under Multimodal Perturbations RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:20:31.091014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T16:33:16.554905Z digest=sha256:e0a48454ef6d25aeed64a4850335a5de95bd7d6294d69ea4c7f69b46404110aa

Observation a6ec3215-5d58-4e78-96af-43ee881f117a · inbound

OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation cites this paper.

OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:20:31.091014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T08:52:07.545191Z digest=sha256:fb3d7abe9b43a7c695d67fd7d117f6832b80bc81730c84a9bf76e000db29e990

Observation 0f2b45cd-7d25-4d2a-9543-c7061f181417 · inbound

Breaking the Epistemic Trap: Active Perception Under Compound Uncertainty cites this paper.

Breaking the Epistemic Trap: Active Perception Under Compound Uncertainty RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-21T02:20:31.091014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T16:08:23.588699Z digest=sha256:1c6ba777fabb0b66c54f673917218afc342e79c5f7aab51e6bc640b0dead57b4

Observation 3d81e558-2497-4c19-9d29-0fa4eff0c2eb · inbound

Breaking the Epistemic Trap: Active Perception Under Compound Uncertainty cites this paper.

Breaking the Epistemic Trap: Active Perception Under Compound Uncertainty RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-21T02:20:31.091014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T11:18:08.362562Z digest=sha256:7f614c31da240baeb7028dcce5ec2df1a02fcc680dabc8fb089e66ef5402d51f

Observation 94ea7ac5-e1cd-45b5-869b-65289d836159 · inbound

How VLAs Fail Differently: Black-Box Action Monitoring Reveals Architecture-Specific Failure Signatures cites this paper.

How VLAs Fail Differently: Black-Box Action Monitoring Reveals Architecture-Specific Failure Signatures RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:20:31.091014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T11:40:15.064339Z digest=sha256:4589ee3762a39eda6aa5de210b3c155a1924e2bbadae849f97e681c2689f1081

Observation 5a1d6dc5-1b70-4d21-a5e9-bf1e01fe7688 · inbound

APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies cites this paper.

APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:20:31.091014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T09:49:56.894300Z digest=sha256:d01056000c014f54f7d8ed0f4382e359c60bd23f227558c787b8e11f6cb9fcb6

Observation 321c6bde-b5df-4f2e-bc80-b8572ddd5ba8 · inbound

Flatness Preserves Instruction Following in Vision-Language-Action Models cites this paper.

Flatness Preserves Instruction Following in Vision-Language-Action Models RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:20:31.091014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T08:06:36.393665Z digest=sha256:fcd456471c4f394bf5cae5c045d64be8d22a79479ea2ba3b7206560a1be8e50f

Observation bcfa979a-1433-48c1-ac78-9dade9708810 · inbound

Sequential Planning via Anchored Robotic Keypoints cites this paper.

Sequential Planning via Anchored Robotic Keypoints RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:20:31.091014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T04:59:12.363425Z digest=sha256:78ce57af11512b33a2a3966f34e3e9af6e328d70dfec93f3b6fc3b211d561910