Pith. sign in

Paper Citation Record · LEDGER

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2307.15818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.15818 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 100 of 363 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:24:59.855729Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T10:07:00.865372Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 96658184-8130-456b-a56a-c52d5da317bc · inbound

Cognitive Architectures for Language Agents cites this paper.

Cognitive Architectures for Language Agents RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:33:44.278357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T19:33:44.146134Z digest=sha256:df495dde77d3e235879174ab41ee8e9cacd13cc1c2e79267d0c24d66c43a68d3

Observation ef194d76-9d83-44ad-b50e-1b4cf63ca8b8 · inbound

GPT-Driver: Learning to Drive with GPT cites this paper.

GPT-Driver: Learning to Drive with GPT RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:05:31.982543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T15:05:31.928650Z digest=sha256:e0de725aa44199e0b93eb31729e6e749534292c3faf9555b8fab723009582b58

Observation f50b1ee2-ebe6-4ea7-b375-ed4d48be7db3 · inbound

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own cites this paper.

Reinforcement Learning with Foundation Priors: Let the Embodied Agent Efficiently Learn on Its Own RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:44:02.621634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T06:40:00.328012Z digest=sha256:0c7f4da929d73787770212a6cbc72ae38714e246da25b2a21dc3bcdcb3361721

Observation 649fbcd2-8058-433b-a436-24504b3c9aae · inbound

Learning Interactive Real-World Simulators cites this paper.

Learning Interactive Real-World Simulators RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:15:18.658514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T02:15:18.265190Z digest=sha256:e4f862c8e1d026840ce1265b90d522a77e056e2f6528dc80abc7201abd5c4a49

Observation 23bba126-a8fd-4f01-b6e9-09e8094c89de · inbound

Open X-Embodiment: Robotic Learning Datasets and RT-X Models cites this paper.

Open X-Embodiment: Robotic Learning Datasets and RT-X Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:23:24.449479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T17:23:24.255829Z digest=sha256:f829d6d5dc437032bcd783b1c048b724a8f1f929a754cd2fddff5cebe4a99f48

Observation 38c972c4-232c-4ccb-9a83-c2208992c99e · inbound

Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models cites this paper.

Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:54:59.000154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T05:54:58.940428Z digest=sha256:80f16d5ded4e0a4e02a342dff0c02ed8b3b0aede85b8a7132953675bf2465e84

Observation 3fd271d6-e777-4d79-844f-8ab2ab44ba53 · inbound

TD-MPC2: Scalable, Robust World Models for Continuous Control cites this paper.

TD-MPC2: Scalable, Robust World Models for Continuous Control RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 156

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T17:27:35.993143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T17:27:35.733800Z digest=sha256:c4115589e678d59b61ee95fe65f97c328c48576ae876c2ea00da5f5bd86a871f

Observation 21e3ab03-a12d-4689-9be7-9794c12146e1 · inbound

Vision-Language Foundation Models as Effective Robot Imitators cites this paper.

Vision-Language Foundation Models as Effective Robot Imitators RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:44:27.605460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:feae8cbe3616b5da58343ee8b1a74b695cd27e622a69fa331959dcb726d14fd2

Observation e5f92c29-2975-4e70-9586-65035dd4c01b · inbound

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation cites this paper.

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T16:32:05.595167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:32:05.538507Z digest=sha256:37c5718f8c2c05889075c336966fa567bcc85350a08a59c36a921e7a68e7e337

Observation 73ce41a8-5064-48c6-80d9-0e4fd8e93486 · inbound

AppAgent: Multimodal Agents as Smartphone Users cites this paper.

AppAgent: Multimodal Agents as Smartphone Users RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T10:16:43.896074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T10:16:43.364787Z digest=sha256:34dc62b319851240cf0cc3e54dd251c1e3ee76762540ae01215a5af78835af14

Observation bef27680-7526-4a59-889c-cf316c3614fa · inbound

Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation cites this paper.

Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:02:55.532844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T22:02:55.240949Z digest=sha256:b730d5af6fc4dd7881eb65fba113229390d69c3d6e0139001c10e1682b83a76d

Observation 916b2d1a-9167-419b-aad0-b5a2b05d2e1a · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 120

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T14:25:59.347132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:f105d31f2d3515e2bb23cfc43cc0b4c3b5705f0bd4c6016563e60ad06c7eaf5a

Observation f022c4e6-8fd8-4677-aec9-aeda741ec060 · inbound

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models cites this paper.

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-12T19:22:35.424710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T19:22:35.305220Z digest=sha256:066d2195224b5aa726656e1dae2b77b7f5596560bde506fe223d99dfce58a0ae

Observation b58cb528-c895-49c0-87a0-a7c36819cb45 · inbound

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation cites this paper.

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:55:20.422893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T04:55:20.362512Z digest=sha256:1bb7928a8811083fd940b59814941a51380d421b54ca8087d06a8046db892bfe

Observation 30ab2bec-9c11-49af-ba4c-a82f90499660 · inbound

RT-H: Action Hierarchies Using Language cites this paper.

RT-H: Action Hierarchies Using Language RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T06:53:27.772992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T06:53:27.642020Z digest=sha256:e6668e51c6398fd5862b9886d4346faefeb7af7bfe55df48fb913b03760a94a1

Observation fb41528e-1b26-46b0-ae0e-888efea28365 · inbound

3D-VLA: A 3D Vision-Language-Action Generative World Model cites this paper.

3D-VLA: A 3D Vision-Language-Action Generative World Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:18:27.264601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:2f8d74d3293adffff27942547ed6b68ea5ced1a13ee81fe018f2a656c90ef2c8

Observation 37a39ef4-f617-4a6d-abd2-6c000a0eaf42 · inbound

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments cites this paper.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:19:32.450810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:580d09c30eeaaee5d5dcc365fd6bab2bfcdf67c09792f8760fe5bf65786237cd

Observation 4882d3c3-7a9d-4410-9107-ade4035e05b6 · inbound

The Platonic Representation Hypothesis cites this paper.

The Platonic Representation Hypothesis RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 219

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:03:56.818142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T06:03:56.328012Z digest=sha256:3ea43e55555d5e23502556e36cd9af7e57d15b53a5895d11f6dbafd58f20723b

Observation cdb17367-46f1-4777-8115-4007acb99935 · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-24T01:25:54.285080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:3e92c17f44262738cf9e1ecfa6ebbf3a1c0fbae34ea27bba2dbadd2a0aaa21fb

Observation 69f71b86-68fc-4236-957a-6ae74cb5af71 · inbound

OpenVLA: An Open-Source Vision-Language-Action Model cites this paper.

OpenVLA: An Open-Source Vision-Language-Action Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:36:04.729348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:46:35.942338Z digest=sha256:a72b8fb8a7d69ad4bc00613cfc56af211da662000eb82e9d7f5f0b0a4271914f

Observation 54bf625c-a741-4d60-81c8-4c46b7298ba1 · inbound

LongVILA: Scaling Long-Context Visual Language Models for Long Videos cites this paper.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.440421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:3ba4c7f3e234089fff949c65ea2d79df758fe6e80489ef2b3e58df6ebdf52548

Observation 8158301a-499c-477b-ba24-38f401b22131 · inbound

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation cites this paper.

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 107

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:25:17.975301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:25:17.847571Z digest=sha256:d032f75b56593c10cd59db6317319fec4a15cddd6d4971584ce2245df66ab379

Observation 6bb2cd43-b995-43ca-9603-ddb731b32873 · inbound

TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation cites this paper.

TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T16:12:26.080139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T16:12:25.980853Z digest=sha256:67df7e4941d7aba023ebdd5eb19513641d179e16228090c96d59c76bd45809ba

Observation a130f85b-dd09-4f7b-9b58-01a4f0dcb0df · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 150

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T12:04:10.698283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:cd1fed4357eb03a4f983a059571372b78e9da965bb9031cd2dd7067481bb56f1

Observation 668a0c42-cfd3-4db6-ab7e-484a94654fc8 · inbound

Towards Robust Surgical Automation via Digital Twin Representations from Foundation Models cites this paper.

Towards Robust Surgical Automation via Digital Twin Representations from Foundation Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:23:24.873272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T20:19:20.382009Z digest=sha256:42206d734c97054f73738357cf0c8220a5a2a7c9b3aae62c06e3150d16686ed0

Observation cb94fd1d-303c-4ee1-8144-d70b041977cc · inbound

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation cites this paper.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:09:33.896519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:66c2204c8935dc705b0342e83b653b6d098cf6cf89522c665fdc2e067391587a

Observation f9ac1071-0598-4bd7-bc05-8314ca9f4cb9 · inbound

Language Conditioned Multi-Finger Dexterous Manipulation Enabled by Physical Compliance and Switching of Controllers cites this paper.

Language Conditioned Multi-Finger Dexterous Manipulation Enabled by Physical Compliance and Switching of Controllers RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:08:21.025254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T19:06:58.600946Z digest=sha256:84e3f5a592d44694aeb941384f56a40c8e7e1c49983ab4b857495cfdd314aa73

Observation 9e164021-00cd-4c88-bc85-f7c0beeb88f3 · inbound

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents cites this paper.

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T09:29:27.362147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T09:29:27.173784Z digest=sha256:15b2da090eac97bf21414a2221f3b5ccc6752797140f18fdbe24afb065ead7f3

Observation 0ead8dff-ba03-423b-94af-1257f7b9ddca · inbound

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control cites this paper.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:36:04.729348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:a49b9ce70bfc64492560de254651ac960d6b9290183580a098ebade10b4ffe39

Observation ea61dd21-5e13-46dc-98ab-88fb0eb5a5a4 · inbound

DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning cites this paper.

DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T16:06:09.618905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T16:06:09.448517Z digest=sha256:de4fb88464cc9bc5df84f77db8a944d57a507d50a8c283f925033be394b16857

Observation afa4cfbc-ac71-4591-86be-8a12bd8caafb · inbound

VeriGraph: Scene Graphs for Execution Verifiable Robot Planning cites this paper.

VeriGraph: Scene Graphs for Execution Verifiable Robot Planning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:13:13.782190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T17:12:31.645346Z digest=sha256:0b7b4f1d2401c72dc21838390890a74369b77273d215a1ebeff9759b5a0c3e47

Observation 929fd749-adcc-45c6-a506-7b00056469fe · inbound

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation cites this paper.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.432800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:469b8218244c46093324ec88fd33138a96e7a6991f8101e997fd0b70db03dc93

Observation 7da82f87-4651-4a8b-8fc2-a7b3a5b42c37 · inbound

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies cites this paper.

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:22:44.360642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-23T08:20:05.898025Z digest=sha256:76c5c755bde3def380f166e807867e059dc52d405ad25505db446d856eabdc7c

Observation abfa9331-448a-47a8-98cf-831b0e642247 · inbound

RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields cites this paper.

RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:52:43.886679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T07:49:37.878395Z digest=sha256:91c5cf813ff32c35bf88328b6d7f482f13756a0ba97da7ea0fcf5487e9fed021

Observation 5457b199-a646-4e3c-8f33-a7ca63a66dbb · inbound

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks cites this paper.

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:51:36.238946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T19:51:36.137985Z digest=sha256:a8d57ac5d3d733bdd2997a191b1deecd88315547f228d115e1255cf0da64362b

Observation f09baac6-b13e-4ac8-9ffd-647ba224fe8e · inbound

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies cites this paper.

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:27:22.867493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T18:27:22.760982Z digest=sha256:712bbfaa31e0bd574220299d81a0c74133f027c9c8b18f0f95b10a512ddd89eb

Observation 8f70d6d8-e7a0-4264-823a-6d06bd4319d0 · inbound

What Matters in Building Vision-Language-Action Models for Generalist Robots cites this paper.

What Matters in Building Vision-Language-Action Models for Generalist Robots RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.706165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:811a92148e1fabd249d028da5b1570553d4274915a33651388579dd61b441c7b

Observation 870883cb-c452-47dd-b7f6-86fa504fe8fa · inbound

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations cites this paper.

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:38:11.219125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T18:38:11.110166Z digest=sha256:211835a3eb1810081471ada408bf4462513dddb604a1a2709a078eee511865ff

Observation e0b41df9-e0d6-4477-9494-b451fa7e1ab9 · inbound

FAST: Efficient Action Tokenization for Vision-Language-Action Models cites this paper.

FAST: Efficient Action Tokenization for Vision-Language-Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:52:31.904673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:87ac21527a01589bcf35d95973bd4da2138bdaa0a659ae2598c9fe4c2972e6f5

Observation bf8d3fb4-5f69-4029-a1e2-c59cca2d2d0c · inbound

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model cites this paper.

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:12:19.863742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T06:12:19.643111Z digest=sha256:e8aa541d09c9c5cdee1e01a2cf822ac63125390847a4ecfb8d20db59571333ee

Observation dc9505dc-693f-4cd5-98f0-83215b576692 · inbound

Large Language Models for Multi-Robot Systems: A Survey cites this paper.

Large Language Models for Multi-Robot Systems: A Survey RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:32:32.087627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T04:32:05.138744Z digest=sha256:c9938a0025f56ecf4ea83fa88847598b47879f914ee2e9721ee8b5518bdfbf04

Observation 7eff8d05-1492-4e42-9627-d79c14e123fd · inbound

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control cites this paper.

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:48:48.944998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T19:48:48.725800Z digest=sha256:ddb51176e930e87a00ea58d82e59975ea78cb943a6d44a9a96ddc520fc90c7a5

Observation 95d55974-f9f6-4a4d-918f-621136103741 · inbound

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models cites this paper.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:53:37.199637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:7c642ef42b2e07822fa3dc6dc9d0b4c05ff2faac2cd047d773f845e9ecae4a6c

Observation e7ea5adf-e6d5-43bd-bb69-c4acd5fa0b3e · inbound

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success cites this paper.

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:35:32.271358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T04:35:31.914360Z digest=sha256:20c77d4c2dc38d7f91543c91d20236be9186f55860d4e4ff7fed99e636e17d1e

Observation 6aa27e55-540d-4995-8c9c-cc0503f68167 · inbound

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning cites this paper.

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:32:22.350668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T01:27:33.123243Z digest=sha256:5f6ccd66d38ccaff7e1213fea11a542b26953014b498266771991391f0a12032

Observation ae82a16e-4262-4e53-90a8-bd0e1ea5d805 · inbound

Kaiwu: A Multimodal Manipulation Dataset and Framework for Robot Learning and Human-Robot Interaction cites this paper.

Kaiwu: A Multimodal Manipulation Dataset and Framework for Robot Learning and Human-Robot Interaction RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:27:21.448877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T01:26:42.319929Z digest=sha256:89551724f33697bf1af977c3c5bcc75e58963cdc50d46b443bd0f17580fb6c0e

Observation cc8628dd-8783-4662-b625-90cf7d12897b · inbound

AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning cites this paper.

AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:06:27.171457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T20:06:27.136345Z digest=sha256:89d4385b1d902c62c271710913070cbb17e5c5d9a4d7890571dc7ebf2e50c028

Observation 436be115-68ae-483b-9d86-890585998ca3 · inbound

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model cites this paper.

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:00:48.944755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T22:00:48.667428Z digest=sha256:9ad0643f4079c082f2a945419dcdcef247afedc4b25bd98c944e4eb955fa1588

Observation 6b34d820-1b21-4719-9585-b85c951bd20b · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-23T00:02:17.887169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:eab2a73fbc88d347a084bcf9601190d834c17076fa5e4e47de88621462deab70

Observation 75830fd6-8b1a-463a-8de5-bd989145fc16 · inbound

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots cites this paper.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:36:04.729348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:a3068992f0f628f4315a5fadc3cce4e58025e204cabd78fcf3632cf33f50955c

Observation 4fe36b50-369c-41ad-9dda-3fffae657425 · inbound

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning cites this paper.

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:47:10.192231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:47:10.146795Z digest=sha256:00b36b8589eae5f22d1275a38f22a5559585347433cd34ca565cb29ee0cd0831

Observation 94ee9ac4-a713-4272-a414-f9028c9d8af9 · inbound

NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks cites this paper.

NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:53:29.438836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T15:53:29.412890Z digest=sha256:ad92dfecbf653d6494a0bb2b6e7189e1069bdea6d050cc1a8869dc16778558b6

Observation 674c5f46-84b4-4d7c-99b7-cb893664acb8 · inbound

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data cites this paper.

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:55:52.197589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:55:52.109166Z digest=sha256:fd09b7efd638336092721f6b8fe796d8a5fb93d8ada51a9fa28e95902e124e11

Observation da7ba968-5944-4f1c-8ea5-dad8ec899c31 · inbound

VLAs are Confined yet Capable of Generalizing to Novel Instructions cites this paper.

VLAs are Confined yet Capable of Generalizing to Novel Instructions RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-22T16:46:47.397622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T16:46:05.993833Z digest=sha256:5d750c30d4cd265fbc14ebb90864fd66cff270480d926e1e32535cb8927b1d7e

Observation 4b1b22d7-c527-41f1-b267-fa64f50eaa04 · inbound

DreamGen: Unlocking Generalization in Robot Learning through Video World Models cites this paper.

DreamGen: Unlocking Generalization in Robot Learning through Video World Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:50:45.573361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:50:45.332466Z digest=sha256:306767a24865a29a560a63c4f2698faef83181a9b44676295b1b022e18e5d993

Observation 50fff180-de6d-44fe-991c-7c4bbfc4cc06 · inbound

Policy Contrastive Decoding for Robotic Foundation Models cites this paper.

Policy Contrastive Decoding for Robotic Foundation Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:11:38.582010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T14:09:48.762737Z digest=sha256:c3ea6e7c61906176a18df6ac0b6cd07749db8ea62fcd247057180f5492b55f87

Observation 6ec1a03f-3de6-410c-94b1-d6626c5d596a · inbound

FLARE: Robot Learning with Implicit World Modeling cites this paper.

FLARE: Robot Learning with Implicit World Modeling RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:59:08.987170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T15:59:08.846629Z digest=sha256:ca1b8b565e3e63b5c71246c4d80a66bf9bab4578d4e89ab2bd1c98e51a8900fe

Observation 9ec798c0-283b-47df-b3c8-0e7e80bd1c3b · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:55:40.502751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:0acdff6e73a3624356a90e6f9da508f532616a5ab3ed60d8f056f9ac4e066df8

Observation d04b9ed3-4832-4b36-a696-5c5eb67e1d1c · inbound

DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion cites this paper.

DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:52:17.951246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T12:50:17.902979Z digest=sha256:2a5ea5590b7dba2ce375ce8ab724aa54aab792280c21b151171f30194805391a

Observation 09d7e6b4-a3a3-49cf-9678-444178251b0e · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:22:37.221359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:66d67ac33ca5fc8831ce94414b30ee47382aabea5719235b63a3a5e23c611bdd

Observation a1c0ce06-e65d-4ac7-91ef-ac9d61fd6c15 · inbound

Real-Time Execution of Action Chunking Flow Policies cites this paper.

Real-Time Execution of Action Chunking Flow Policies RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:18:51.761539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T14:18:51.613045Z digest=sha256:095d4b60846aaf124c73111b9484dc727d425f910840635ccc51bda0576fb409

Observation c6b04c45-8f01-4280-97ac-3488f0dd72d2 · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:50.699817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:3096e30c304e846762f8fe452ac92c35a343611f7229f8d8e736e36bfc2e2bdf

Observation 4eaf2be8-1aa8-49b4-92eb-4ea2fc502cd1 · inbound

Block-wise Adaptive Caching for Accelerating Diffusion Policy cites this paper.

Block-wise Adaptive Caching for Accelerating Diffusion Policy RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:37:13.994419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T09:36:09.790248Z digest=sha256:0b07039eb867627b2265b0a094f2dbe21dc20b26ced5fdf7cd187d12a3208978

Observation f11fb299-de8e-418b-baec-7c9deb9959a6 · inbound

GRaD-Nav++: Vision-Language Model Enabled Visual Drone Navigation with Gaussian Radiance Fields and Differentiable Dynamics cites this paper.

GRaD-Nav++: Vision-Language Model Enabled Visual Drone Navigation with Gaussian Radiance Fields and Differentiable Dynamics RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:30:49.385527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T00:26:28.696429Z digest=sha256:e740e6c6a915c8c33e1af43b127d0691b9780cebce29c940251e5a649b3e60fe

Observation 9d2cf2ee-fa5c-4141-afe1-6dd8233bd2d9 · inbound

RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation cites this paper.

RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:40:27.339354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T06:40:27.206337Z digest=sha256:490479e972efc4bce3f64f6915537aa8aafd88b5c89f4fcb0e2d6b83d5fd037f

Observation 3d6a7b7e-3f2a-4e8a-a616-d3b28e46cabd · inbound

WorldVLA: Towards Autoregressive Action World Model cites this paper.

WorldVLA: Towards Autoregressive Action World Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:57:08.017187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T22:57:07.883617Z digest=sha256:ea38486f593202312f7fa418fa250a1c41721c846fbc09d68d6860e60a463442

Observation a89c0052-ad6d-47a9-b50b-dbf080a2c966 · inbound

A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation cites this paper.

A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:32:56.726297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T04:32:56.397350Z digest=sha256:7618f0d9970f57863f2fce546b8516da5ec7c5e3e4116959fe7585d771d88bf6

Observation 03dcca80-50b5-4ba6-9b3b-85dfc3804b4a · inbound

Preemptive Solving of Future Problems: Multitask Preplay in Humans and Machines cites this paper.

Preemptive Solving of Future Problems: Multitask Preplay in Humans and Machines RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:42:07.621610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T06:38:33.969927Z digest=sha256:2c5717f57bd5f089d24d46f0862c40d6428efda98a4b3453554ffe30ed2106bb

Observation 53bfa4d4-aa9d-4ef3-9eb8-5fa31b68ff16 · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-17T08:04:12.654077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:ee0de188a37a352b41e1a01c56767758d8613993e3ed133996bb0cddb6f353f2

Observation 402f1164-5b32-4826-a52b-22a45bf3ff43 · inbound

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning cites this paper.

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:22:00.904165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T03:18:14.655384Z digest=sha256:a2c9776afdf7f6530ac1497bf2e68822bb3ea9bbbceef3025604c0662c5983f5

Observation 24f98853-d3f5-47ab-b2f1-9d88375305fb · inbound

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation cites this paper.

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:28:41.975598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T21:28:41.904725Z digest=sha256:d0c0e0fb3683b11acc5908ff7e14d254351333a8f10da8fb64e57a254ecb6448

Observation 8b306df5-cd1a-445c-a1ac-a4a7c46ced50 · inbound

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation cites this paper.

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:43:24.448144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T20:43:24.417901Z digest=sha256:595185f397787bd3665ad58b93caafb0990a0d4656212ae0d4ba0256b7fd7430

Observation 34e88aee-4055-4ac5-99ea-018ca00f7785 · inbound

Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI cites this paper.

Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 225

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:01:14.210554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T09:56:36.716680Z digest=sha256:64e48cf5cf857c0d0cb05c5ad92983792017e77d5a96f8dfd032a7cea6b2e931

Observation 8518c6a5-a2ee-4014-98c3-1e5baad3ec48 · inbound

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation cites this paper.

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T08:51:09.086024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T08:46:51.278024Z digest=sha256:cbab7e5451a53fea28bdf99df93f85975a9a1090f58802e03f5b74c0c43ee12e

Observation f9786ea3-b935-4cfa-bd83-263aff97709f · inbound

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models cites this paper.

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T10:24:59.855729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:24:59.855729Z digest=sha256:53eee0fa8d6e482bc60117e1abc54db82be43e767e6ae5e4e0b05b7aa7c3517f

Observation b4c6ec9b-26e4-4e2a-84d8-dfc3a381bddc · inbound

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy cites this paper.

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:09:39.735048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T20:09:39.677347Z digest=sha256:1cbf2f51b5209c2687bf967f23bb1ec9cb86ac8f84ef1990590bc3b1d1c5e9ce

Observation 28da32b0-07fa-4681-b657-bf4255060496 · inbound

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models cites this paper.

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:40.820281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:34:40.820281Z digest=sha256:702ddf18af79cde0bb31c956308f3da00f85495d9573b484668ec5057958b459

Observation 90529829-446a-46ba-b581-2aed5cffbc7e · inbound

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model cites this paper.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:45.804069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:45.804069Z digest=sha256:bf9641e77dbbc9bf97b89594a51d73c3dbb89be843d67f71513e0a805354eee5

Observation c5071b48-8ac2-4f67-a6f6-98e51df95086 · inbound

LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation cites this paper.

LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:35:27.926770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T07:34:55.240907Z digest=sha256:279da6c812a667ae79b1e5c02b4f12e277e049eef099da5e9a7d8776a670a0bc

Observation a35b5dd0-f523-48e5-9031-b0825f01733f · inbound

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models cites this paper.

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:30:18.350850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T21:28:18.630934Z digest=sha256:50d0b06ea44306da3aec7d4338b13266ddb69588b6d3dd62c4ca3a62509fc819

Observation bf80d64c-1d5d-4bbf-aa01-24c17c134a5f · inbound

DynaMimicGen: A Data Generation Framework for Robot Learning of Dynamic Tasks cites this paper.

DynaMimicGen: A Data Generation Framework for Robot Learning of Dynamic Tasks RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T21:14:41.516853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:14:41.516853Z digest=sha256:0b41dfa4cd2a95a7e87564fc4d711821cdaef8efe55733ca24728c65ab537a90

Observation 30cf9be9-17be-4f22-a6b1-808061085006 · inbound

DASIP: Dynamic Test-Time Compute Scaling for Robot Control with Stochastic Interpolant Policies cites this paper.

DASIP: Dynamic Test-Time Compute Scaling for Robot Control with Stochastic Interpolant Policies RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T20:13:26.448498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:13:26.448498Z digest=sha256:5720cc203357be072e79d02a138f311863a251c99eb535b95d808fe897df4b15

Observation 7bfd0808-ce72-4877-b94c-a3803f1582f2 · inbound

SimScale: Learning to Drive via Real-World Simulation at Scale cites this paper.

SimScale: Learning to Drive via Real-World Simulation at Scale RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-17T04:34:01.922251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T04:33:03.629533Z digest=sha256:8ec46ba5068cf018950c372966aca568c1aa7fe2a566674d15e8676927ffb993

Observation e32151b2-88c8-4d91-aa2f-8c8a6e415793 · inbound

Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning cites this paper.

Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T17:30:50.131213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:30:50.131213Z digest=sha256:0e7854830e3f761b7cd0e1b05469c4459816be7d7377d25184fae902fdd30a0d

Observation 47318a33-5b9f-4103-9de0-5b38181a5ee9 · inbound

Large Video Planner Enables Generalizable Robot Control cites this paper.

Large Video Planner Enables Generalizable Robot Control RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:28:33.925089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:26:32.048309Z digest=sha256:63dbb868a17a9b00c4dba282d757e5dd64c41bd4a55021062ce73256eebd9378

Observation d8b81cd4-22d4-4f2b-a90c-d2801d265914 · inbound

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models cites this paper.

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T13:53:24.245885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:53:24.245885Z digest=sha256:e6f1c43462535a4f0d56bca701efb6d71e84f77327d33338ccdb82359c1797d5

Observation 4ec77fc6-5cf2-45eb-a555-a5c37206250b · inbound

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation cites this paper.

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T18:41:10.946746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T18:39:59.449746Z digest=sha256:2353d477f0525fe51c3e08bd1cc26ab960072cda195195cc62d60032c7297cfb

Observation 0a4bbb9e-f51d-4272-a56c-0a5d670723be · inbound

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation cites this paper.

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T13:35:54.050075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:35:54.050075Z digest=sha256:8b17348e35b78393da3eebc14f92407c0ad0243ced99cabfab200802ccc4f53e

Observation 62233dc9-936e-419b-a6f3-c38e0f58c4dc · inbound

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models cites this paper.

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T12:30:33.750549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:30:33.750549Z digest=sha256:3fa228e1b5dd6778247b988a84d9d5296510b687aab4e53fe7ad51168568e59c

Observation e64fe6ad-2e91-488a-839a-5a187154cd2b · inbound

State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space cites this paper.

State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T12:18:21.634748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:18:21.634748Z digest=sha256:f4c9d8ce3de75d164ed86471d5191e68b88a08901db73b28a6dbbb3e5695e235

Observation dcec3e9b-37a2-47e4-8bac-8cbafd3baa8f · inbound

Real2Sim via Active Perception with Behavior Trees Automatically Generated by VLMs cites this paper.

Real2Sim via Active Perception with Behavior Trees Automatically Generated by VLMs RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-21T15:10:16.661647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T15:07:54.930273Z digest=sha256:a6b58430c6db90912c976f028a69ad1b6866ce209989dd441228a777564e0eb0

Observation 7c612aea-1fe0-47d7-bedf-09063fbbef2c · inbound

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning cites this paper.

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T14:50:12.856956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T14:50:12.804707Z digest=sha256:c241547a0748e396b68cfcd4b46bf3224ee62534badc8898e5fbde4fb6890df6

Observation 6cba2652-a325-4b4d-acc6-69686cbb1b06 · inbound

Eval-Actions: Fine-Grained Execution Quality Evaluation for Robotic Manipulation cites this paper.

Eval-Actions: Fine-Grained Execution Quality Evaluation for Robotic Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T08:00:25.395013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:00:25.395013Z digest=sha256:e90f0ff1ee8b638a4c9a7d60e9bb174f7593b41cdebb4ec6e2be8de15f7e6643

Observation f6e86cb1-2d70-4e36-b77e-547a39ef03ea · inbound

How Users Understand Robot Foundation Model Performance through Task Success Rates and Beyond cites this paper.

How Users Understand Robot Foundation Model Performance through Task Success Rates and Beyond RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2023

Resolution
malformed identifier
no resolver link, observed 2026-08-03T04:53:23.674341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:53:23.674341Z digest=sha256:4f0cf3a722892d10b2dcb8a6b25228c3d5443573420daaa06bd8efe2a424a0cb

Observation 1f131e54-7d80-47a3-914f-60f2e8965678 · inbound

ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs cites this paper.

ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:00:40.541325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T06:00:02.043029Z digest=sha256:4e6f0cd40bd1ddae85b8d645f79ad7813f62f51b6f0a7efa904ea4cc5c3af6b9

Observation 52dc38f3-28cf-4a64-b69c-f63ec1c7d087 · inbound

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation cites this paper.

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:14:10.845690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:13:53.818915Z digest=sha256:50904eb3529d16a7f091e280398512feec7e0a88c24dc05a56fa0037bab3164e

Observation 33d214d7-c209-4613-86ea-8c58d81925b5 · inbound

PRISM-XR: Empowering Privacy-Aware XR Collaboration with Multimodal Large Language Models cites this paper.

PRISM-XR: Empowering Privacy-Aware XR Collaboration with Multimodal Large Language Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:07:20.550707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T05:05:41.709469Z digest=sha256:def1a1e83a6273ca38ddc90f6b13fcbd48c43a6246337108788e42a026ad3445

Observation 7f0aa1ec-7573-4c8c-89bd-9d7c10351990 · inbound

World Action Models are Zero-shot Policies cites this paper.

World Action Models are Zero-shot Policies RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:18:15.836058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T16:18:15.003371Z digest=sha256:4368340e141d1b9f92cc67708cc90b514a130956a07332b24604b13b823b9419

Observation f2996fc3-11d9-4bfe-bdbd-062bea046405 · inbound

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs cites this paper.

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T22:14:26.017939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:14:26.017939Z digest=sha256:52ae4711ca718d808e72c365e6ee95428b0dacbb74dea9615ec4885f5e3a28b4

Observation c20da66f-f8b7-4a19-a200-b6c04932699b · inbound

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models cites this paper.

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:20:17.538528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T20:18:31.988002Z digest=sha256:d1f13de9be32be94a8c988de1ab89cddf4fe54cf4cec0b5694d87738225365fd