Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:04.383484Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 20 inbound Pith citation observations for arXiv:2505.21906.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:04.383484Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:08:10.115509Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:49:46.316884Z
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6d21ea7c-2866-435c-973f-a3dd8583ba81 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d0d4bab-577a-4f16-9761-867515298871 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb6be59-bc76-4caf-92a2-b347b7b136a4 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7ff553a-31a7-429b-a864-bd232afe9cda · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8acab03-d2af-4d4d-9344-82a6bb1694e4 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge PaliGemma: A versatile 3B VLM for transfer
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e18183a-1248-4276-82f8-ec2f5c3291ae · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 226371c4-8467-4e3d-a7c1-017f7d6f1c8c · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2c7eacb-e3b7-463d-8742-a214e14fb596 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36:74952–74965, 2023
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6947f35-87a2-46e3-a53d-f8d9265340f5 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Reasoning Models Can Be Effective Without Thinking
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33193132-135f-45ea-9ddc-72f63e163e57 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Openvla: An open-source vision-language-action model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 55827294-2694-4182-b4dd-68de7dac08c1 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 548aeeac-794b-4726-977f-088f427409c4 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Visual reinforcement learning with self-supervised 3d representations.IEEE Robotics and Automation Letters, 8(5):2890–2897, 2023
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 82d8ebb0-3fe7-4a9c-8a86-7b1672f396dc · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Discrete Policy: Learning Disentangled Action Space for Multi-Task Robotic Manipulation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 746f057e-c8de-4848-a8d1-5f557a1b7376 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21d30c9c-8d72-4997-8f6e-6fced690ce5d · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Any2policy: Learning visuomotor policy with any-modality.Advances in Neural Information Processing Systems, 37:133518–133540, 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0fc88331-e5bd-4b89-86b6-6607624c4309 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Any-point Trajectory Modeling for Policy Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 695fd35d-8c9e-4a2a-9aff-31d95383513d · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Retrieval-augmented embodied agents
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52a0c22d-9591-43c9-b18b-14c2a596b3f8 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge RT-1: Robotics Transformer for Real-World Control at Scale
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91f2e76f-3719-4539-ae8b-4878063b1871 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Rt-affordance: Affordances are versatile intermediate representations for robot manipulation, 2024
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a6aaff8d-86cc-47f4-928e-d0788e7ca553 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dab7c23-ac78-4632-82d9-6c8438b81bed · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8c7d9c8-7e39-4594-be27-af871733f963 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adcd24dd-49c7-46ca-b096-e49bb0e0f592 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Mail: Improving imitation learning with selective state space models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a59c795f-e8ad-4785-b17f-5edc3faddf78 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b8e407a-0420-49b2-a79f-b764eee9cfff · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Vision-Language Foundation Models as Effective Robot Imitators
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6abf7fac-59dc-445e-9a07-4e67f2b4dee9 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge An embodied generalist agent in 3d world
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 81db2c8e-c22d-4e86-922d-84ed3e3cf046 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36c10730-9e37-41bf-b368-2dfca4b73650 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge FAST: Efficient Action Tokenization for Vision-Language-Action Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b796f4bb-a1f0-4644-a7fa-f65ecb5fcd9a · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge π0: A vision-language-action flow model for general robot control, 2024
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff4ebc0e-ca74-4e34-8d29-d2164948b1de · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge OpenVLA: An Open-Source Vision-Language-Action Model
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c4177d1-3793-4444-824e-03d8bf31e882 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1ddc19d-7136-4002-968a-9f7bb5a87959 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfc82a78-bea8-472f-b61d-6ecc9759d7a4 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c055702b-42d5-45a5-8da6-aebad243287c · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Training Diffusion Models with Reinforcement Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7560aca-d665-465d-8b2a-f544313c6dcc · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be950bdd-6c22-4ced-8164-733187ea9338 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge The Ingredients for Robotic Diffusion Transformers
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e91f4925-a14b-4183-be26-c6967b0680d2 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Data scaling laws in imitation learning for robotic manipulation, 2024
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 20673b07-97a8-47c7-8149-8c6dcc97bac2 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Multimodal diffusion transformer: Learning versatile behavior from multimodal goals
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c40028a2-2dd3-4dc3-9500-aae8fa2059bb · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Aloha unleashed: A simple recipe for robot dexterity
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 33c41de6-8225-42ea-b84a-a0d2332fb78f · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Fine-Tuning of Continuous-Time Diffusion Models as Entropy-Regularized Control
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a657ba2-5406-4fa1-a63d-7a7ebe6c9e00 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Feedback Efficient Online Fine-Tuning of Diffusion Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17512fd7-14bc-4572-bf21-f5f0627d553b · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Gemini Robotics: Bringing AI into the Physical World
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5564302-9a1c-45ad-88cd-5fc42a2d1342 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 370bbd6c-6ff8-4ac2-b488-f84ef934979a · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83bca668-757d-4639-8e0b-39db868c162f · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Quar-vla: Vision-language-action model for quadruped robots
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd9caa4d-7dad-4310-b342-57f2e7c467e9 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b01c3a5a-9fba-4a2f-9007-4fd83e3b3deb · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3795b51a-6f3b-4b44-b2ac-45e454f1f574 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13683616-8f83-449b-9bcf-61f3e9db0953 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Robomamba: Efficient vision- language-action model for robotic reasoning and manipulation.Advances in Neural Information Processing Systems, 37:40085–40110, 2024
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aebf1c98-049a-45c9-abfe-6ebad3420eea · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Deer-vla: Dynamic inference of multimodal large language models for efficient robot execution.Advances in Neural Information Processing Systems, 37:56619–56643, 2024
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e2c2f74-258b-4012-91f4-fda7d7938d15 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d02f426d-edf6-4864-a9e6-ae8c32b90383 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bee1918-1f07-4b6a-b3bd-1b580917c248 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc454b8e-d899-488d-9009-45f1e0e41b7c · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Chain-of-thought prompting elicits reasoning in large language models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 483ba450-c5cf-441a-b022-761b421380c2 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Robotic Control via Embodied Chain-of-Thought Reasoning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e40c096e-f28b-4109-bf45-bfaee6bd7b19 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f450f948-70e7-4198-8484-78173268b6a4 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b71bc9a-7055-4762-b6ba-40edfcfacf0e · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2f290be-f8a1-46c1-8e1c-654bbba7905d · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1a9f423-8651-4a84-80f2-4494d6b44898 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4971d5ef-374a-449f-8a17-49f31795baf1 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic Manipulation
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c35bd5b3-604d-40bf-8062-347a8b121928 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcd4bf35-08f6-4584-8d82-ba1130d069e3 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge PointVLA: Injecting the 3D World into Vision-Language-Action Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98c2a8d3-5339-4cca-98f9-0c01eec39b24 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 282edde4-87b7-421b-92d2-41e24090bf53 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Microsoft coco: Common objects in context
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 04eea78b-1e08-4379-92fc-403f3015ca81 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Towards vqa models that can read
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 61f29ec0-7e88-4d1d-985f-f887751adcd3 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f543ec4-7f8f-4b01-ba79-aae7b6ba8950 · outbound
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Octo: An open-source generalist robot policy
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7b90b889-cb7a-4be7-bbaf-b74925dd9442 · inbound
What Matters in Building Vision-Language-Action Models for Generalist Robots ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a0e2b8c1-28e1-46ef-bed6-456df92c9d11 · inbound
SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3d798a09-016e-4b4a-badf-e51fd4ed9225 · inbound
RationalVLA: A Rational Vision-Language-Action Model with Dual System ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91bd68e6-817b-4383-90fe-56f24360044f · inbound
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 131
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6cb31c1a-69ee-47cd-914e-be08a3b32bd5 · inbound
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a9e238f-c452-44a9-ac56-732d17bd6da2 · inbound
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4cfccbd2-71cf-4a53-9f38-40d77c92705b · inbound
GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 51b4255c-f5d1-4777-b590-5064be053c6e · inbound
Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 20f8ca47-5f32-4d55-89c2-651193f0b6ff · inbound
Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1eb1fddc-ed36-443c-8628-0548704f57c9 · inbound
Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 999a750d-1f10-4cbf-95c6-c0e76e5ec01a · inbound
VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c4fedb81-ba07-4d9a-9aef-731750d8ad25 · inbound
PhysBrain 1.0 Technical Report ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 39aa038f-dc04-48fb-a73a-f3a8e39bc85b · inbound
DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e79f361a-6613-4805-b7cc-5660a41a7c0e · inbound
Continuous Reasoning for Vision-Language-Action ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 13aab588-7370-48d8-8881-0bc4931c2e1e · inbound
Policy-based Foveated Imaging and Perception ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 12d49e0d-37ca-4ab3-97e5-d7c91875184a · inbound
OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 668abe37-031f-4b57-9a78-5655687e2540 · inbound
Bridging Semantics and Kinematics: A Modular Framework for Zero-Shot Robotic Manipulation ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 89823bb4-2bfc-4161-b2cc-78d9df11fd8a · inbound
Last-Meter Precision Navigation for UAVs: A Diffusion-Refined Aerial Visual Servoing Approach ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 731e90b9-b286-469b-b613-037f883a850c · inbound
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7182c33b-5725-4bcf-a650-ddc8ef32db45 · inbound
SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.