REVIEW 19 cited by
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Humans possess a unified cognitive ability to perceive, comprehend, and interact with the physical world. Why can't large language models replicate this holistic understanding? Through a systematic analysis of existing training paradigms in vision-language-action models (VLA), we identify two key challenges: spurious forgetting, where robot training overwrites crucial visual-text alignments, and task interference, where competing control and understanding tasks degrade performance when trained jointly. To overcome these limitations, we propose ChatVLA, a novel framework featuring Phased Alignment Training, which incrementally integrates multimodal data after initial control mastery, and a Mixture-of-Experts architecture to minimize task interference. ChatVLA demonstrates competitive performance on visual question-answering datasets and significantly surpasses state-of-the-art vision-language-action (VLA) methods on multimodal understanding benchmarks. Notably, it achieves a six times higher performance on MMMU and scores 47.2% on MMStar with a more parameter-efficient design than ECoT. Furthermore, ChatVLA demonstrates superior performance on 25 real-world robot manipulation tasks compared to existing VLA methods like OpenVLA. Our findings highlight the potential of our unified framework for achieving both robust multimodal understanding and effective robot control.
Forward citations
Cited by 19 Pith papers
-
RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.
-
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment
Preserving pretrained VLM features with layer-wise distillation plus supervising the language head on discretized action directions improves OOD generalization of VLA policies on LIBERO, CALVIN, and a real xArm7.
-
Nautilus: From One Prompt to Plug-and-Play Robot Learning
NAUTILUS is a prompt-driven harness that automates plug-and-play adapters, typed contracts, and validation for policies, benchmarks, and robots in learning research.
-
UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation
UniSteer unifies human corrective actions and noise-space RL for VLA adaptation by inverting actions to noise targets, raising success rates from 20% to 90% in 66 minutes across four real-world manipulation tasks.
-
Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
A 6K-question embodied-reasoning benchmark plus a flow-matching action tokenizer let one 3B vision-language model reason and manipulate better than continuous- or discrete-action VLA baselines.
-
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
An open benchmark with 170 graded manipulation tasks shows current VLA robot policies memorize their training settings, degrade sharply under visual shifts, ignore safety constraints, and fail to compose long-horizon skills.
-
Contrastive Representation Regularization for Vision-Language-Action Models
Adding a robot-state-aware contrastive loss to VLA training improves manipulation success on RoboCasa-Kitchen and real-robot tasks.
-
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
A VLM trained on auto-generated failure trajectories with executable correction actions helps VLA models detect and fix manipulation errors, lifting success rates by up to 22.6 percentage points.
-
LLaDA-VLA: Vision Language Diffusion Action Models
LLaDA-VLA applies a masked diffusion vision-language model to robot control with localized action-token classification and hierarchical decoding, achieving SOTA success rates on SimplerEnv, CALVIN, and real-robot tasks.
-
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
ROSA trains a VLA model jointly on expert actions and automatically recorded robot states, improving success rates and generalization, particularly with few demonstrations.
-
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
VeBrain unifies perception, spatial reasoning, and robot control in one MLLM by representing control as keypoint detection plus skill selection, with a robotic adapter for deployment.
-
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
ChatVLA-2 uses dynamic mixture-of-experts and a two-stage training recipe to let a vision-language-action model retain pretrained reasoning while following robot instructions.
-
Choose What to Manipulate: Revealing Data Scaling Laws in Bounding-Box Guided Policies for Semantic Manipulation
A bounding-box-conditioned diffusion policy shows a power-law improvement with the number of object classes in training data, reaching about 85% success on four semantic manipulation tasks.
-
RynnVLA-002: A Unified Vision-Language-Action and World Model
A single model that jointly predicts robot actions and future images outperforms separate action-only and video-only models on LIBERO and real SO100 manipulation tasks.
-
Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
A plug-and-play fine-tuning method using two VAEs and a latent-distance guidance loss improves cross-embodiment and cross-task success rates of diffusion- and flow-based VLA policies.
-
ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning
Fine-tuning a vision-language-action robot model on teacher-generated reasoning rationales raises average simulated manipulation success by up to 8.6 percentage points over the SpatialVLA baseline.
-
Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges
A survey that taxonomizes robotic manipulation policies trained by imitation learning, traces their evolution, and compiles benchmark comparisons.
-
Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
A survey that categorizes continual learning methods for generative models into architecture-based, regularization-based, and replay-based paradigms across four model families.
-
Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents
An agentic framework using a GPT-4o planner, an OpenVLA executor, and a LoRA-fine-tuned Qwen2.5-VL verifier achieves 79.6% average success on LIBERO by decomposing and verifying subgoals.
Discussion (0). Continue with ORCID to comment.