Pith. sign in

REVIEW 19 cited by

ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.14420 v2 pith:W4XUSKKN submitted 2025-02-20 cs.RO cs.CVcs.LG

classification cs.ROcs.CVcs.LG
keywords understandingchatvlacontrolmultimodalperformancerobottrainingunified
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans possess a unified cognitive ability to perceive, comprehend, and interact with the physical world. Why can't large language models replicate this holistic understanding? Through a systematic analysis of existing training paradigms in vision-language-action models (VLA), we identify two key challenges: spurious forgetting, where robot training overwrites crucial visual-text alignments, and task interference, where competing control and understanding tasks degrade performance when trained jointly. To overcome these limitations, we propose ChatVLA, a novel framework featuring Phased Alignment Training, which incrementally integrates multimodal data after initial control mastery, and a Mixture-of-Experts architecture to minimize task interference. ChatVLA demonstrates competitive performance on visual question-answering datasets and significantly surpasses state-of-the-art vision-language-action (VLA) methods on multimodal understanding benchmarks. Notably, it achieves a six times higher performance on MMMU and scores 47.2% on MMStar with a more parameter-efficient design than ECoT. Furthermore, ChatVLA demonstrates superior performance on 25 real-world robot manipulation tasks compared to existing VLA methods like OpenVLA. Our findings highlight the potential of our unified framework for achieving both robust multimodal understanding and effective robot control.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.

  2. Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Preserving pretrained VLM features with layer-wise distillation plus supervising the language head on discretized action directions improves OOD generalization of VLA policies on LIBERO, CALVIN, and a real xArm7.

  3. Nautilus: From One Prompt to Plug-and-Play Robot Learning

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    NAUTILUS is a prompt-driven harness that automates plug-and-play adapters, typed contracts, and validation for policies, benchmarks, and robots in learning research.

  4. UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    UniSteer unifies human corrective actions and noise-space RL for VLA adaptation by inverting actions to noise targets, raising success rates from 20% to 90% in 66 minutes across four real-world manipulation tasks.

  5. Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

    cs.RO 2025-12 conditional novelty 6.0 of 10

    A 6K-question embodied-reasoning benchmark plus a flow-matching action tokenizer let one 3B vision-language model reason and manipulate better than continuous- or discrete-action VLA baselines.

  6. VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models

    cs.RO 2025-12 conditional novelty 6.0 of 10

    An open benchmark with 170 graded manipulation tasks shows current VLA robot policies memorize their training settings, degrade sharply under visual shifts, ignore safety constraints, and fail to compose long-horizon skills.

  7. Contrastive Representation Regularization for Vision-Language-Action Models

    cs.RO 2025-10 conditional novelty 6.0 of 10

    Adding a robot-state-aware contrastive loss to VLA training improves manipulation success on RoboCasa-Kitchen and real-robot tasks.

  8. FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models

    cs.RO 2025-10 conditional novelty 6.0 of 10

    A VLM trained on auto-generated failure trajectories with executable correction actions helps VLA models detect and fix manipulation errors, lifting success rates by up to 22.6 percentage points.

  9. LLaDA-VLA: Vision Language Diffusion Action Models

    cs.RO 2025-09 conditional novelty 6.0 of 10

    LLaDA-VLA applies a masked diffusion vision-language model to robot control with localized action-token classification and hierarchical decoding, achieving SOTA success rates on SimplerEnv, CALVIN, and real-robot tasks.

  10. ROSA: Harnessing Robot States for Vision-Language and Action Alignment

    cs.RO 2025-06 conditional novelty 6.0 of 10

    ROSA trains a VLA model jointly on expert actions and automatically recorded robot states, improving success rates and generalization, particularly with few demonstrations.

  11. Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

    cs.CV 2025-05 reject novelty 6.0 of 10

    VeBrain unifies perception, spatial reasoning, and robot control in one MLLM by representing control as keypoint detection plus skill selection, with a robotic adapter for deployment.

  12. ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

    cs.RO 2025-05 conditional novelty 6.0 of 10

    ChatVLA-2 uses dynamic mixture-of-experts and a two-stage training recipe to let a vision-language-action model retain pretrained reasoning while following robot instructions.

  13. Choose What to Manipulate: Revealing Data Scaling Laws in Bounding-Box Guided Policies for Semantic Manipulation

    cs.RO 2026-02 conditional novelty 5.0 of 10

    A bounding-box-conditioned diffusion policy shows a power-law improvement with the number of object classes in training data, reaching about 85% success on four semantic manipulation tasks.

  14. RynnVLA-002: A Unified Vision-Language-Action and World Model

    cs.RO 2025-11 conditional novelty 5.0 of 10

    A single model that jointly predicts robot actions and future images outperforms separate action-only and video-only models on LIBERO and real SO100 manipulation tasks.

  15. Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance

    cs.RO 2025-09 conditional novelty 5.0 of 10

    A plug-and-play fine-tuning method using two VAEs and a latent-distance guidance loss improves cross-embodiment and cross-task success rates of diffusion- and flow-based VLA policies.

  16. ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning

    cs.RO 2025-05 conditional novelty 5.0 of 10

    Fine-tuning a vision-language-action robot model on teacher-generated reasoning rationales raises average simulated manipulation success by up to 8.6 percentage points over the SpatialVLA baseline.

  17. Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges

    cs.RO 2025-08 conditional novelty 4.0 of 10

    A survey that taxonomizes robotic manipulation policies trained by imitation learning, traces their evolution, and compiles benchmark comparisons.

  18. Continual Learning for Generative AI: From LLMs to MLLMs and Beyond

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A survey that categorizes continual learning methods for generative models into architecture-based, regularization-based, and replay-based paradigms across four model families.

  19. Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents

    cs.RO 2025-05 conditional novelty 4.0 of 10

    An agentic framework using a GPT-4o planner, an OpenVLA executor, and a LoRA-fine-tuned Qwen2.5-VL verifier achieves 79.6% average success on LIBERO by decomposing and verifying subgoals.

Pith tools