Pith. sign in

REVIEW 5 cited by

VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.13508 v2 pith:KYSFS7XY submitted 2025-02-19 cs.RO

classification cs.RO
keywords speechvlasrobotmodelcustomizedinstructionsinteractionmanipulation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-language-action models (VLAs) have become increasingly popular in robot manipulation for their end-to-end design and remarkable performance. However, existing VLAs rely heavily on vision-language models (VLMs) that only support text-based instructions, neglecting the more natural speech modality for human-robot interaction. Traditional speech integration methods usually involves a separate speech recognition system, which complicates the model and introduces error propagation. Moreover, the transcription procedure would lose non-semantic information in the raw speech, such as voiceprint, which may be crucial for robots to successfully complete customized tasks. To overcome above challenges, we propose VLAS, a novel end-to-end VLA that integrates speech recognition directly into the robot policy model. VLAS allows the robot to understand spoken commands through inner speech-text alignment and produces corresponding actions to fulfill the task. We also present two new datasets, SQA and CSI, to support a three-stage tuning process for speech instructions, which empowers VLAS with the ability of multimodal interaction across text, image, speech, and robot actions. Taking a step further, a voice retrieval-augmented generation (RAG) paradigm is designed to enable our model to effectively handle tasks that require individual-specific knowledge. Our extensive experiments show that VLAS can effectively accomplish robot manipulation tasks with diverse speech commands, offering a seamless and customized interaction experience.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Monotonic Progress: Retry-Supervised Value Learning for Robot Imitation

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    Sparse retry keypoints plus pairwise preference learning yield mistake-sensitive values that reweight mixed-quality demos and raise real-robot imitation success over progress-based baselines.

  2. UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    UniSteer unifies human corrective actions and noise-space RL for VLA adaptation by inverting actions to noise targets, raising success rates from 20% to 90% in 66 minutes across four real-world manipulation tasks.

  3. Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting

    cs.RO 2025-12 conditional novelty 6.0 of 10

    A frozen vision-language-action robot policy can manipulate a user-specific object when the object is grounded from a few reference photos, highlighted in the camera view, and the instruction is rewritten to name the ...

  4. Differential Amplifier-Inspired AmpAttention for Multi-View Robotic Manipulation

    cs.RO 2026-07 conditional novelty 5.0 of 10

    AmpAttention and RVAF raise multi-view robotic manipulation success and cut training time by suppressing attention noise with a differential-amplifier-style mechanism plus a CMRR loss.

  5. VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A convolutional residual VQ-VAE action tokenizer trained on over 100x more data than prior work improves OpenVLA success rates and inference speed on several manipulation tasks.

Pith tools