Pith. sign in

REVIEW 6 cited by

A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.17418 v2 pith:D4FDP6BI submitted 2024-05-27 cs.CV

classification cs.CV
keywords systemmanipulationfastactionsfailuremodelslowtasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, some studies have integrated Multimodal Large Language Models into robotic manipulation, constructing vision-language-action models (VLAs) to interpret multimodal information and predict SE(3) poses. While VLAs have shown promising progress, they may suffer from failures when faced with novel and complex tasks. To emulate human-like reasoning for more robust manipulation, we propose the self-corrected (SC-)VLA framework, which integrates fast system for directly predicting actions and slow system for reflecting on failed actions within a single VLA policy. For the fast system, we incorporate parameter-efficient fine-tuning to equip the model with pose prediction capabilities while preserving the inherent reasoning abilities of MLLMs. For the slow system, we propose a Chain-of-Thought training strategy for failure correction, designed to mimic human reflection after a manipulation failure. Specifically, our model learns to identify the causes of action failures, adaptively seek expert feedback, reflect on the current failure scenario, and iteratively generate corrective actions, step by step. Furthermore, a continuous policy learning method is designed based on successfully corrected samples, enhancing the fast system's adaptability to the current configuration. We compare SC-VLA with the previous SOTA VLA in both simulation and real-world tasks, demonstrating an efficient correction process and improved manipulation accuracy on both seen and unseen tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Token-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Streaming intermediate hidden states from a slow VLM to a fast flow-matching planner, token by token, improves dynamic social VLN success and reduces observation staleness.

  2. CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding

    cs.RO 2026-01 reject novelty 6.0 of 10

    CycleVLA adds progress-triggered VLM failure checks, subtask backtracking, and MBR consensus decoding to VLAs, raising LIBERO average success from 89.3% to 95.3% and claiming 91% real-robot success.

  3. Reinforcement Learning for Flow-Matching Policies

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Reward-weighted flow matching and GRPO with a learned reward surrogate both improve flow-matching policies beyond a suboptimal demonstrator on simulated unicycle tasks.

  4. The Latent Color Subspace: Emergent Order in High-Dimensional Chaos

    cs.LG 2026-03 unverdicted novelty 5.0 of 10

    FLUX.1’s VAE latent space contains an interpretable Hue–Saturation–Lightness structure that enables training-free color prediction and control via closed-form latent edits.

  5. Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning

    cs.RO 2025-06 conditional novelty 5.0 of 10

    FiS-VLA embeds a diffusion-based action module into the final transformer blocks of a vision-language model, achieving 69% mean success on RLBench and a claimed 117.7 Hz control frequency.

  6. ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models

    cs.RO 2025-05 conditional novelty 5.0 of 10

    ManipLVM-R1 applies RLVR with IoU and trajectory-distance rewards to train a 3B VLM for affordance perception and trajectory prediction, claiming better performance and generalization than SFT on 50% of the data.

Pith tools