Pith. sign in

REVIEW 4 cited by

Automating Robot Failure Recovery Using Vision-Language Models With Optimized Prompts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.03966 v1 pith:GUPWJFLM submitted 2024-09-06 cs.RO

classification cs.RO
keywords promptsrecoveryfailuresvlmsreasoningfailuremodelsmotion-level
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current robot autonomy struggles to operate beyond the assumed Operational Design Domain (ODD), the specific set of conditions and environments in which the system is designed to function, while the real-world is rife with uncertainties that may lead to failures. Automating recovery remains a significant challenge. Traditional methods often rely on human intervention to manually address failures or require exhaustive enumeration of failure cases and the design of specific recovery policies for each scenario, both of which are labor-intensive. Foundational Vision-Language Models (VLMs), which demonstrate remarkable common-sense generalization and reasoning capabilities, have broader, potentially unbounded ODDs. However, limitations in spatial reasoning continue to be a common challenge for many VLMs when applied to robot control and motion-level error recovery. In this paper, we investigate how optimizing visual and text prompts can enhance the spatial reasoning of VLMs, enabling them to function effectively as black-box controllers for both motion-level position correction and task-level recovery from unknown failures. Specifically, the optimizations include identifying key visual elements in visual prompts, highlighting these elements in text prompts for querying, and decomposing the reasoning process for failure detection and control generation. In experiments, prompt optimizations significantly outperform pre-trained Vision-Language-Action Models in correcting motion-level position errors and improve accuracy by 65.78% compared to VLMs with unoptimized prompts. Additionally, for task-level failures, optimized prompts enhanced the success rate by 5.8%, 5.8%, and 7.5% in VLMs' abilities to detect failures, analyze issues, and generate recovery plans, respectively, across a wide range of unknown errors in Lego assembly.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models

    cs.RO 2025-10 conditional novelty 6.0 of 10

    A VLM trained on auto-generated failure trajectories with executable correction actions helps VLA models detect and fix manipulation errors, lifting success rates by up to 22.6 percentage points.

  2. Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A self-critique, revision, and verification loop makes small vision-language models produce more detailed and more executable robot plans, beating their own baselines and, on the paper's judge-based evaluation, plans ...

  3. NeSyPack: A Neuro-Symbolic Framework for Bimanual Logistics Packing

    cs.RO 2025-06 conditional novelty 5.0 of 10

    NeSyPack, a hierarchical neuro-symbolic controller, achieved high packing success rates and won the WBCD competition at ICRA 2025.

  4. LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks

    cs.RO 2025-05 conditional novelty 4.0 of 10

    A unified vision-language-action model that emits a sub-task description followed by a discrete action token outperforms modular and action-only baselines on simulated long-horizon tabletop tasks.

Pith tools