Pith. sign in

REVIEW 15 cited by

Yell At Your Robot: Improving On-the-Fly from Language Corrections

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.12910 v1 pith:ELOZOK6E submitted 2024-03-19 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords feedbackhigh-levellanguagecorrectionslong-horizonlow-levelpoliciesrobot
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hierarchical policies that combine language and low-level control have been shown to perform impressively long-horizon robotic tasks, by leveraging either zero-shot high-level planners like pretrained language and vision-language models (LLMs/VLMs) or models trained on annotated robotic demonstrations. However, for complex and dexterous skills, attaining high success rates on long-horizon tasks still represents a major challenge -- the longer the task is, the more likely it is that some stage will fail. Can humans help the robot to continuously improve its long-horizon task performance through intuitive and natural feedback? In this paper, we make the following observation: high-level policies that index into sufficiently rich and expressive low-level language-conditioned skills can be readily supervised with human feedback in the form of language corrections. We show that even fine-grained corrections, such as small movements ("move a bit to the left"), can be effectively incorporated into high-level policies, and that such corrections can be readily obtained from humans observing the robot and making occasional suggestions. This framework enables robots not only to rapidly adapt to real-time language feedback, but also incorporate this feedback into an iterative training scheme that improves the high-level policy's ability to correct errors in both low-level execution and high-level decision-making purely from verbal feedback. Our evaluation on real hardware shows that this leads to significant performance improvement in long-horizon, dexterous manipulation tasks without the need for any additional teleoperation. Videos and code are available at https://yay-robot.github.io/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Zero2Skill: Bootstrapping Robot Skills through Autonomous Data Collection, Training, and Deployment

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Corrective Memory lets a robot data collector reuse natural-language corrections across rounds, cutting human time to 16% of teleoperation while matching its success rate and downstream policy performance.

  2. FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models

    cs.RO 2025-10 conditional novelty 6.0 of 10

    A VLM trained on auto-generated failure trajectories with executable correction actions helps VLA models detect and fix manipulation errors, lifting success rates by up to 22.6 percentage points.

  3. Robix: A Unified Model for Robot Interaction, Reasoning and Planning

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A three-stage-trained VLM unifies robot planning and dialogue, and beats commercial VLMs on the authors' interactive-task benchmarks.

  4. CueLearner: Bootstrapping and local policy adaptation from relative feedback

    cs.RO 2025-07 conditional novelty 6.0 of 10

    CueLearner learns a relative-feedback model from a small number of human directional corrections and uses it to guide off-policy RL exploration or refine a deployed policy.

  5. FEAST: A Flexible Mealtime-Assistance System Towards In-the-Wild Personalization

    cs.RO 2025-06 conditional novelty 6.0 of 10

    FEAST is a mealtime assistance robot that uses LLM-editable behavior trees and modular tools to let care recipients personalize feeding, drinking, and mouth wiping in real home settings.

  6. SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models

    cs.RO 2025-06 conditional novelty 6.0 of 10

    SwitchVLA trains a vision-language-action policy to handle mid-execution instruction changes by conditioning on contact state and a three-way behavior mode, using only existing single-task demonstrations.

  7. Learning Compositional Behaviors from Demonstration and Language

    cs.RO 2025-05 conditional novelty 6.0 of 10

    BLADE learns structured, planable action representations from language-annotated demonstrations and composes them with a symbolic planner, outperforming latent and LLM/VLM baselines on new manipulation tasks.

  8. A Few Words Go a Long Way: Language Guided Robot Policy Synthesis

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Interactive LLM program synthesis plus a persistent skill library from natural-language corrections outperforms zero-shot VLAs and one-shot code policies on complex real-robot manipulation.

  9. Scene Graph-Guided Proactive Replanning for Failure-Resilient Embodied Agent

    cs.RO 2025-08 conditional novelty 5.0 of 10

    A robot replanner that compares scene graphs to successful demonstrations before each subtask, triggering LLM-based replanning on mismatch, raises task success in AI2-THOR.

  10. A Human-in-the-loop Approach to Robot Action Replanning through LLM Common-Sense Reasoning

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A human-in-the-loop system lets users refine vision-generated robot behavior trees through natural-language requests to GPT-4o, correcting errors and adapting plans before execution.

  11. Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization

    cs.RO 2025-07 conditional novelty 5.0 of 10

    Tactile-VLA fuses tactile sensing into a vision-language-action model so force-related instructions and corrective reasoning transfer to new contact-rich tasks with few demonstrations.

  12. ProVox: Personalization and Proactive Planning for Situated Human-Robot Collaboration

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A personalized, proactive LLM planner suggests helpful next actions during human-robot lunch packing, reporting 38.7% faster task execution at the cost of an extra 5.6-minute setup phase.

  13. Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning

    cs.RO 2025-08 reject novelty 4.0 of 10

    A review that categorizes large-model-empowered embodied AI into hierarchical and end-to-end decision-making, imitation and reinforcement learning, and world models.

  14. Steering Robots with Inference-Time Interactions

    cs.RO 2025-06 conditional novelty 4.0 of 10

    Frozen imitation policies can be steered at inference time via user interactions, with a diffusion-sampling method and a constraint-enforcing framework that provides formal task guarantees.

  15. LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks

    cs.RO 2025-05 conditional novelty 4.0 of 10

    A unified vision-language-action model that emits a sub-task description followed by a discrete action token outperforms modular and action-only baselines on simulated long-horizon tabletop tasks.

Pith tools