Pith. sign in

REVIEW 3 cited by

AVID: Learning Multi-Stage Tasks via Pixel-Level Translation of Human Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.04443 v3 pith:SEG56LOS submitted 2019-12-10 cs.RO cs.CVcs.LG

classification cs.ROcs.CVcs.LG
keywords humanlearningrobottasksautomatedimagetaskthen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Robotic reinforcement learning (RL) holds the promise of enabling robots to learn complex behaviors through experience. However, realizing this promise for long-horizon tasks in the real world requires mechanisms to reduce human burden in terms of defining the task and scaffolding the learning process. In this paper, we study how these challenges can be alleviated with an automated robotic learning framework, in which multi-stage tasks are defined simply by providing videos of a human demonstrator and then learned autonomously by the robot from raw image observations. A central challenge in imitating human videos is the difference in appearance between the human and robot, which typically requires manual correspondence. We instead take an automated approach and perform pixel-level image translation via CycleGAN to convert the human demonstration into a video of a robot, which can then be used to construct a reward function for a model-based RL algorithm. The robot then learns the task one stage at a time, automatically learning how to reset each stage to retry it multiple times without human-provided resets. This makes the learning process largely automatic, from intuitive task specification via a video to automated training with minimal human intervention. We demonstrate that our approach is capable of learning complex tasks, such as operating a coffee machine, directly from raw image observations, requiring only 20 minutes to provide human demonstrations and about 180 minutes of robot interaction.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

    cs.RO 2026-04 unverdicted novelty 6.5 of 10

    EgoVerse releases 1,362 hours of standardized egocentric human data across 1,965 tasks and shows via multi-lab experiments that robot policy performance scales with human data volume when the data aligns with robot ob...

  2. Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A two-stage pipeline trains a robot policy that accepts a human demonstration video as a prompt and generalizes beyond its robot training tasks, with success rates of up to 79 percent on known task variations and unde...

  3. Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning

    cs.LG 2025-02 conditional novelty 5.0 of 10

    PrefVLM combines VLM-generated trajectory preferences with selective human feedback and inverse-dynamics VLM adaptation, matching PEBBLE on five Meta-World tasks with up to 2x fewer human labels.

Pith tools