Pith. sign in

REVIEW 7 cited by

GenAug: Retargeting behaviors to unseen situations via Generative Augmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.06671 v2 pith:7UHR7KAH submitted 2023-02-13 cs.RO

classification cs.RO
keywords datageneralizationgenerativemodelsrobotgenauglearningpre-trained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Robot learning methods have the potential for widespread generalization across tasks, environments, and objects. However, these methods require large diverse datasets that are expensive to collect in real-world robotics settings. For robot learning to generalize, we must be able to leverage sources of data or priors beyond the robot's own experience. In this work, we posit that image-text generative models, which are pre-trained on large corpora of web-scraped data, can serve as such a data source. We show that despite these generative models being trained on largely non-robotics data, they can serve as effective ways to impart priors into the process of robot learning in a way that enables widespread generalization. In particular, we show how pre-trained generative models can serve as effective tools for semantically meaningful data augmentation. By leveraging these pre-trained models for generating appropriate "semantic" data augmentations, we propose a system GenAug that is able to significantly improve policy generalization. We apply GenAug to tabletop manipulation tasks, showing the ability to re-target behavior to novel scenarios, while only requiring marginal amounts of real-world data. We demonstrate the efficacy of this system on a number of object manipulation problems in the real world, showing a 40% improvement in generalization to novel scenes and objects.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  2. DynamicManip: Enabling Dynamic Manipulation from a Single Static Demonstration

    cs.RO 2026-08 conditional novelty 6.0 of 10

    DynamicManip synthesizes diverse dynamic manipulation demonstrations from one static demonstration and uses stage-aware adaptive inference to improve success rates and reduce latency.

  3. Affordance-Based Manipulation Planning with Text Goals and Sim-to-Real Generalisation via Real-to-Sim Image Conversion

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Affordance recognition, multi-step visual effect prediction, and multimodal text matching produce robot plans that handle occlusion; real-to-sim conversion enables hardware transfer.

  4. VLM-TDP: VLM-guided Trajectory-conditioned Diffusion Policy for Robust Long-Horizon Manipulation

    cs.RO 2025-07 conditional novelty 6.0 of 10

    VLM-TDP guides a diffusion-based robot policy with VLM-generated voxel trajectories, improving success rates by roughly 30-44% and adding robustness to noise and scene changes.

  5. RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A generative model and wrist camera turn human hand videos into robot gripper demonstrations that train manipulation policies at success rates close to those trained on real gripper data.

  6. Video2Policy: Scaling up Manipulation Tasks in Simulation through Internet Videos

    cs.RO 2025-02 conditional novelty 6.0 of 10

    Video2Policy automatically builds simulated manipulation tasks from internet videos and trains RL policies using LLM-generated rewards, reaching 88% average simulation success and 47% sim2real success.

  7. ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents

    cs.CV 2025-07 conditional novelty 5.0 of 10

    ERMV edits 4D multi-view robot videos from one edited frame plus robot states, and VLA policies trained on the edited data show higher success rates in simulation and real robot tests.

Pith tools