Pith. sign in

REVIEW 6 cited by

CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.05711 v2 pith:A335MGIC submitted 2022-12-12 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords learningrobotcactidatatrainingframeworkkitchenperform
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale training have propelled significant progress in various sub-fields of AI such as computer vision and natural language processing. However, building robot learning systems at a comparable scale remains challenging. To develop robots that can perform a wide range of skills and adapt to new scenarios, efficient methods for collecting vast and diverse amounts of data on physical robot systems are required, as well as the capability to train high-capacity policies using such datasets. In this work, we propose a framework for scaling robot learning, with specific focus on multi-task and multi-scene manipulation in kitchen environments, both in simulation and in the real world. Our proposed framework, CACTI, comprises four stages that separately handle: data collection, data augmentation, visual representation learning, and imitation policy training, to enable scalability in robot learning . We make use of state-of-the-art generative models as part of the data augmentation stage, and use pre-trained out-of-domain visual representations to improve training efficiency. Experimental results demonstrate the effectiveness of our approach. On a real robot setup, CACTI enables efficient training of a single policy that can perform 10 manipulation tasks involving kitchen objects, and is robust to varying layouts of distractors. In a simulated kitchen environment, CACTI trains a single policy to perform 18 semantic tasks across 100 layout variations for each individual task. We will release the simulation task benchmark and augmented datasets in both real and simulated environments to facilitate future research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  2. RoboLight: A Dataset with Linearly Composable Illumination for Robotic Manipulation

    cs.RO 2026-03 conditional novelty 6.0 of 10

    A dataset that records identical robot manipulation tasks under 14 controlled lighting conditions and uses HDR linearity to synthesize 196,000 additional lighting-varied episodes.

  3. OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction

    cs.RO 2025-09 conditional novelty 6.0 of 10

    An interaction-mesh retargeting pipeline with hard constraints generates robot training references that preserve object/terrain contacts, enabling long-horizon humanoid loco-manipulation with minimal rewards.

  4. VLM-TDP: VLM-guided Trajectory-conditioned Diffusion Policy for Robust Long-Horizon Manipulation

    cs.RO 2025-07 conditional novelty 6.0 of 10

    VLM-TDP guides a diffusion-based robot policy with VLM-generated voxel trajectories, improving success rates by roughly 30-44% and adding robustness to noise and scene changes.

  5. RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A generative model and wrist camera turn human hand videos into robot gripper demonstrations that train manipulation policies at success rates close to those trained on real gripper data.

  6. ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents

    cs.CV 2025-07 conditional novelty 5.0 of 10

    ERMV edits 4D multi-view robot videos from one edited frame plus robot states, and VLA policies trained on the edited data show higher success rates in simulation and real robot tests.

Pith tools