Pith. sign in

REVIEW 5 cited by

Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.14957 v2 pith:7IKPJQIX submitted 2024-12-19 cs.RO cs.CV

classification cs.ROcs.CV
keywords worldrobotmodelsactionsdigitaldremalearningtwins
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A world model provides an agent with a representation of its environment, enabling it to predict the causal consequences of its actions. Current world models typically cannot directly and explicitly imitate the actual environment in front of a robot, often resulting in unrealistic behaviors and hallucinations that make them unsuitable for real-world robotics applications. To overcome those challenges, we propose to rethink robot world models as learnable digital twins. We introduce DreMa, a new approach for constructing digital twins automatically using learned explicit representations of the real world and its dynamics, bridging the gap between traditional digital twins and world models. DreMa replicates the observed world and its structure by integrating Gaussian Splatting and physics simulators, allowing robots to imagine novel configurations of objects and to predict the future consequences of robot actions thanks to its compositionality. We leverage this capability to generate new data for imitation learning by applying equivariant transformations to a small set of demonstrations. Our evaluations across various settings demonstrate significant improvements in accuracy and robustness by incrementing actions and object distributions, reducing the data needed to learn a policy and improving the generalization of the agents. As a highlight, we show that a real Franka Emika Panda robot, powered by DreMa's imagination, can successfully learn novel physical tasks from just a single example per task variation (one-shot policy learning). Our project page can be found in: https://dreamtomanipulate.github.io/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DemoBridge: A Simulation-in-the-Loop Toolkit for Single-View Human Demonstration Retargeting

    cs.RO 2026-07 conditional novelty 6.0 of 10

    DemoBridge retargets single-view human hand demonstrations into physics-validated, collision-aware robot trajectories via whole-trajectory optimization and simulation-in-the-loop re-planning.

  2. DynaMimicGen: A Data Generation Framework for Robot Learning of Dynamic Tasks

    cs.RO 2025-11 conditional novelty 6.0 of 10

    DynaMimicGen generates large robot-training datasets from one or two demonstrations by adapting DMP-based trajectories in real time to moving object poses, improving downstream imitation-learning policies over MimicGen.

  3. Calib3R: A 3D Foundation Model for Multi-Camera to Robot Calibration and 3D Metric-Scaled Scene Reconstruction

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A patternless joint optimization of camera-to-robot calibration and metric-scaled 3D reconstruction, built on MASt3R pointmaps and per-camera scale factors.

  4. Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A robot policy trained on one real demonstration plus AI-generated 3D views succeeds from novel initial poses, including opposite-side starts, across six real manipulation tasks.

  5. DSG-World: Learning a 3D Gaussian World Model from Dual State Videos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    DSG-World builds two segmented 3D Gaussian fields from two scene states and trains them with mutual consistency, enabling novel-state simulation without inpainting or dense capture.

Pith tools