Pith. sign in

REVIEW 16 cited by

Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.16862 v1 pith:4FABLZ2E submitted 2024-06-24 cs.RO cs.CV

classification cs.ROcs.CV
keywords learningpolicyvideoallowsexecutiongenerativehumanmodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

A key challenge in manipulation is learning a policy that can robustly generalize to diverse visual environments. A promising mechanism for learning robust policies is to leverage video generative models, which are pretrained on large-scale datasets of internet videos. In this paper, we propose a visuomotor policy learning framework that fine-tunes a video diffusion model on human demonstrations of a given task. At test time, we generate an example of an execution of the task conditioned on images of a novel scene, and use this synthesized execution directly to control the robot. Our key insight is that using common tools allows us to effortlessly bridge the embodiment gap between the human hand and the robot manipulator. We evaluate our approach on four tasks of increasing complexity and demonstrate that harnessing internet-scale generative models allows the learned policy to achieve a significantly higher degree of generalization than existing behavior cloning approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Pre-training a VLA model on 100k hours of auto-labeled UMI trajectories, then post-training on robot data, yields SOTA simulated manipulation and data-efficient fine-tuning.

  2. RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A tri-branch diffusion model co-generates RGB, depth, and optical flow from a single RGB-D image, and an inverse dynamics head on its internal latents achieves state-of-the-art bimanual manipulation success rates.

  3. MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A geometry-consistent multi-view RGBD 4D world model for robot manipulation, whose actions are recovered by test-time optimization of a learned trajectory latent, outperforming single- and dual-view world-model baseli...

  4. On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning

    cs.LG 2026-02 conditional novelty 6.0 of 10

    VM-IDM and IDM labeling coincide at infinite unlabeled data, and IDM learning is more label-efficient than BC because the ground-truth IDM is typically less complex and less stochastic than the expert policy.

  5. Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation

    cs.RO 2025-08 conditional novelty 6.0 of 10

    GVF-TAPE predicts future RGB-D frames from an image and text, then extracts end-effector poses to control a robot, achieving strong success rates without action-labeled data.

  6. AMPLIFY: Actionless Motion Priors for Robot Learning from Videos

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A three-stage pipeline that turns keypoint tracks into discrete motion tokens, predicts them from action-free video, and decodes them into actions yields large few-shot and zero-shot policy improvements in robot manipulation.

  7. Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A two-stage pipeline trains a robot policy that accepts a human demonstration video as a prompt and generalizes beyond its robot training tasks, with success rates of up to 79 percent on known task variations and unde...

  8. TesserAct: Learning 4D Embodied World Models

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A 4D embodied world model that generates RGB-depth-normal videos from an image and instruction, reconstructs the scene as point clouds, and uses those point clouds to train better robot manipulation policies.

  9. Solving New Tasks by Adapting Internet Video Knowledge

    cs.LG 2025-04 conditional novelty 6.0 of 10

    Inverse Probabilistic Adaptation, a score-composition variant that keeps the large video model as the base and consults a small in-domain model, achieves 68.3% average success on MetaWorld policy supervision and stays...

  10. VILP: Imitation Learning with Latent Video Planning

    cs.RO 2025-02 conditional novelty 6.0 of 10

    VILP generates multi-view future videos in a compressed latent space and converts them into robot actions with a lightweight policy, achieving real-time receding horizon control on tested manipulation tasks.

  11. FLIP: Flow-Centric Generative Planning as General-Purpose Manipulation World Model

    cs.RO 2024-12 conditional novelty 6.0 of 10

    FLIP combines flow generation, flow-conditioned video prediction, and a learned value function to plan long robot manipulations from only image and language inputs.

  12. ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A video diffusion model generates scene-conditioned, step-by-step visual instructions from an input image and text prompts, trained on a new 0.6M-sequence dataset.

  13. Distilling Physical Priors into Streaming World Models

    cs.CV 2026-08 conditional novelty 5.0 of 10

    PhyS adds physics-aware video data, teacher distillation, and windowed reward routing to make streaming world models generate more physically plausible long rollouts.

  14. PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration

    cs.CV 2026-07 conditional novelty 5.0 of 10

    PAVXploreRL post-trains action-conditioned world models with VJEPA-2 latent rewards and perturbed 'OOD' actions, reporting a 5.6% average gain and lowered policy-overestimation bias.

  15. GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning

    cs.RO 2025-08 conditional novelty 5.0 of 10

    GenFlowRL converts generated 2D object keypoint flows into a compact delta-flow that shapes dense RL rewards, improving robot manipulation performance on 10 simulation and real-world probe tasks.

  16. A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges

    cs.CV 2025-01 reject novelty 2.0 of 10

    A survey that catalogs large vision-language models, their alignment methods, benchmarks, and challenges, but is compromised by inconsistent counts and misclassified entries.

Pith tools