Pith. sign in

REVIEW 3 cited by

AVID: Adapting Video Diffusion Models to World Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.12822 v2 pith:H7TWH3BF submitted 2024-10-01 cs.CV cs.LG

classification cs.CVcs.LG
keywords modelsaviddiffusionmodelpretrainedvideovideosworld
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale generative models have achieved remarkable success in a number of domains. However, for sequential decision-making problems, such as robotics, action-labelled data is often scarce and therefore scaling-up foundation models for decision-making remains a challenge. A potential solution lies in leveraging widely-available unlabelled videos to train world models that simulate the consequences of actions. If the world model is accurate, it can be used to optimize decision-making in downstream tasks. Image-to-video diffusion models are already capable of generating highly realistic synthetic videos. However, these models are not action-conditioned, and the most powerful models are closed-source which means they cannot be finetuned. In this work, we propose to adapt pretrained video diffusion models to action-conditioned world models, without access to the parameters of the pretrained model. Our approach, AVID, trains an adapter on a small domain-specific dataset of action-labelled videos. AVID uses a learned mask to modify the intermediate outputs of the pretrained model and generate accurate action-conditioned videos. We evaluate AVID on video game and real-world robotics data, and show that it outperforms existing baselines for diffusion model adaptation.1 Our results demonstrate that if utilized correctly, pretrained video models have the potential to be powerful tools for embodied AI.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction

    cs.CV 2025-05 conditional novelty 6.0 of 10

    CogRobot uses optical flow as an intermediate variable to fine-tune a text-to-video model for predicting bimanual robot trajectories, then maps those predictions to actions with a goal-conditioned diffusion policy.

  2. Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression

    cs.RO 2025-02 conditional novelty 6.0 of 10

    HMA is a masked autoregressive transformer that predicts future video and actions across many robot embodiments, running up to 15x faster than prior diffusion-based video simulators while matching or improving visual ...

  3. Pre-Trained Video Generative Models as World Simulators

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A lightweight action-conditioning module and a motion-reinforced loss convert pre-trained video generators into action-following world simulators that also speed up model-based reinforcement learning.

Pith tools