Pith. sign in

REVIEW 6 cited by

Diffusion Models for Video Prediction and Infilling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.07696 v3 pith:PC4C364T submitted 2022-06-15 cs.CV cs.LGstat.ML

classification cs.CVcs.LGstat.ML
keywords videodiffusionmodelspredictionableconditioninggenerativeinfilling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Predicting and anticipating future outcomes or reasoning about missing information in a sequence are critical skills for agents to be able to make intelligent decisions. This requires strong, temporally coherent generative capabilities. Diffusion models have shown remarkable success in several generative tasks, but have not been extensively explored in the video domain. We present Random-Mask Video Diffusion (RaMViD), which extends image diffusion models to videos using 3D convolutions, and introduces a new conditioning technique during training. By varying the mask we condition on, the model is able to perform video prediction, infilling, and upsampling. Due to our simple conditioning scheme, we can utilize the same architecture as used for unconditional training, which allows us to train the model in a conditional and unconditional fashion at the same time. We evaluate RaMViD on two benchmark datasets for video prediction, on which we achieve state-of-the-art results, and one for video generation. High-resolution videos are provided at https://sites.google.com/view/video-diffusion-prediction.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ELT: Elastic Looped Transformers for Visual Generation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Weight-shared looped transformers trained with intra-loop self-distillation match MaskGIT-class FID/FVD at roughly 4x fewer parameters and support any-time inference across loop counts.

  2. Programmatic Video Prediction Using Large Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ProgGen uses language-model-written programs for perception, dynamics, and rendering to predict future video frames from about ten training examples, beating large diffusion baselines on two synthetic benchmarks.

  3. Masked Generative Nested Transformers with Decode Time Scaling

    cs.CV 2025-02 conditional novelty 6.0 of 10

    MaGNeTS schedules progressively larger nested transformer sub-models over decode iterations and caches key-value pairs of unmasked tokens, achieving 2.5-3.7x compute reduction with competitive FID/FVD.

  4. CPKD: Clinical Prior Knowledge-Constrained Diffusion Models for Surgical Phase Recognition in Endoscopic Submucosal Dissection

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A diffusion-based generative model with training-time masking and clinical logic constraints achieves state-of-the-art surgical phase recognition on ESD videos and a small gain on cholecystectomy videos.

  5. Multi-View Face and Gesture Animation with Dynamic Gaussians

    cs.CV 2026-08 conditional novelty 4.0 of 10

    Combining separate face and hand models with a parametric body and Gaussian splatting enables multi-view-consistent upper-body avatars that can be re-animated with new expressions and gestures.

  6. On the Benefits of Instance Decomposition in Video Prediction Models

    cs.CV 2025-01 reject novelty 4.0 of 10

    Explicit instance decomposition with per-class shared weights improves latent-transformer video prediction in the paper's experiments, but the claimed advantage is weakened by mismatched parameter counts and test-set ...

Pith tools