Pith. sign in

REVIEW 8 cited by

Photorealistic Video Generation with Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.06662 v1 pith:YVYQU3QO submitted 2023-12-11 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords generationvideodiffusionmodelsapproachdecisionsdesignlatent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We present W.A.L.T, a transformer-based approach for photorealistic video generation via diffusion modeling. Our approach has two key design decisions. First, we use a causal encoder to jointly compress images and videos within a unified latent space, enabling training and generation across modalities. Second, for memory and training efficiency, we use a window attention architecture tailored for joint spatial and spatiotemporal generative modeling. Taken together these design decisions enable us to achieve state-of-the-art performance on established video (UCF-101 and Kinetics-600) and image (ImageNet) generation benchmarks without using classifier free guidance. Finally, we also train a cascade of three models for the task of text-to-video generation consisting of a base latent video diffusion model, and two video super-resolution diffusion models to generate videos of $512 \times 896$ resolution at $8$ frames per second.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ActionParty: Multi-Subject Action Binding in Generative Video Games

    cs.CV 2026-04 conditional novelty 7.0 of 10

    ActionParty binds discrete actions to individual subjects in a single generated video by jointly modeling subject state tokens and video latents, controlling up to seven players across 46 Melting Pot games.

  2. WaiT for the Signal: Simple Frequency-Aware Flow-Matching

    cs.CV 2026-07 conditional novelty 6.0 of 10

    WaiT delays high-frequency wavelet bands in flow-matching image generation until coarse structure emerges, improving quality and cutting compute, with a reported SOTA FID of 1.30 on ImageNet 512.

  3. Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments

    cs.LG 2026-01 conditional novelty 6.0 of 10

    Flow equivariant world models use a latent memory that shifts with the agent and with inferred object motion, giving stable long-horizon prediction under partial observability.

  4. MambaVideo for Discrete Video Tokenization with Channel-Split Quantization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A Mamba-based hierarchical video tokenizer with channel-split quantization achieves state-of-the-art reconstruction and generation scores while preserving token count.

  5. Populate-A-Scene: Affordance-Aware Human Video Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A fine-tuned text-to-video model inserts a person into a scene and generates an interaction video without bounding boxes or pose input, and its attention maps reveal a latent sense of affordance.

  6. HiWave: Training-Free High-Resolution Image Generation via Wavelet-Based Diffusion Sampling

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A training-free diffusion sampling pipeline combines patch-wise DDIM inversion with wavelet-domain frequency guidance to generate coherent 4096x4096 images from SDXL.

  7. VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment

    cs.CV 2026-03 conditional novelty 5.5 of 10

    A VGGT-based pointwise reprojection reward with geometry-aware sampling improves video geometric consistency via SFT/DPO and causal test-time search.

  8. Leveraging Pre-Trained Visual Models for AI-Generated Video Detection

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Pre-trained SigLIP/VideoMAE features with a linear probe or nearest-neighbor distance separate real videos from text-to-video model outputs, reaching about 90% average F1 on the new VID-AID benchmark, but much lower o...

Pith tools