Pith. sign in

REVIEW 21 cited by

Lumiere: A Space-Time Diffusion Model for Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.12945 v2 pith:BDCF3H52 submitted 2024-01-23 cs.CV

classification cs.CV
keywords videomodeltemporaldiffusiongenerationspace-timeintroducelumiere
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Lumiere -- a text-to-video diffusion model designed for synthesizing videos that portray realistic, diverse and coherent motion -- a pivotal challenge in video synthesis. To this end, we introduce a Space-Time U-Net architecture that generates the entire temporal duration of the video at once, through a single pass in the model. This is in contrast to existing video models which synthesize distant keyframes followed by temporal super-resolution -- an approach that inherently makes global temporal consistency difficult to achieve. By deploying both spatial and (importantly) temporal down- and up-sampling and leveraging a pre-trained text-to-image diffusion model, our model learns to directly generate a full-frame-rate, low-resolution video by processing it in multiple space-time scales. We demonstrate state-of-the-art text-to-video generation results, and show that our design easily facilitates a wide range of content creation tasks and video editing applications, including image-to-video, video inpainting, and stylized generation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Using a synthetic alphabet video testbed, balanced data mixing and caption precision are shown to dominate T2V model quality, while CFG and fine-tuning only partially compensate for corrupted captions.

  2. GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video Diffusion

    cs.CV 2025-05 conditional novelty 7.0 of 10

    GeoMan predicts temporally consistent depth and normals for human videos by conditioning an image-to-video diffusion model on first-frame geometry and using a root-relative depth representation.

  3. OSVE: One Step Video Editing with One Step Diffusion Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    OSVE performs text-guided video editing with one-step diffusion models by training a single-pass inversion encoder and unifying frame latents for cross-frame attention, achieving quality comparable to multi-step metho...

  4. PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Iteratively simulating, verifying, and repairing physics programs gives video generation more reliable fine-grained control over object motion than one-shot configuration.

  5. Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Single-image novel view synthesis is decomposed into panorama outpainting plus keyframe-conditioned video diffusion, producing loop-consistent scene tours.

  6. CineScale: Free Lunch in High-Resolution Cinematic Visual Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    CineScale extends pre-trained diffusion models to 8k image and 4k video generation with mostly tuning-free inference plus a small LoRA adaptation for video.

  7. Vid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free Simulation

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Vid2Sim recovers 3D geometry, appearance, and elastic material parameters from multi-view videos using a feed-forward network plus a fast refinement, enabling mesh-free reduced-order simulation.

  8. Humanoid World Models: Open World Foundation Models for Humanoid Robotics

    cs.RO 2025-06 conditional novelty 6.0 of 10

    Masked-transformers trained on humanoid video forecast future frames with better FID than flow-matching models, and parameter sharing cut model size 33-53% with minimal quality loss.

  9. FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FlowMo reduces temporal artifacts in video generation by guiding the denoising process to lower the maximum patch-wise variance of consecutive-frame differences in the latent space.

  10. MOVi: Training-free Text-conditioned Multi-Object Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MOVi improves multi-object video generation without retraining by using LLM-planned trajectories to reinitialize the diffusion noise and by reweighting attention to stop objects from mixing together.

  11. MotionPro: A Precise Motion Controller for Image-to-Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MotionPro uses region-wise trajectories and a motion mask to control object and camera motion in image-to-video generation, reporting improved trajectory alignment over prior methods.

  12. HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

    cs.CV 2025-05 conditional novelty 6.0 of 10

    HunyuanVideo-Avatar is an audio-driven video generator that enables emotion-controllable and multi-character animation by injecting character images, routing audio via face masks, and transferring emotion from referen...

  13. ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ProphetDWM is a one-stage diffusion world model that jointly predicts future driving video and low-level actions from a current frame and a short action sequence.

  14. Self Gradient Forcing: Native Long Video Extrapolation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Self Gradient Forcing lets later-video losses train how earlier generated video latents are encoded into the causal KV cache, improving measured long-video consistency over plain Self Forcing.

  15. OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    OmniV2V is one diffusion-transformer model that performs eight video generation and editing tasks by combining mask, pose, image, and text-instruction conditions.

  16. Multi-View Face and Gesture Animation with Dynamic Gaussians

    cs.CV 2026-08 conditional novelty 4.0 of 10

    Combining separate face and hand models with a parametric body and Gaussian splatting enables multi-view-consistent upper-body avatars that can be re-animated with new expressions and gestures.

  17. Advances in 4D Representation: Geometry, Motion, and Interaction

    cs.CV 2025-10 conditional novelty 4.0 of 10

    A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.

  18. SplitMeanFlow: Interval Splitting Consistency in Few-Step Generative Modeling

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A purely algebraic interval-splitting consistency objective trains few-step generative models without JVP computations and recovers MeanFlow's differential identity as a special limit.

  19. Leveraging Pre-Trained Visual Models for AI-Generated Video Detection

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Pre-trained SigLIP/VideoMAE features with a linear probe or nearest-neighbor distance separate real videos from text-to-video model outputs, reaching about 90% average F1 on the new VID-AID benchmark, but much lower o...

  20. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

  21. A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality

    cs.CV 2025-07 conditional novelty 3.0 of 10

    A survey of 32 long-video generation papers, presenting a taxonomy and component recommendations for backbones, text encoders, objectives, and positional encodings.

Pith tools