REVIEW 21 cited by
Lumiere: A Space-Time Diffusion Model for Video Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce Lumiere -- a text-to-video diffusion model designed for synthesizing videos that portray realistic, diverse and coherent motion -- a pivotal challenge in video synthesis. To this end, we introduce a Space-Time U-Net architecture that generates the entire temporal duration of the video at once, through a single pass in the model. This is in contrast to existing video models which synthesize distant keyframes followed by temporal super-resolution -- an approach that inherently makes global temporal consistency difficult to achieve. By deploying both spatial and (importantly) temporal down- and up-sampling and leveraging a pre-trained text-to-image diffusion model, our model learns to directly generate a full-frame-rate, low-resolution video by processing it in multiple space-time scales. We demonstrate state-of-the-art text-to-video generation results, and show that our design easily facilitates a wide range of content creation tasks and video editing applications, including image-to-video, video inpainting, and stylized generation.
Forward citations
Cited by 21 Pith papers
-
Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
Using a synthetic alphabet video testbed, balanced data mixing and caption precision are shown to dominate T2V model quality, while CFG and fine-tuning only partially compensate for corrupted captions.
-
GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video Diffusion
GeoMan predicts temporally consistent depth and normals for human videos by conditioning an image-to-video diffusion model on first-frame geometry and using a root-relative depth representation.
-
OSVE: One Step Video Editing with One Step Diffusion Models
OSVE performs text-guided video editing with one-step diffusion models by training a single-pass inversion encoder and unifying frame latents for cross-frame attention, achieving quality comparable to multi-step metho...
-
PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation
Iteratively simulating, verifying, and repairing physics programs gives video generation more reliable fine-grained control over object motion than one-shot configuration.
-
Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion
Single-image novel view synthesis is decomposed into panorama outpainting plus keyframe-conditioned video diffusion, producing loop-consistent scene tours.
-
CineScale: Free Lunch in High-Resolution Cinematic Visual Generation
CineScale extends pre-trained diffusion models to 8k image and 4k video generation with mostly tuning-free inference plus a small LoRA adaptation for video.
-
Vid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free Simulation
Vid2Sim recovers 3D geometry, appearance, and elastic material parameters from multi-view videos using a feed-forward network plus a fast refinement, enabling mesh-free reduced-order simulation.
-
Humanoid World Models: Open World Foundation Models for Humanoid Robotics
Masked-transformers trained on humanoid video forecast future frames with better FID than flow-matching models, and parameter sharing cut model size 33-53% with minimal quality loss.
-
FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation
FlowMo reduces temporal artifacts in video generation by guiding the denoising process to lower the maximum patch-wise variance of consecutive-frame differences in the latent space.
-
MOVi: Training-free Text-conditioned Multi-Object Video Generation
MOVi improves multi-object video generation without retraining by using LLM-planned trajectories to reinitialize the diffusion noise and by reweighting attention to stop objects from mixing together.
-
MotionPro: A Precise Motion Controller for Image-to-Video Generation
MotionPro uses region-wise trajectories and a motion mask to control object and camera motion in image-to-video generation, reporting improved trajectory alignment over prior methods.
-
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
HunyuanVideo-Avatar is an audio-driven video generator that enables emotion-controllable and multi-character animation by injecting character images, routing audio via face masks, and transferring emotion from referen...
-
ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos
ProphetDWM is a one-stage diffusion world model that jointly predicts future driving video and low-level actions from a current frame and a short action sequence.
-
Self Gradient Forcing: Native Long Video Extrapolation
Self Gradient Forcing lets later-video losses train how earlier generated video latents are encoded into the causal KV cache, improving measured long-video consistency over plain Self Forcing.
-
OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation
OmniV2V is one diffusion-transformer model that performs eight video generation and editing tasks by combining mask, pose, image, and text-instruction conditions.
-
Multi-View Face and Gesture Animation with Dynamic Gaussians
Combining separate face and hand models with a parametric body and Gaussian splatting enables multi-view-consistent upper-body avatars that can be re-animated with new expressions and gestures.
-
Advances in 4D Representation: Geometry, Motion, and Interaction
A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.
-
SplitMeanFlow: Interval Splitting Consistency in Few-Step Generative Modeling
A purely algebraic interval-splitting consistency objective trains few-step generative models without JVP computations and recovers MeanFlow's differential identity as a special limit.
-
Leveraging Pre-Trained Visual Models for AI-Generated Video Detection
Pre-trained SigLIP/VideoMAE features with a linear probe or nearest-neighbor distance separate real videos from text-to-video model outputs, reaching about 90% average F1 on the new VID-AID benchmark, but much lower o...
-
Reinforcement Learning: From Algorithms To Foundation Models
A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.
-
A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality
A survey of 32 long-video generation papers, presenting a taxonomy and component recommendations for backbones, text encoders, objectives, and positional encodings.
Discussion (0). Sign in to comment.