Pith. sign in

REVIEW 18 cited by

LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.15103 v2 pith:COBW72NZ submitted 2023-09-26 cs.CV

classification cs.CV
keywords videomodellaviegenerationhigh-qualitymodelspre-trainedtemporal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work aims to learn a high-quality text-to-video (T2V) generative model by leveraging a pre-trained text-to-image (T2I) model as a basis. It is a highly desirable yet challenging task to simultaneously a) accomplish the synthesis of visually realistic and temporally coherent videos while b) preserving the strong creative generation nature of the pre-trained T2I model. To this end, we propose LaVie, an integrated video generation framework that operates on cascaded video latent diffusion models, comprising a base T2V model, a temporal interpolation model, and a video super-resolution model. Our key insights are two-fold: 1) We reveal that the incorporation of simple temporal self-attentions, coupled with rotary positional encoding, adequately captures the temporal correlations inherent in video data. 2) Additionally, we validate that the process of joint image-video fine-tuning plays a pivotal role in producing high-quality and creative outcomes. To enhance the performance of LaVie, we contribute a comprehensive and diverse video dataset named Vimeo25M, consisting of 25 million text-video pairs that prioritize quality, diversity, and aesthetic appeal. Extensive experiments demonstrate that LaVie achieves state-of-the-art performance both quantitatively and qualitatively. Furthermore, we showcase the versatility of pre-trained LaVie models in various long video generation and personalized video synthesis applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Adding persistently updated, supervised world-state register tokens to streaming multi-agent diffusion improves cross-agent consistency and visual quality in two-agent Minecraft generation.

  2. TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion Model

    eess.IV 2026-07 conditional novelty 6.0 of 10

    TCAM-Diff fits 3D medical volumes into triplane features via a decoder-only autoencoder and generates new volumes with a cross-attention diffusion model, reporting better reconstruction and generation scores than VAE/...

  3. CineScale: Free Lunch in High-Resolution Cinematic Visual Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    CineScale extends pre-trained diffusion models to 8k image and 4k video generation with mostly tuning-free inference plus a small LoRA adaptation for video.

  4. Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    EVS combines a text-to-image and a text-to-video diffusion model in a single denoising pass, improving frame quality and temporal consistency without retraining.

  5. "PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    PhyWorldBench evaluates 12 text-to-video models on 1,050 physics prompts; the best model passes both semantic adherence and physical commonsense checks in only 26.2% of videos.

  6. FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FreeLong++ extends short-video diffusion models to 4x to 8x longer clips, without retraining, by fusing multiple windowed attention branches through frequency-domain filters and a spectral noise initialization.

  7. VMoBA: Mixture-of-Block Attention for Video Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    VMoBA is a sparse attention mechanism for video diffusion models that combines cyclic 1D-2D-3D block partitioning with global and threshold-based block selection to reduce training FLOPs while keeping generation quality.

  8. GenWorld: Towards Detecting AI-generated Real-world Simulation Videos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GenWorld is a 100k real-world-simulation video forgery benchmark, and SpannDetector uses multi-view 3D consistency to detect AI-generated videos, especially world-model outputs that fool existing detectors.

  9. FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FlowMo reduces temporal artifacts in video generation by guiding the denoising process to lower the maximum patch-wise variance of consecutive-frame differences in the latent space.

  10. Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A three-stage pipeline combining sparse 'tile' attention with multi-step consistency distillation makes Open-Sora-Plan video generation up to 7.8x faster while keeping the aggregate VBench final score within 1%.

  11. LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion

    cs.CV 2025-07 conditional novelty 5.0 of 10

    LiON-LoRA adds a learned scaling token to video-diffusion LoRA adapters, enabling linear and independent control of camera trajectory and object motion strength.

  12. VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A new benchmark, VF-Eval, measures how well multimodal LLMs check, detect, and reason about errors in AI-generated videos, and shows frontier models remain far below human performance.

  13. RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Metric-scale depth alignment plus scene-constrained noise shaping improves camera controllability and video quality for image-to-video generation on RealEstate10K.

  14. Light-A-Video: Training-free Video Relighting via Progressive Light Fusion

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Light-A-Video relights videos without training by injecting image relight results into a video diffusion model's denoising loop with cross-frame attention and progressive blending.

  15. Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation

    cs.CL 2026-07 reject novelty 4.0 of 10

    Text2Sign generates smooth 64×64 signer clips from text on one GPU, but its denoising audit shows the clips hardly depend on prompt identity.

  16. LIA-X: Interpretable Latent Portrait Animator

    cs.CV 2025-08 conditional novelty 4.0 of 10

    Adding an L1 sparsity penalty to the motion dictionary of the LIA portrait animator produces disentangled, human-interpretable motion vectors that support controllable image and video editing and scale to roughly one ...

  17. DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion

    cs.CV 2025-06 conditional novelty 4.0 of 10

    DiffuseSlide boosts the frame rate of latent diffusion videos via latent interpolation, noise re-injection, and sliding-window denoising, reporting better FVD, PSNR, and SSIM than several baselines on WebVid-10M.

  18. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

Pith tools