Pith. sign in

REVIEW 10 cited by

VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.09047 v1 pith:34XB6TJB submitted 2024-01-17 cs.CV

classification cs.CV
keywords high-qualitymodelsvideomodulesvideoslow-qualityspatialtemporal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-video generation aims to produce a video based on a given prompt. Recently, several commercial video models have been able to generate plausible videos with minimal noise, excellent details, and high aesthetic scores. However, these models rely on large-scale, well-filtered, high-quality videos that are not accessible to the community. Many existing research works, which train models using the low-quality WebVid-10M dataset, struggle to generate high-quality videos because the models are optimized to fit WebVid-10M. In this work, we explore the training scheme of video models extended from Stable Diffusion and investigate the feasibility of leveraging low-quality videos and synthesized high-quality images to obtain a high-quality video model. We first analyze the connection between the spatial and temporal modules of video models and the distribution shift to low-quality videos. We observe that full training of all modules results in a stronger coupling between spatial and temporal modules than only training temporal modules. Based on this stronger coupling, we shift the distribution to higher quality without motion degradation by finetuning spatial modules with high-quality images, resulting in a generic high-quality video model. Evaluations are conducted to demonstrate the superiority of the proposed method, particularly in picture quality, motion, and concept composition.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Using a synthetic alphabet video testbed, balanced data mixing and caption precision are shown to dominate T2V model quality, while CFG and fine-tuning only partially compensate for corrupted captions.

  2. LayerFlow: A Unified Model for Layer-aware Video Generation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    LayerFlow is a unified diffusion-transformer model that generates transparent foreground, background, and blended video layers from per-layer prompts, and supports decomposition and conditioned generation in one framework.

  3. VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An MLLM extracts transferable physical cues from a reference video and conditions a pretrained I2V generator so new scenes follow that physics without exhaustive prompts.

  4. Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Single-image novel view synthesis is decomposed into panorama outpainting plus keyframe-conditioned video diffusion, producing loop-consistent scene tours.

  5. MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MedVideoCap-55K, a 55,803-clip caption-rich medical video dataset, enables MedGen, a LoRA fine-tune of HunyuanVideo that reports top open-source scores and near-commercial quality on medical video benchmarks.

  6. VMoBA: Mixture-of-Block Attention for Video Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    VMoBA is a sparse attention mechanism for video diffusion models that combines cyclic 1D-2D-3D block partitioning with global and threshold-based block selection to reduce training FLOPs while keeping generation quality.

  7. BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    The paper introduces a 3,254-video benchmark with pixel-level artifact masks for AI-generated video, and reports that fine-tuning on it improves artifact localization.

  8. GenWorld: Towards Detecting AI-generated Real-world Simulation Videos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GenWorld is a 100k real-world-simulation video forgery benchmark, and SpannDetector uses multi-view 3D consistency to detect AI-generated videos, especially world-model outputs that fool existing detectors.

  9. Retrieval-Driven Training-Free AI-Generated Video Attribution

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A training-free retrieval pipeline using adaptive color transforms, multi-scale quantized residuals, and temporal aggregation attributes AI-generated videos to one of eight generators with 84.6% Rank-1 and 78.3% mAP o...

  10. RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images

    cs.CV 2026-07 conditional novelty 4.0 of 10

    Reversing bit-plane weights ('bit-reversed image') plus a gradient-selected 32×32 patch lets a small ResNet detect AI-generated images with state-of-the-art accuracy on many benchmarks.

Pith tools