Pith. sign in

REVIEW 6 cited by

A Good Image Generator Is What You Need for High-Resolution Video Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.15069 v1 pith:KYFWRRQ5 submitted 2021-04-30 cs.CV

classification cs.CV
keywords videoimagecontentmotionsynthesisframeworkgeneratorhigh-resolution
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image and video synthesis are closely related areas aiming at generating content from noise. While rapid progress has been demonstrated in improving image-based models to handle large resolutions, high-quality renderings, and wide variations in image content, achieving comparable video generation results remains problematic. We present a framework that leverages contemporary image generators to render high-resolution videos. We frame the video synthesis problem as discovering a trajectory in the latent space of a pre-trained and fixed image generator. Not only does such a framework render high-resolution videos, but it also is an order of magnitude more computationally efficient. We introduce a motion generator that discovers the desired trajectory, in which content and motion are disentangled. With such a representation, our framework allows for a broad range of applications, including content and motion manipulation. Furthermore, we introduce a new task, which we call cross-domain video synthesis, in which the image and motion generators are trained on disjoint datasets belonging to different domains. This allows for generating moving objects for which the desired video data is not available. Extensive experiments on various datasets demonstrate the advantages of our methods over existing video generation techniques. Code will be released at https://github.com/snap-research/MoCoGAN-HD.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Interspatial Attention for Efficient 4D Human Video Generation

    cs.CV 2025-05 conditional novelty 7.0 of 10

    A new symmetric 3D-to-2D attention mechanism with relative positional encodings, plus a motion-tuned video VAE, improves controllable 4D human video generation.

  2. Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A temporal attention purification and skip-connection rerouting method that separates motion learning from appearance learning in text-to-video diffusion model customization.

  3. Taming Teacher Forcing for Masked Autoregressive Video Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Complete Teacher Forcing, conditioning masked frames on complete previous frames instead of masked ones, substantially improves frame-level autoregressive video generation quality and temporal coherence.

  4. DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A 1B-parameter autoregressive world model generates over 40 seconds of controllable driving video with near-SOTA FVD, but key claims rest on inconsistent and cross-paper comparisons.

  5. CFSynthesis: Controllable and Free-view 3D Human Video Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    CFSynthesis generates free-view human videos from one reference image by conditioning a diffusion model on a textured SMPL body model and separately encoded foreground and background.

  6. FlipSketch: Flipping Static Drawings to Text-Guided Sketch Animations

    cs.GR 2024-11 conditional novelty 6.0 of 10

    A text-guided system that animates a single raster sketch into a dynamic video by fine-tuning a text-to-video diffusion model and steering its attention maps with the input drawing.

Pith tools