Pith. sign in

REVIEW 9 cited by

Generating Videos with Dynamics-aware Implicit Generative Adversarial Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.10571 v1 pith:K6G54EWZ submitted 2022-02-21 cs.CV cs.LG

classification cs.CVcs.LG
keywords videovideosadversarialdigangenerationgenerativeimplicitlong
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the deep learning era, long video generation of high-quality still remains challenging due to the spatio-temporal complexity and continuity of videos. Existing prior works have attempted to model video distribution by representing videos as 3D grids of RGB values, which impedes the scale of generated videos and neglects continuous dynamics. In this paper, we found that the recent emerging paradigm of implicit neural representations (INRs) that encodes a continuous signal into a parameterized neural network effectively mitigates the issue. By utilizing INRs of video, we propose dynamics-aware implicit generative adversarial network (DIGAN), a novel generative adversarial network for video generation. Specifically, we introduce (a) an INR-based video generator that improves the motion dynamics by manipulating the space and time coordinates differently and (b) a motion discriminator that efficiently identifies the unnatural motions without observing the entire long frame sequences. We demonstrate the superiority of DIGAN under various datasets, along with multiple intriguing properties, e.g., long video synthesis, video extrapolation, and non-autoregressive video generation. For example, DIGAN improves the previous state-of-the-art FVD score on UCF-101 by 30.7% and can be trained on 128 frame videos of 128x128 resolution, 80 frames longer than the 48 frames of the previous state-of-the-art method.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Every Image Listens, Every Image Dances: Music-Driven Image Animation

    cs.CV 2025-01 conditional novelty 7.0 of 10

    MuseDance animates a reference image into a music-synchronized dance video conditioned only on the audio track and a text description, and contributes a new 2,904-video dataset.

  2. ELT: Elastic Looped Transformers for Visual Generation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Weight-shared looped transformers trained with intra-loop self-distillation match MaskGIT-class FID/FVD at roughly 4x fewer parameters and support any-time inference across loop counts.

  3. Taming Teacher Forcing for Masked Autoregressive Video Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Complete Teacher Forcing, conditioning masked frames on complete previous frames instead of masked ones, substantially improves frame-level autoregressive video generation quality and temporal coherence.

  4. Ingredients: Blending Custom Photos with Video Diffusion Transformers

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Ingredients adds a mask-supervised identity router to a video diffusion transformer, enabling multi-person videos from a few reference photos without per-identity fine-tuning.

  5. DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A 1B-parameter autoregressive world model generates over 40 seconds of controllable driving video with near-SOTA FVD, but key claims rest on inconsistent and cross-paper comparisons.

  6. Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling

    cs.CV 2024-11 conditional novelty 6.0 of 10

    STG boosts video diffusion sample quality by guiding away from a self-produced weak model obtained by skipping spatiotemporal layers, with no extra training.

  7. Towards Precise Scaling Laws for Video Diffusion Transformers

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Video diffusion transformers follow scaling laws, and tuning learning rate and batch size per model and data size makes those laws precise enough to predict loss and optimal model size at larger scales.

  8. HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment

    cs.CV 2025-02 reject novelty 4.0 of 10

    HuViDPO claims the first DPO-based alignment for text-to-video generation, but its loss reduces to the known DPO-SDXL objective and the evaluation is not reproducible.

  9. Video Diffusion Transformers are In-Context Learners

    cs.CV 2024-12 conditional novelty 4.0 of 10

    Concatenating multiple video clips into one input and fine-tuning a LoRA adapter lets a pretrained video diffusion transformer produce consistent multi-scene videos from a single prompt.

Pith tools