REVIEW 9 cited by
Generating Videos with Dynamics-aware Implicit Generative Adversarial Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In the deep learning era, long video generation of high-quality still remains challenging due to the spatio-temporal complexity and continuity of videos. Existing prior works have attempted to model video distribution by representing videos as 3D grids of RGB values, which impedes the scale of generated videos and neglects continuous dynamics. In this paper, we found that the recent emerging paradigm of implicit neural representations (INRs) that encodes a continuous signal into a parameterized neural network effectively mitigates the issue. By utilizing INRs of video, we propose dynamics-aware implicit generative adversarial network (DIGAN), a novel generative adversarial network for video generation. Specifically, we introduce (a) an INR-based video generator that improves the motion dynamics by manipulating the space and time coordinates differently and (b) a motion discriminator that efficiently identifies the unnatural motions without observing the entire long frame sequences. We demonstrate the superiority of DIGAN under various datasets, along with multiple intriguing properties, e.g., long video synthesis, video extrapolation, and non-autoregressive video generation. For example, DIGAN improves the previous state-of-the-art FVD score on UCF-101 by 30.7% and can be trained on 128 frame videos of 128x128 resolution, 80 frames longer than the 48 frames of the previous state-of-the-art method.
Forward citations
Cited by 9 Pith papers
-
Every Image Listens, Every Image Dances: Music-Driven Image Animation
MuseDance animates a reference image into a music-synchronized dance video conditioned only on the audio track and a text description, and contributes a new 2,904-video dataset.
-
ELT: Elastic Looped Transformers for Visual Generation
Weight-shared looped transformers trained with intra-loop self-distillation match MaskGIT-class FID/FVD at roughly 4x fewer parameters and support any-time inference across loop counts.
-
Taming Teacher Forcing for Masked Autoregressive Video Generation
Complete Teacher Forcing, conditioning masked frames on complete previous frames instead of masked ones, substantially improves frame-level autoregressive video generation quality and temporal coherence.
-
Ingredients: Blending Custom Photos with Video Diffusion Transformers
Ingredients adds a mask-supervised identity router to a video diffusion transformer, enabling multi-person videos from a few reference photos without per-identity fine-tuning.
-
DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
A 1B-parameter autoregressive world model generates over 40 seconds of controllable driving video with near-SOTA FVD, but key claims rest on inconsistent and cross-paper comparisons.
-
Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling
STG boosts video diffusion sample quality by guiding away from a self-produced weak model obtained by skipping spatiotemporal layers, with no extra training.
-
Towards Precise Scaling Laws for Video Diffusion Transformers
Video diffusion transformers follow scaling laws, and tuning learning rate and batch size per model and data size makes those laws precise enough to predict loss and optimal model size at larger scales.
-
HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment
HuViDPO claims the first DPO-based alignment for text-to-video generation, but its loss reduces to the known DPO-SDXL objective and the evaluation is not reproducible.
-
Video Diffusion Transformers are In-Context Learners
Concatenating multiple video clips into one input and fine-tuning a LoRA adapter lets a pretrained video diffusion transformer produce consistent multi-scene videos from a single prompt.
Discussion (0). Continue with ORCID to comment.