REVIEW 2 cited by
Improved Conditional VRNNs for Video Prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Predicting future frames for a video sequence is a challenging generative modeling task. Promising approaches include probabilistic latent variable models such as the Variational Auto-Encoder. While VAEs can handle uncertainty and model multiple possible future outcomes, they have a tendency to produce blurry predictions. In this work we argue that this is a sign of underfitting. To address this issue, we propose to increase the expressiveness of the latent distributions and to use higher capacity likelihood models. Our approach relies on a hierarchy of latent variables, which defines a family of flexible prior and posterior distributions in order to better model the probability of future sequences. We validate our proposal through a series of ablation experiments and compare our approach to current state-of-the-art latent variable models. Our method performs favorably under several metrics in three different datasets.
Forward citations
Cited by 2 Pith papers
-
Efficient Continuous Video Flow Model for Video Prediction
The paper adapts the authors' prior continuous-video-process framework to latent space, reporting state-of-the-art FVD on KTH, BAIR, Human3.6M, and UCF101 with fewer parameters and sampling steps.
-
Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
CVP trains a network to reverse a continuous interpolation between past and future frames, reporting competitive FVD scores and 25-step sampling on KTH, BAIR, Human3.6M, and UCF101.
Discussion (0). Continue with ORCID to comment.