Pith. sign in

REVIEW 4 cited by

Diverse Video Generation using a Gaussian Process Trigger

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.04619 v1 pith:GCYBU4WS submitted 2021-07-09 cs.CV cs.AIcs.LGcs.RO

classification cs.CVcs.AIcs.LGcs.RO
keywords futurediversegenerationgivenstatesvideodistributiondiversity
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Generating future frames given a few context (or past) frames is a challenging task. It requires modeling the temporal coherence of videos and multi-modality in terms of diversity in the potential future states. Current variational approaches for video generation tend to marginalize over multi-modal future outcomes. Instead, we propose to explicitly model the multi-modality in the future outcomes and leverage it to sample diverse futures. Our approach, Diverse Video Generator, uses a Gaussian Process (GP) to learn priors on future states given the past and maintains a probability distribution over possible futures given a particular sample. In addition, we leverage the changes in this distribution over time to control the sampling of diverse future states by estimating the end of ongoing sequences. That is, we use the variance of GP over the output function space to trigger a change in an action sequence. We achieve state-of-the-art results on diverse future frame generation in terms of reconstruction quality and diversity of the generated sequences.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Video Decomposition Prior: A Methodology to Decompose Videos into Layers

    cs.CV 2024-12 conditional novelty 5.0 of 10

    VDP decomposes a single test video into layers and opacity maps via two U-Nets, achieving strong unsupervised video object segmentation, dehazing, and relighting, though the relighting model reduces to gamma correction.

  2. Obstacle-aware Gaussian Process Regression

    cs.LG 2024-12 reject novelty 4.0 of 10

    GP-ND adds a log-KL divergence penalty between a GP's predictive distribution and Gaussian blobs placed on negative data pairs, aiming to fit positive points while avoiding obstacles.

  3. Efficient Continuous Video Flow Model for Video Prediction

    cs.CV 2024-12 conditional novelty 4.0 of 10

    The paper adapts the authors' prior continuous-video-process framework to latent space, reporting state-of-the-art FVD on KTH, BAIR, Human3.6M, and UCF101 with fewer parameters and sampling steps.

  4. Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction

    cs.CV 2024-12 conditional novelty 4.0 of 10

    CVP trains a network to reverse a continuous interpolation between past and future frames, reporting competitive FVD scores and 25-step sampling on KTH, BAIR, Human3.6M, and UCF101.

Pith tools