Pith. sign in

REVIEW 1 cited by

Unsupervised Video Prediction from a Single Frame by Estimating 3D Dynamic Scene Structure

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.09051 v1 pith:SR4OH4PK submitted 2021-06-16 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords framesegmentationstructureunsupervisedcameraframesfuturegenerate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Our goal in this work is to generate realistic videos given just one initial frame as input. Existing unsupervised approaches to this task do not consider the fact that a video typically shows a 3D environment, and that this should remain coherent from frame to frame even as the camera and objects move. We address this by developing a model that first estimates the latent 3D structure of the scene, including the segmentation of any moving objects. It then predicts future frames by simulating the object and camera dynamics, and rendering the resulting views. Importantly, it is trained end-to-end using only the unsupervised objective of predicting future frames, without any 3D information nor segmentation annotations. Experiments on two challenging datasets of natural videos show that our model can estimate 3D structure and motion segmentation from a single frame, and hence generate plausible and varied predictions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Benefits of Instance Decomposition in Video Prediction Models

    cs.CV 2025-01 reject novelty 4.0 of 10

    Explicit instance decomposition with per-class shared weights improves latent-transformer video prediction in the paper's experiments, but the claimed advantage is weakened by mismatched parameter counts and test-set ...

Pith tools