Pith. sign in

REVIEW 1 cited by

Learning Semantic-Aware Dynamics for Video Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.09762 v1 pith:P4Q37ECH submitted 2021-04-20 cs.CV

classification cs.CV
keywords videomotionpredictedregionssceneexplicitlyframeslayout
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose an architecture and training scheme to predict video frames by explicitly modeling dis-occlusions and capturing the evolution of semantically consistent regions in the video. The scene layout (semantic map) and motion (optical flow) are decomposed into layers, which are predicted and fused with their context to generate future layouts and motions. The appearance of the scene is warped from past frames using the predicted motion in co-visible regions; dis-occluded regions are synthesized with content-aware inpainting utilizing the predicted scene layout. The result is a predictive model that explicitly represents objects and learns their class-specific motion, which we evaluate on video prediction benchmarks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Benefits of Instance Decomposition in Video Prediction Models

    cs.CV 2025-01 reject novelty 4.0 of 10

    Explicit instance decomposition with per-class shared weights improves latent-transformer video prediction in the paper's experiments, but the claimed advantage is weakened by mismatched parameter counts and test-set ...

Pith tools