Pith. sign in

REVIEW 2 cited by

Transformation-Based Models of Video Sequences

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1701.08435 v3 pith:PTX5EIIX submitted 2017-01-29 cs.LG cs.CV

classification cs.LGcs.CV
keywords frameframesmodelspredictionsequencesvideoapproachgiven
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work we propose a simple unsupervised approach for next frame prediction in video. Instead of directly predicting the pixels in a frame given past frames, we predict the transformations needed for generating the next frame in a sequence, given the transformations of the past frames. This leads to sharper results, while using a smaller prediction model. In order to enable a fair comparison between different video frame prediction models, we also propose a new evaluation protocol. We use generated frames as input to a classifier trained with ground truth sequences. This criterion guarantees that models scoring high are those producing sequences which preserve discriminative features, as opposed to merely penalizing any deviation, plausible or not, from the ground truth. Our proposed approach compares favourably against more sophisticated ones on the UCF-101 data set, while also being more efficient in terms of the number of parameters and computational cost.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FIction: 4D Future Interaction Prediction from Video

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FICTION predicts future 3D interaction locations and body poses up to three minutes ahead from egocentric video and a 3D scene map, and claims substantial gains over prior methods on a new Ego-Exo4D benchmark.

  2. Everything is a Video: Unifying Modalities through Next-Frame Prediction

    cs.CV 2024-11 conditional novelty 5.0 of 10

    The paper reformulates text, image, video, and audio tasks as next-frame video prediction by rendering everything into 64x64 frames, and shows a 41M-parameter transformer can solve them without pretrained encoders.

Pith tools