Pith. sign in

REVIEW 2 cited by

Implicit Stacked Autoregressive Model for Video Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.07849 v1 pith:BDUHPONE submitted 2023-03-14 cs.CV

classification cs.CV
keywords autoregressivepredictionmethodsstackedmodelframefutureimplicit
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Future frame prediction has been approached through two primary methods: autoregressive and non-autoregressive. Autoregressive methods rely on the Markov assumption and can achieve high accuracy in the early stages of prediction when errors are not yet accumulated. However, their performance tends to decline as the number of time steps increases. In contrast, non-autoregressive methods can achieve relatively high performance but lack correlation between predictions for each time step. In this paper, we propose an Implicit Stacked Autoregressive Model for Video Prediction (IAM4VP), which is an implicit video prediction model that applies a stacked autoregressive method. Like non-autoregressive methods, stacked autoregressive methods use the same observed frame to estimate all future frames. However, they use their own predictions as input, similar to autoregressive methods. As the number of time steps increases, predictions are sequentially stacked in the queue. To evaluate the effectiveness of IAM4VP, we conducted experiments on three common future frame prediction benchmark datasets and weather\&climate prediction benchmark datasets. The results demonstrate that our proposed model achieves state-of-the-art performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. STLight: a Fully Convolutional Approach for Efficient Predictive Learning by Spatio-Temporal joint Processing

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A fully convolutional spatio-temporal patch mixer, STLight, matches or beats RNN-based baselines on predictive learning benchmarks at a fraction of the parameter and FLOP budget.

  2. Everything is a Video: Unifying Modalities through Next-Frame Prediction

    cs.CV 2024-11 conditional novelty 5.0 of 10

    The paper reformulates text, image, video, and audio tasks as next-frame video prediction by rendering everything into 64x64 frames, and shows a 41M-parameter transformer can solve them without pretrained encoders.

Pith tools