Pith. sign in

REVIEW 6 cited by

Scaling Autoregressive Video Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.02634 v3 pith:LEJ2YQDK submitted 2019-06-06 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords videomodelshighcomplexcontinuationsoftenresultsautoregressive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Due to the statistical complexity of video, the high degree of inherent stochasticity, and the sheer amount of data, generating natural video remains a challenging task. State-of-the-art video generation models often attempt to address these issues by combining sometimes complex, usually video-specific neural network architectures, latent variable models, adversarial training and a range of other methods. Despite their often high complexity, these approaches still fall short of generating high quality video continuations outside of narrow domains and often struggle with fidelity. In contrast, we show that conceptually simple autoregressive video generation models based on a three-dimensional self-attention mechanism achieve competitive results across multiple metrics on popular benchmark datasets, for which they produce continuations of high fidelity and realism. We also present results from training our models on Kinetics, a large scale action recognition dataset comprised of YouTube videos exhibiting phenomena such as camera movement, complex object interactions and diverse human movement. While modeling these phenomena consistently remains elusive, we hope that our results, which include occasional realistic continuations encourage further research on comparatively complex, large scale datasets such as Kinetics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation

    cs.CV 2026-02 conditional novelty 7.0 of 10

    Causal Forcing uses an autoregressive teacher for ODE initialization in diffusion distillation to close the causal attention gap and deliver better real-time video generation than Self Forcing.

  2. Energy-Based Transformers are Scalable Learners and Thinkers

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Energy-Based Transformers learn to predict by gradient-descent minimization of a learned energy function, and the paper reports faster pretraining scaling and inference-time thinking gains over Transformer++ and Diffu...

  3. Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new story-completion benchmark, StoryEval, shows that 11 current text-to-video models complete fewer than half of the consecutive events in short story prompts.

  4. VRAG: Learning World Models for Interactive Video Generation

    cs.CV 2025-05 unverdicted novelty 5.0 of 10

    VRAG improves long-horizon interactive video generation by conditioning autoregressive diffusion on retrieved historical frames and explicit global state, outperforming long-context baselines on the tested Minecraft a...

  5. An Empirical Study of Autoregressive Pre-training from Videos

    cs.CV 2025-01 conditional novelty 5.0 of 10

    Autoregressive next-token prediction on video and image tokens yields competitive visual representations across recognition, tracking, and robotics benchmarks, with scaling laws that are slower than those of language models.

  6. UniCP: A Unified Caching and Pruning Framework for Efficient Video Generation

    cs.CV 2025-02 reject novelty 4.0 of 10

    UniCP combines error-aware caching and PCA-based pruning to speed up diffusion-transformer video generation by about 1.6x.

Pith tools