REVIEW 5 cited by
LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present LARP, a novel video tokenizer designed to overcome limitations in current video tokenization methods for autoregressive (AR) generative models. Unlike traditional patchwise tokenizers that directly encode local visual patches into discrete tokens, LARP introduces a holistic tokenization scheme that gathers information from the visual content using a set of learned holistic queries. This design allows LARP to capture more global and semantic representations, rather than being limited to local patch-level information. Furthermore, it offers flexibility by supporting an arbitrary number of discrete tokens, enabling adaptive and efficient tokenization based on the specific requirements of the task. To align the discrete token space with downstream AR generation tasks, LARP integrates a lightweight AR transformer as a training-time prior model that predicts the next token on its discrete latent space. By incorporating the prior model during training, LARP learns a latent space that is not only optimized for video reconstruction but is also structured in a way that is more conducive to autoregressive generation. Moreover, this process defines a sequential order for the discrete tokens, progressively pushing them toward an optimal configuration during training, ensuring smoother and more accurate AR generation at inference time. Comprehensive experiments demonstrate LARP's strong performance, achieving state-of-the-art FVD on the UCF101 class-conditional video generation benchmark. LARP enhances the compatibility of AR models with videos and opens up the potential to build unified high-fidelity multimodal large language models (MLLMs).
Forward citations
Cited by 5 Pith papers
-
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
Frozen video foundation features can be compressed into reconstruction-capable, generation-friendly latents that improve video generation quality and convergence speed.
-
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
A ViT-based visual tokenizer study shows latent code size drives reconstruction quality, encoder scaling gives little benefit for generation, and ViTok reaches competitive or state-of-the-art results with fewer FLOPs.
-
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
SweetTok compresses 17-frame 256x256 videos into 1,280 tokens via decoupled spatial-temporal query autoencoding and a parts-of-speech language codebook, reporting an rFVD of 20.46 on UCF-101 versus 35.15 for LARP.
-
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
Diffusion-conditioned denoising trains a video tokenizer whose features support both video question answering and text-to-video generation in a single LLM.
-
Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction
CoordTok encodes a 128-frame video into three 2D triplane latents (1280 tokens total) and reconstructs randomly sampled patch coordinates, enabling efficient long-video tokenization and 128-frame generation.
Discussion (0). Continue with ORCID to comment.