Pith. sign in

REVIEW 3 cited by

ElasticTok: Adaptive Tokenization for Image and Video

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08368 v2 pith:AXQWGILU submitted 2024-10-10 cs.LG

classification cs.LG
keywords tokensvideoelastictokmodelsnumberdataefficientencoding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Efficient video tokenization remains a key bottleneck in learning general purpose vision models that are capable of processing long video sequences. Prevailing approaches are restricted to encoding videos to a fixed number of tokens, where too few tokens will result in overly lossy encodings, and too many tokens will result in prohibitively long sequence lengths. In this work, we introduce ElasticTok, a method that conditions on prior frames to adaptively encode a frame into a variable number of tokens. To enable this in a computationally scalable way, we propose a masking technique that drops a random number of tokens at the end of each frames's token encoding. During inference, ElasticTok can dynamically allocate tokens when needed -- more complex data can leverage more tokens, while simpler data only needs a few tokens. Our empirical evaluations on images and video demonstrate the effectiveness of our approach in efficient token usage, paving the way for future development of more powerful multimodal models, world models, and agents.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A point-prompted segmentation model gains efficiency by foveated tokenization, cutting tokens from 4096 to 172 while staying competitive on mIoU benchmarks.

  2. ELT: Elastic Looped Transformers for Visual Generation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Weight-shared looped transformers trained with intra-loop self-distillation match MaskGIT-class FID/FVD at roughly 4x fewer parameters and support any-time inference across loop counts.

  3. HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    HumanGenesis is a claimed state-of-the-art framework that couples 3D Gaussian reconstruction, LLM-based critique, pose guidance, and diffusion-based harmonization for synthetic human video generation; unverified in th...

Pith tools