REVIEW 9 cited by
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Tokenizer, serving as a translator to map the intricate visual data into a compact latent space, lies at the core of visual generative models. Based on the finding that existing tokenizers are tailored to image or video inputs, this paper presents OmniTokenizer, a transformer-based tokenizer for joint image and video tokenization. OmniTokenizer is designed with a spatial-temporal decoupled architecture, which integrates window and causal attention for spatial and temporal modeling. To exploit the complementary nature of image and video data, we further propose a progressive training strategy, where OmniTokenizer is first trained on image data on a fixed resolution to develop the spatial encoding capacity and then jointly trained on image and video data on multiple resolutions to learn the temporal dynamics. OmniTokenizer, for the first time, handles both image and video inputs within a unified framework and proves the possibility of realizing their synergy. Extensive experiments demonstrate that OmniTokenizer achieves state-of-the-art (SOTA) reconstruction performance on various image and video datasets, e.g., 1.11 reconstruction FID on ImageNet and 42 reconstruction FVD on UCF-101, beating the previous SOTA methods by 13% and 26%, respectively. Additionally, we also show that when integrated with OmniTokenizer, both language model-based approaches and diffusion models can realize advanced visual synthesis performance, underscoring the superiority and versatility of our method. Code is available at https://github.com/FoundationVision/OmniTokenizer.
Forward citations
Cited by 9 Pith papers
-
MambaVideo for Discrete Video Tokenization with Channel-Split Quantization
A Mamba-based hierarchical video tokenizer with channel-split quantization achieves state-of-the-art reconstruction and generation scores while preserving token count.
-
Taming Teacher Forcing for Masked Autoregressive Video Generation
Complete Teacher Forcing, conditioning masked frames on complete previous frames instead of masked ones, substantially improves frame-level autoregressive video generation quality and temporal coherence.
-
Parallelized Autoregressive Visual Generation
Grouping spatially distant visual tokens into parallel prediction steps reduces autoregressive generation steps by 3.9x to 11.3x with modest FID/FVD loss.
-
Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
A bitwise next-scale autoregressive model with a 2^32-vocabulary classifier and bitwise self-correction reaches top text-to-image benchmark scores at 2B parameters.
-
Scaling Image Tokenizers with Grouped Spherical Quantization
GSQ combines spherical codebook initialization, normalized lookup, and group-wise latent decomposition to achieve strong reconstruction quality at 16x spatial downsampling in far fewer training steps than prior tokenizers.
-
Factorized Visual Tokenization and Generation
A factorized quantizer with disentanglement and semantic supervision achieves state-of-the-art reconstruction FID, 0.24 at 8x downsample on ImageNet, and improves autoregressive image generation compared to VQ baselines.
-
One-D-Piece: Image Tokenizer Meets Quality-Controllable Compression
One-D-Piece trains a 1D discrete image tokenizer with randomized tail truncation so that reconstruction quality can be controlled by choosing the number of tokens.
-
VidTok: A Versatile and Open-Source Video Tokenizer
VidTok reports state-of-the-art video reconstruction accuracy among open and published video tokenizers, using FSQ for discrete tokens, 2D+1D convolutions, and a two-stage low-resolution-to-high-resolution training recipe.
-
LaVin-DiT: Large Vision Diffusion Transformer
A single diffusion transformer with a spatial-temporal VAE and in-context example pairs unifies over 20 vision tasks, with strong scores on several image benchmarks and uneven evidence on video tasks.
Discussion (0). Continue with ORCID to comment.