Pith. sign in

REVIEW 9 cited by

OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.09399 v1 pith:N34MAQ6U submitted 2024-06-13 cs.CV

classification cs.CV
keywords omnitokenizerimagevideodatavisualreconstructiontokenizerfirst
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Tokenizer, serving as a translator to map the intricate visual data into a compact latent space, lies at the core of visual generative models. Based on the finding that existing tokenizers are tailored to image or video inputs, this paper presents OmniTokenizer, a transformer-based tokenizer for joint image and video tokenization. OmniTokenizer is designed with a spatial-temporal decoupled architecture, which integrates window and causal attention for spatial and temporal modeling. To exploit the complementary nature of image and video data, we further propose a progressive training strategy, where OmniTokenizer is first trained on image data on a fixed resolution to develop the spatial encoding capacity and then jointly trained on image and video data on multiple resolutions to learn the temporal dynamics. OmniTokenizer, for the first time, handles both image and video inputs within a unified framework and proves the possibility of realizing their synergy. Extensive experiments demonstrate that OmniTokenizer achieves state-of-the-art (SOTA) reconstruction performance on various image and video datasets, e.g., 1.11 reconstruction FID on ImageNet and 42 reconstruction FVD on UCF-101, beating the previous SOTA methods by 13% and 26%, respectively. Additionally, we also show that when integrated with OmniTokenizer, both language model-based approaches and diffusion models can realize advanced visual synthesis performance, underscoring the superiority and versatility of our method. Code is available at https://github.com/FoundationVision/OmniTokenizer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MambaVideo for Discrete Video Tokenization with Channel-Split Quantization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A Mamba-based hierarchical video tokenizer with channel-split quantization achieves state-of-the-art reconstruction and generation scores while preserving token count.

  2. Taming Teacher Forcing for Masked Autoregressive Video Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Complete Teacher Forcing, conditioning masked frames on complete previous frames instead of masked ones, substantially improves frame-level autoregressive video generation quality and temporal coherence.

  3. Parallelized Autoregressive Visual Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Grouping spatially distant visual tokens into parallel prediction steps reduces autoregressive generation steps by 3.9x to 11.3x with modest FID/FVD loss.

  4. Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A bitwise next-scale autoregressive model with a 2^32-vocabulary classifier and bitwise self-correction reaches top text-to-image benchmark scores at 2B parameters.

  5. Scaling Image Tokenizers with Grouped Spherical Quantization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GSQ combines spherical codebook initialization, normalized lookup, and group-wise latent decomposition to achieve strong reconstruction quality at 16x spatial downsampling in far fewer training steps than prior tokenizers.

  6. Factorized Visual Tokenization and Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A factorized quantizer with disentanglement and semantic supervision achieves state-of-the-art reconstruction FID, 0.24 at 8x downsample on ImageNet, and improves autoregressive image generation compared to VQ baselines.

  7. One-D-Piece: Image Tokenizer Meets Quality-Controllable Compression

    cs.CV 2025-01 conditional novelty 5.0 of 10

    One-D-Piece trains a 1D discrete image tokenizer with randomized tail truncation so that reconstruction quality can be controlled by choosing the number of tokens.

  8. VidTok: A Versatile and Open-Source Video Tokenizer

    cs.CV 2024-12 conditional novelty 5.0 of 10

    VidTok reports state-of-the-art video reconstruction accuracy among open and published video tokenizers, using FSQ for discrete tokens, 2D+1D convolutions, and a two-stage low-resolution-to-high-resolution training recipe.

  9. LaVin-DiT: Large Vision Diffusion Transformer

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A single diffusion transformer with a spatial-temporal VAE and in-context example pairs unifies over 20 vision tasks, with strong scores on several image benchmarks and uneven evidence on video tasks.

Pith tools