Pith. sign in

REVIEW 13 cited by

Image and Video Tokenization with Binary Spherical Quantization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.07548 v1 pith:5CB5BO4F submitted 2024-06-11 cs.CV cs.ITcs.LGeess.IVmath.IT

classification cs.CVcs.ITcs.LGeess.IVmath.IT
keywords videoimagebinarybsq-vitquantizationvisualachievescompression
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We propose a new transformer-based image and video tokenizer with Binary Spherical Quantization (BSQ). BSQ projects the high-dimensional visual embedding to a lower-dimensional hypersphere and then applies binary quantization. BSQ is (1) parameter-efficient without an explicit codebook, (2) scalable to arbitrary token dimensions, and (3) compact: compressing visual data by up to 100$\times$ with minimal distortion. Our tokenizer uses a transformer encoder and decoder with simple block-wise causal masking to support variable-length videos as input. The resulting BSQ-ViT achieves state-of-the-art visual reconstruction quality on image and video reconstruction benchmarks with 2.4$\times$ throughput compared to the best prior methods. Furthermore, by learning an autoregressive prior for adaptive arithmetic coding, BSQ-ViT achieves comparable results on video compression with state-of-the-art video compression standards. BSQ-ViT also enables masked language models to achieve competitive image synthesis quality to GAN- and diffusion-based methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units

    cs.RO 2026-08 conditional novelty 7.0 of 10

    DigitCode tokenizes hand motion by anatomical units, showing the token span (bone/finger/hand) matters more than the quantizer family, and reduces symbolic reconstruction error by about three quarters.

  2. ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    ScaleMoGen introduces a scale-wise autoregressive framework that quantizes motions into hierarchical discrete tokens and predicts next-scale maps to achieve SOTA FID 0.030 on HumanML3D and text-guided editing.

  3. Generative Refinement Networks for Visual Synthesis

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Hierarchical Binary Quantization plus global refinement AR yields 0.56 rFID reconstruction and 1.81 gFID class-conditional generation on ImageNet, with competitive T2I/T2V at 2B scale.

  4. ELT: Elastic Looped Transformers for Visual Generation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Weight-shared looped transformers trained with intra-loop self-distillation match MaskGIT-class FID/FVD at roughly 4x fewer parameters and support any-time inference across loop counts.

  5. WindFM: An Open-Source Foundation Model for Zero-Shot Wind Power Forecasting

    cs.LG 2025-09 conditional novelty 6.0 of 10

    A wind-specific foundation model, WindFM, uses hierarchical tokenization and autoregressive pre-training on the NREL WIND Toolkit to achieve state-of-the-art zero-shot wind power forecasts.

  6. Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new interactive segmentation decoder that routes computation to boundary regions, using binary quantization attention and mixture-of-experts, achieves state-of-the-art accuracy with CPU-friendly latency.

  7. MambaVideo for Discrete Video Tokenization with Channel-Split Quantization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A Mamba-based hierarchical video tokenizer with channel-split quantization achieves state-of-the-art reconstruction and generation scores while preserving token count.

  8. multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data

    cs.LG 2025-05 conditional novelty 6.0 of 10

    multivariateGPT extends next-token prediction to jointly predict the class and continuous value of mixed categorical and numeric time series, with Gaussian uncertainty, and outperforms discrete-token baselines on clin...

  9. TokBench: Evaluating Your Visual Tokenizer before Visual Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    TokBench measures text recognition accuracy and face similarity on reconstructed images and videos across 16 tokenizers, showing small text and faces are poorly preserved and often missed by traditional metrics.

  10. UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    UniGen shows a 1.5B model trained on open data can beat larger systems on image understanding and generation once it verifies its own outputs with chain-of-thought and Best-of-N selection.

  11. QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    QLIP trains a quantized image autoencoder with both reconstruction and text-alignment losses, yielding a tokenizer that supports multimodal understanding and text-to-image generation in one model.

  12. Wireless TokenCom: RL-Based Tokenizer Agreement for Multi-User Wireless Token Communications

    cs.LG 2026-02 conditional novelty 5.0 of 10

    Joint tokenizer/codebook selection, subchannel assignment, and beamforming for multi-user video TokenCom is posed as an MDP and solved by DQN for discrete choices and DDPG for beamforming, with simulated gains over H.265.

  13. LGQ: Learnable Geometric Quantization for Image Tokenization

    cs.CV 2026-02 reject novelty 4.0 of 10

    LGQ reports better ImageNet reconstruction FID than FSQ/SimVQ using soft-to-hard learnable-codebook quantization, but its abstract's generation and utilization claims are contradicted by the body.

Pith tools