REVIEW 13 cited by
Image and Video Tokenization with Binary Spherical Quantization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We propose a new transformer-based image and video tokenizer with Binary Spherical Quantization (BSQ). BSQ projects the high-dimensional visual embedding to a lower-dimensional hypersphere and then applies binary quantization. BSQ is (1) parameter-efficient without an explicit codebook, (2) scalable to arbitrary token dimensions, and (3) compact: compressing visual data by up to 100$\times$ with minimal distortion. Our tokenizer uses a transformer encoder and decoder with simple block-wise causal masking to support variable-length videos as input. The resulting BSQ-ViT achieves state-of-the-art visual reconstruction quality on image and video reconstruction benchmarks with 2.4$\times$ throughput compared to the best prior methods. Furthermore, by learning an autoregressive prior for adaptive arithmetic coding, BSQ-ViT achieves comparable results on video compression with state-of-the-art video compression standards. BSQ-ViT also enables masked language models to achieve competitive image synthesis quality to GAN- and diffusion-based methods.
Forward citations
Cited by 13 Pith papers
-
DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units
DigitCode tokenizes hand motion by anatomical units, showing the token span (bone/finger/hand) matters more than the quantizer family, and reduces symbolic reconstruction error by about three quarters.
-
ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation
ScaleMoGen introduces a scale-wise autoregressive framework that quantizes motions into hierarchical discrete tokens and predicts next-scale maps to achieve SOTA FID 0.030 on HumanML3D and text-guided editing.
-
Generative Refinement Networks for Visual Synthesis
Hierarchical Binary Quantization plus global refinement AR yields 0.56 rFID reconstruction and 1.81 gFID class-conditional generation on ImageNet, with competitive T2I/T2V at 2B scale.
-
ELT: Elastic Looped Transformers for Visual Generation
Weight-shared looped transformers trained with intra-loop self-distillation match MaskGIT-class FID/FVD at roughly 4x fewer parameters and support any-time inference across loop counts.
-
WindFM: An Open-Source Foundation Model for Zero-Shot Wind Power Forecasting
A wind-specific foundation model, WindFM, uses hierarchical tokenization and autoregressive pre-training on the NREL WIND Toolkit to achieve state-of-the-art zero-shot wind power forecasts.
-
Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive
A new interactive segmentation decoder that routes computation to boundary regions, using binary quantization attention and mixture-of-experts, achieves state-of-the-art accuracy with CPU-friendly latency.
-
MambaVideo for Discrete Video Tokenization with Channel-Split Quantization
A Mamba-based hierarchical video tokenizer with channel-split quantization achieves state-of-the-art reconstruction and generation scores while preserving token count.
-
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
multivariateGPT extends next-token prediction to jointly predict the class and continuous value of mixed categorical and numeric time series, with Gaussian uncertainty, and outperforms discrete-token baselines on clin...
-
TokBench: Evaluating Your Visual Tokenizer before Visual Generation
TokBench measures text recognition accuracy and face similarity on reconstructed images and videos across 16 tokenizers, showing small text and faces are poorly preserved and often missed by traditional metrics.
-
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
UniGen shows a 1.5B model trained on open data can beat larger systems on image understanding and generation once it verifies its own outputs with chain-of-thought and Best-of-N selection.
-
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
QLIP trains a quantized image autoencoder with both reconstruction and text-alignment losses, yielding a tokenizer that supports multimodal understanding and text-to-image generation in one model.
-
Wireless TokenCom: RL-Based Tokenizer Agreement for Multi-User Wireless Token Communications
Joint tokenizer/codebook selection, subchannel assignment, and beamforming for multi-user video TokenCom is posed as an MDP and solved by DQN for discrete choices and DDPG for beamforming, with simulated gains over H.265.
-
LGQ: Learnable Geometric Quantization for Image Tokenization
LGQ reports better ImageNet reconstruction FID than FSQ/SimVQ using soft-to-hard learnable-codebook quantization, but its abstract's generation and utilization claims are contradicted by the body.
Discussion (0). Continue with ORCID to comment.