REVIEW 8 cited by
Scaling Vision Transformers to 22 Billion Parameters
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Vision Transformers (ViT) have introduced the same architecture to image and video modelling, but these have not yet been successfully scaled to nearly the same degree; the largest dense ViT contains 4B parameters (Chen et al., 2022). We present a recipe for highly efficient and stable training of a 22B-parameter ViT (ViT-22B) and perform a wide variety of experiments on the resulting model. When evaluated on downstream tasks (often with a lightweight linear model on frozen features), ViT-22B demonstrates increasing performance with scale. We further observe other interesting benefits of scale, including an improved tradeoff between fairness and performance, state-of-the-art alignment to human visual perception in terms of shape/texture bias, and improved robustness. ViT-22B demonstrates the potential for "LLM-like" scaling in vision, and provides key steps towards getting there.
Forward citations
Cited by 8 Pith papers
-
One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse
Different low-precision errors converge on the same query-key spectral runaway, entry is gated by temporal sign-coherence, and a dormant query-key normalization guard contains it.
-
Predict before you train: Scaling Laws for particle physics foundation models
A Chinchilla-style law fit on ParticleViT runs below 10^19 FLOPs predicts held-out pretraining loss within ~1% at >100× compute and tracks downstream jet-tagging rejection.
-
Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors
Splitting weight matrices into a fixed-norm direction and learnable per-row/column magnitudes improves LLM training over AdamW/Muon, removes weight decay and warmup, and transfers the optimal LR across width.
-
Distributed Cross-Channel Hierarchical Aggregation for Foundation Models
Distributed Cross-Channel Hierarchical Aggregation (D-CHAG) reduces memory and boosts throughput for multi-channel vision foundation models by spreading tokenization and channel fusion across GPUs with only a small qu...
-
Adversarial Attacks on Robotic Vision Language Action Models
Text-based adversarial suffixes can make OpenVLA robot policies elicit chosen target actions with over 90% success on one-hot targets and persist across rollout steps.
-
Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models
Direct Ascent Synthesis generates recognizable images from CLIP embeddings by optimizing a sum of multi-resolution image components, requiring no generative training.
-
Opt.Gear Technical Report
Opt.Gear is a family of efficient on-device language models using a ConvKV-gated mixer with sparse attention, trained on 0.5T tokens without distillation, claiming up to 4.9x NPU speedups and 20 TPS on a Cortex-M7.
-
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
An open-source image-generation family shows that agentic prompt rewriting and a stronger text encoder can lift quality to near closed-source levels with only 208.62M images and about $400K of training compute.
Discussion (0). Continue with ORCID to comment.