Pith. sign in

REVIEW 8 cited by

Scaling Vision Transformers to 22 Billion Parameters

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.05442 v1 pith:VRCOCLKD submitted 2023-02-10 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords parametersscalingtransformersvisionvit-22bdemonstratesimprovedlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Vision Transformers (ViT) have introduced the same architecture to image and video modelling, but these have not yet been successfully scaled to nearly the same degree; the largest dense ViT contains 4B parameters (Chen et al., 2022). We present a recipe for highly efficient and stable training of a 22B-parameter ViT (ViT-22B) and perform a wide variety of experiments on the resulting model. When evaluated on downstream tasks (often with a lightweight linear model on frozen features), ViT-22B demonstrates increasing performance with scale. We further observe other interesting benefits of scale, including an improved tradeoff between fairness and performance, state-of-the-art alignment to human visual perception in terms of shape/texture bias, and improved robustness. ViT-22B demonstrates the potential for "LLM-like" scaling in vision, and provides key steps towards getting there.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 118 citations worldwide. Full citation record

  1. One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse

    cs.LG 2026-08 conditional novelty 7.0 of 10

    Different low-precision errors converge on the same query-key spectral runaway, entry is gated by temporal sign-coherence, and a dormant query-key normalization guard contains it.

  2. Predict before you train: Scaling Laws for particle physics foundation models

    hep-ex 2026-07 conditional novelty 7.0 of 10

    A Chinchilla-style law fit on ParticleViT runs below 10^19 FLOPs predicts held-out pretraining loss within ~1% at >100× compute and tracks downstream jet-tagging rejection.

  3. Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Splitting weight matrices into a fixed-norm direction and learnable per-row/column magnitudes improves LLM training over AdamW/Muon, removes weight decay and warmup, and transfers the optimal LR across width.

  4. Distributed Cross-Channel Hierarchical Aggregation for Foundation Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Distributed Cross-Channel Hierarchical Aggregation (D-CHAG) reduces memory and boosts throughput for multi-channel vision foundation models by spreading tokenization and channel fusion across GPUs with only a small qu...

  5. Adversarial Attacks on Robotic Vision Language Action Models

    cs.RO 2025-06 conditional novelty 6.0 of 10

    Text-based adversarial suffixes can make OpenVLA robot policies elicit chosen target actions with over 90% success on one-hot targets and persist across rollout steps.

  6. Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Direct Ascent Synthesis generates recognizable images from CLIP embeddings by optimizing a sum of multi-resolution image components, requiring no generative training.

  7. Opt.Gear Technical Report

    cs.CL 2026-08 conditional novelty 4.0 of 10

    Opt.Gear is a family of efficient on-device language models using a ConvKV-gated mixer with sparse attention, trained on 0.5T tokens without distillation, claiming up to 4.9x NPU speedups and 20 TPS on a Cortex-M7.

  8. Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget

    cs.CV 2026-07 conditional novelty 4.0 of 10

    An open-source image-generation family shows that agentic prompt rewriting and a stronger text encoder can lift quality to near closed-source levels with only 208.62M images and about $400K of training compute.

Pith tools