Pith. sign in

REVIEW 2 cited by

Going deeper with Image Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.17239 v2 pith:6CXRGECZ submitted 2021-03-31 cs.CV

classification cs.CV
keywords transformersimageaccuracyarchitecturebeenclassificationdatadeeper
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformers have been recently adapted for large scale image classification, achieving high scores shaking up the long supremacy of convolutional neural networks. However the optimization of image transformers has been little studied so far. In this work, we build and optimize deeper transformer networks for image classification. In particular, we investigate the interplay of architecture and optimization of such dedicated transformers. We make two transformers architecture changes that significantly improve the accuracy of deep transformers. This leads us to produce models whose performance does not saturate early with more depth, for instance we obtain 86.5% top-1 accuracy on Imagenet when training with no external data, we thus attain the current SOTA with less FLOPs and parameters. Moreover, our best model establishes the new state of the art on Imagenet with Reassessed labels and Imagenet-V2 / match frequency, in the setting with no additional training data. We share our code and models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FourCastNet 3: A geometric approach to probabilistic machine-learning weather forecasting at scale

    cs.LG 2025-07 conditional novelty 7.0 of 10

    A purely convolutional, spherical-geometry weather model trained with a combined spatial and spectral CRPS loss delivers GenCast-level skill, IFS-beating accuracy, and stable spectra out to 60 days.

  2. NOVO: Unlearning-Compliant Vision Transformers

    cs.CV 2025-07 conditional novelty 6.0 of 10

    NOVO is a vision transformer that forgets classes at inference time by removing learned class keys, trained with simulated unlearning to generalize to any forget set.

Pith tools