Pith. sign in

REVIEW 13 cited by

Training data-efficient image transformers & distillation through attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.12877 v2 pith:3DW57JXI submitted 2020-12-23 cs.CV

classification cs.CV
keywords attentiondistillationimageimagenettransformersaccuracycompetitivetasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, neural networks purely based on attention were shown to address image understanding tasks such as image classification. However, these visual transformers are pre-trained with hundreds of millions of images using an expensive infrastructure, thereby limiting their adoption. In this work, we produce a competitive convolution-free transformer by training on Imagenet only. We train them on a single computer in less than 3 days. Our reference vision transformer (86M parameters) achieves top-1 accuracy of 83.1% (single-crop evaluation) on ImageNet with no external data. More importantly, we introduce a teacher-student strategy specific to transformers. It relies on a distillation token ensuring that the student learns from the teacher through attention. We show the interest of this token-based distillation, especially when using a convnet as a teacher. This leads us to report results competitive with convnets for both Imagenet (where we obtain up to 85.2% accuracy) and when transferring to other tasks. We share our code and models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learn from A Rationalist: Distilling Intermediate Interpretable Rationales

    cs.LG 2026-01 conditional novelty 6.0 of 10

    Distilling a teacher rationale-extraction model's feature selections and predictions into smaller students improves student accuracy by up to ~14 points on CIFAR-10 while keeping the same rationale sparsity.

  2. MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs

    cs.CR 2025-08 reject novelty 6.0 of 10

    MoEcho claims to compromise user privacy in MoE LLMs and VLMs via four CPU and GPU side channels, but the provided manuscript body contains no supporting content.

  3. Fine-Grained Image Recognition from Scratch with Teacher-Guided Data Augmentation

    cs.CV 2025-07 reject novelty 6.0 of 10

    TGDA trains fine-grained-recognition students from scratch using attention maps and soft labels from a fine-tuned, pretrained teacher, reporting large gains on three benchmarks.

  4. AdvMIM: Adversarial Masked Image Modeling for Semi-Supervised Medical Image Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AdvMIM trains a transformer to segment masked images and adversarially aligns original and masked domains, improving semi-supervised medical image segmentation.

  5. From Pixels to Components: Eigenvector Masking for Visual Representation Learning

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Masking principal components instead of pixel patches in masked autoencoders yields better image classification representations across CIFAR10, TinyImageNet, and three MedMNIST datasets.

  6. Efficient Learned Image Compression Through Knowledge Distillation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    Knowledge-distilled students with 64 or more channels match the rate-distortion performance of a 128-channel teacher while cutting memory by 68% and energy by 34%.

  7. I-Segmenter: Integer-Only Vision Transformer for Efficient Semantic Segmentation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    I-Segmenter is an integer-only Vision Transformer for semantic segmentation that keeps mIoU within roughly 5 points of the FP32 baseline while cutting model size by up to 3.8x.

  8. SAAT: Synergistic Alternating Aggregation Transformer for Image Super-Resolution

    cs.CV 2025-06 reject novelty 4.0 of 10

    SAAT, an alternating channel-spatial-window attention Transformer, reports slight PSNR/SSIM gains over SwinIR and HAT, but the evidence is weakened by inconsistent baselines and test-set tuning.

  9. In Context Learning with Vision Transformers: Case Study

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A decoder-only transformer with a CNN or ViT image encoder learns random linear, convolutional, and ViT functions on 8x8 CIFAR-10 images in-context from a handful of examples.

  10. ViSIR: Vision Transformer Single Image Reconstruction Method for Earth System Models

    cs.CV 2025-02 reject novelty 4.0 of 10

    ViSIR, a Vision Transformer with a SIREN head, is claimed to improve super-resolution quality on E3SM climate images by 2 to 8 dB over four baselines, though the experimental setup is incomplete.

  11. The evolution of AI from image interpretation toward scientific inference in nanoparticle electron microscopy

    cond-mat.mtrl-sci 2026-07 conditional novelty 3.5 of 10

    AI for nanoparticle TEM/STEM has progressed from detection and segmentation to physics-informed restoration, 2D-to-3D inference, and spatiotemporal analysis of in situ dynamics, with remaining gaps in benchmarking and...

  12. Kolmogorov-Arnold Fourier Networks

    cs.LG 2025-02 reject novelty 3.0 of 10

    KAF is a Kolmogorov-Arnold-style network that uses Random Fourier Features and a hybrid GELU activation, claiming better efficiency and high-frequency accuracy, though the parameter reduction and sigma=1.64 'derivatio...

  13. Performance Analysis of Traditional VQA Models Under Limited Computational Resources

    cs.CV 2025-02 reject novelty 2.0 of 10

    An empirical comparison claims BidGRU with embedding size 300 and vocabulary 3000 is the best resource-constrained VQA configuration, but the paper lacks dataset and statistical details.

Pith tools