REVIEW 13 cited by
Training data-efficient image transformers & distillation through attention
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recently, neural networks purely based on attention were shown to address image understanding tasks such as image classification. However, these visual transformers are pre-trained with hundreds of millions of images using an expensive infrastructure, thereby limiting their adoption. In this work, we produce a competitive convolution-free transformer by training on Imagenet only. We train them on a single computer in less than 3 days. Our reference vision transformer (86M parameters) achieves top-1 accuracy of 83.1% (single-crop evaluation) on ImageNet with no external data. More importantly, we introduce a teacher-student strategy specific to transformers. It relies on a distillation token ensuring that the student learns from the teacher through attention. We show the interest of this token-based distillation, especially when using a convnet as a teacher. This leads us to report results competitive with convnets for both Imagenet (where we obtain up to 85.2% accuracy) and when transferring to other tasks. We share our code and models.
Forward citations
Cited by 13 Pith papers
-
Learn from A Rationalist: Distilling Intermediate Interpretable Rationales
Distilling a teacher rationale-extraction model's feature selections and predictions into smaller students improves student accuracy by up to ~14 points on CIFAR-10 while keeping the same rationale sparsity.
-
MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs
MoEcho claims to compromise user privacy in MoE LLMs and VLMs via four CPU and GPU side channels, but the provided manuscript body contains no supporting content.
-
Fine-Grained Image Recognition from Scratch with Teacher-Guided Data Augmentation
TGDA trains fine-grained-recognition students from scratch using attention maps and soft labels from a fine-tuned, pretrained teacher, reporting large gains on three benchmarks.
-
AdvMIM: Adversarial Masked Image Modeling for Semi-Supervised Medical Image Segmentation
AdvMIM trains a transformer to segment masked images and adversarially aligns original and masked domains, improving semi-supervised medical image segmentation.
-
From Pixels to Components: Eigenvector Masking for Visual Representation Learning
Masking principal components instead of pixel patches in masked autoencoders yields better image classification representations across CIFAR10, TinyImageNet, and three MedMNIST datasets.
-
Efficient Learned Image Compression Through Knowledge Distillation
Knowledge-distilled students with 64 or more channels match the rate-distortion performance of a 128-channel teacher while cutting memory by 68% and energy by 34%.
-
I-Segmenter: Integer-Only Vision Transformer for Efficient Semantic Segmentation
I-Segmenter is an integer-only Vision Transformer for semantic segmentation that keeps mIoU within roughly 5 points of the FP32 baseline while cutting model size by up to 3.8x.
-
SAAT: Synergistic Alternating Aggregation Transformer for Image Super-Resolution
SAAT, an alternating channel-spatial-window attention Transformer, reports slight PSNR/SSIM gains over SwinIR and HAT, but the evidence is weakened by inconsistent baselines and test-set tuning.
-
In Context Learning with Vision Transformers: Case Study
A decoder-only transformer with a CNN or ViT image encoder learns random linear, convolutional, and ViT functions on 8x8 CIFAR-10 images in-context from a handful of examples.
-
ViSIR: Vision Transformer Single Image Reconstruction Method for Earth System Models
ViSIR, a Vision Transformer with a SIREN head, is claimed to improve super-resolution quality on E3SM climate images by 2 to 8 dB over four baselines, though the experimental setup is incomplete.
-
The evolution of AI from image interpretation toward scientific inference in nanoparticle electron microscopy
AI for nanoparticle TEM/STEM has progressed from detection and segmentation to physics-informed restoration, 2D-to-3D inference, and spatiotemporal analysis of in situ dynamics, with remaining gaps in benchmarking and...
-
Kolmogorov-Arnold Fourier Networks
KAF is a Kolmogorov-Arnold-style network that uses Random Fourier Features and a hybrid GELU activation, claiming better efficiency and high-frequency accuracy, though the parameter reduction and sigma=1.64 'derivatio...
-
Performance Analysis of Traditional VQA Models Under Limited Computational Resources
An empirical comparison claims BidGRU with embedding size 300 and vocabulary 3000 is the best resource-constrained VQA configuration, but the paper lacks dataset and statistical details.
Discussion (0). Continue with ORCID to comment.