Pith. sign in

REVIEW 4 cited by

How to Train Vision Transformer on Small-scale Datasets?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.07240 v1 pith:PF2TUW54 submitted 2022-10-13 cs.CV

classification cs.CV
keywords visiondatasetstransformermodelssmall-scaletrainarchitecturebiases
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Vision Transformer (ViT), a radically different architecture than convolutional neural networks offers multiple advantages including design simplicity, robustness and state-of-the-art performance on many vision tasks. However, in contrast to convolutional neural networks, Vision Transformer lacks inherent inductive biases. Therefore, successful training of such models is mainly attributed to pre-training on large-scale datasets such as ImageNet with 1.2M or JFT with 300M images. This hinders the direct adaption of Vision Transformer for small-scale datasets. In this work, we show that self-supervised inductive biases can be learned directly from small-scale datasets and serve as an effective weight initialization scheme for fine-tuning. This allows to train these models without large-scale pre-training, changes to model architecture or loss functions. We present thorough experiments to successfully train monolithic and non-monolithic Vision Transformers on five small datasets including CIFAR10/100, CINIC10, SVHN, Tiny-ImageNet and two fine-grained datasets: Aircraft and Cars. Our approach consistently improves the performance of Vision Transformers while retaining their properties such as attention to salient regions and higher robustness. Our codes and pre-trained models are available at: https://github.com/hananshafi/vits-for-small-scale-datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Benign Overfitting Does Not Occur in Diffusion Models

    stat.ML 2026-07 conditional novelty 7.0 of 10

    Benign overfitting and double descent do not occur in diffusion models: population and empirical score-matching losses cannot both be small without exponentially many samples.

  2. Cyborg Insect Factory: Automatic Assembly System to Build up Insect-computer Hybrid Robot Based on Vision-guided Robotic Arm Manipulation of Custom Bipolar Electrodes

    cs.RO 2024-11 conditional novelty 7.0 of 10

    An automatic vision-guided robotic assembly system builds steerable and decelerable cyborg cockroaches in 68 seconds with control performance comparable to manual assembly.

  3. Learning to Adapt to Position Bias in Vision Transformer Classifiers

    cs.CV 2025-05 reject novelty 6.0 of 10

    A vision transformer can learn a single scalar that gates its position embedding, and a new SHAP-based metric measures when that gate should be small or large.

  4. Powerful Design of Small Vision Transformer on CIFAR10

    cs.LG 2025-01 conditional novelty 3.0 of 10

    A Tiny ViT on CIFAR-10 reaches 93.95% accuracy with two classifier tokens at reduced width, while low-rank query compression causes only a small accuracy drop.

Pith tools