Pith. sign in

REVIEW 4 cited by

Understanding Why ViT Trains Badly on Small Datasets: An Intuitive Perspective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.03751 v1 pith:SUTFO3WK submitted 2023-02-07 cs.CV cs.LG

classification cs.CVcs.LG
keywords datasetssmalltrainedresnet-18accuracyattentionperformanceperspective
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision transformer (ViT) is an attention neural network architecture that is shown to be effective for computer vision tasks. However, compared to ResNet-18 with a similar number of parameters, ViT has a significantly lower evaluation accuracy when trained on small datasets. To facilitate studies in related fields, we provide a visual intuition to help understand why it is the case. We first compare the performance of the two models and confirm that ViT has less accuracy than ResNet-18 when trained on small datasets. We then interpret the results by showing attention map visualization for ViT and feature map visualization for ResNet-18. The difference is further analyzed through a representation similarity perspective. We conclude that the representation of ViT trained on small datasets is hugely different from ViT trained on large datasets, which may be the reason why the performance drops a lot on small datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients

    cs.LG 2025-09 conditional novelty 6.0 of 10

    A warm-up phase of ordinary federated training followed by zeroth-order forward-pass-only updates lets low-resource clients participate in federated pre-training from random initialization.

  2. eMamba: Efficient Acceleration Framework for Mamba Models in Edge Computing

    cs.LG 2025-08 conditional novelty 6.0 of 10

    An end-to-end Mamba edge accelerator using hardware-friendly approximations, INT8 quantization, and NAS achieves 4.95x-5.62x lower latency and 1.63x-19.9x smaller models than ViT/CNN baselines.

  3. Efficient optimization of expensive black-box simulators via marginal means, with application to neutrino detector design

    stat.ML 2025-08 unverdicted novelty 6.0 of 10

    A new estimator, BOMM, uses marginal mean functions to propose optimizer candidates beyond evaluated simulator runs, with proven consistency and improved high-dimensional rates under an additive model.

  4. Self-Supervised Ultrasound-Video Segmentation with Feature Prediction and 3D Localised Loss

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A 3D relative-localisation auxiliary loss improves V-JEPA pre-training for cardiac ultrasound video segmentation, with larger gains in low-label regimes.

Pith tools