Pith. sign in

REVIEW 1 cited by

Pre-training of Lightweight Vision Transformers on Small Datasets with Minimally Scaled Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.03752 v1 pith:DTLUHBY2 submitted 2024-02-06 cs.CV cs.LG

classification cs.CVcs.LG
keywords datasetslightweightsmallimagesperformancecifar-10cifar-100image
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Can a lightweight Vision Transformer (ViT) match or exceed the performance of Convolutional Neural Networks (CNNs) like ResNet on small datasets with small image resolutions? This report demonstrates that a pure ViT can indeed achieve superior performance through pre-training, using a masked auto-encoder technique with minimal image scaling. Our experiments on the CIFAR-10 and CIFAR-100 datasets involved ViT models with fewer than 3.65 million parameters and a multiply-accumulate (MAC) count below 0.27G, qualifying them as 'lightweight' models. Unlike previous approaches, our method attains state-of-the-art performance among similar lightweight transformer-based architectures without significantly scaling up images from CIFAR-10 and CIFAR-100. This achievement underscores the efficiency of our model, not only in handling small datasets but also in effectively processing images close to their original scale.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-modal transformer for signal classification in nanopore blockade experiments

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A multi-modal transformer fusing raw current traces, wavelet images, and catch22 features classifies nanopore peptide events with 92.6% macro accuracy—more than 10 points above the best single-modality baseline.

Pith tools