Pith. sign in

REVIEW 32 cited by

ImageNet-21K Pretraining for the Masses

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.10972 v4 pith:522CHFUV submitted 2021-04-22 cs.CV cs.LG

classification cs.CVcs.LG
keywords pretrainingimagenet-21kmodelsavailabledatasetefficienttaskstraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

ImageNet-1K serves as the primary dataset for pretraining deep learning models for computer vision tasks. ImageNet-21K dataset, which is bigger and more diverse, is used less frequently for pretraining, mainly due to its complexity, low accessibility, and underestimation of its added value. This paper aims to close this gap, and make high-quality efficient pretraining on ImageNet-21K available for everyone. Via a dedicated preprocessing stage, utilization of WordNet hierarchical structure, and a novel training scheme called semantic softmax, we show that various models significantly benefit from ImageNet-21K pretraining on numerous datasets and tasks, including small mobile-oriented models. We also show that we outperform previous ImageNet-21K pretraining schemes for prominent new models like ViT and Mixer. Our proposed pretraining pipeline is efficient, accessible, and leads to SoTA reproducible results, from a publicly available dataset. The training code and pretrained models are available at: https://github.com/Alibaba-MIIL/ImageNet21K

Discussion (0). Sign in to comment.

Forward citations

Cited by 32 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CanViT: Toward Active-Vision Foundation Models

    cs.CV 2026-03 conditional novelty 8.0 of 10

    CanViT is the first task- and policy-agnostic AVFM pretrained via passive-to-active dense latent distillation on 13.2M scenes and 1B random glimpses, achieving 38.5% ADE20K mIoU in one glimpse and 84.5% ImageNet-1k to...

  2. Condensing Large-Scale Datasets Directly with Minimal Information Loss

    cs.CV 2026-07 unverdicted novelty 7.0 of 10

    CIM directly aligns data distributions to condense large-scale datasets with minimal information loss, achieving new SOTA results on ImageNet-1K distillation at IPC=10.

  3. Identifying Latent Concepts and Structures for Generalized Category Discovery

    cs.CV 2026-07 unverdicted novelty 7.0 of 10

    CPF-GCD enforces low-rank compositional structure on vision backbone features via spatial primitive fields so that novel categories emerge as new activation patterns over a shared vocabulary of reusable visual primitives.

  4. GVCCS: A Dataset for Contrail Identification and Tracking on Visible Whole Sky Camera Sequences

    cs.CV 2025-07 conditional novelty 7.0 of 10

    GVCCS is the first open dataset of ground-based visible all-sky camera video with instance-level contrail masks, temporal tracking, and flight IDs, plus Mask2Former baselines.

  5. The Edge-on Galaxies in the DESI survey (EGIDE): sample building and photometry

    astro-ph.GA 2026-06 unverdicted novelty 6.0 of 10

    The EGIDE project releases a tenfold larger catalogue of edge-on galaxies with griz photometry, stellar masses, redshifts and star formation rates, finding that red-sequence galaxies are thicker than blue-cloud ones a...

  6. The Edge-on Galaxies in the DESI survey (EGIDE): sample building and photometry

    astro-ph.GA 2026-06 conditional novelty 6.0 of 10

    EGIDE provides 149,215 edge-on galaxy candidates from DESI DR10 with homogeneous photometry, masses, and redshifts, ten times larger than EGIPS.

  7. ExDet: Open-Domain Open-Vocabulary Detection with Cross-modal Extrapolation and Rectification

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    ExDet proposes a lightweight framework using text-guided extrapolation, detector-compatible rectification, and recalibrated proposals to achieve state-of-the-art open-domain open-vocabulary detection on multiple benchmarks.

  8. VISReg: Variance-Invariance-Sketching Regularization for JEPA training

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    VISReg replaces covariance in VICReg-style objectives with sliced-Wasserstein sketching for JEPA training, claiming better OOD performance and resilience to collapse.

  9. LV-OSD: Language-Vision-Complementary Open-Set Object Detection

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    The authors define the LV-OSD problem and introduce the LVDor dual-branch framework with TPDW dynamic weighting and PRM masking to align multimodal prompts for open-set detection.

  10. Energy-Structured Low-Rank Adaptation for Continual Learning

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    E²-LoRA structures low-rank adaptations by energy concentration and ordering with dynamic rank allocation to achieve state-of-the-art continual learning.

  11. Weierstrass Positional Encoding for Vision Transformers

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    WePE encodes 2D patch positions in Vision Transformers via Weierstrass elliptic functions on the complex plane to exploit double periodicity and derive relative positions algebraically.

  12. Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Exploiting linear structure in VLM embeddings, a synthetic-data pre-training method yields background-invariant representations that exceed 90% worst-group accuracy on Waterbirds even under 100% spurious correlation w...

  13. Parameter-Efficient Adaptation of Pre-Trained Vision Foundation Models for Active and Passive Seismic Data Denoising

    physics.geo-ph 2026-04 conditional novelty 6.0 of 10

    Adapting vision foundation models with LoRA and kurtosis-guided unsupervised test-time adaptation matches or exceeds domain-specific models for seismic denoising across multiple sites and unseen data.

  14. Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Using contrastive examples with vision-language models and a new CLIP-based scoring method called CSP produces more faithful and granular neuron labels than prior activation-only approaches.

  15. StableTTA: Improving Vision Model Performance by Training-free Test-Time Adaptation Methods

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    StableTTA improves ImageNet-1K accuracy across 71 vision models by stabilizing logit aggregation under coherent-batch inference and enabling efficient single-forward-pass adaptation.

  16. MePo: Meta Post-Refinement for Rehearsal-Free General Continual Learning

    cs.AI 2026-02 unverdicted novelty 6.0 of 10

    MePo refines pretrained backbones via meta-learning on constructed pseudo tasks and initializes a meta covariance matrix to enable robust second-order alignment, yielding 12-15% gains on CIFAR-100, ImageNet-R and CUB-...

  17. Page image classification for content-specific data processing

    cs.IR 2025-07 conditional novelty 6.0 of 10

    Fine-tuned image classifiers, led by RegNetY-16GF at about 99% top-1 accuracy, can sort historical archaeological page scans into 11 content categories, and the authors release the dataset, code, and model weights.

  18. Towards Continuous Home Cage Monitoring: An Evaluation of Tracking and Identification Strategies for Laboratory Mice

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A real-time mouse tracking and ear-tag identity pipeline reports 95.28% identification accuracy and fewer ID switches than SLEAP and DeepLabCut on a 100-minute home-cage dataset.

  19. Opto-ViT: Architecting a Near-Sensor Region of Interest-Aware Vision Transformer Accelerator with Silicon Photonics

    cs.AR 2025-07 conditional novelty 6.0 of 10

    A near-sensor vision transformer accelerator combines VCSEL-microring photonic matrix multiplication with region-of-interest patch pruning, reporting 100.4 KFPS/W and up to 84% energy savings.

  20. Revisiting Audio-Visual Segmentation with Vision-Centric Transformer

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Vision-derived object queries with prototype prompting achieve new state-of-the-art results on AVSBench audio-visual segmentation.

  21. LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

    cs.CV 2023-10 unverdicted novelty 6.0 of 10

    LanguageBind aligns video, infrared, depth, and audio to a frozen language encoder via contrastive learning on the new VIDAL-10M dataset, extending video-language pretraining to N modalities.

  22. JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    JetViT uses post-training attention search to hybridize full-attention ViTs with linear and window attention blocks, achieving up to 1.79x throughput gains on high-res images while preserving accuracy on DINOv3 and De...

  23. H3Former: Hypergraph-based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A hypergraph-based token-to-region aggregation plus a hyperbolic hierarchical contrastive loss yields reported state-of-the-art fine-grained classification accuracy on four benchmarks.

  24. Revisiting Deepfake Detection: Chronological Continual Learning and the Limits of Generalization

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A chronological continual learning study finds deepfake detectors retain past knowledge but generalize to future generators at near-random AUC around 0.5.

  25. Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset

    cs.CV 2025-07 conditional novelty 5.0 of 10

    The ODOR dataset contributes 38,116 fine-grained object annotations over 4,712 artworks, benchmarked with five detector families, to stress-test object detection on dense, occluded, and off-centre objects in historica...

  26. MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

    cs.CV 2023-12 unverdicted novelty 5.0 of 10

    MobileVLM achieves on-par performance with much larger vision-language models on standard benchmarks while delivering state-of-the-art inference speeds of 21.5 tokens per second on Snapdragon 888 CPU and 65.3 on Jetso...

  27. Page image classifier fine-tuned on century-spanning archives of scanned documents for further content-specific processing

    cs.CV 2026-05 unverdicted novelty 4.0 of 10

    Fine-tuned RegNetY-16GF reaches 99.16% accuracy classifying 48k century-old Czech archival pages into 11 visual content types, beating a 75% hand-crafted feature baseline, with models and data released publicly.

  28. HandyLabel: Towards Post-Processing to Real-Time Annotation Using Skeleton Based Hand Gesture Recognition

    cs.HC 2025-11 unverdicted novelty 4.0 of 10

    HandyLabel enables real-time data annotation by mapping hand gestures to labels, with ResNet50 on skeleton-preprocessed HaGRID data reaching 0.923 F1-score and 88.9% of 46 study participants preferring it to tradition...

  29. CascadeFormer: A Family of Two-stage Cascading Transformers for Skeleton-based Human Action Recognition

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A masked-pretrained skeleton transformer with a second fine-tuning transformer and cross-attention fusion reaches 94.66% on Penn Action, 91.16% on N-UCLA, and 81.01%/88.17% on NTU RGB+D 60 cross-subject/cross-view.

  30. Page image classification for content-specific data processing

    cs.IR 2025-07 unverdicted novelty 3.0 of 10

    Develops an image classification system for historical document pages to support content-specific analysis pipelines.

  31. 4th Workshop on Maritime Computer Vision (MaCVi): Challenge Overview

    cs.CV 2026-04 unverdicted novelty 2.0 of 10

    The report overviews five maritime computer vision benchmark challenges, their datasets, protocols, quantitative results, and top team approaches from the MaCVi 2026 workshop.

  32. ViT-ProtoNet for Few-Shot Image Classification: A Multi-Benchmark Evaluation

    cs.CV 2025-07 reject novelty 2.0 of 10

    ViT-ProtoNet, a Prototypical Network with a ViT-Small encoder, is reported to reach 95-97% 5-shot accuracy on three benchmarks and 81.88% on FC100, but the evaluation lacks critical baselines.

Pith tools