Pith. sign in

REVIEW 29 cited by

A ConvNet for the 2020s

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.03545 v2 pith:KOEQNG3K submitted 2022-01-10 cs.CV

classification cs.CV
keywords transformersconvnetvisionstandardaccuracyconvnetsdesigndetection
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The "Roaring 20s" of visual recognition began with the introduction of Vision Transformers (ViTs), which quickly superseded ConvNets as the state-of-the-art image classification model. A vanilla ViT, on the other hand, faces difficulties when applied to general computer vision tasks such as object detection and semantic segmentation. It is the hierarchical Transformers (e.g., Swin Transformers) that reintroduced several ConvNet priors, making Transformers practically viable as a generic vision backbone and demonstrating remarkable performance on a wide variety of vision tasks. However, the effectiveness of such hybrid approaches is still largely credited to the intrinsic superiority of Transformers, rather than the inherent inductive biases of convolutions. In this work, we reexamine the design spaces and test the limits of what a pure ConvNet can achieve. We gradually "modernize" a standard ResNet toward the design of a vision Transformer, and discover several key components that contribute to the performance difference along the way. The outcome of this exploration is a family of pure ConvNet models dubbed ConvNeXt. Constructed entirely from standard ConvNet modules, ConvNeXts compete favorably with Transformers in terms of accuracy and scalability, achieving 87.8% ImageNet top-1 accuracy and outperforming Swin Transformers on COCO detection and ADE20K segmentation, while maintaining the simplicity and efficiency of standard ConvNets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 29 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 178 citations worldwide. Full citation record

  1. AIMIP Phase 1: systematic evaluations of AI weather and climate models

    physics.ao-ph 2026-05 unverdicted novelty 7.0 of 10

    Under one protocol, most AI climate models reproduce historical climatology and ENSO response as well as a CMIP6 model, but some underestimate warming trends and all diverge on +2/+4K SST experiments.

  2. FourCastNet 3: A geometric approach to probabilistic machine-learning weather forecasting at scale

    cs.LG 2025-07 conditional novelty 7.0 of 10

    A purely convolutional, spherical-geometry weather model trained with a combined spatial and spectral CRPS loss delivers GenCast-level skill, IFS-beating accuracy, and stable spectra out to 60 days.

  3. ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A single-step IMLE generator with per-stage supervision and a robust loss reports FID 2.56 on ImageNet-256 by filtering ~5% of samples at test time.

  4. C3DIR: A Deep Learning 3-Dimensional Cloud Property Retrieval Scheme for Passive Satellite Imagers

    physics.ao-ph 2026-07 conditional novelty 6.0 of 10

    C3DIR is a single multi-sensor deep-learning model that retrieves 3-D ice, liquid, and rain water content from passive imagers using voxel-to-voxel collocation with active-sensor profiles.

  5. Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A continuously refreshed, incentive-driven deepfake detector beats static detectors on in-the-wild benchmarks and improves on post-export AI-generated media.

  6. Geographically-Weighted Weakly Supervised Bayesian High-Resolution Transformer for 200m Resolution Pan-Arctic Sea Ice Concentration Mapping and Uncertainty Estimation using Sentinel-1, RCM, and AMSR2 Data

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A 200m pan-Arctic sea ice concentration model using a Bayesian Transformer with a geographically-weighted weakly supervised loss; reported 0.70 feature detection accuracy and R²=0.90 vs ASI.

  7. Model soups need only one ingredient

    cs.LG 2026-02 conditional novelty 6.0 of 10

    A single checkpoint, edited by splitting each layer's update with SVD and reweighting the high- and low-energy parts, reaches soup-level OOD robustness without multi-model training.

  8. Khana: A Comprehensive Indian Cuisine Dataset

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Khana is an Indian cuisine image dataset with 131K images across 80 dish categories and baseline classification accuracies up to 86.7% top-1 with ConvNeXT.

  9. RRTO: A High-Performance Transparent Offloading System for Model Inference in Mobile Edge Computing

    cs.NI 2025-07 conditional novelty 6.0 of 10

    RRTO identifies static inference operator sequences from CUDA call logs alone and replays them on an edge GPU, cutting transparent-offloading communication to 11 RPCs per inference instead of thousands, with performan...

  10. Beginning with You: Perceptual-Initialization Improves Vision-Language Representation and Alignment

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Perceptually initializing a CLIP vision encoder with NIGHTS triplet judgments before YFCC15M contrastive training improves zero-shot accuracy and retrieval over an identical random-start baseline.

  11. DiffEx: Explaining a Classifier with Diffusion Models to Identify Microscopic Cellular Variations

    cs.CV 2025-02 conditional novelty 6.0 of 10

    DiffEx builds a classifier-aware latent space with a diffusion model, finds contrastive directions in it, and ranks them to produce visual explanations that reveal cellular phenotype changes.

  12. Small, Bias-Free, Blind and Convolutional Denoiser: A compact ConvNeXt U-Net for blind Gaussian color-image denoising

    eess.IV 2026-07 conditional novelty 5.0 of 10

    BF-ConvUNeXt, a 0.82M-parameter bias-free ConvNeXt U-Net, is degree-1 homogeneous and matches DnCNN/FFDNet in blind color denoising, extrapolating smoothly beyond its training noise range.

  13. REVELIO -- Universal Multimodal Task Load Estimation for Cross-Domain Generalization

    cs.LG 2025-09 conditional novelty 5.0 of 10

    A new cognitive-load dataset with driving, gaming, and n-back tasks shows multimodal models beat unimodal ones, but cross-domain transfer remains poor.

  14. StixelNExT++: Lightweight Monocular Scene Segmentation and Representation for Collective Perception

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A monocular neural network predicts 3D Stixels directly from RGB images in about 10 ms, with a self-defined Waymo evaluation showing competitive performance within 30 m.

  15. A Controlled Visual-Backbone Benchmark for Multimodal Short-Term Solar Irradiance Forecasting

    eess.IV 2026-07 conditional novelty 4.0 of 10

    Under one fixed multimodal 10-min irradiance forecaster, VMamba-S and Swin-B nearly tie best Folsom RMSE (~65.4 W/m²) over smart persistence, while NREL’s tiny matched split favors persistence.

  16. Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning

    cs.LG 2026-07 conditional novelty 4.0 of 10

    On a large World of Tanks dataset, a shared multi-task model with equal weighting or PCGrad outperforms single-task models on average, and task/map pre-training helps most in low-data regimes.

  17. Revisiting Simple Baselines for In-The-Wild Deepfake Detection

    cs.CV 2025-09 conditional novelty 4.0 of 10

    Finetuned CLIP-pretrained ConvNeXt-base and ViT-b32 classifiers reach 81% accuracy on Deepfake-Eval-2024, within noise of the leading commercial detector's 82%.

  18. Classifying Mitotic Figures in the MIDOG25 Challenge with Deep Ensemble Learning and Rule Based Refinement

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A ConvNeXt ensemble achieves 84% balanced accuracy on MIDOG25 atypical mitotic figure classification, while a rule-based refinement module trades sensitivity for specificity.

  19. Beyond Linear Bottlenecks: Spline-Based Knowledge Distillation for Culturally Diverse Art Style Classification

    cs.CV 2025-07 reject novelty 4.0 of 10

    Replacing MLP projection heads with Kolmogorov-Arnold Network heads in a dual-teacher self-supervised art-style classifier yields Top-1 accuracy gains of around 0.2 to 1.0 percentage points on WikiArt and Pandora18k, ...

  20. Faithful, Interpretable Chest X-ray Diagnosis with Anti-Aliased B-cos Networks

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Combining B-cos networks with anti-aliasing pooling (FLC or BlurPool) reduces grid artifacts in chest X-ray explanation maps while keeping diagnostic accuracy close to baseline networks.

  21. Applying multimodal learning to Classify transient Detections Early (AppleCiDEr) I: Data set, methods, and infrastructure

    astro-ph.IM 2025-07 conditional novelty 4.0 of 10

    AppleCiDEr combines photometry, images, metadata, and spectra in one deep learning pipeline to classify ZTF transients and variable stars, with high accuracy on common classes but poor performance on tidal disruption events.

  22. First-of-its-kind AI model for bioacoustic detection using a lightweight associative memory Hopfield neural network

    cs.LG 2025-07 reject novelty 4.0 of 10

    A Hopfield neural network trained on two bat calls classifies 10,384 recordings in 5.4 seconds with claimed accuracy up to 80%, though the headline numbers hinge on removing ambiguous calls.

  23. Bridging the gap in FER: addressing age bias in deep learning

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Age-aware loss reweighting and age-conditioned training reduce elderly recognition errors in facial expression models, even when age labels are automatically estimated, but the evidence rests on a small elderly benchmark.

  24. Comparison of ConvNeXt and Vision-Language Models for Breast Density Assessment in Screening Mammography

    eess.IV 2025-06 conditional novelty 4.0 of 10

    Fine-tuned ConvNeXt achieves 0.73 accuracy and 0.78 F1, beating BioMedCLIP linear probe (0.64/0.63) and zero-shot (0.47/0.31) for BI-RADS breast density classification.

  25. Predicting Genetic Mutations from Single-Cell Bone Marrow Images in Acute Myeloid Leukemia Using Noise-Robust Deep Learning Models

    eess.IV 2025-06 reject novelty 4.0 of 10

    A deep learning pipeline claims 85% accuracy for classifying four AML-associated mutations from single-cell blood and bone marrow images, but the evaluation has unresolved validity issues.

  26. CASE: Contrastive Activation for Saliency Estimation

    cs.CV 2025-06 conditional novelty 4.0 of 10

    CASE removes gradient components shared with confused classes to produce more class-distinct saliency maps, validated on a top-k overlap diagnostic where many existing methods show class-insensitive behavior.

  27. Automated Fetal Biometry Assessment with Deep Ensembles using Sparse-Sampling of 2D Intrapartum Ultrasound Images

    eess.IV 2025-05 reject novelty 3.0 of 10

    A fetal biometry pipeline using sparse sampling and deep ensembles reports high accuracy, but its published measurement errors conflict with its own results tables.

  28. A Survey on Training-free Open-Vocabulary Semantic Segmentation

    cs.CV 2025-05 conditional novelty 2.0 of 10

    A structured review of over 30 training-free open-vocabulary semantic segmentation methods, organized by whether they rely on CLIP alone, auxiliary visual foundation models, or generative models.

  29. SIM-Net: A Multimodal Fusion Network Using Inferred 3D Object Shape Point Clouds from RGB Images for 2D Classification

    cs.CV 2025-06

Pith tools