REVIEW 29 cited by
A ConvNet for the 2020s
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The "Roaring 20s" of visual recognition began with the introduction of Vision Transformers (ViTs), which quickly superseded ConvNets as the state-of-the-art image classification model. A vanilla ViT, on the other hand, faces difficulties when applied to general computer vision tasks such as object detection and semantic segmentation. It is the hierarchical Transformers (e.g., Swin Transformers) that reintroduced several ConvNet priors, making Transformers practically viable as a generic vision backbone and demonstrating remarkable performance on a wide variety of vision tasks. However, the effectiveness of such hybrid approaches is still largely credited to the intrinsic superiority of Transformers, rather than the inherent inductive biases of convolutions. In this work, we reexamine the design spaces and test the limits of what a pure ConvNet can achieve. We gradually "modernize" a standard ResNet toward the design of a vision Transformer, and discover several key components that contribute to the performance difference along the way. The outcome of this exploration is a family of pure ConvNet models dubbed ConvNeXt. Constructed entirely from standard ConvNet modules, ConvNeXts compete favorably with Transformers in terms of accuracy and scalability, achieving 87.8% ImageNet top-1 accuracy and outperforming Swin Transformers on COCO detection and ADE20K segmentation, while maintaining the simplicity and efficiency of standard ConvNets.
Forward citations
Cited by 29 Pith papers
-
AIMIP Phase 1: systematic evaluations of AI weather and climate models
Under one protocol, most AI climate models reproduce historical climatology and ENSO response as well as a CMIP6 model, but some underestimate warming trends and all diverge on +2/+4K SST experiments.
-
FourCastNet 3: A geometric approach to probabilistic machine-learning weather forecasting at scale
A purely convolutional, spherical-geometry weather model trained with a combined spatial and spectral CRPS loss delivers GenCast-level skill, IFS-beating accuracy, and stable spectra out to 60 days.
-
ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling
A single-step IMLE generator with per-stage supervision and a robust loss reports FID 2.56 on ImageNet-256 by filtering ~5% of samples at test time.
-
C3DIR: A Deep Learning 3-Dimensional Cloud Property Retrieval Scheme for Passive Satellite Imagers
C3DIR is a single multi-sensor deep-learning model that retrieves 3-D ice, liquid, and rain water content from passive imagers using voxel-to-voxel collocation with active-sensor profiles.
-
Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System
A continuously refreshed, incentive-driven deepfake detector beats static detectors on in-the-wild benchmarks and improves on post-export AI-generated media.
-
Geographically-Weighted Weakly Supervised Bayesian High-Resolution Transformer for 200m Resolution Pan-Arctic Sea Ice Concentration Mapping and Uncertainty Estimation using Sentinel-1, RCM, and AMSR2 Data
A 200m pan-Arctic sea ice concentration model using a Bayesian Transformer with a geographically-weighted weakly supervised loss; reported 0.70 feature detection accuracy and R²=0.90 vs ASI.
-
Model soups need only one ingredient
A single checkpoint, edited by splitting each layer's update with SVD and reweighting the high- and low-energy parts, reaches soup-level OOD robustness without multi-model training.
-
Khana: A Comprehensive Indian Cuisine Dataset
Khana is an Indian cuisine image dataset with 131K images across 80 dish categories and baseline classification accuracies up to 86.7% top-1 with ConvNeXT.
-
RRTO: A High-Performance Transparent Offloading System for Model Inference in Mobile Edge Computing
RRTO identifies static inference operator sequences from CUDA call logs alone and replays them on an edge GPU, cutting transparent-offloading communication to 11 RPCs per inference instead of thousands, with performan...
-
Beginning with You: Perceptual-Initialization Improves Vision-Language Representation and Alignment
Perceptually initializing a CLIP vision encoder with NIGHTS triplet judgments before YFCC15M contrastive training improves zero-shot accuracy and retrieval over an identical random-start baseline.
-
DiffEx: Explaining a Classifier with Diffusion Models to Identify Microscopic Cellular Variations
DiffEx builds a classifier-aware latent space with a diffusion model, finds contrastive directions in it, and ranks them to produce visual explanations that reveal cellular phenotype changes.
-
Small, Bias-Free, Blind and Convolutional Denoiser: A compact ConvNeXt U-Net for blind Gaussian color-image denoising
BF-ConvUNeXt, a 0.82M-parameter bias-free ConvNeXt U-Net, is degree-1 homogeneous and matches DnCNN/FFDNet in blind color denoising, extrapolating smoothly beyond its training noise range.
-
REVELIO -- Universal Multimodal Task Load Estimation for Cross-Domain Generalization
A new cognitive-load dataset with driving, gaming, and n-back tasks shows multimodal models beat unimodal ones, but cross-domain transfer remains poor.
-
StixelNExT++: Lightweight Monocular Scene Segmentation and Representation for Collective Perception
A monocular neural network predicts 3D Stixels directly from RGB images in about 10 ms, with a self-defined Waymo evaluation showing competitive performance within 30 m.
-
A Controlled Visual-Backbone Benchmark for Multimodal Short-Term Solar Irradiance Forecasting
Under one fixed multimodal 10-min irradiance forecaster, VMamba-S and Swin-B nearly tie best Folsom RMSE (~65.4 W/m²) over smart persistence, while NREL’s tiny matched split favors persistence.
-
Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning
On a large World of Tanks dataset, a shared multi-task model with equal weighting or PCGrad outperforms single-task models on average, and task/map pre-training helps most in low-data regimes.
-
Revisiting Simple Baselines for In-The-Wild Deepfake Detection
Finetuned CLIP-pretrained ConvNeXt-base and ViT-b32 classifiers reach 81% accuracy on Deepfake-Eval-2024, within noise of the leading commercial detector's 82%.
-
Classifying Mitotic Figures in the MIDOG25 Challenge with Deep Ensemble Learning and Rule Based Refinement
A ConvNeXt ensemble achieves 84% balanced accuracy on MIDOG25 atypical mitotic figure classification, while a rule-based refinement module trades sensitivity for specificity.
-
Beyond Linear Bottlenecks: Spline-Based Knowledge Distillation for Culturally Diverse Art Style Classification
Replacing MLP projection heads with Kolmogorov-Arnold Network heads in a dual-teacher self-supervised art-style classifier yields Top-1 accuracy gains of around 0.2 to 1.0 percentage points on WikiArt and Pandora18k, ...
-
Faithful, Interpretable Chest X-ray Diagnosis with Anti-Aliased B-cos Networks
Combining B-cos networks with anti-aliasing pooling (FLC or BlurPool) reduces grid artifacts in chest X-ray explanation maps while keeping diagnostic accuracy close to baseline networks.
-
Applying multimodal learning to Classify transient Detections Early (AppleCiDEr) I: Data set, methods, and infrastructure
AppleCiDEr combines photometry, images, metadata, and spectra in one deep learning pipeline to classify ZTF transients and variable stars, with high accuracy on common classes but poor performance on tidal disruption events.
-
First-of-its-kind AI model for bioacoustic detection using a lightweight associative memory Hopfield neural network
A Hopfield neural network trained on two bat calls classifies 10,384 recordings in 5.4 seconds with claimed accuracy up to 80%, though the headline numbers hinge on removing ambiguous calls.
-
Bridging the gap in FER: addressing age bias in deep learning
Age-aware loss reweighting and age-conditioned training reduce elderly recognition errors in facial expression models, even when age labels are automatically estimated, but the evidence rests on a small elderly benchmark.
-
Comparison of ConvNeXt and Vision-Language Models for Breast Density Assessment in Screening Mammography
Fine-tuned ConvNeXt achieves 0.73 accuracy and 0.78 F1, beating BioMedCLIP linear probe (0.64/0.63) and zero-shot (0.47/0.31) for BI-RADS breast density classification.
-
Predicting Genetic Mutations from Single-Cell Bone Marrow Images in Acute Myeloid Leukemia Using Noise-Robust Deep Learning Models
A deep learning pipeline claims 85% accuracy for classifying four AML-associated mutations from single-cell blood and bone marrow images, but the evaluation has unresolved validity issues.
-
CASE: Contrastive Activation for Saliency Estimation
CASE removes gradient components shared with confused classes to produce more class-distinct saliency maps, validated on a top-k overlap diagnostic where many existing methods show class-insensitive behavior.
-
Automated Fetal Biometry Assessment with Deep Ensembles using Sparse-Sampling of 2D Intrapartum Ultrasound Images
A fetal biometry pipeline using sparse sampling and deep ensembles reports high accuracy, but its published measurement errors conflict with its own results tables.
-
A Survey on Training-free Open-Vocabulary Semantic Segmentation
A structured review of over 30 training-free open-vocabulary semantic segmentation methods, organized by whether they rely on CLIP alone, auxiliary visual foundation models, or generative models.
- SIM-Net: A Multimodal Fusion Network Using Inferred 3D Object Shape Point Clouds from RGB Images for 2D Classification
Discussion (0). Continue with ORCID to comment.