REVIEW 14 cited by
ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Driven by improved architectures and better representation learning frameworks, the field of visual recognition has enjoyed rapid modernization and performance boost in the early 2020s. For example, modern ConvNets, represented by ConvNeXt, have demonstrated strong performance in various scenarios. While these models were originally designed for supervised learning with ImageNet labels, they can also potentially benefit from self-supervised learning techniques such as masked autoencoders (MAE). However, we found that simply combining these two approaches leads to subpar performance. In this paper, we propose a fully convolutional masked autoencoder framework and a new Global Response Normalization (GRN) layer that can be added to the ConvNeXt architecture to enhance inter-channel feature competition. This co-design of self-supervised learning techniques and architectural improvement results in a new model family called ConvNeXt V2, which significantly improves the performance of pure ConvNets on various recognition benchmarks, including ImageNet classification, COCO detection, and ADE20K segmentation. We also provide pre-trained ConvNeXt V2 models of various sizes, ranging from an efficient 3.7M-parameter Atto model with 76.7% top-1 accuracy on ImageNet, to a 650M Huge model that achieves a state-of-the-art 88.9% accuracy using only public training data.
Forward citations
Cited by 14 Pith papers
-
MVGBench: Comprehensive Benchmark for Multi-view Generation Models
MVGBench evaluates multi-view generators through self-consistency of 3D reconstructions and uses this protocol to rank 12 models and build a better one.
-
LaPrune: Controllable Differentiable Sparsity at Million Scale
A differentiable top-k mask layer that enforces an exact selection budget and uses a normalized hardness parameter to interpolate from equal-weight masks to hard binary masks, with saturation theory and million-scale results.
-
AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
A new 5K Arabic meme benchmark with fine-grained hate-type labels and a 66K silver-labeled auxiliary set shows fine-tuned VLMs beat zero-shot ones but all models lag on rare hate categories.
-
GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification
GenSyn10 provides 60k CIFAR-10-aligned images from FLUX.2, HunyuanImage-3.0, and Qwen-Image-2512, showing detectors lose 4–18 points of accuracy on an unseen generator.
-
AI-guided stimuli discovery and generation to optimize facial emotion perception studies in autism
Behavior-aligned ANNs prospectively select diagnostic facial expressions that enlarge autistic–neurotypical emotion-judgment gaps, and GAN-guided transforms of those faces reduce the gaps under phenotype-matched validation.
-
Mahalanobis++: Improving OOD Detection via Feature Normalization
L2-normalizing pre-logit features improves Mahalanobis-based out-of-distribution detection across 44 ImageNet models, reducing the average false positive rate by 7.6 percentage points.
-
Disentanglement and Assessment of Shortcuts in Ophthalmological Retinal Imaging Exams
On mBRSET, disentangling sensitive attributes from DR predictions improved DINOv2 AUROC by 2 points but dropped ConvNeXt V2 by 7 and Swin V2 by 3, with no consistent fairness gain.
-
QuarterMap: Efficient Post-Training Token Pruning for Visual State Space Models
QuarterMap prunes spatial activations before VMamba's four-directional scan and upsamples after, yielding up to 1.11x throughput with under 1% accuracy loss on ImageNet classification.
-
On the rankability of visual embeddings
Visual embeddings from CLIP and other vision encoders encode ordinal attributes along linear directions, recoverable from as few as two extreme reference images, without full supervision.
-
Harmonized Interpretable ECG Waveform Features for Robust Cross-Dataset Clinical Prediction
Harmonized waveform-derived ECG features retain ≥90% of internal AUROC under bidirectional MIMIC↔Alberta transfer while matching vendor machine-measurement models within ~2.5% internally.
-
Cross-Modal Fusion of OCT and OCT angiography enface for Improved Diagnostics of Diabetic Retinopathy
Bidirectional cross-modal attention fusion of OCT B-scans with single-channel enface OCTA (real or diffusion-translated) consistently beats a ConvNeXt V2 OCT-only baseline for binary DR classification across two cohor...
-
AeroLite-MDNet: Lightweight Multi-task Deviation Detection Network for UAV Landing
A YOLO-style detection-plus-segmentation network warns of UAV landing deviations with 98.6% accuracy on a new 9,142-image dataset.
-
WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration
WaveLLDM, a lightweight latent diffusion model with a neural codec, achieves low spectral distortion (LSD 0.48-0.60) on speech restoration but scores far below SOTA on PESQ and STOI.
- SIM-Net: A Multimodal Fusion Network Using Inferred 3D Object Shape Point Clouds from RGB Images for 2D Classification
Discussion (0). Continue with ORCID to comment.