Pith. sign in

REVIEW 14 cited by

ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.00808 v1 pith:LVEW4VQE submitted 2023-01-02 cs.CV

classification cs.CV
keywords convnextlearningperformanceconvnetsimagenetmaskedmodelvarious
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Driven by improved architectures and better representation learning frameworks, the field of visual recognition has enjoyed rapid modernization and performance boost in the early 2020s. For example, modern ConvNets, represented by ConvNeXt, have demonstrated strong performance in various scenarios. While these models were originally designed for supervised learning with ImageNet labels, they can also potentially benefit from self-supervised learning techniques such as masked autoencoders (MAE). However, we found that simply combining these two approaches leads to subpar performance. In this paper, we propose a fully convolutional masked autoencoder framework and a new Global Response Normalization (GRN) layer that can be added to the ConvNeXt architecture to enhance inter-channel feature competition. This co-design of self-supervised learning techniques and architectural improvement results in a new model family called ConvNeXt V2, which significantly improves the performance of pure ConvNets on various recognition benchmarks, including ImageNet classification, COCO detection, and ADE20K segmentation. We also provide pre-trained ConvNeXt V2 models of various sizes, ranging from an efficient 3.7M-parameter Atto model with 76.7% top-1 accuracy on ImageNet, to a 650M Huge model that achieves a state-of-the-art 88.9% accuracy using only public training data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 65 citations worldwide. Full citation record

  1. MVGBench: Comprehensive Benchmark for Multi-view Generation Models

    cs.GR 2025-06 conditional novelty 7.0 of 10

    MVGBench evaluates multi-view generators through self-consistency of 3D reconstructions and uses this protocol to rank 12 models and build a better one.

  2. LaPrune: Controllable Differentiable Sparsity at Million Scale

    cs.LG 2026-08 accept novelty 6.0 of 10

    A differentiable top-k mask layer that enforces an exact selection budget and uses a normalized hardness parameter to interpolate from equal-weight masks to hard binary masks, with saturation theory and million-scale results.

  3. AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A new 5K Arabic meme benchmark with fine-grained hate-type labels and a 66K silver-labeled auxiliary set shows fine-tuned VLMs beat zero-shot ones but all models lag on rare hate categories.

  4. GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification

    cs.CV 2026-07 conditional novelty 6.0 of 10

    GenSyn10 provides 60k CIFAR-10-aligned images from FLUX.2, HunyuanImage-3.0, and Qwen-Image-2512, showing detectors lose 4–18 points of accuracy on an unseen generator.

  5. AI-guided stimuli discovery and generation to optimize facial emotion perception studies in autism

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Behavior-aligned ANNs prospectively select diagnostic facial expressions that enlarge autistic–neurotypical emotion-judgment gaps, and GAN-guided transforms of those faces reduce the gaps under phenotype-matched validation.

  6. Mahalanobis++: Improving OOD Detection via Feature Normalization

    cs.LG 2025-05 conditional novelty 6.0 of 10

    L2-normalizing pre-logit features improves Mahalanobis-based out-of-distribution detection across 44 ImageNet models, reducing the average false positive rate by 7.6 percentage points.

  7. Disentanglement and Assessment of Shortcuts in Ophthalmological Retinal Imaging Exams

    cs.CV 2025-07 conditional novelty 5.0 of 10

    On mBRSET, disentangling sensitive attributes from DR predictions improved DINOv2 AUROC by 2 points but dropped ConvNeXt V2 by 7 and Swin V2 by 3, with no consistent fairness gain.

  8. QuarterMap: Efficient Post-Training Token Pruning for Visual State Space Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    QuarterMap prunes spatial activations before VMamba's four-directional scan and upsamples after, yielding up to 1.11x throughput with under 1% accuracy loss on ImageNet classification.

  9. On the rankability of visual embeddings

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Visual embeddings from CLIP and other vision encoders encode ordinal attributes along linear directions, recoverable from as few as two extreme reference images, without full supervision.

  10. Harmonized Interpretable ECG Waveform Features for Robust Cross-Dataset Clinical Prediction

    cs.LG 2026-07 conditional novelty 4.0 of 10

    Harmonized waveform-derived ECG features retain ≥90% of internal AUROC under bidirectional MIMIC↔Alberta transfer while matching vendor machine-measurement models within ~2.5% internally.

  11. Cross-Modal Fusion of OCT and OCT angiography enface for Improved Diagnostics of Diabetic Retinopathy

    eess.IV 2026-07 conditional novelty 4.0 of 10

    Bidirectional cross-modal attention fusion of OCT B-scans with single-channel enface OCTA (real or diffusion-translated) consistently beats a ConvNeXt V2 OCT-only baseline for binary DR classification across two cohor...

  12. AeroLite-MDNet: Lightweight Multi-task Deviation Detection Network for UAV Landing

    cs.RO 2025-06 conditional novelty 4.0 of 10

    A YOLO-style detection-plus-segmentation network warns of UAV landing deviations with 98.6% accuracy on a new 9,142-image dataset.

  13. WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration

    cs.SD 2025-08 conditional novelty 3.0 of 10

    WaveLLDM, a lightweight latent diffusion model with a neural codec, achieves low spectral distortion (LSD 0.48-0.60) on speech restoration but scores far below SOTA on PESQ and STOI.

  14. SIM-Net: A Multimodal Fusion Network Using Inferred 3D Object Shape Point Clouds from RGB Images for 2D Classification

    cs.CV 2025-06

Pith tools