Pith. sign in

REVIEW 4 cited by

Conv2Former: A Simple Transformer-Style ConvNet for Visual Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.11943 v1 pith:EC7AYF2R submitted 2022-11-22 cs.CV

classification cs.CV
keywords convolutionalconv2formerconvnetssimpledesignmodulationrecognitiontransformers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper does not attempt to design a state-of-the-art method for visual recognition but investigates a more efficient way to make use of convolutions to encode spatial features. By comparing the design principles of the recent convolutional neural networks ConvNets) and Vision Transformers, we propose to simplify the self-attention by leveraging a convolutional modulation operation. We show that such a simple approach can better take advantage of the large kernels (>=7x7) nested in convolutional layers. We build a family of hierarchical ConvNets using the proposed convolutional modulation, termed Conv2Former. Our network is simple and easy to follow. Experiments show that our Conv2Former outperforms existent popular ConvNets and vision Transformers, like Swin Transformer and ConvNeXt in all ImageNet classification, COCO object detection and ADE20k semantic segmentation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ClawEnvKit: Automatic Environment Generation for Claw-Like Agents

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    EVT improves the RMT backbone by using Euclidean-distance attention decay and 1D token grouping, achieving 86.6% top-1 on ImageNet-1K at 384×384 resolution.

  2. Conceptualizing Multi-scale Wavelet Attention and Ray-based Encoding for Human-Object Interaction Detection

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A wavelet backbone with ray-origin attention encoding improves efficiency and self-comparison accuracy for human-object interaction detection, but remains below the FGAHOI baseline in accuracy despite fewer parameters.

  3. All-in-One Image Compression and Restoration

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A single learned image codec performs joint compression and all-in-one restoration for haze, snow, rain, and Gaussian noise with one set of weights, preserving clean-image performance.

  4. The model is the message: Lightweight convolutional autoencoders applied to noisy imaging data for planetary science and astrobiology

    astro-ph.EP 2025-07 conditional novelty 4.0 of 10

    A simple convolutional autoencoder reconstructs planetary images with up to 99% pixel loss, and the author argues its latent space could be a more efficient data product than raw imagery.

Pith tools