REVIEW 4 cited by
Conv2Former: A Simple Transformer-Style ConvNet for Visual Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper does not attempt to design a state-of-the-art method for visual recognition but investigates a more efficient way to make use of convolutions to encode spatial features. By comparing the design principles of the recent convolutional neural networks ConvNets) and Vision Transformers, we propose to simplify the self-attention by leveraging a convolutional modulation operation. We show that such a simple approach can better take advantage of the large kernels (>=7x7) nested in convolutional layers. We build a family of hierarchical ConvNets using the proposed convolutional modulation, termed Conv2Former. Our network is simple and easy to follow. Experiments show that our Conv2Former outperforms existent popular ConvNets and vision Transformers, like Swin Transformer and ConvNeXt in all ImageNet classification, COCO object detection and ADE20k semantic segmentation.
Forward citations
Cited by 4 Pith papers
-
ClawEnvKit: Automatic Environment Generation for Claw-Like Agents
EVT improves the RMT backbone by using Euclidean-distance attention decay and 1D token grouping, achieving 86.6% top-1 on ImageNet-1K at 384×384 resolution.
-
Conceptualizing Multi-scale Wavelet Attention and Ray-based Encoding for Human-Object Interaction Detection
A wavelet backbone with ray-origin attention encoding improves efficiency and self-comparison accuracy for human-object interaction detection, but remains below the FGAHOI baseline in accuracy despite fewer parameters.
-
All-in-One Image Compression and Restoration
A single learned image codec performs joint compression and all-in-one restoration for haze, snow, rain, and Gaussian noise with one set of weights, preserving clean-image performance.
-
The model is the message: Lightweight convolutional autoencoders applied to noisy imaging data for planetary science and astrobiology
A simple convolutional autoencoder reconstructs planetary images with up to 99% pixel loss, and the author argues its latent space could be a more efficient data product than raw imagery.
Discussion (0). Continue with ORCID to comment.