Pith. sign in

REVIEW 2 cited by

Refiner: Refining Self-attention for Vision Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.03714 v1 pith:5HHPNCLS submitted 2021-06-07 cs.CV

classification cs.CV
keywords refinervitsattentionself-attentionmapsworksaccuracyaggregated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision Transformers (ViTs) have shown competitive accuracy in image classification tasks compared with CNNs. Yet, they generally require much more data for model pre-training. Most of recent works thus are dedicated to designing more complex architectures or training methods to address the data-efficiency issue of ViTs. However, few of them explore improving the self-attention mechanism, a key factor distinguishing ViTs from CNNs. Different from existing works, we introduce a conceptually simple scheme, called refiner, to directly refine the self-attention maps of ViTs. Specifically, refiner explores attention expansion that projects the multi-head attention maps to a higher-dimensional space to promote their diversity. Further, refiner applies convolutions to augment local patterns of the attention maps, which we show is equivalent to a distributed local attention features are aggregated locally with learnable kernels and then globally aggregated with self-attention. Extensive experiments demonstrate that refiner works surprisingly well. Significantly, it enables ViTs to achieve 86% top-1 classification accuracy on ImageNet with only 81M parameters.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. STeInFormer: Spatial-Temporal Interaction Transformer Architecture for Remote Sensing Change Detection

    cs.CV 2024-12 conditional novelty 6.0 of 10

    STeInFormer enhances remote sensing change detection by interacting bi-temporal features during feature extraction and using fixed DCT frequency components as a parameter-light token mixer.

  2. Unified Local and Global Attention Interaction Modeling for Vision Transformers

    cs.CV 2024-12 conditional novelty 4.0 of 10

    Adding local and global token interactions before self-attention, via an aggressive convolution-pooling block and a concept-attention block, improves RetinaNet object detection mAP over non-pretrained ViT, Swin, and D...

Pith tools