Pith. sign in

REVIEW 2 cited by

Skip-Attention: Improving Vision Transformers by Paying Less Attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.02240 v2 pith:7ZUJ6DJM submitted 2023-01-05 cs.CV

classification cs.CV
keywords layersself-attentionacrossattentioncomputationallydenoisingimagemethod
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This work aims to improve the efficiency of vision transformers (ViT). While ViTs use computationally expensive self-attention operations in every layer, we identify that these operations are highly correlated across layers -- a key redundancy that causes unnecessary computations. Based on this observation, we propose SkipAt, a method to reuse self-attention computation from preceding layers to approximate attention at one or more subsequent layers. To ensure that reusing self-attention blocks across layers does not degrade the performance, we introduce a simple parametric function, which outperforms the baseline transformer's performance while running computationally faster. We show the effectiveness of our method in image classification and self-supervised learning on ImageNet-1K, semantic segmentation on ADE20K, image denoising on SIDD, and video denoising on DAVIS. We achieve improved throughput at the same-or-higher accuracy levels in all these tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Surprising Effectiveness of Attention Transfer for Vision Transformers

    cs.LG 2024-11 conditional novelty 7.0 of 10

    Transferring only the attention maps of a pre-trained ViT recovers the accuracy gain of full fine-tuning on ImageNet-1K classification.

  2. Memory Efficient Matting with Adaptive Token Routing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Adaptive token routing with a lightweight refinement branch lets a ViT matting model run on full-resolution high-res images at about 12% of the memory of the ViTMatte baseline with only a small accuracy drop on Compos...

Pith tools