Pith. sign in

REVIEW 2 cited by

SpectFormer: Frequency and Attention is what you need in a Vision Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.06446 v2 pith:DRRLDFPJ submitted 2023-04-13 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords attentioncitespectrallayersmulti-headedtransformertransformersfurther
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision transformers have been applied successfully for image recognition tasks. There have been either multi-headed self-attention based (ViT \cite{dosovitskiy2020image}, DeIT, \cite{touvron2021training}) similar to the original work in textual models or more recently based on spectral layers (Fnet\cite{lee2021fnet}, GFNet\cite{rao2021global}, AFNO\cite{guibas2021efficient}). We hypothesize that both spectral and multi-headed attention plays a major role. We investigate this hypothesis through this work and observe that indeed combining spectral and multi-headed attention layers provides a better transformer architecture. We thus propose the novel Spectformer architecture for transformers that combines spectral and multi-headed attention layers. We believe that the resulting representation allows the transformer to capture the feature representation appropriately and it yields improved performance over other transformer representations. For instance, it improves the top-1 accuracy by 2\% on ImageNet compared to both GFNet-H and LiT. SpectFormer-S reaches 84.25\% top-1 accuracy on ImageNet-1K (state of the art for small version). Further, Spectformer-L achieves 85.7\% that is the state of the art for the comparable base version of the transformers. We further ensure that we obtain reasonable results in other scenarios such as transfer learning on standard datasets such as CIFAR-10, CIFAR-100, Oxford-IIIT-flower, and Standford Car datasets. We then investigate its use in downstream tasks such of object detection and instance segmentation on the MS-COCO dataset and observe that Spectformer shows consistent performance that is comparable to the best backbones and can be further optimized and improved. Hence, we believe that combined spectral and attention layers are what are needed for vision transformers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scalable Event Cloud Network for Event-based Classification

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A frequency-aware network operating on raw-like event clouds matches or beats prior event-based models on nine benchmarks while using roughly 0.1 G MACs, far below frame and voxel baselines.

  2. FSTA-SNN:Frequency-based Spatial-Temporal Attention Module for Spiking Neural Networks

    cs.NE 2024-12 conditional novelty 6.0 of 10

    FSTA-SNN introduces a frequency-based spatial-temporal attention module for spiking neural networks, cutting spike firing rate by about 34 percent and improving accuracy on CIFAR-10/100, ImageNet, and CIFAR10-DVS.

Pith tools