REVIEW 4 major objections 5 minor 56 references
Combining Aggregated Attention and Transformer Architecture for Accurate and Efficient Performance of Spiking Neural Networks
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read SAFormer claims top SNN accuracy using attention without a value matrix.
desk verdict A plausible spiking attention variant with a real reproducibility gap: the forward pass is underspecified where reduced-length attention meets full-length depthwise features, so the headline numbers can't be verified as written. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Spike Aggregated Self-Attention (SASA) mechanism. Query and key are obtained by applying adaptive-average-pooling aggregation $\mathrm{AG}(\cdot)$ to floating-point projections, then binarized through spike neurons, yielding matrices $Q, K \in \mathbb{R}^{T \times n \times D}$ with $n \ll N$. The attention map is $\mathrm{SASA}'(Q, K) = \mathrm{SN}(\mathrm{SUM}_c(Q \otimes K))$, a Hadamard product followed by column-wise summation with no value matrix. In parallel, the full-resolution key projection passes through a depthwise convolution to produce $K_D$, and the final output is $\mathrm{BN}(\mathrm{Linear}(\mathrm{SN}(K_D \oplus \mathrm{SASA}'(Q, K))))$. The machinery works by using aggregation to reduce sparsity and computational cost, using key-only attention to eliminate the value pathway, and using the depthwise convolution to restore feature diversity.
What would settle it
Run a concrete forward pass through Eq. 17 with $n < N$ and inspect the shapes: if the implementation must broadcast or reshape $\mathrm{SASA}'(Q, K)$ to length $N$ to add $K_D$, the equations as written are incomplete. A more decisive test is to remove the SASA attention term entirely, keeping only the depthwise-convolved key path, and measure accuracy; if accuracy stays near the reported levels, the attention mechanism is not carrying the claimed load. A further check is to replace the learned $W_Q$ and $W_K$ with fixed random binary projections of the same sparsity and observe whether the reported accuracy persists.
Extended reading notes
Core claim
SAFormer establishes that a spiking self-attention mechanism can omit the value matrix entirely and compute attention from aggregated query and key spikes, then add a depthwise-convolved key feature to the attention map; this SASA mechanism is claimed to combine linear complexity with accuracy that exceeds existing spiking Transformers. The paper reports accuracy of 95.8% on CIFAR-10, 79.07% on CIFAR-100, 81.3% on CIFAR10-DVS, and 98.3% on DVS128-Gesture, with estimated theoretical energies of 0.49 mJ, 0.58 mJ, 1.67 mJ, and 1.23 mJ respectively. The core discovery is that attention weights alone, modulated by features drawn from the key matrix, can carry enough information for classification in the sparse spike domain, provided the attention map is computed on denser downsampled aggregates and supplemented by local depthwise features.
Load-bearing premise
The load-bearing premise is that the attention map computed from the downsampled $n$-length query and key can be combined by element-wise addition with the full-resolution $N$-length depthwise-convolved key features, and that attention weights alone carry enough information for classification; the paper does not specify how the $n$ and $N$ shapes are reconciled, so if that reduction or broadcast fails, or the information content is insufficient, the architecture does not work as stated.
Editorial extensions
If this is right
- Attention in spiking Transformers can be made linear in sequence length without giving up accuracy, by computing attention weights on downsampled query and key matrices.
- Removing the value matrix and replacing value-weighted summation with key-derived depthwise features lowers theoretical energy, with reported reductions of about 90% versus Spikformer and about 6% versus S-Transformer on CIFAR-10.
- The architecture remains accurate even with an extremely small aggregated sequence length: at $n=1$, accuracy is comparable to the S-Transformer baseline, suggesting the attention signal can be highly compressed.
- SAFormer reaches state-of-the-art accuracy with only 2-4 encoder blocks and 4-16 time steps, which supports deployment on resource-constrained hardware.
Reading between the lines
- The paper does not specify how the $n$-length attention map in Eq. 14 is reconciled with the $N$-length depthwise-convolved feature in Eq. 17; if the implementation broadcasts or reshapes the attention map back to length $N$, that hidden operation is essential to the reported efficiency and should be verified in code.
- All reported energies are theoretical estimates from Eq. 19, not hardware measurements; actual gains on neuromorphic chips could differ, so a measured comparison would be the natural next test.
- The success of dropping the value matrix suggests spiking Transformers may benefit more from feature-diversity mechanisms such as convolutions than from learned value projections, which could be explored on larger-scale vision or language tasks.
- The robustness at $n=1$ hints that SASA might be capturing a global spike-rate statistic rather than fine-grained spatial attention; an ablation that replaces the attention term with a fixed statistic could clarify what is actually doing the work.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SAFormer, a spiking Transformer architecture whose core Spike Aggregated Self-Attention (SASA) mechanism computes attention from downsampled query and key spike matrices via a Hadamard product and column-wise summation, omits the value matrix, and augments the attention map with a depthwise convolution (DWC) branch. The authors report state-of-the-art accuracy on CIFAR-10 (95.8%), CIFAR-100 (79.07%), CIFAR10-DVS (81.3%), and DVS128-Gesture (98.3%), with claimed theoretical energies of 0.49 mJ and 0.58 mJ on CIFAR-10 and CIFAR-100, respectively, together with ablations, time-step studies, and complexity/energy analyses.
Significance. If the claims hold, SAFormer offers a plausible route to linear-complexity spiking attention with competitive accuracy and low energy, and the comparison against external Spikformer and S-Transformer baselines using standard Horowitz coefficients is a strength. The ablation in Table 3 supports the qualitative benefit of both the aggregation function and the DWC module, and the conceptual idea of removing the value matrix while enriching features with depthwise convolution is clearly motivated. However, the forward pass as written cannot be reimplemented because of a shape inconsistency between the downsampled attention output and the full-resolution DWC branch, and the accuracy and energy evidence is presented without variance or absolute baseline energy values. No code or weights are released, so the central 'outperforms' claim is not independently verifiable in the current manuscript.
major comments (4)
- [§3.4, Eqs. (12)–(17)] The forward pass is dimensionally undefined. Equations (12) and (13) set Q,K ∈ R^{T×n×D}; Equation (14) then defines SASA'(Q,K) = SN(SUMc(Q⊗K)), whose output is either R^{T×n} or R^{T×n×D} depending on the interpretation of 'column-wise summation'. Equation (17) requires SASA(Q,K) = BN(Linear(SN(KD ⊕ SASA'(Q,K)))) with KD ∈ R^{T×N×D} and ⊕ declared element-wise addition. No reduction, broadcast, upsampling, or reshaping rule that reconciles n and N is specified anywhere in §3.4. Because every encoder block feeds this output into the MLP and residual stream, the architecture cannot be instantiated from the manuscript alone, and the reported accuracy and energy numbers cannot be checked. Please specify the exact output shape of SUMc and the precise rule for combining the two branches.
- [Tables 1 and 2; Figure 2 caption] Figure 2 states that 'the averages and standard deviations are calculated over three independent runs', but Tables 1 and 2 report only single accuracy values with no variance or significance information. The claimed improvements over Spikformer (0.3% on CIFAR-10, 0.87% on CIFAR-100, 0.4% on CIFAR10-DVS) and S-Transformer (0.2% on CIFAR-10, 0.67% on CIFAR-100, while SAFormer is actually 1.0% lower on DVS128-Gesture) are small enough that run-to-run variation could change the reported ranking. Report mean±standard deviation or confidence intervals for at least the compared Transformer models, and state whether the margins are statistically meaningful.
- [§4.1, Eqs. (19)–(20)] The energy comparison is not auditable as presented. No table reports absolute energy consumption for the baseline models; the text gives only relative reductions such as '90.49% reduction' and '5.8% reduction', which cannot be checked without the underlying mJ values. Furthermore, Equation (20) defines SP as the sum of spike-based operations over convolutional, fully connected, and SASA layers, but it does not show whether the DWC branch's depthwise convolutions are included in SOP_SASA or in the layer-wise counts, and no per-layer SOP values are given. Provide a complete energy table with per-model absolute mJ values and a detailed SOP accounting that explicitly includes or excludes every operation in the forward pass.
- [§4.5 and Appendix A.1, Table A.4] The reduction from N to n is central to the claimed linear complexity and energy savings, but the paper never states the value of n used for the main results in Tables 1 and 2; Figure 6 varies n over a range, yet the configuration that produces the headline accuracy is not identified. Without this value, the O(nD) complexity and the numerical entries in Table A.5 cannot be quantified. Additionally, Table A.4's O(nD) excludes the DWC module, and the assertion in Appendix A.1 that DWC 'does not significantly affect the linear time complexity' is an unquantified assumption; the full-resolution depthwise convolution has a cost of O(T·N·D·k) that should be reported explicitly.
minor comments (5)
- [Figure 4 caption] The word 'colume' should be 'column'; please also clarify how the attention maps are pooled across the T dimension for visualization.
- [Table 3] The header 'w\o D' is likely a typographical artifact for 'w/o D' (without DWC); please correct the formatting for readability.
- [Appendix A.1] The sentence 'although convolution operations are generally nonlinear' is incorrect as stated: convolution is a linear operation, and depthwise convolution is also linear. If the intended meaning concerns nonlinear activations or the nonlinear behavior of spiking layers, please rephrase.
- [§4.4] The sentence 'providing that the performance decline is due to the smaller number of categories' should read 'proving' or 'indicating'; please reword.
- [Eq. (20)] The symbol SP is used both as the summed spike-operation count and as the upper index P in the sum; please define the notation to avoid ambiguity.
Circularity Check
No circularity: SAFormer's accuracy and energy claims rest on external benchmark comparisons and published Horowitz energy constants, not on fitted parameters or a self-citation chain.
full rationale
The paper does not contain any load-bearing step that reduces to its own inputs by construction. The SASA attention formula SASA'(Q,K)=SN(SUMc(Q⊗K)) is introduced as a definition of the proposed mechanism, not derived from the mechanism itself, and the subsequent DWC combination in Eq. 17 is a design choice rather than a predicted consequence. Energy consumption is computed via Eq. 19 using externally published Horowitz coefficients (EMAC=4.6 pJ, EAC=0.9 pJ) and measured or estimated firing rates; no parameter is fitted to a subset of the accuracy or energy data and then renamed as a prediction. The reported accuracies are empirical comparisons against Spikformer, S-Transformer, and other external baselines, so the central performance claim is self-contained benchmark evidence. The authors' citations to prior Transformer/SNN work are used for architectural conventions and baseline values, not as an unverified uniqueness theorem or as the sole support for the paper's own results. The dimensional ambiguity between KD∈R^{T×N×D} and SASA' derived from Q,K∈R^{T×n×D} in Eq. 17 is a specification and reproducibility concern, but it is not a circularity: it does not show that any claimed result is equivalent to its input by definition or by fit.
Assumptions & free parameters
free parameters (2)
- n (aggregated sequence length) =
chosen by hand; no single fitted value
- simulation time step T =
T=4 for CIFAR, T=16 for DVS datasets
assumptions (5)
- domain assumption The LIF neuron model with Heaviside firing and reset adequately models the spiking dynamics for the reported accuracy results.
- domain assumption Adaptive average pooling applied to floating-point query/key features before spike encoding preserves the information needed for attention and reduces sparsity without loss of accuracy.
- domain assumption The value matrix is redundant in spiking self-attention because sparse activation makes it largely redundant.
- domain assumption Energy per MAC (4.6 pJ) and per AC (0.9 pJ) from Horowitz's 45nm model accurately reflects the energy cost of all operations in SAFormer and the compared baselines.
- ad hoc to paper Including the DWC module does not significantly change the linear time complexity of SASA.
Cite this review
Pith. "Pith review of Combining Aggregated Attention and Transformer Architecture for Accurate and Efficient Performance of Spiking Neural Networks." pith.science (2026). https://pith.science/paper/V6XG6RCZ
@misc{pith2026241213553,
author = {Pith},
title = {Pith review of: Combining Aggregated Attention and Transformer Architecture for Accurate and Efficient Performance of Spiking Neural Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/V6XG6RCZ}},
note = {Machine review of arXiv:2412.13553}
}
read the original abstract
Spiking Neural Networks have attracted significant attention in recent years due to their distinctive low-power characteristics. Meanwhile, Transformer models, known for their powerful self-attention mechanisms and parallel processing capabilities, have demonstrated exceptional performance across various domains, including natural language processing and computer vision. Despite the significant advantages of both SNNs and Transformers, directly combining the low-power benefits of SNNs with the high performance of Transformers remains challenging. Specifically, while the sparse computing mode of SNNs contributes to reduced energy consumption, traditional attention mechanisms depend on dense matrix computations and complex softmax operations. This reliance poses significant challenges for effective execution in low-power scenarios. Given the tremendous success of Transformers in deep learning, it is a necessary step to explore the integration of SNNs and Transformers to harness the strengths of both. In this paper, we propose a novel model architecture, Spike Aggregation Transformer (SAFormer), that integrates the low-power characteristics of SNNs with the high-performance advantages of Transformer models. The core contribution of SAFormer lies in the design of the Spike Aggregated Self-Attention (SASA) mechanism, which significantly simplifies the computation process by calculating attention weights using only the spike matrices query and key, thereby effectively reducing energy consumption. Additionally, we introduce a Depthwise Convolution Module (DWC) to enhance the feature extraction capabilities, further improving overall accuracy. We evaluated and demonstrated that SAFormer outperforms state-of-the-art SNNs in both accuracy and energy consumption, highlighting its significant advantages in low-power and high-performance computing.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Spike-train level backpropagation for training deep recurrent spiking neural networks
Wenrui Zhang and Peng Li. Spike-train level backpropagation for training deep recurrent spiking neural networks. Advances in neural information processing systems, 32, 2019
work page 2019
-
[2]
Spiking deep residual networks
Yangfan Hu, Huajin Tang, and Gang Pan. Spiking deep residual networks. IEEE Transactions on Neural Networks and Learning Systems, 34(8):5200–5205, 2021
work page 2021
-
[3]
Dynamic spiking graph neural net- works
Nan Yin, Mengzhu Wang, Zhenghan Chen, Giulia De Masi, Huan Xiong, and Bin Gu. Dynamic spiking graph neural net- works. In Proceedings of the AAAI Conference on Artificial In- telligence, volume 38, pages 16495–16503, 2024
work page 2024
-
[4]
Spikegpt: Generative pre-trained language model with spiking neural networks
Rui-Jie Zhu, Qihang Zhao, Guoqi Li, and Jason K Eshraghian. Spikegpt: Generative pre-trained language model with spiking neural networks. arXiv preprint arXiv:2302.13939, 2023
arXiv 2023
-
[5]
Spikformer: When spiking neural network meets transformer
Zhaokun Zhou, Yuesheng Zhu, Chao He, Yaowei Wang, Shuicheng Y AN, Yonghong Tian, and Li Yuan. Spikformer: When spiking neural network meets transformer. In The Eleventh International Conference on Learning Representa- tions, 2023
work page 2023
-
[6]
Man Yao, Jiakui Hu, Zhaokun Zhou, Li Yuan, Yonghong Tian, Bo Xu, and Guoqi Li. Spike-driven transformer. Advances in Neural Information Processing Systems, 36, 2024
work page 2024
-
[7]
Spikingformer: Spike-driven residual learning for transformer-based spiking neural network
Chenlin Zhou, Liutao Yu, Zhaokun Zhou, Zhengyu Ma, Han Zhang, Huihui Zhou, and Yonghong Tian. Spikingformer: Spike-driven residual learning for transformer-based spiking neural network. arXiv preprint arXiv:2304.11954, 2023
arXiv 2023
-
[8]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polo- sukhin. Attention is all you need. Advances in neural informa- tion processing systems, 30, 2017
work page 2017
Show all 56 references
-
[9]
Qkformer: Hierarchical spiking transformer using qk attention
Chenlin Zhou, Han Zhang, Zhaokun Zhou, Liutao Yu, Liwei Huang, Xiaopeng Fan, Li Yuan, Zhengyu Ma, Huihui Zhou, and Yonghong Tian. Qkformer: Hierarchical spiking transformer using qk attention. arXiv preprint arXiv:2403.16552, 2024
2024 arXiv
-
[10]
Masked spiking transformer
Ziqing Wang, Yuetong Fang, Jiahang Cao, Qiang Zhang, Zhon- grui Wang, and Renjing Xu. Masked spiking transformer. In Proceedings of the IEEE /CVF International Conference on Computer Vision, pages 1761–1771, 2023
2023
-
[11]
Attention-free spikformer: Mixing spike sequences with simple linear transforms
Qingyu Wang, Duzhen Zhang, Tielin Zhang, and Bo Xu. Attention-free spikformer: Mixing spike sequences with simple linear transforms. arXiv preprint arXiv:2308.02557, 2023
2023 arXiv
-
[12]
An image is worth 16x16 words: Transformers for image recog- nition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa De- hghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recog- nition at scale. arXiv preprint a...
2010 arXiv
-
[13]
Cf-vit: A general coarse-to- fine method for vision transformer
Mengzhao Chen, Mingbao Lin, Ke Li, Yunhang Shen, Yongjian Wu, Fei Chao, and Rongrong Ji. Cf-vit: A general coarse-to- fine method for vision transformer. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 37, pages 7042– 7052, 2023
2023
-
[14]
Scaling vision transformers to 22 billion parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Peter Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, et al. Scaling vision transformers to 22 billion parameters. In Inter- national Conference on Machine L...
2023
-
[15]
Fq-vit: Post-training quantization for fully quantized vi- sion transformer
Yang Lin, Tianyu Zhang, Peiqin Sun, Zheng Li, and Shuchang Zhou. Fq-vit: Post-training quantization for fully quantized vi- sion transformer. arXiv preprint arXiv:2111.13824, 2021
2021 arXiv
-
[16]
Top-down visual attention from analysis by synthesis
Baifeng Shi, Trevor Darrell, and Xin Wang. Top-down visual attention from analysis by synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition, pages 2102–2112, 2023
2023
-
[17]
Riformer: Keep your vision backbone effective but removing token mixer
Jiahao Wang, Songyang Zhang, Yong Liu, Taiqiang Wu, Yujiu Yang, Xihui Liu, Kai Chen, Ping Luo, and Dahua Lin. Riformer: Keep your vision backbone effective but removing token mixer. In Proceedings of the IEEE /CVF Conference on Computer Vi- sion and Pattern Recognition, pages ...
2023
-
[18]
A closer look at self-supervised lightweight vision transformers
Shaoru Wang, Jin Gao, Zeming Li, Xiaoqin Zhang, and Weim- ing Hu. A closer look at self-supervised lightweight vision transformers. In International Conference on Machine Learn- ing, pages 35624–35641. PMLR, 2023
2023
-
[19]
E fficientvit: Memory e fficient vision transformer with cascaded group attention
Xinyu Liu, Houwen Peng, Ningxin Zheng, Yuqing Yang, Han Hu, and Yixuan Yuan. E fficientvit: Memory e fficient vision transformer with cascaded group attention. In Proceedings of the IEEE /CVF Conference on Computer Vision and Pattern Recognition, pages 14420–14430, 2023
2023
-
[20]
Spikeformer: A novel architecture for training high-performance low-latency spiking neural network
Yudong Li, Yunlin Lei, and Xu Yang. Spikeformer: A novel architecture for training high-performance low-latency spiking neural network. arXiv preprint arXiv:2211.10686, 2022
2022 arXiv
-
[21]
Deep residual learning in spik- ing neural networks
Wei Fang, Zhaofei Yu, Yanqi Chen, Tiejun Huang, Timoth ´ee Masquelier, and Yonghong Tian. Deep residual learning in spik- ing neural networks. Advances in Neural Information Process- ing Systems, 34:21056–21069, 2021
2021
-
[22]
Enhanc- ing the performance of transformer-based spiking neural net- works by improved downsampling with precise gradient back- propagation
Chenlin Zhou, Han Zhang, Zhaokun Zhou, Liutao Yu, Zhengyu Ma, Huihui Zhou, Xiaopeng Fan, and Yonghong Tian. Enhanc- ing the performance of transformer-based spiking neural net- works by improved downsampling with precise gradient back- propagation. arXiv preprint arXiv:2305.05...
2023 arXiv
-
[23]
Temporal-wise attention spiking neural networks for event streams classification
Man Yao, Huanhuan Gao, Guangshe Zhao, Dingheng Wang, Yi- han Lin, Zhaoxu Yang, and Guoqi Li. Temporal-wise attention spiking neural networks for event streams classification. InPro- ceedings of the IEEE /CVF International Conference on Com- puter Vision, pages 10221–10230, 2021
2021
-
[24]
Tcja-snn: Temporal-channel joint attention for spiking neural networks
Rui-Jie Zhu, Malu Zhang, Qihang Zhao, Haoyu Deng, Yule Duan, and Liang-Jian Deng. Tcja-snn: Temporal-channel joint attention for spiking neural networks. IEEE Transactions on Neural Networks and Learning Systems, 2024
2024
-
[25]
Spatial-temporal self-attention for asyn- chronous spiking neural networks
Yuchen Wang, Kexin Shi, Chengzhuo Lu, Yuguo Liu, Malu Zhang, and Hong Qu. Spatial-temporal self-attention for asyn- chronous spiking neural networks. In Proceedings of the Thirty- Second International Joint Conference on Artificial Intelligence, IJCAI-23, volume 8, pages 3085–...
2023
-
[26]
Attention spiking neural networks
Man Yao, Guangshe Zhao, Hengyu Zhang, Yifan Hu, Lei Deng, Yonghong Tian, Bo Xu, and Guoqi Li. Attention spiking neural networks. IEEE transactions on pattern analysis and machine intelligence, 2023
2023
-
[27]
Stsc-snn: Spatio-temporal synaptic connection with temporal convolution and attention for spiking neural net- works
Chengting Yu, Zheming Gu, Da Li, Gaoang Wang, Aili Wang, and Erping Li. Stsc-snn: Spatio-temporal synaptic connection with temporal convolution and attention for spiking neural net- works. Frontiers in Neuroscience, 16:1079357, 2022
2022
-
[28]
Dista: De- noising spiking transformer with intrinsic plasticity and spa- tiotemporal attention
Boxun Xu, Hejia Geng, Yuxuan Yin, and Peng Li. Dista: De- noising spiking transformer with intrinsic plasticity and spa- tiotemporal attention. arXiv preprint arXiv:2311.09376, 2023
2023 arXiv
-
[29]
A spatial–channel– temporal-fused attention for spiking neural networks
Wuque Cai, Hongze Sun, Rui Liu, Yan Cui, Jun Wang, Yang Xia, Dezhong Yao, and Daqing Guo. A spatial–channel– temporal-fused attention for spiking neural networks. IEEE transactions on Neural Networks and Learning Systems, 2023
2023
-
[30]
Simple model of spiking neurons
Eugene M Izhikevich. Simple model of spiking neurons. IEEE Transactions on neural networks, 14(6):1569–1572, 2003
2003
-
[31]
A quantitative de- scription of membrane current and its application to conduction and excitation in nerve
Alan L Hodgkin and Andrew F Huxley. A quantitative de- scription of membrane current and its application to conduction and excitation in nerve. The Journal of physiology, 117(4):500, 1952
1952
-
[32]
A novel image denoising algorithm combining attention mechanism and residual unet network
Shifei Ding, Qidong Wang, Lili Guo, Jian Zhang, and Ling Ding. A novel image denoising algorithm combining attention mechanism and residual unet network. Knowledge and Infor- mation Systems, 66(1):581–611, 2024. 13
2024
-
[33]
Agent attention: On the integration of softmax and linear attention
Dongchen Han, Tianzhu Ye, Yizeng Han, Zhuofan Xia, Shiji Song, and Gao Huang. Agent attention: On the integration of softmax and linear attention. arXiv preprint arXiv:2312.08874, 2023
2023 arXiv
-
[34]
Flatten transformer: Vision transformer using focused linear attention
Dongchen Han, Xuran Pan, Yizeng Han, Shiji Song, and Gao Huang. Flatten transformer: Vision transformer using focused linear attention. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5961–5971, 2023
2023
-
[35]
Learning multiple lay- ers of features from tiny images
Alex Krizhevsky, Geo ffrey Hinton, et al. Learning multiple lay- ers of features from tiny images. 2009
2009
-
[36]
Cifar10-dvs: an event-stream dataset for object classifica- tion
Hongmin Li, Hanchao Liu, Xiangyang Ji, Guoqi Li, and Luping Shi. Cifar10-dvs: an event-stream dataset for object classifica- tion. Frontiers in neuroscience, 11:244131, 2017
2017
-
[37]
A low power, fully event-based gesture recognition system
Arnon Amir, Brian Taba, David Berg, Timothy Melano, Jef- frey McKinstry, Carmelo Di Nolfo, Tapan Nayak, Alexander Andreopoulos, Guillaume Garreau, Marcela Mendoza, et al. A low power, fully event-based gesture recognition system. InPro- ceedings of the IEEE conference on compu...
2017
-
[38]
Enabling deep spiking neural networks with hybrid conversion and spike timing dependent backpropagation
Nitin Rathi, Gopalakrishnan Srinivasan, Priyadarshini Panda, and Kaushik Roy. Enabling deep spiking neural networks with hybrid conversion and spike timing dependent backpropagation. arXiv preprint arXiv:2005.01807, 2020
2005 arXiv
-
[39]
Optimizing deeper spiking neural networks for dynamic vision sensing
Youngeun Kim and Priyadarshini Panda. Optimizing deeper spiking neural networks for dynamic vision sensing. Neural Networks, 144:686–698, 2021
2021
-
[40]
Incorporating learnable membrane time constant to enhance learning of spiking neural networks
Wei Fang, Zhaofei Yu, Yanqi Chen, Timoth ´ee Masquelier, Tiejun Huang, and Yonghong Tian. Incorporating learnable membrane time constant to enhance learning of spiking neural networks. In Proceedings of the IEEE /CVF international con- ference on computer vision, pages 2661–2671, 2021
2021
-
[41]
Spatio- temporal backpropagation for training high-performance spik- ing neural networks
Yujie Wu, Lei Deng, Guoqi Li, Jun Zhu, and Luping Shi. Spatio- temporal backpropagation for training high-performance spik- ing neural networks. Frontiers in neuroscience, 12:331, 2018
2018
-
[42]
Direct training for spiking neural networks: Faster, larger, better
Yujie Wu, Lei Deng, Guoqi Li, Jun Zhu, Yuan Xie, and Luping Shi. Direct training for spiking neural networks: Faster, larger, better. In Proceedings of the AAAI conference on artificial intel- ligence, volume 33, pages 1311–1318, 2019
2019
-
[43]
Temporal spike sequence learn- ing via backpropagation for deep spiking neural networks
Wenrui Zhang and Peng Li. Temporal spike sequence learn- ing via backpropagation for deep spiking neural networks. Ad- vances in neural information processing systems , 33:12022– 12033, 2020
2020
-
[44]
Di fferentiable spike: Rethinking gradient-descent for training spiking neural networks
Yuhang Li, Yufei Guo, Shanghang Zhang, Shikuang Deng, Yongqing Hai, and Shi Gu. Di fferentiable spike: Rethinking gradient-descent for training spiking neural networks. Advances in Neural Information Processing Systems , 34:23426–23439, 2021
2021
-
[45]
Go- ing deeper with directly-trained larger spiking neural networks
Hanle Zheng, Yujie Wu, Lei Deng, Yifan Hu, and Guoqi Li. Go- ing deeper with directly-trained larger spiking neural networks. In Proceedings of the AAAI conference on artificial intelligence, volume 35, pages 11062–11070, 2021
2021
-
[46]
Temporal efficient training of spiking neural network via gra- dient re-weighting
Shikuang Deng, Yuhang Li, Shanghang Zhang, and Shi Gu. Temporal efficient training of spiking neural network via gra- dient re-weighting. arXiv preprint arXiv:2202.11946, 2022
2022 arXiv
-
[47]
Diet-snn: A low-latency spik- ing neural network with direct input encoding and leakage and threshold optimization
Nitin Rathi and Kaushik Roy. Diet-snn: A low-latency spik- ing neural network with direct input encoding and leakage and threshold optimization. IEEE Transactions on Neural Networks and Learning Systems, 34(6):3174–3182, 2021
2021
-
[48]
Ad- vancing spiking neural networks toward deep residual learning
Yifan Hu, Lei Deng, Yujie Wu, Man Yao, and Guoqi Li. Ad- vancing spiking neural networks toward deep residual learning. IEEE Transactions on Neural Networks and Learning Systems , 2024
2024
-
[49]
Training high-performance low-latency spiking neural networks by differentiation on spike representation
Qingyan Meng, Mingqing Xiao, Shen Yan, Yisen Wang, Zhouchen Lin, and Zhi-Quan Luo. Training high-performance low-latency spiking neural networks by differentiation on spike representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pa...
2022
-
[50]
Fixing weight decay regu- larization in adam
Ilya Loshchilov, Frank Hutter, et al. Fixing weight decay regu- larization in adam. arXiv preprint arXiv:1711.05101, 5, 2017
2017 arXiv
-
[51]
Swin transformer: Hier- archical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hier- archical vision transformer using shifted windows. In Proceed- ings of the IEEE/CVF international conference on computer vi- sion, pages 10012–10022, 2021
2021
-
[52]
Neuronal dynamics: From single neurons to networks and models of cognition
Wulfram Gerstner, Werner M Kistler, Richard Naud, and Liam Paninski. Neuronal dynamics: From single neurons to networks and models of cognition. Cambridge University Press, 2014
2014
-
[53]
1.1 computing’s energy problem (and what we can do about it)
Mark Horowitz. 1.1 computing’s energy problem (and what we can do about it). In 2014 IEEE international solid-state cir- cuits conference digest of technical papers (ISSCC) , pages 10–
2014
-
[54]
Liaf-net: Leaky integrate and analog fire net- work for lightweight and e fficient spatiotemporal information processing
Zhenzhi Wu, Hehui Zhang, Yihan Lin, Guoqi Li, Meng Wang, and Ye Tang. Liaf-net: Leaky integrate and analog fire net- work for lightweight and e fficient spatiotemporal information processing. IEEE Transactions on Neural Networks and Learn- ing Systems, 33(11):6249–6262, 2021
2021
-
[55]
E fficient processing of spatio-temporal data streams with spiking neural networks
Alexander Kugele, Thomas Pfeil, Michael Pfei ffer, and Elis- abetta Chicca. E fficient processing of spatio-temporal data streams with spiking neural networks. Frontiers in neuro- science, 14:512192, 2020
2020
-
[56]
Synaptic plasticity dynamics for deep continuous local learning (decolle)
Jacques Kaiser, Hesham Mostafa, and Emre Neftci. Synaptic plasticity dynamics for deep continuous local learning (decolle). Frontiers in Neuroscience, 14:424, 2020. 14
2020
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.