REVIEW 4 major objections 5 minor 1 cited by
DA-LIF: Dual Adaptive Leaky Integrate-and-Fire Model for Deep Spiking Neural Networks
T0 review · 4 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read Giving each layer of a spiking network its own learnable input gain and membrane leak—instead of a single shared time constant—lets the DA-LIF neuron reach state-of-the-art accuracy on static and neuromorphic benchmarks with fewer…
desk verdict A clean, plausible SNN neuron tweak with solid ablations, but the 'independent decays' claim needs a PLIF-style control and the SOTA comparisons are under-specified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the dual-adaptive LIF neuron, a discrete-time spiking neuron whose membrane update is $V^{t,n} = \beta_n H^{t-1,n} + \alpha_n X^{t,n}$, with $\alpha_n$ and $\beta_n$ learned by gradient descent independently for each layer $n$ and shared across timesteps within a layer. It replaces the single factor $1/\tau_m$ from the vanilla LIF with two separate learnable factors, so the neuron can amplify or attenuate incoming spikes while independently deciding how quickly to forget its own state. The training machinery is spatio-temporal backpropagation with a rectangular surrogate gradient, which lets the non-differentiable spike function pass gradients to $\alpha$ and $\beta$. This combination is what lets the network assign different spatio-temporal filtering roles to different layers.
What would settle it
Train two identical ResNet spiking networks on the same dataset, differing only in whether the membrane update uses two independent learnable decays (DA-LIF) or a single shared learnable decay, with every other hyperparameter and augmentation step fixed; if the accuracy gap disappears, the dual-decay mechanism is not the cause of the reported gains.
Extended reading notes
Core claim
The paper's central claim is that the LIF neuron's decay should be split into two independent learnable coefficients: a spatial coefficient $\alpha$ that controls how much of the current input enters the membrane potential, and a temporal coefficient $\beta$ that controls how much of the previous potential is retained. Starting from the continuous LIF equation and giving separate time constants to the leak and input terms, the discretized update becomes $V^{t,n} = \beta_n H^{t-1,n} + \alpha_n X^{t,n}$, with both coefficients shared across timesteps within a layer but free to differ across layers. Trained with spatio-temporal backpropagation and surrogate gradients, this neuron model reaches 96.72% on CIFAR-10, 80.59% on CIFAR-100, 70.58% on ImageNet, 82.42% on CIFAR10-DVS, and 98.61% on DVS128 Gesture, matching or exceeding earlier methods while using fewer timesteps and adding only negligible parameters.
Load-bearing premise
The reported accuracy gains are attributed to the two independent decays rather than to differences in the training recipe, but the implementation details omit batch size, data augmentation, and learning-rate schedule, so the comparison against prior methods is not fully controlled.
Editorial extensions
If this is right
- Because the two decays are scalars shared within each layer, the accuracy gains come with almost no additional parameters and no additional spike computations.
- The model works across static image benchmarks and neuromorphic event-stream benchmarks without changing the neuron formulation, suggesting the decoupling is dataset-agnostic.
- Reported accuracies at smaller timesteps imply lower latency and lower energy per inference on spike-based hardware.
- Ablations indicate the two decays are complementary: independently tuning $\alpha$ alone gives +0.76% and $\beta$ alone +0.37%, while tuning both gives +1.28% on CIFAR-100.
- The learned $\alpha$ and $\beta$ values drift apart across layers, with shallow layers leaning temporal and deeper layers spatial, indicating the model discovers a layer-specific division of labor.
Reading between the lines
- The paper leaves implicit that sharing $\alpha$ and $\beta$ within a layer is a design choice; making them per-channel or per-neuron is a natural next step that would trade a few more parameters for finer filtering, and the reported energy analysis suggests the added cost would still be modest.
- The decoupling principle is not tied to the LIF equation specifically; applying separate spatial and temporal learnable coefficients to gated or adaptive-threshold neuron models could yield similar expressiveness gains in other spiking architectures.
- The paper's ablations tune one decay at a time from a fixed baseline; comparing against a learned single-decay neuron would separate the benefit of learning decays from the benefit of decoupling them, a distinction the reported numbers do not fully resolve.
- If the observed divergence of $\alpha$ and $\beta$ across layers is consistent across runs, the learned values could serve as a diagnostic for which layers need more temporal memory, guiding layer-specific architectural choices.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DA-LIF, a spiking neuron model in which the spatial decay α_n and temporal decay β_n of the LIF membrane update are learned independently per layer (Eq. (9)). The authors train deep SNNs end-to-end with surrogate gradients and evaluate on CIFAR-10, CIFAR-100, ImageNet, CIFAR10-DVS, and DVS128 Gesture, reporting state-of-the-art or competitive accuracies with small timesteps. They also provide ablations isolating the contribution of α and β, an analysis of the learned decay distributions, and an energy-consumption table.
Significance. If the results hold, DA-LIF is a simple and parameter-efficient modification of the LIF neuron that improves accuracy across static and neuromorphic benchmarks, while the ablation separating α-only, β-only, and combined conditions gives some evidence that both decay mechanisms matter. The method is trained end-to-end against external classification benchmarks, so the improvements are not forced by construction. The main limitations are that the claimed superiority over prior work is not consistent in all configurations of Table I, the energy-efficiency claim lacks a comparative baseline, and the closest prior model (PLIF, a single learnable decay) is not included as a controlled baseline, leaving the specific benefit of independent versus coupled decays undemonstrated.
major comments (4)
- [Section IV-B, Table I] The claim that the model "consistently outperforms SOTA methods" is contradicted by the paper's own table: on ResNet-20 with T=4 on CIFAR-10, DA-LIF reaches 94.16%, below MPBN's 94.28% under the same architecture and timestep. Because the comparison also spans unstated training protocols, the general superiority claim is not established. Please restrict the claim to configurations where DA-LIF actually outperforms prior methods, or provide a matched-protocol comparison.
- [Section IV-D.1] The ablations compare a fixed LIF baseline, α-only learnable, β-only learnable, and both learnable, but no PLIF-style single shared learnable decay (e.g., β = 1 − α) is included as a baseline. Since PLIF already learns a membrane decay, the specific novelty of DA-LIF—independent rather than coupled spatial/temporal decays—needs to be tested against such a control under the same training protocol to show that the independence itself is beneficial.
- [Section IV-C, Table III] The energy-efficiency table reports only DA-LIF values for ACs, MACs, FLOPs, parameters, and energy; no ANN or reference SNN baseline with the same architecture is provided. The statement that "SNNs show greater energy efficiency" cannot be evaluated from Table III. Please add comparator columns or remove the energy-efficiency claim.
- [Section IV-A.2] Reproducibility and causal attribution are undermined by omitted standard hyperparameters: batch size, data-augmentation pipeline, learning-rate schedule, and normalization details are not reported. The implementation paragraph also appears internally inconsistent, stating that CIFAR10, CIFAR100, and CIFAR10-DVS are trained for 1000 epochs while "the DVS CIFAR-10 dataset is trained for 200 epochs," which conflates CIFAR10-DVS with DVS CIFAR-10. Please state the exact epoch counts and all training hyperparameters unambiguously.
minor comments (5)
- [Section I and Section IV-D.1] The Introduction states that the experimental results show a maximum improvement of 1.84%, but the ablation in Section IV-D.1 reports a combined improvement of 1.28% on CIFAR-100; please reconcile these numbers or specify which experiment the 1.84% refers to.
- [Section III-B, Eq. (8)] The stated range [-1, 1] for α and β is not accompanied by a description of how it is enforced during training; please specify the parameterization, clamping, or projection used.
- [Section IV-D.2] There are typographical issues such as "T anh" and "T anhwere"; additionally, it is unclear what activation functions are being compared and what is initialized to 1, so please clarify the experimental setup.
- [Tables I and II] The number of runs used to compute the reported standard deviations is not stated; please specify the number of seeds for the mean and error bars.
- [General] There is no code or data availability statement; including one would improve reproducibility and help readers verify the reported benchmark numbers.
Circularity Check
No significant circularity: DA-LIF is an empirical model proposal trained end-to-end on external benchmarks, with no self-citation chain or fitted-input-called-prediction step.
full rationale
The paper's derivation chain is a parameterization, not a circular reduction. Starting from the continuous LIF equation, the authors introduce variable decays and arrive at the update V_t,n = beta_n H_{t-1,n} + alpha_n X_t,n (Eq. 9). This is a model definition: alpha multiplies the spatial input current and beta multiplies the temporal membrane potential, so the labels 'spatial' and 'temporal' are descriptive of where the parameters act, not a conclusion derived from the model. The reported accuracies are measured outcomes on external datasets (CIFAR, ImageNet, DVS), not quantities forced by the definition of alpha and beta. The learnable decays are trained by gradient descent against classification losses, and the ablation compares fixed decays against learnable decays; no accuracy number is obtained by fitting a parameter to the test set and then re-reporting it as a prediction. The paper contains no load-bearing self-citations: the references are to prior SNN and neuroscience work by other groups, and the authors do not invoke any 'uniqueness theorem' from their own prior papers. The skeptical concern about missing controlled baselines and omitted training-protocol details (batch size, augmentation, schedule) is a legitimate experimental-rigor criticism, but it is not circularity: an uncontrolled comparison can weaken causal attribution without making the result equivalent to its inputs by construction. Accordingly, the circularity score is 0.
Assumptions & free parameters
free parameters (4)
- alpha (spatial decay) =
learned; initialized 1.0, range [-1, 1]
- beta (temporal decay) =
learned; initialized 1.0, range [-1, 1]
- firing threshold vth =
1.0
- surrogate gradient width a =
1
assumptions (3)
- standard math The discretized LIF recurrence Eq. 9 is a valid approximation of the continuous neuron dynamics in Eq. 5.
- domain assumption The surrogate gradient in Eq. 11 gives a useful training signal for the non-differentiable Heaviside spike function.
- domain assumption Published baseline accuracies in Tables I and II are accurate and were obtained under comparable conditions.
Cite this review
Pith. "Pith review of DA-LIF: Dual Adaptive Leaky Integrate-and-Fire Model for Deep Spiking Neural Networks." pith.science (2026). https://pith.science/paper/BWLDXR4Q
@misc{pith2026250210422,
author = {Pith},
title = {Pith review of: DA-LIF: Dual Adaptive Leaky Integrate-and-Fire Model for Deep Spiking Neural Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/BWLDXR4Q}},
note = {Machine review of arXiv:2502.10422}
}
read the original abstract
Spiking Neural Networks (SNNs) are valued for their ability to process spatio-temporal information efficiently, offering biological plausibility, low energy consumption, and compatibility with neuromorphic hardware. However, the commonly used Leaky Integrate-and-Fire (LIF) model overlooks neuron heterogeneity and independently processes spatial and temporal information, limiting the expressive power of SNNs. In this paper, we propose the Dual Adaptive Leaky Integrate-and-Fire (DA-LIF) model, which introduces spatial and temporal tuning with independently learnable decays. Evaluations on both static (CIFAR10/100, ImageNet) and neuromorphic datasets (CIFAR10-DVS, DVS128 Gesture) demonstrate superior accuracy with fewer timesteps compared to state-of-the-art methods. Importantly, DA-LIF achieves these improvements with minimal additional parameters, maintaining low energy consumption. Extensive ablation studies further highlight the robustness and effectiveness of the DA-LIF model.
Figures
Forward citations
Cited by 1 Pith paper
-
Edge Intelligence with Spiking Neural Networks
A comprehensive review of spiking neural networks for edge computing, covering neuron models, learning algorithms, hardware, deployment, security, and evaluation, with a claim to be the first survey on this specific i...
Reference graph
Works this paper leans on
-
[1]
Networks of spiking neurons: The third generation of neural network models,
W. Maass, “Networks of spiking neurons: The third generation of neural network models,” Neural Networks, vol. 10, no. 9, pp. 1659–1671, Dec. 1997
work page 1997
-
[2]
Homeostatic plasticity in the developing nervous system,
G. G. Turrigiano and S. B. Nelson, “Homeostatic plasticity in the developing nervous system,” Nature Reviews Neuroscience, vol. 5, no. 2, pp. 97–107, Feb. 2004
work page 2004
-
[3]
Beyond plasticity: the dynamic impact of electrical synapses on neural circuits,
P. Alcam ´ı and A. E. Pereda, “Beyond plasticity: the dynamic impact of electrical synapses on neural circuits,” Nature Reviews Neuroscience, vol. 20, no. 5, pp. 253–271, May 2019
work page 2019
-
[4]
Lapicque’s introduction of the integrate-and-fire model neuron (1907),
L. F. Abbott, “Lapicque’s introduction of the integrate-and-fire model neuron (1907),” Brain Research Bulletin, vol. 50, no. 5-6, pp. 303–304, Nov. 1999
work page 1907
-
[5]
Going Deeper With Directly-Trained Larger Spiking Neural Networks,
H. Zheng, Y . Wu, L. Deng, Y . Hu, and G. Li, “Going Deeper With Directly-Trained Larger Spiking Neural Networks,” Pro- ceedings of AAAI , vol. 35, pp. 11 062–11 070, May 2021, number: 12
work page 2021
-
[6]
Membrane Potential Batch Normal- ization for Spiking Neural Networks,
Y . Guo, Y . Zhang, Y . Chen, W. Peng, X. Liu, L. Zhang, X. Huang, and Z. Ma, “Membrane Potential Batch Normal- ization for Spiking Neural Networks,” in Proceedings of ICCV, Oct. 2023, pp. 19 420–19 430
work page 2023
-
[7]
Deep Residual Learning in Spiking Neural Networks,
W. Fang, Z. Yu, Y . Chen, T. Huang, T. Masquelier, and Y . Tian, “Deep Residual Learning in Spiking Neural Networks,” Proceedings of NeurIPS , vol. 34, pp. 21 056–21 069, 2021
work page 2021
-
[8]
L. Feng, Q. Liu, H. Tang, D. Ma, and G. Pan, “Multi-Level Firing with Spiking DS-ResNet: Enabling Better and Deeper Directly-Trained Spiking Neural Networks,” in Proceedings of IJCAI, Jul. 2022, pp. 2471–2477
work page 2022
Show all 35 references
-
[9]
Advancing Spiking Neural Networks Toward Deep Residual Learning,
Y . Hu, L. Deng, Y . Wu, M. Yao, and G. Li, “Advancing Spiking Neural Networks Toward Deep Residual Learning,” TNNLS, pp. 1–15, 2024
2024
-
[10]
IM-Loss: Information Maximization Loss for Spiking Neural Networks,
Y . Guo, Y . Chen, L. Zhang, X. Liu, Y . Wang, X. Huang, and Z. Ma, “IM-Loss: Information Maximization Loss for Spiking Neural Networks,” Proceedings of NeurIPS , vol. 35, pp. 156– 166, Dec. 2022
2022
-
[11]
Constructing Deep Spiking Neural Networks From Artificial Neural Networks With Knowledge Distillation,
Q. Xu, Y . Li, J. Shen, J. K. Liu, H. Tang, and G. Pan, “Constructing Deep Spiking Neural Networks From Artificial Neural Networks With Knowledge Distillation,” in Proceedings of CVPR, Jun. 2023, pp. 7886–7895
2023
-
[12]
Gradual Surrogate Gra- dient Learning in Deep Spiking Neural Networks,
Y . Chen, S. Zhang, S. Ren, and H. Qu, “Gradual Surrogate Gra- dient Learning in Deep Spiking Neural Networks,” in ICASSP, May 2022, pp. 8927–8931, iSSN: 2379-190X
2022
-
[13]
Event- Based Multimodal Spiking Neural Network with Attention Mechanism,
Q. Liu, D. Xing, L. Feng, H. Tang, and G. Pan, “Event- Based Multimodal Spiking Neural Network with Attention Mechanism,” in ICASSP, May 2022, pp. 8922–8926, iSSN: 2379-190X
2022
-
[14]
Long short-term memory and Learning-to-learn in networks of spiking neurons,
G. Bellec, D. Salaj, A. Subramoney, R. Legenstein, and W. Maass, “Long short-term memory and Learning-to-learn in networks of spiking neurons,” Proceedings of NeurIPS, vol. 31, 2018
2018
-
[15]
LTMD: Learn- ing Improvement of Spiking Neural Networks with Learnable Thresholding Neurons and Moderate Dropout,
S. Wang, T. H. Cheng, and M.-H. Lim, “LTMD: Learn- ing Improvement of Spiking Neural Networks with Learnable Thresholding Neurons and Moderate Dropout,” Proceedings of NeurIPS, vol. 35, pp. 28 350–28 362, Dec. 2022
2022
-
[16]
Incorporating Learnable Membrane Time Constant To Enhance Learning of Spiking Neural Networks,
W. Fang, Z. Yu, Y . Chen, T. Masquelier, T. Huang, and Y . Tian, “Incorporating Learnable Membrane Time Constant To Enhance Learning of Spiking Neural Networks,” in Proceedings of ICCV, 2021, pp. 2661–2671
2021
-
[17]
DIET-SNN: A Low-Latency Spiking Neural Network With Direct Input Encoding and Leakage and Threshold Optimization,
N. Rathi and K. Roy, “DIET-SNN: A Low-Latency Spiking Neural Network With Direct Input Encoding and Leakage and Threshold Optimization,” IEEE TNNLS , vol. 34, no. 6, pp. 3174–3182, Jun. 2023
2023
-
[18]
Biologically Inspired Dynamic Thresholds for Spik- ing Neural Networks,
J. Ding, B. Dong, F. Heide, Y . Ding, Y . Zhou, B. Yin, and X. Yang, “Biologically Inspired Dynamic Thresholds for Spik- ing Neural Networks,” Proceedings of NeurIPS , vol. 35, pp. 6090–6103, Dec. 2022
2022
-
[19]
Rsnn: Recurrent spiking neural networks for dynamic spatial- temporal information processing,
Q. Xu, X. Fang, Y . Li, J. Shen, D. Ma, Y . Xu, and G. Pan, “Rsnn: Recurrent spiking neural networks for dynamic spatial- temporal information processing,” in ACM Multimedia 2024
2024
-
[20]
Spatio-Temporal Backpropagation for Training High-Performance Spiking Neu- ral Networks,
Y . Wu, L. Deng, G. Li, J. Zhu, and L. Shi, “Spatio-Temporal Backpropagation for Training High-Performance Spiking Neu- ral Networks,” Frontiers in Neuroscience , vol. 12, p. 323875, 2018
2018
-
[21]
Cifar-10 (canadian institute for advanced research),
A. Krizhevsky, V . Nair, and G. Hinton, “Cifar-10 (canadian institute for advanced research),” vol. 5, no. 4, p. 1, 2010
2010
-
[22]
ImageNet: A large-scale hierarchical image database,
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei- Fei, “ImageNet: A large-scale hierarchical image database,” in Proceedings of CVPR , Jun. 2009, pp. 248–255, iSSN: 1063- 6919
2009
-
[23]
CIFAR10-DVS: An Event-Stream Dataset for Object Classification,
H. Li, H. Liu, X. Ji, G. Li, and L. Shi, “CIFAR10-DVS: An Event-Stream Dataset for Object Classification,” Frontiers in Neuroscience, vol. 11, 2017
2017
-
[24]
A Low Power, Fully Event-Based Gesture Recognition System,
A. Amir, B. Taba, D. Berg, T. Melano, J. McKinstry, C. Di Nolfo, T. Nayak, A. Andreopoulos, G. Garreau, M. Men- doza, J. Kusnitz, M. Debole, S. Esser, T. Delbruck, M. Flickner, and D. Modha, “A Low Power, Fully Event-Based Gesture Recognition System,” in Proceedings of CVPR, 2...
2017
-
[25]
RecDis-SNN: Rectifying Membrane Potential Dis- tribution for Directly Training Spiking Neural Networks,
Y . Guo, X. Tong, Y . Chen, L. Zhang, X. Liu, Z. Ma, and X. Huang, “RecDis-SNN: Rectifying Membrane Potential Dis- tribution for Directly Training Spiking Neural Networks,” in Proceedings of CVPR , Jun. 2022, pp. 326–335
2022
-
[26]
GLIF: A Unified Gated Leaky Integrate-and-Fire Neuron for Spiking Neural Networks,
X. Yao, F. Li, Z. Mo, and J. Cheng, “GLIF: A Unified Gated Leaky Integrate-and-Fire Neuron for Spiking Neural Networks,” Proceedings of NeurIPS, vol. 35, pp. 32 160–32 171, Dec. 2022
2022
-
[27]
Temporal Efficient Train- ing of Spiking Neural Network via Gradient Re-weighting,
S. Deng, Y . Li, S. Zhang, and S. Gu, “Temporal Efficient Train- ing of Spiking Neural Network via Gradient Re-weighting,” in Proceedings of ICLR , Oct. 2021
2021
-
[28]
Learnable Surrogate Gradient for Direct Training Spiking Neural Networks,
S. Lian, J. Shen, Q. Liu, Z. Wang, R. Yan, and H. Tang, “Learnable Surrogate Gradient for Direct Training Spiking Neural Networks,” in Proceedings of IJCAI , Aug. 2023, pp. 3002–3010
2023
-
[29]
Tensor decomposition based attention module for spiking neu- ral networks,
H. Deng, R. Zhu, X. Qiu, Y . Duan, M. Zhang, and L.-J. Deng, “Tensor decomposition based attention module for spiking neu- ral networks,” Knowledge-Based Systems, vol. 295, p. 111780, 2024
2024
-
[30]
IM-LIF: Improved Neuronal Dynamics With Attention Mechanism for Direct Training Deep Spiking Neural Network,
S. Lian, J. Shen, Z. Wang, and H. Tang, “IM-LIF: Improved Neuronal Dynamics With Attention Mechanism for Direct Training Deep Spiking Neural Network,” IEEE TETCI, pp. 1– 11, 2024
2024
-
[31]
Spatial- Temporal Self-Attention for Asynchronous Spiking Neural Net- works,
Y . Wang, K. Shi, C. Lu, Y . Liu, M. Zhang, and H. Qu, “Spatial- Temporal Self-Attention for Asynchronous Spiking Neural Net- works,” in Proceedings of IJCAI , Aug. 2023, pp. 3085–3093
2023
-
[32]
Attention Spiking Neural Networks,
M. Yao, G. Zhao, H. Zhang, Y . Hu, L. Deng, Y . Tian, B. Xu, and G. Li, “Attention Spiking Neural Networks,” IEEE TPAMI, vol. 45, no. 8, pp. 9393–9410, Aug. 2023
2023
-
[33]
Inherent Redundancy in Spiking Neural Networks,
M. Yao, J. Hu, G. Zhao, Y . Wang, Z. Zhang, B. Xu, and G. Li, “Inherent Redundancy in Spiking Neural Networks,” in Proceedings of the ICCV , 2023, pp. 16 924–16 934
2023
-
[34]
Temporal-Wise Attention Spiking Neural Networks for Event Streams Classification,
M. Yao, H. Gao, G. Zhao, D. Wang, Y . Lin, Z. Yang, and G. Li, “Temporal-Wise Attention Spiking Neural Networks for Event Streams Classification,” in Proceedings of ICCV , 2021, pp. 10 221–10 230
2021
-
[35]
Recent Advances in Scalable Energy-Efficient and Trustworthy Spiking Neural Networks: from Algorithms to Technology,
S. Kundu, R.-J. Zhu, A. Jaiswal, and P. A. Beerel, “Recent Advances in Scalable Energy-Efficient and Trustworthy Spiking Neural Networks: from Algorithms to Technology,” in ICASSP, Apr. 2024, pp. 13 256–13 260, iSSN: 2379-190X
2024
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.