Pith. sign in

REVIEW 2 minor 78 references

DBHN-Net: Dual-Branch Hybrid Neural Network For Low-Complexity Monaural Speech Enhancement

T0 review · 0 major / 2 minor · reviewed 2026-06-27 · grok-4.3

Pith's one-line read A hybrid neural network for speech enhancement pairs spiking and conventional branches to cut computational complexity by a factor of 7.5 while preserving performance.

desk verdict DBHN-Net gives a concrete hybrid ANN-SNN design that trades off power for performance in speech enhancement and backs the trade-off with results on three datasets. read the letter →

arxiv 2606.05911 v1 pith:GYH4PVVW submitted 2026-06-04 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords speechenhancementspikingneuralnetworkshybridnetworklowcomplexitymonauraldual-brancharchitecturetime-frequencyfusioncomputationalefficiency
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces DBHN-Net, a dual-branch architecture that integrates spiking neural networks for reduced power use with artificial neural networks to recover lost information from binary activations. The design incorporates specialized modules like BandSplit, TF-Mamba, SFEG, ITB, an Interaction module, and TF-Cross Attention-Fusion to enable effective feature compression, refinement, and cross-branch information exchange in the time-frequency domain. By addressing the information loss in SNNs through guided fusion, the model aims to make high-performance speech enhancement feasible on low-power devices. Results indicate it outperforms baselines on three datasets at significantly lower complexity.

What carries the argument

The dual-branch hybrid network with Interaction module and TF-Cross Attention-Fusion module that facilitates information exchange between the ANN branch and the SNN branch to compensate for spiking information loss.

What would settle it

An ablation study removing the TF-Cross Attention-Fusion module and measuring if speech enhancement metrics on the datasets fall to levels comparable to pure SNN models without the claimed performance retention.

Watch

Extended reading notes

Core claim

The DBHN-Net maintains superior performance across three public datasets while achieving an average 7.5 fold reduction in computational complexity compared to baseline models by using a dual-branch ANN-SNN structure with inter-branch fusion mechanisms that allow the SNN branch to retain critical information despite its discrete activations.

Load-bearing premise

The Interaction module and TF-Cross Attention-Fusion module can successfully guide the SNN branch to retain critical information and thereby offset the information loss inherent to discrete spiking activations.

Editorial extensions

If this is right

  • The model can be deployed in resource-constrained devices for real-time speech enhancement due to its lower complexity.
  • SNN branches in hybrid setups reduce energy consumption while the ANN branch preserves accuracy in audio processing tasks.
  • The fusion modules enable the SNN branch to retain critical time-frequency features despite binary spiking activations.
  • Residual connections in SFEG and ITB components further refine representations and mitigate information loss across branches.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • This hybrid design could be tested on other audio tasks like source separation or dereverberation to check if the complexity savings generalize.
  • The inter-branch fusion approach may suggest ways to combine discrete and continuous activations in non-audio domains such as sensor data processing.
  • Scaling the BandSplit and TF-Mamba modules to longer audio sequences could reveal limits on the claimed efficiency gains.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 2 minor

Summary. The paper proposes DBHN-Net, a dual-branch hybrid neural network combining ANN and SNN branches for monaural speech enhancement. It introduces BandSplit and TF-Mamba modules, SFEG and ITB components with residual connections, an Interaction module for inter-branch exchange, and a TF-Cross Attention-Fusion module for time-frequency domain fusion to guide the SNN branch. The central empirical claim is that the model achieves superior performance across three public datasets while delivering an average 7.5-fold reduction in computational complexity relative to baselines.

Significance. If the reported results hold, the work is significant for enabling low-power, high-performance speech enhancement on edge devices. The hybrid ANN-SNN design with explicit fusion mechanisms directly tackles SNN information loss while leveraging SNN efficiency, and the provision of dataset results, complexity metrics, and module diagrams supplies falsifiable empirical support for the performance-complexity trade-off.

minor comments (2)
  1. [Abstract] Abstract: the claim of results on 'three public datasets' is not accompanied by dataset names or any quantitative metrics (PESQ, STOI, complexity figures); while the full manuscript supplies these, the abstract should include at least one sentence naming the datasets and summarizing the key numbers to allow immediate evaluation.
  2. [TF-Cross Attention-Fusion module description] The description of the TF-Cross Attention-Fusion module (section describing inter-branch fusion) states it 'data-adaptively guid[es] the SNN branch to retain more critical information,' but does not specify the exact attention formulation or loss terms used to enforce this guidance; adding the precise equations would strengthen reproducibility.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for their positive assessment of DBHN-Net and the recommendation for minor revision. The review correctly identifies the core contributions of the dual-branch hybrid architecture, the BandSplit/TF-Mamba modules, SFEG/ITB components, and the cross-attention fusion mechanism, as well as the empirical results on three datasets showing maintained performance at substantially reduced complexity. We will prepare a revised manuscript addressing any minor points.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity

full rationale

The paper proposes an empirical neural network architecture (DBHN-Net) for speech enhancement and validates it through experiments on three public datasets. No derivation chain, first-principles predictions, fitted parameters renamed as outputs, or self-referential equations appear in the abstract or architecture description. Claims of performance and complexity reduction rest on reported metrics rather than reducing to inputs by construction. No load-bearing self-citations or uniqueness theorems are invoked. The work is self-contained as an engineering contribution with external empirical benchmarks.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The central claim rests on the unstated assumptions that standard supervised training of hybrid networks is stable, that the chosen public datasets are representative, and that the proposed modules function as described; none of these are evidenced in the abstract.

assumptions (2)
  • domain assumption Hybrid ANN-SNN training converges to a useful operating point
    Implicit in any claim that the dual-branch design works
  • domain assumption The three public datasets are appropriate benchmarks for the task
    Required for the performance claim to be meaningful

how reviews work

0 comments
Cite this review

Pith. "Pith review of DBHN-Net: Dual-Branch Hybrid Neural Network For Low-Complexity Monaural Speech Enhancement." pith.science (2026). https://pith.science/paper/GYH4PVVW

@misc{pith2026260605911,
  author       = {Pith},
  title        = {Pith review of: DBHN-Net: Dual-Branch Hybrid Neural Network For Low-Complexity Monaural Speech Enhancement},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GYH4PVVW}},
  note         = {Machine review of arXiv:2606.05911}
}
read the original abstract

Although artificial neural network (ANN) based speech enhancement (SE) methods demonstrate excellent performance, the high computational complexity and high energy consumption hinder their deployment in practical front-end processing tasks.} Currently, the spiking neural networks (SNNs) have shown potential in reducing power consumption. However, the discrete binary activation and complex spatio-temporal dynamics of SNNs often result in information loss. The current challenge therefore focuses on how to maintain performance and reduce computational complexity. To address this issue, this work propose a Dual-Branch Hybrid Neural (DBHN) Network. 1) In terms of network architecture: A dual-branch network integrating ANN and SNN was designed, where the SNN branch reduces power consumption while the ANN branch addresses information loss; The BandSplit and Time-Frequency (TF) -Mamba modules were developed to simultaneously compress energy consumption and enhance model performance; Spiking Feature Extraction Group (SFEG) and Information Transformation Block (ITB) components were implemented with residual connections to mitigate information loss while further refining feature representations. 2) To facilitate inter-branch information fusion: An Interaction module was designed to promote information exchange at various stages of the dual-branch network; A TF-Cross Attention-Fusion module was designed to perform time-frequency domain fusion of dual-branch information while data-adaptively guiding the SNN branch to retain more critical information. Results show that the proposed model maintains superior performance across three public datasets while achieving an average 7.5 fold reduction in computational complexity compared to baseline models.

Figures

Figures reproduced from arXiv: 2606.05911 by the authors.

Figure 1
Figure 1. (a) The proposed Dual-Branch Hybrid Neural Network takes complex spectra as input. The upper branch is the ANN pathway, primarily consisting of a BandSplit module and a TF-Mamba sequential modeling module. The lower branch represents the SNN pathway, mainly comprising a Spiking Feature Extraction Group and an Information Transformation Block. (b) The proposed Spiking Feature Extraction Block primarily employs LIF un… view at source ↗
Figure 3
Figure 3. (a) The Band-Split module, operating at the initial stage of the ANN branch, decomposes the complex spectrum along the frequency dimension. (b) The Band-Merge module, functioning at the final stage of the ANN branch, reconstructs the segmented complex spectra. domain and frequency-domain fusion, generating enhanced complex spectra that are finally transformed into enhanced speech via ISTFT. C. ANN Branch Architectur… view at source ↗
Figure 4
Figure 4. The proposed Information Transformation Block operates at the [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figures from the paper (1 more)
Figure 5
Figure 5. Figure 5: The figure shows the visualization results of ablation experiments. The models were trained on the WSJ0-SI84 and DNS-Challenge datasets, and [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

78 extracted references · 3 canonical work pages

  1. [1]

    Validity and robustness of denoisers: A proof of concept in speech denoising,

    M. Ebrahimi, Q. Alfalouji, and M. Basirat, “Validity and robustness of denoisers: A proof of concept in speech denoising,”IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 650–665, 2025

  2. [2]

    Dubbing movies via hierarchical phoneme modeling and acoustic diffusion denoising,

    L. Li, G. Cong, Y . Qi, Z.-J. Zha, Q. Wu, Q. Z. Sheng, Q. Huang, and M.-H. Yang, “Dubbing movies via hierarchical phoneme modeling and acoustic diffusion denoising,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 11, pp. 10 361–10 377, 2025. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 13

  3. [3]

    Bsdb-net: Band-split dual-branch network with selective state spaces mechanism for monaural speech enhancement,

    C. Fan, E. Liu, A. Li, J. Tao, J. Zhou, J. Li, C. Zheng, and Z. Lv, “Bsdb-net: Band-split dual-branch network with selective state spaces mechanism for monaural speech enhancement,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 22, 2025, pp. 23 850–23 858

  4. [4]

    Cross-modal knowledge distillation with multi-stage adaptive feature fusion for speech separa- tion,

    C. Fan, W. Xiang, J. Tao, J. Yi, and Z. Lv, “Cross-modal knowledge distillation with multi-stage adaptive feature fusion for speech separa- tion,”IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 935–948, 2025

  5. [5]

    Waveform-domain speech enhancement using spectrogram encoding for robust speech recognition,

    H. Shi, M. Mimura, and T. Kawahara, “Waveform-domain speech enhancement using spectrogram encoding for robust speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, pp. 3049–3060, 2024

  6. [6]

    Automatic speech recognition: A survey of deep learning techniques and approaches,

    H. Ahlawat, N. Aggarwal, and D. Gupta, “Automatic speech recognition: A survey of deep learning techniques and approaches,”International Journal of Cognitive Computing in Engineering, 2025

  7. [7]

    Seeing helps hearing: A multi-modal dataset and a mamba- based dual branch parallel network for auditory attention decoding,

    C. Fan, H. Zhang, Q. Ni, J. Zhang, J. Tao, J. Zhou, J. Yi, Z. Lv, and X. Wu, “Seeing helps hearing: A multi-modal dataset and a mamba- based dual branch parallel network for auditory attention decoding,” Information Fusion, vol. 118, pp. 102 946–102 957, 2025

  8. [8]

    An overview of deep-learning-based audio-visual speech en- hancement and separation,

    D. Michelsanti, Z.-H. Tan, S.-X. Zhang, Y . Xu, M. Yu, D. Yu, and J. Jensen, “An overview of deep-learning-based audio-visual speech en- hancement and separation,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 1368–1396, 2021

Show all 78 references
  1. [9]

    Sse-net: Towards low-power-consumption spiking neural network for monaural speech enhancement,

    E. Liu, A. Li, C. Fan, C. Zheng, J. Yi, R. Fu, X. Li, J. Zhou, and Z. Lv, “Sse-net: Towards low-power-consumption spiking neural network for monaural speech enhancement,”IEEE Transactions on Audio, Speech and Language Processing, pp. 1–13, 2026

  2. [10]

    Ai flow: Perspectives, scenarios, and approaches,

    H. An, S. Huang, S. Huang, R. Li, Y . Liang, J. Shao, Z. Wang, C. Yuan, C. Zhang, H. Zhang, W. Zhuang, and X. Li, “Ai flow: Perspectives, scenarios, and approaches,”arXiv preprint arXiv:2506.12479, 2025

  3. [11]

    Compact deep neural networks for real-time speech enhancement on resource-limited devices,

    F. E. Wahab, Z. Ye, N. Saleem, and R. Ullah, “Compact deep neural networks for real-time speech enhancement on resource-limited devices,” Speech Communication, vol. 156, pp. 103 008–103 019, 2024

  4. [12]

    Dccrn: Deep complex convolution recurrent network for phase-aware speech enhancement,

    Y . Hu, Y . Liu, S. Lv, M. Xing, S. Zhang, Y . Fu, J. Wu, B. Zhang, and L. Xie, “Dccrn: Deep complex convolution recurrent network for phase-aware speech enhancement,” ininterspeech, 2020, pp. 380–390

  5. [13]

    Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation,

    Y . Luo, Z. Chen, and T. Yoshioka, “Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation,” inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 46–50

  6. [14]

    Large-scale training to increase speech intelligibility for hearing-impaired listeners in novel noises,

    J. Chen, Y . Wang, S. E. Yoho, D. Wang, and E. W. Healy, “Large-scale training to increase speech intelligibility for hearing-impaired listeners in novel noises,”The Journal of the Acoustical Society of America, vol. 139, no. 5, pp. 2604–2612, 2016

  7. [15]

    Two heads are better than one: A two-stage complex spectral mapping approach for monaural speech enhancement,

    A. Li, W. Liu, C. Zheng, C. Fan, and X. Li, “Two heads are better than one: A two-stage complex spectral mapping approach for monaural speech enhancement,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 1829–1843, 2021

  8. [16]

    Dbt-net: Dual-branch federative magnitude and phase estimation with attention- in-attention transformer for monaural speech enhancement,

    G. Yu, A. Li, H. Wang, Y . Wang, Y . Ke, and C. Zheng, “Dbt-net: Dual-branch federative magnitude and phase estimation with attention- in-attention transformer for monaural speech enhancement,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, pp. 2629–2...

  9. [17]

    A convolutional recurrent neural network for real-time speech enhancement

    K. Tan and D. Wang, “A convolutional recurrent neural network for real-time speech enhancement.” inInterspeech, 3229–3233, p. 2018

  10. [18]

    Learning complex spectral mapping with gated convolutional recurrent networks for monaural speech enhancement,

    ——, “Learning complex spectral mapping with gated convolutional recurrent networks for monaural speech enhancement,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 28, pp. 380–390, 2019

  11. [19]

    On the compensation between magnitude and phase in speech separation,

    Z.-Q. Wang, G. Wichern, and J. Le Roux, “On the compensation between magnitude and phase in speech separation,”IEEE Signal Processing Letters, vol. 28, pp. 2018–2022, 2021

  12. [20]

    Seeing helps hearing: A multi-modal dataset and a mamba- based dual branch parallel network for auditory attention decoding,

    C. Fan, H. Zhang, Q. Ni, J. Zhang, J. Tao, J. Zhou, J. Yi, Z. Lv, and X. Wu, “Seeing helps hearing: A multi-modal dataset and a mamba- based dual branch parallel network for auditory attention decoding,” Information Fusion, vol. 118, p. 102946, 2025. [Online]. Available: https...

  13. [21]

    Glance and gaze: A collaborative learning framework for single-channel speech enhancement,

    A. Li, C. Zheng, L. Zhang, and X. Li, “Glance and gaze: A collaborative learning framework for single-channel speech enhancement,”Applied Acoustics, vol. 187, pp. 108 499–108 508, 2022

  14. [22]

    Fullsubnet: A full-band and sub- band fusion model for real-time single-channel speech enhancement,

    X. Hao, X. Su, R. Horaud, and X. Li, “Fullsubnet: A full-band and sub- band fusion model for real-time single-channel speech enhancement,” inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 6633–6637

  15. [23]

    A low-power streaming speech enhance- ment accelerator for edge devices,

    C.-H. Wu and T.-S. Chang, “A low-power streaming speech enhance- ment accelerator for edge devices,”IEEE Open Journal of Circuits and Systems, vol. 5, pp. 128–140, 2024

  16. [24]

    Flowse: Flow matching- based speech enhancement,

    S. Lee, S. Cheong, S. Han, and J. W. Shin, “Flowse: Flow matching- based speech enhancement,” inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5

  17. [25]

    Toward ultralow- power neuromorphic speech enhancement with spiking-fullsubnet,

    X. Hao, C. Ma, Q. Yang, J. Wu, and K. C. Tan, “Toward ultralow- power neuromorphic speech enhancement with spiking-fullsubnet,” IEEE Transactions on Neural Networks and Learning Systems, vol. 36, no. 9, pp. 17 350–17 364, 2025

  18. [26]

    Speech emotion recognition based on spiking neural network and convolutional neural network,

    C. Du, F. Liu, B. Kang, and T. Hou, “Speech emotion recognition based on spiking neural network and convolutional neural network,” Engineering Applications of Artificial Intelligence, vol. 147, p. 110314, 2025

  19. [27]

    A hybrid ann- snn architecture for low-power and low-latency visual perception,

    A. Aydin, M. Gehrig, D. Gehrig, and D. Scaramuzza, “A hybrid ann- snn architecture for low-power and low-latency visual perception,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 5701–5711

  20. [28]

    Spiking neural networks on fpga: A survey of methodologies and recent advancements,

    M. Karamimanesh, E. Abiri, M. Shahsavari, K. Hassanli, A. van Schaik, and J. Eshraghian, “Spiking neural networks on fpga: A survey of methodologies and recent advancements,”Neural Networks, p. 107256, 2025

  21. [29]

    The intel neuromorphic dns challenge,

    J. Timcheck, S. B. Shrestha, D. B. D. Rubin, A. Kupryjanow, G. Orchard, L. Pindor, T. Shea, and M. Davies, “The intel neuromorphic dns challenge,”Neuromorphic Computing and Engineering, vol. 3, no. 3, p. 034005, 2023

  22. [30]

    A hybrid ann- snn architecture for low-power and low-latency visual perception,

    A. Aydin, M. Gehrig, D. Gehrig, and D. Scaramuzza, “A hybrid ann- snn architecture for low-power and low-latency visual perception,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 5701–5711

  23. [31]

    Hynita: A neuromorphic inference and training accelerator for hybrid ann-snn fusion models,

    Y . Zhong, L. Lun, Z. Wang, J. Ruan, Y . Gao, X. Cui, X. Zhang, and Y . Wang, “Hynita: A neuromorphic inference and training accelerator for hybrid ann-snn fusion models,” in2025 IEEE International Symposium on Circuits and Systems (ISCAS). IEEE, 2025, pp. 1–5

  24. [32]

    Minimizing informa- tion loss reduces spiking neuronal networks to differential equations,

    J. Chang, Z. Li, Z. Wang, L. Tao, and Z.-C. Xiao, “Minimizing informa- tion loss reduces spiking neuronal networks to differential equations,” Journal of Computational Physics, p. 114117, 2025

  25. [33]

    Ai flow at the network edge,

    J. Shao and X. Li, “Ai flow at the network edge,”IEEE Network, 2025, early Access

  26. [34]

    Naturalspeech: End-to-end text-to-speech synthesis with human-level quality,

    X. Tan, J. Chen, H. Liu, J. Cong, C. Zhang, Y . Liu, X. Wang, Y . Leng, Y . Yi, L. He, S. Zhao, T. Qin, F. Soong, and T.-Y . Liu, “Naturalspeech: End-to-end text-to-speech synthesis with human-level quality,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. ...

  27. [35]

    Boss: Beyond-semantic speech,

    Q. Wang, Z. Li, H. Lv, H. Chen, Y . Song, J. Kang, J. Lian, J. Li, Y . Li, Z. Heet al., “Boss: Beyond-semantic speech,”arXiv preprint arXiv:2507.17563, 2025

  28. [36]

    Fullsubnet+: Channel attention fullsubnet with complex spectrograms for speech enhancement,

    J. Chen, Z. Wang, D. Tuo, Z. Wu, S. Kang, and H. Meng, “Fullsubnet+: Channel attention fullsubnet with complex spectrograms for speech enhancement,” inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 7857–7861

  29. [37]

    Taylor, can you hear me now? a taylor-unfolding framework for monaural speech enhancement,

    A. Li, S. You, G. Yu, C. Zheng, and X. Li, “Taylor, can you hear me now? a taylor-unfolding framework for monaural speech enhancement,” inProceedings of the Thirty-First International Joint Conference on Artificial Intelligence, L. D. Raedt, Ed. International Joint Conferences...

  30. [38]

    Comp- net: Complementary network for single-channel speech enhancement,

    C. Fan, H. Zhang, A. Li, W. Xiang, C. Zheng, Z. Lv, and X. Wu, “Comp- net: Complementary network for single-channel speech enhancement,” Neural Networks, vol. 168, pp. 508–517, 2023

  31. [39]

    Learning a spiking neural network for efficient image deraining,

    T. Song, G. Jin, P. Li, K. Jiang, X. Chen, and J. Jin, “Learning a spiking neural network for efficient image deraining,” inProceedings of the Thirty-Third International Joint Conference on Artificial Intelligence. International Joint Conferences on Artificial Intelligence Org...

  32. [40]

    Adaptation and learning of spatio-temporal thresholds in spiking neural networks,

    J. Fu, S. Gou, P. Wang, L. Jiao, Z. Guo, J. Li, and R. Liu, “Adaptation and learning of spatio-temporal thresholds in spiking neural networks,” Neurocomputing, p. 130423, 2025

  33. [41]

    Enhancing representation of spiking neural networks via similarity- sensitive contrastive learning,

    Y . Zhang, X. Liu, Y . Chen, W. Peng, Y . Guo, X. Huang, and Z. Ma, “Enhancing representation of spiking neural networks via similarity- sensitive contrastive learning,” inAAAI Conference on Artificial Intelli- gence, vol. 38, no. 15, 2024, pp. 16 926–16 934

  34. [42]

    Spikingbert: Distilling bert to train spiking language models using implicit differentiation,

    M. Bal and A. Sengupta, “Spikingbert: Distilling bert to train spiking language models using implicit differentiation,” inAAAI Conference on Artificial Intelligence, vol. 38, no. 10, 2024, pp. 10 998–11 006

  35. [43]

    Tc-lif: A two- compartment spiking neuron model for long-term sequential modelling,

    S. Zhang, Q. Yang, C. Ma, J. Wu, H. Li, and K. C. Tan, “Tc-lif: A two- compartment spiking neuron model for long-term sequential modelling,” JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 14 inAAAI Conference on Artificial Intelligence, vol. 38, no. 15, 2024, pp. 16...

  36. [44]

    Learning a spiking neural network for efficient image deraining,

    T. Song, G. Jin, P. Li, K. Jiang, X. Chen, and J. Jin, “Learning a spiking neural network for efficient image deraining,” inInternational Joint Conference on Artificial Intelligence, 2024

  37. [45]

    Spikelm: Towards general spike-driven language modeling via elastic bi-spiking mechanisms,

    X. Xing, Z. Zhang, Z. Ni, S. Xiao, Y . Ju, S. Fan, Y . Wang, J. Zhang, and G. Li, “Spikelm: Towards general spike-driven language modeling via elastic bi-spiking mechanisms,” inInternational Conference on Machine Learning, 2024, pp. 54 698–54 714

  38. [46]

    Dpsnn: Spiking neural network for low- latency streaming speech enhancement,

    T. Sun and S. M. Bohte, “Dpsnn: Spiking neural network for low- latency streaming speech enhancement,”Neuromorphic Computing and Engineering, 2024

  39. [47]

    Temporally dynamic spiking transformer network for speech enhancement,

    M. A. Alohali, N. Saleem, D. Rhouma, M. Medani, H. Elmannai, and S. Bourouis, “Temporally dynamic spiking transformer network for speech enhancement,”IEEE Access, 2024

  40. [48]

    Single channel speech enhancement using u-net spiking neural networks,

    A. Riahi and ´E. Plourde, “Single channel speech enhancement using u-net spiking neural networks,” in2023 IEEE Canadian Conference on Electrical and Computer Engineering. IEEE, 2023, pp. 111–116

  41. [49]

    Gerstner and W

    W. Gerstner and W. M. Kistler,Spiking neuron models: Single neurons, populations, plasticity. Cambridge university press, 2002

  42. [50]

    Rmp-loss: Regularizing membrane potential distribution for spiking neural networks,

    Y . Guo, X. Liu, Y . Chen, L. Zhang, W. Peng, Y . Zhang, X. Huang, and Z. Ma, “Rmp-loss: Regularizing membrane potential distribution for spiking neural networks,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 17 391–17 401

  43. [51]

    The design for the wall street journal-based csr corpus,

    D. B. Paul and J. Baker, “The design for the wall street journal-based csr corpus,” inSpeech and Natural Language: Proceedings of a Workshop Held at Harriman, New York, February 23-26, 1992, 1992

  44. [52]

    The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results,

    C. K. Reddy, V . Gopal, R. Cutler, E. Beyrami, R. Cheng, H. Dubey, S. Matusevych, R. Aichner, A. Aazami, S. Braunet al., “The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results,”Interspeech 2020, pp. 83–93, 2020

  45. [53]

    Assessment for automatic speech recog- nition: Ii. noisex-92: A database and an experiment to study the effect of additive noise on speech recognition systems,

    A. Varga and H. J. Steeneken, “Assessment for automatic speech recog- nition: Ii. noisex-92: A database and an experiment to study the effect of additive noise on speech recognition systems,”Speech communication, vol. 12, no. 3, pp. 247–251, 1993

  46. [54]

    The voice bank corpus: Design, collection and data analysis of a large regional accent speech database,

    C. Veaux, J. Yamagishi, and S. King, “The voice bank corpus: Design, collection and data analysis of a large regional accent speech database,” in2013 international conference oriental COCOSDA held jointly with 2013 conference on Asian spoken language research and evaluation (O...

  47. [55]

    Investigating rnn-based speech enhancement methods for noise-robust text-to-speech,

    C. V . Botinhao, X. Wang, S. Takaki, and J. Yamagishi, “Investigating rnn-based speech enhancement methods for noise-robust text-to-speech,” in9th ISCA Speech Synthesis Workshop, 2016, pp. 159–165

  48. [56]

    On the importance of power compression and phase estimation in monaural speech dereverberation,

    A. Li, C. Zheng, R. Peng, and X. Li, “On the importance of power compression and phase estimation in monaural speech dereverberation,” JASA express letters, vol. 1, no. 1, 2021

  49. [57]

    Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation,

    Y . Luo and N. Mesgarani, “Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation,”IEEE/ACM trans- actions on audio, speech, and language processing, vol. 27, no. 8, pp. 1256–1266, 2019

  50. [58]

    Segan: Speech enhancement generative adversarial network,

    S. Pascual, A. Bonafonte, and J. Serr `a, “Segan: Speech enhancement generative adversarial network,” inInterspeech, 2017, pp. 3642–3646

  51. [59]

    Time-frequency masking-based speech enhancement using generative adversarial network,

    M. H. Soni, N. Shah, and H. A. Patil, “Time-frequency masking-based speech enhancement using generative adversarial network,” inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 5039–5043

  52. [60]

    Metricgan: Generative adversarial networks based black-box metric scores optimization for speech enhancement,

    S.-W. Fu, C.-F. Liao, Y . Tsao, and S.-D. Lin, “Metricgan: Generative adversarial networks based black-box metric scores optimization for speech enhancement,” inInternational Conference on Machine Learn- ing. PmLR, 2019, pp. 2031–2041

  53. [61]

    Wavenet: A generative model for raw audio,

    S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalch- brenner, A. Senior, K. Kavukcuogluet al., “Wavenet: A generative model for raw audio,”arXiv preprint arXiv:1609.03499, vol. 12, 2016

  54. [62]

    Srtnet: Time domain speech enhancement via stochastic refinement,

    Z. Qiu, M. Fu, Y . Yu, L. Yin, F. Sun, and H. Huang, “Srtnet: Time domain speech enhancement via stochastic refinement,” inIEEE Interna- tional Conference on Acoustics, Speech and Signal Processing(ICASSP), 2023, pp. 1–5

  55. [63]

    Phasen: A phase-and- harmonics-aware speech enhancement network,

    D. Yin, C. Luo, Z. Xiong, and W. Zeng, “Phasen: A phase-and- harmonics-aware speech enhancement network,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 05, 2020, pp. 9458–9465

  56. [64]

    Speech enhancement using self-adaptation and multi-head self- attention,

    Y . Koizumi, K. Yatabe, M. Delcroix, Y . Masuyama, and D. Takeuchi, “Speech enhancement using self-adaptation and multi-head self- attention,” inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 181–185

  57. [65]

    Tstnn: Two-stage transformer based neural network for speech enhancement in the time domain,

    K. Wang, B. He, and W.-P. Zhu, “Tstnn: Two-stage transformer based neural network for speech enhancement in the time domain,” pp. 7098– 7102, 2021

  58. [66]

    A multi-dimensional deep structured state space approach to speech enhancement using small- footprint models,

    K. P-J, C. Yang, S. Siniscalchiet al., “A multi-dimensional deep structured state space approach to speech enhancement using small- footprint models,” inintrespeech, 2023, pp. 2453–2457

  59. [67]

    A two-stage framework in cross-spectrum domain for real-time speech enhancement,

    Y . Zhang, H. Zou, and J. Zhu, “A two-stage framework in cross-spectrum domain for real-time speech enhancement,” inIEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 12 587–12 591

  60. [68]

    Dual-signal transformation lstm network for real-time noise suppression,

    N. L. Westhausen and B. T. Meyer, “Dual-signal transformation lstm network for real-time noise suppression,” inInterspeech, 2021, pp. 280– 290

  61. [69]

    Iifc-net: A monaural speech enhancement network with high-order information interaction and fea- ture calibration,

    W. Wei, Y . Hu, H. Huang, and L. He, “Iifc-net: A monaural speech enhancement network with high-order information interaction and fea- ture calibration,”IEEE Signal Processing Letters, vol. 31, pp. 196–200, 2024

  62. [70]

    A mask free neural network for monaural speech enhancement,

    L. Liu, H. Guan, J. Ma, W. Dai, G. Wang, and S. Ding, “A mask free neural network for monaural speech enhancement,” inInterspeech, 2023, pp. 2468–2472

  63. [71]

    Sicrn: Advancing speech enhancement through state space model and inplace convolution techniques,

    C. Zhao, S. He, and X. Zhang, “Sicrn: Advancing speech enhancement through state space model and inplace convolution techniques,” inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 10 506–10 510

  64. [72]

    Exploiting bispectral features for single-channel speech enhancement,

    V . Parvathala, R. Gundluru, S. Sankala, and S. R. M. Kodukula, “Exploiting bispectral features for single-channel speech enhancement,” inInterspeech, 2025, pp. 2385–2389

  65. [73]

    Tsdt-net: Ultra- low-complexity two-stage model combining dual-path-transformer and transform-average-concatenate network for speech enhancement,

    Y . Gao, H. Chen, S. Zhang, Q. Yang, and J. Chen, “Tsdt-net: Ultra- low-complexity two-stage model combining dual-path-transformer and transform-average-concatenate network for speech enhancement,” in Interspeech, 2025, pp. 71–75

  66. [74]

    Per- ceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,

    A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Per- ceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in2001 IEEE international conference on acoustics, speech, and signal processing. Procee...

  67. [75]

    An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,

    J. Jensen and C. H. Taal, “An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,”IEEE/ACM Transac- tions on Audio, Speech, and Language Processing, vol. 24, no. 11, pp. 2009–2022, 2016

  68. [76]

    Evaluation of objective quality measures for speech enhancement,

    Y . Hu and P. C. Loizou, “Evaluation of objective quality measures for speech enhancement,”IEEE Transactions on audio, speech, and language processing, vol. 16, no. 1, pp. 229–238, 2007

  69. [77]

    Dccrn+: Channel-wise subband dccrn with snr estimation for speech enhancement,

    S. Lv, Y . Hu, S. Zhang, and L. Xie, “Dccrn+: Channel-wise subband dccrn with snr estimation for speech enhancement,” inInterspeech, 2021, pp. 2385–2389

  70. [78]

    Densely connected neural network with dilated convolutions for real-time speech enhancement in the time domain,

    A. Pandey and D. Wang, “Densely connected neural network with dilated convolutions for real-time speech enhancement in the time domain,” inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 6629–6633. Cunhang Fan(Member, IEEE) received...

Pith tools

Reviewed June 27, 2026 · model on record in the stance chip above.