REVIEW 3 major objections 4 minor 52 references
Neural Kalman Filters for Acoustic Echo Cancellation
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Neural Kalman filters beat classic FDKF on echo and reconvergence.
desk verdict Useful controlled comparison of neural Kalman AEC variants, but the headline claim overreaches what the statistics support. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the acoustic state-space model, a first-order Markov process $h(n+1)=a\,h(n)+\Delta h(n)$ for the room impulse response combined with the observation $y(n)=s(n)+n(n)+x^{\mathsf T}(n)h(n)$, whose optimal recursive estimator is the frequency-domain adaptive Kalman filter. The FDKF updates a per-bin filter state $\hat{H}_{\ell,k}=A\hat{H}_{\ell-1,k}+A K_{\ell,k} E_{\ell,k}$ with a Kalman gain $K_{\ell,k}$ that balances echo-path tracking against observation noise; its practical weakness is that the process and observation covariances must be supplied a priori. The machinery is to keep this update loop intact and replace individual blocks with DNNs -- the Kalman gain, the transition factor $A_\ell$, the filter-state postprocessing, or the reference signal (distortion model) -- and the comparison is made meaningful by training all variants in one framework with the same log-MSE echo loss and equalized algorithmic delay.
What would settle it
Run the four original released implementations (or author-provided code) with their original filter lengths, DFT sizes, losses, and training procedures on the paper's Dtest and DNL test sets and compare ERLE and (re)convergence against FDKF; if FDKF matches or beats the neural variants on echo reduction or double-talk near-end measures under those settings, the paper's ranking and design recommendations would not transfer to the published methods.
Extended reading notes
Core claim
The central claim is that the classical FDKF's limitations -- model linearity and the need to hand-set noise covariances -- can be relieved by letting DNNs take over one of four roles: estimating the Kalman gain, estimating the state-transition factor, postprocessing the filter-state update, or modeling loudspeaker distortion as a learned reference signal. Under a fair common framework (same delay, same effective reference length, same data and loss), each of the four neural Kalman filters reaches higher ERLE and faster (re)convergence than the FDKF on linear and nonlinear echo paths, and the gain-estimating variants match or exceed the FDKF's double-talk near-end preservation as measured by PESQ, STOI, and LPS. The authors further observe that methods with a learned distortion model behave like mask-based echo suppressors rather than subtractive cancellers, which explains their aggressive echo reduction but weaker near-end quality; per-bin processing with shared weights gives smaller parameter counts and flexible filter lengths, while fully connected hybrids are outpaced by an end-to-end network in most metrics and resources.
Load-bearing premise
The ranking assumes that the authors' re-implementations of DLAC-Kalman, NKF, NeuralKalman, and DeepAdaptive faithfully represent the original published methods, since every model was retrained from scratch with a common loss, modified filter lengths, DFT sizes, and (for NeuralKalman) a changed kernel, any of which could alter the outcome.
Editorial extensions
If this is right
- If the central claim is right, a small DNN estimating the Kalman gain (DLAC-Kalman or NKF) can improve FDKF's echo reduction and convergence without harming near-end speech, making it a practical low-parameter upgrade.
- The learned-distortion hybrids (NeuralKalman, DeepAdaptive) behave as mask-based echo suppressors, so they should be evaluated and deployed with that behavior in mind, not as pure system identifiers.
- The FDKF remains the minimal-complexity choice when compute or parameters are the binding constraint, while fully connected neural Kalman filters are dominated by end-to-end networks in most metrics and resources.
- Per-bin architectures decouple frame rate from update rate and can reuse one trained model at different filter lengths, which suits low-latency or variable-latency operation.
- Residual echo from nonlinearities and long echo tails remains, so a postfilter is still beneficial after neural Kalman filters.
Reading between the lines
- The paper's equalization of algorithmic delay and effective reference length, though fair, may favor OLS-based FDKF and penalize OLA-based hybrids as originally tuned; if original authors' settings change results, the ranking is implementation-sensitive.
- Because the trained distortion models can absorb the echo path (e.g., learning $\hat{H}=1$), the Kalman loop may decouple from the physical echo path; a testable extension is to constrain the distortion model with an auxiliary loss to keep the filter identifiable.
- The per-bin weight-sharing designs suggest a path to few-shot or variable-length AEC, where one compact network trained at a fixed DFT size operates at different filter lengths, an extension the paper notes but does not test.
- The observation that the AECMOS 'other degradation' metric underrates near-end distortion at low SER suggests that objective metrics may mislead hybrid-AEC comparisons; an ASR-based or listening-test evaluation on the paper's test sets would strengthen the near-end preservation ranking.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper revisits the frequency-domain adaptive Kalman filter (FDKF) for acoustic echo cancellation and presents a unified framework for four neural Kalman filter variants (DLAC-Kalman, NKF, NeuralKalman, DeepAdaptive) that replace or augment components such as Kalman gain, state transition, distortion modeling, or filter-state update with DNNs. It then reports a comparative evaluation against FDKF and an end-to-end neural baseline (CGGN16) under linear and nonlinear loudspeaker conditions, including double-talk, RIR switches, and WGN excitation. The central claims are that neural Kalman filters achieve faster (re)convergence, better echo cancellation, and in some cases better near-end speech preservation than FDKF, and that the comparison provides design guidance (e.g., per-bin vs. fully connected processing).
Significance. If the central claims are established, this is a valuable contribution: it provides a common mathematical notation and experimental framework for a recently emerging class of hybrid DNN/Kalman AEC methods, includes a careful attempt to equalize algorithmic delay and effective reference length, and releases code and evaluation scripts. The inclusion of an Oracle-FDKF upper bound, a fully data-driven CGGN16 baseline, complexity and parameter counts, and a language-independent LPS metric are commendable and strengthen the paper's utility as a benchmark. However, the comparative ranking that carries the conclusions is not statistically supported because the authors explicitly omit standard deviations, and the re-implementations of prior methods involve nontrivial modifications, so the significance of the specific rank order remains to be established.
major comments (3)
- [§IV-C, Fig. 7] The paper states on p. 14 that 'the broad range of SER values naturally leads to large standard deviations in most metrics. Accordingly, we do not report standard deviations.' This is load-bearing for the central claim because the ranking of methods is based on mean metric differences over 60 files per condition, and Fig. 7 presents bar charts with no error bars or significance tests. For example, the PESQ/STOI/LPS differences between DLAC-Kalman and FDKF in Fig. 7 appear small, and without confidence intervals the observed ordering cannot be distinguished from sampling noise. I request the authors to either report standard deviations or confidence intervals, or perform pairwise significance tests (e.g., matched-pairs bootstrap on the 60 files), and to adjust the conclusions to the strength of the statistical evidence.
- [§III-A–D and §IV-A] The manuscript states on p. 9 that the authors 'replicate the authors' methods as closely as possible,' but then lists departures: NeuralKalman is modified to use a 3x3 causal kernel, DLAC-Kalman is implemented in a narrowband variant, all models are retrained from scratch in a common PyTorch framework with the same loss, and DFT sizes and filter lengths are changed (K=512, 1408, 896) to equalize delay and effective reference length. These changes mean the reported ranks and design recommendations apply to the authors' re-implementations rather than necessarily to the original published methods. The paper should either validate the re-implementations against the original code/results or explicitly state in the conclusions that the ranking is specific to the re-implemented versions used here.
- [§IV-D, Fig. 6; §V] The conclusion that 'neural Kalman filters reveal better echo performance than FDKF along with better (re)convergence' is contradicted by the paper's own WGN results: in Fig. 6, DeepAdaptive and NeuralKalman are limited to roughly 10 dB ERLE while FDKF reaches a much higher final ERLE, and NKF/DLAC-Kalman also appear close to FDKF. The claim is therefore not uniformly supported across test sets. The conclusions should be qualified to the speech-excitation conditions, or the paper should provide a mechanism (e.g., an interaction analysis of method × excitation) that reconciles this condition with the overall claim.
minor comments (4)
- [§IV-F] The 'informal subjective listening' section reports impressions from 106 files but provides no details on the number of listeners, the rating scale, or any statistical treatment; it is anecdotal and should be labeled as such rather than being used to support the comparative ranking.
- [Throughout] The name 'DeepAdaptive' is occasionally written as 'Deep Adaptive' (e.g., p. 16, Fig. 5 caption); please standardize the spelling.
- [Eq. (13)] The observation noise variance smoothing factor β=0.5 is introduced without comment; a sentence or citation explaining this choice would improve reproducibility.
- [§IV-A] The definition of the effective reference input length M = K + (L−1)·R is helpful, but the notation R is used both for frame shift and for RIR length in the introduction (Eq. (1) uses N for RIR length); please ensure the symbols do not clash in the reader's mind.
Circularity Check
No significant circularity; comparative empirical study with self-citations that are not load-bearing, while missing significance testing is a statistical limitation, not circular reasoning.
full rationale
This paper is a comparative benchmark, not a derivation from first principles. The central claim that neural Kalman filters outperform FDKF is an empirical observation over test sets Dtest, DWGN_test, and DNL_test, each with 60 files drawn from TIMIT, the Aachen Impulse Response database, and ETSI noise. The compared methods (DLAC-Kalman [32], NKF [33], NeuralKalman [34], DeepAdaptive [35]) are external methods reimplemented in a common PyTorch framework; training uses the CSTR-VCTK corpus with a log-MSE AEC loss (Eq. 21) and early stopping on Ddev, with no test-set fitting. The FDKF baseline is classical and cited to [5]; although one co-author (Enzner) is an originator of FDKF, the baseline equations are standard and the claimed improvements are measured against independent external implementations, so the self-citation is not load-bearing. There is an explicit limitation: 'Note that the broad range of SER values naturally leads to large standard deviations in most metrics. Accordingly, we do not report standard deviations.' This weakens the statistical support for the headline ranking and belongs under correctness/statistical rigor, not circularity. Likewise, the statements 'we replicate the authors’ methods as closely as possible' and the described modifications to kernel size, DFT size, and filter taps raise external-validity concerns, but they do not make the result equal to its inputs by construction. No equation in the paper defines the predicted metric in terms of a fitted parameter or a self-cited uniqueness theorem; no circular step can be exhibited.
Assumptions & free parameters
assumptions (4)
- domain assumption First-order Markov model for the time-varying RIR (Eq. 4)
- standard math DFT-domain diagonalization approximation
- ad hoc to paper Re-implementations faithfully represent the original published methods
- domain assumption Training and test data represent real-world AEC conditions
Cite this review
Pith. "Pith review of Neural Kalman Filters for Acoustic Echo Cancellation." pith.science (2026). https://pith.science/paper/CZEENVZN
@misc{pith2026250116367,
author = {Pith},
title = {Pith review of: Neural Kalman Filters for Acoustic Echo Cancellation},
year = {2026},
howpublished = {\url{https://pith.science/paper/CZEENVZN}},
note = {Machine review of arXiv:2501.16367}
}
read the original abstract
Kalman filtering is a powerful approach to adaptive filtering for various problems in signal processing. The frequency-domain adaptive Kalman filter (FDKF), based on the concept of the acoustic state space, provides a unifying solution to the adaptive filter update and the related stepsize control. It was conceived for the problem of acoustic echo cancellation and, as such, is frequently applied in hands-free systems. This article motivates and briefly recapitulates the linear FDKF and investigates how it can be further supported by deep neural networks (DNNs) in various ways, specifically to overcome the challenges and limitations related to the usually required estimation of process and observation noise covariances for the Kalman filter. While the mere FDKF comes with very low computational complexity, its neural Kalman filter variants may deliver faster (re)convergence, better echo cancellation, and even exceed the FDKF in its excellent double-talk near-end speech preservation both under linear and nonlinear loudspeaker conditions. To provide a synopsis of the state of the art, this article contributes a comparison of a range of DNN-based extensions of FDKF in the same training framework and using the same data.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
E. Hänsler and G. Schmidt,Acoustic Echo and Noise Control: A Practical Approach. Hoboken, NJ, USA: John Wiley & Sons, Ltd, 2004
work page 2004
-
[2]
Analysis and Design of Multirate Systems for Cancellation of Acoustical Echoes,
W. Kellermann, “Analysis and Design of Multirate Systems for Cancellation of Acoustical Echoes,” inProc. of ICASSP, (New York, NY, USA), pp. 2570–2573, Apr. 1988
work page 1988
-
[3]
Acoustic Echo Control: An Application of Very-High-Order Adaptive Filters,
C. Breining, P. Dreiseitel, E. Hänsler, A. Mader, B. Nitsch, H. Puder, T. Schertler, G. Schmidt, and J. Tilp, “Acoustic Echo Control: An Application of Very-High-Order Adaptive Filters,” IEEE Signal Process. Mag. , vol. 16, pp. 42–69, July 1999
work page 1999
-
[4]
J. Benesty, T. Gänsler, D. Morgan, M. Sondhi, and S. Gay, Advances in Network and Acoustic Echo Cancellation . Berlin, Germany: Springer, 2001
work page 2001
-
[5]
Frequency-Domain Adaptive Kalman Filter for Acoustic Echo Control in Hands-Free Telephones,
G. Enzner and P. Vary, “Frequency-Domain Adaptive Kalman Filter for Acoustic Echo Control in Hands-Free Telephones,” Signal Processing, vol. 86, pp. 1140–1156, June 2006
work page 2006
-
[6]
Audio Signal Processing in the 21st Century: The Important Outcomes of the Past 25 Years,
G. Richard, P. Smaragdis, S. Gannot, P. A. Naylor, S. Makino, W. Kellermann, and A. Sugiyama, “Audio Signal Processing in the 21st Century: The Important Outcomes of the Past 25 Years,” IEEE Signal Process. Mag. , vol. 40, no. 5, pp. 12–26, 2023
work page 2023
-
[7]
AEC in A Netshell: On Target and Topology Choices for FCRN Acoustic Echo Cancellation,
J. Franzen, E. Seidel, and T. Fingscheidt, “AEC in A Netshell: On Target and Topology Choices for FCRN Acoustic Echo Cancellation,” in Proc. of ICASSP , (Toronto, ON, Canada), pp. 156–160, June 2021
work page 2021
-
[8]
Y2-Net FCRN for Acoustic Echo and Noise Suppression,
E. Seidel, J. Franzen, M. Strake, and T. Fingscheidt, “Y2-Net FCRN for Acoustic Echo and Noise Suppression,” inProc. of Interspeech, (Brno, Czech Republic), pp. 4763–4767, Oct. 2021
work page 2021
Show all 52 references
-
[9]
Task Splitting for DNN-Based Acoustic Echo and Noise Removal,
S. Braun and M. L. Valero, “Task Splitting for DNN-Based Acoustic Echo and Noise Removal,” in Proc. of IW AENC , (Bamberg, Germany), pp. 386–390, Sept. 2022
2022
-
[10]
Efficient Deep Acoustic Echo Suppression with Condition-Aware Training,
E. Seidel, P. Mowlaee, and T. Fingscheidt, “Efficient Deep Acoustic Echo Suppression with Condition-Aware Training,” in Proc. of W ASPAA, (New Paltz, NY, USA), pp. 1–5, Oct. 2023
2023
-
[11]
Deep Learning for Joint Acoustic Echo and Noise Cancellation with Nonlinear Distor- tions,
H. Zhang, K. Tan, and D. Wang, “Deep Learning for Joint Acoustic Echo and Noise Cancellation with Nonlinear Distor- tions,” in Proc. of Interspeech, (Graz, Austria), pp. 4255–4259, Sept. 2019
2019
-
[12]
Low-Complexity, Real-Time Joint Neural Echo Con- trol and Speech Enhancement Based On PercepNet,
J.-M. Valin, S. Tenneti, K. Helwani, U. Isik, and A. Krish- naswamy, “Low-Complexity, Real-Time Joint Neural Echo Con- trol and Speech Enhancement Based On PercepNet,” inProc. of ICASSP, (Toronto, ON, Canada), pp. 7133–7137, June 2021
2021
-
[13]
Combining Adaptive Filtering and Complex- Valued Deep Postfiltering for Acoustic Echo Cancellation,
M. Halimeh, T. Haubner, A. Briegleb, A. Schmidt, and W. Kellermann, “Combining Adaptive Filtering and Complex- Valued Deep Postfiltering for Acoustic Echo Cancellation,” in Proc. of ICASSP , (Toronto, ON, Canada), pp. 121–125, June 2021
2021
-
[14]
Deep Residual Echo Sup- pression with a Tunable Tradeoff Between Signal Distortion and Echo Suppression,
A. Ivry, I. Cohen, and B. Berdugo, “Deep Residual Echo Sup- pression with a Tunable Tradeoff Between Signal Distortion and Echo Suppression,” in Proc. of ICASSP , (Toronto, ON, Canada), pp. 126–130, June 2021
2021
-
[15]
KalmanNet: Neural Network Aided Kalman Filtering for Partially Known Dynamics,
G. Revach, N. Shlezinger, X. Ni, A. L. Escoriza, R. J. G. van Sloun, and Y. C. Eldar, “KalmanNet: Neural Network Aided Kalman Filtering for Partially Known Dynamics,”IEEE Trans. Sig. Proc., vol. 70, p. 1532–1547, Jan. 2022
2022
-
[16]
Meta-AF: Meta- Learning for Adaptive Filters,
J. Casebeer, N. J. Bryan, and P. Smaragdis, “Meta-AF: Meta- Learning for Adaptive Filters,”IEEE T-ASLP, vol. 31, pp. 355– 370, 2023
2023
-
[17]
Bayesian Inference Model for Applications of Time- Varying Acoustic System Identification,
G. Enzner, “Bayesian Inference Model for Applications of Time- Varying Acoustic System Identification,” inProc. of EUSIPCO, pp. 2126–2130, Aug. 2010
2010
-
[18]
Study of the General Kalman Filter for Echo Cancellation,
C. Paleologu, J. Benesty, and S. Ciochina, “Study of the General Kalman Filter for Echo Cancellation,”IEEE T-ASLP, vol. 21, pp. 1539–1549, Aug. 2013
2013
-
[19]
An Automotive Wideband Stereo Acoustic Echo Canceler Using Frequency- Domain Adaptive Filtering,
M.-A. Jung, S. Elshamy, and T. Fingscheidt, “An Automotive Wideband Stereo Acoustic Echo Canceler Using Frequency- Domain Adaptive Filtering,” in Proc. of EUSIPCO , (Lisbon, Portugal), pp. 1452–1456, Sept. 2014
2014
-
[20]
Frequency-Domain Adap- tive Kalman Filter with Fast Recovery of Abrupt Echo-Path Changes,
F. Yang, G. Enzner, and J. Yang, “Frequency-Domain Adap- tive Kalman Filter with Fast Recovery of Abrupt Echo-Path Changes,” IEEE SP Letters, vol. 24, pp. 1778–1782, June 2017
2017
-
[21]
Efficient Nonlinear Acoustic Echo Cancellation by Dual- Stage Multi-Channel Kalman Filtering,
M. Schrammen, S. Kühl, S. Markovich-Golan, and P. Jax, “Efficient Nonlinear Acoustic Echo Cancellation by Dual- Stage Multi-Channel Kalman Filtering,” inProc. of ICASSP , (Brighton, UK), pp. 975–979, May 2019
2019
-
[22]
ICASSP 2023 Acoustic Echo Cancellation Challenge,
R. Cutler, A. Saabas, T. Parnamaa, M. Purin, E. Indenbom, N.- C. Ristea, J. Guzhvin, H. Gamper, S. Braun, and R. Aichner, “ICASSP 2023 Acoustic Echo Cancellation Challenge,”arXiv preprint:2309.12553, Sept. 2023
2023 arXiv
-
[23]
An Efficient Residual Echo SupressionforMulti-ChannelAcousticEchoCancellationBased on the Frequency-Domain Adaptive Kalman Filter,
J. Franzen and T. Fingscheidt, “An Efficient Residual Echo SupressionforMulti-ChannelAcousticEchoCancellationBased on the Frequency-Domain Adaptive Kalman Filter,” inProc. of ICASSP, (Calgary, Canada), pp. 226–230, Apr. 2018
2018
-
[24]
Convergence and Performance Analysis of Classical, Hybrid, and Deep Acoustic Echo Control,
E. Seidel, P. Mowlaee, and T. Fingscheidt, “Convergence and Performance Analysis of Classical, Hybrid, and Deep Acoustic Echo Control,” IEEE T-ASLP, vol. 32, pp. 2857–2870, 2024
2024
-
[25]
Haykin, Adaptive Filter Theory
S. Haykin, Adaptive Filter Theory . Hoboken, NJ, USA: Prentice-Hall, 2002. 23
2002
-
[26]
S. L. Gay and J. Benesty, eds.,Acoustic Signal Processing for Telecommunication. USA: Kluwer Academic Publishers, 2000
2000
-
[27]
Vary and R
P. Vary and R. Martin,Digital Speech Transmission. Hoboken, NJ, USA: John Wiley & Sons, Ltd, 2006
2006
-
[28]
L. L. Scharf, Statistical Signal Processing . Addison-Wesley Publishing Company, 1991
1991
-
[29]
Recursive Bayesian Control of Multi- channel Acoustic Echo Cancellation,
S. Malik and G. Enzner, “Recursive Bayesian Control of Multi- channel Acoustic Echo Cancellation,”IEEE SP Letters, vol. 18, pp. 619–622, Nov. 2011
2011
-
[30]
Fast Implementation of LMS Adaptive Filters,
E. Ferrara, “Fast Implementation of LMS Adaptive Filters,” vol. 28, pp. 474–475, August 1980
1980
-
[31]
State-Space Architec- ture of the Partitioned-Block-Based Acoustic Echo Controller,
F. Kuech, E. Mabande, and G. Enzner, “State-Space Architec- ture of the Partitioned-Block-Based Acoustic Echo Controller,” in Proc. of ICASSP,(Florence,Italy),pp.1295–1299,May2014
-
[32]
End-to-End Deep Learning-Based Adaptation Control for Linear Acoustic Echo Cancellation,
T. Haubner, A. Brendel, and W. Kellermann, “End-to-End Deep Learning-Based Adaptation Control for Linear Acoustic Echo Cancellation,”IEEE T-ASLP, vol. 32, pp. 227–238, 2024
2024
-
[33]
Low- Complexity Acoustic Echo Cancellation with Neural Kalman Filtering,
D. Yang, F. Jiang, W. Wu, X. Fang, and M. Cao, “Low- Complexity Acoustic Echo Cancellation with Neural Kalman Filtering,” in Proc. of ICASSP , (Rhodes Island, Greece), pp. 7846–7850, June 2023
2023
-
[34]
NeuralKal- man: A Learnable Kalman Filter for Acoustic Echo Cancella- tion,
Y. Zhang, M. Yu, H. Zhang, D. Yu, and D. Wang, “NeuralKal- man: A Learnable Kalman Filter for Acoustic Echo Cancella- tion,” arXiv preprint:2301.12363, Jan. 2023
2023 arXiv
-
[35]
DeepAdaptiveAEC:HybridofDeepLearning and Adaptive Acoustic Echo Cancellation,
H. Zhang, S. Kandadai, H. Rao, M. Kim, T. Pruthi, and T.Kristjansson,“DeepAdaptiveAEC:HybridofDeepLearning and Adaptive Acoustic Echo Cancellation,” inProc. of ICASSP, (Singapore), pp. 756–760, May 2022
2022
-
[36]
Deep Filtering: Signal Ex- traction and Reconstruction Using Complex Time-Frequency Filters,
W. Mack and E. A. P. Habets, “Deep Filtering: Signal Ex- traction and Reconstruction Using Complex Time-Frequency Filters,” IEEE Signal Processing Letters , vol. 27, pp. 61–65, 2020
2020
-
[37]
AECMOS: A Speech Quality Assessment Metric for Echo Impairment,
M. Purin, S. Sootla, M. Sponza, A. Saabas, and R. Cutler, “AECMOS: A Speech Quality Assessment Metric for Echo Impairment,” in Proc. of ICASSP , (Singapore), pp. 901–905, May 2022
2022
-
[38]
P.862.2 Corrigendum 1: Wideband Extension to Rec
ITU-T Rec. P.862.2 Corrigendum 1: Wideband Extension to Rec. P.862 for the Assessment of Wideband Telephone Networks and Speech Codecs, Oct. 2017
2017
-
[39]
An Algorithm for Intelligibility Prediction of Time–Frequency Weighted Noisy Speech,
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “An Algorithm for Intelligibility Prediction of Time–Frequency Weighted Noisy Speech,” IEEE T-ASLP , vol. 19, no. 7, pp. 2125–2136, 2011
2011
-
[40]
PyTorch: An Imperative Style, High- Performance Deep Learning Library,
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, et al. , “PyTorch: An Imperative Style, High- Performance Deep Learning Library,” in Proc. of NeurIPS , (Vancouver, BC, Canada), pp. 8024–8035, Dec. 2019
2019
-
[41]
CSTR VCTK Corpus: English Multi-Speaker Corpus for CSTR Voice Cloning Toolkit
J. Yamagishi, C. Veaux, and K. MacDonald, “CSTR VCTK Corpus: English Multi-Speaker Corpus for CSTR Voice Cloning Toolkit.” University of Edinburgh. The Centre for Speech Tech- nology Research, 2017
2017
-
[42]
Tutorial: Loudspeaker Nonlinearities – Causes, Parameters, Symptoms,
W. Klippel, “Tutorial: Loudspeaker Nonlinearities – Causes, Parameters, Symptoms,” Journal of the Audio Engineering Society, vol. 54, pp. 907–939, Oct. 2006
2006
-
[43]
The Diverse Environ- ments Multi-Channel Acoustic Noise Database: A Database of Multichannel Environmental Noise Recordings,
J. Thiemann, N. Ito, and E. Vincent, “The Diverse Environ- ments Multi-Channel Acoustic Noise Database: A Database of Multichannel Environmental Noise Recordings,”The Journal of the Acoustical Society of America,vol.133,no.5,pp.3591–3591, 2013
2013
-
[44]
The QUT-NOISE-TIMIT Corpus for the Evaluation of Voice Activ- ity Detection Algorithms,
D. B. Dean, S. Sridharan, R. J. Vogt, and M. W. Mason, “The QUT-NOISE-TIMIT Corpus for the Evaluation of Voice Activ- ity Detection Algorithms,” inProc. of Interspeech, (Makuhari, Japan), p. 3110–3113, Sept. 2010
2010
-
[45]
Adam: A Method for Stochastic Optimization,
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” inProc. of ICLR, (San Diego, CA, USA), pp. 1– 15, May 2015
2015
-
[46]
TIMIT Acoustic-Phonetic Continous Speech Corpus
J.S.Garofolo,L.F.Lamel,W.M.Fisher,J.G.Fiscus,andD.S. Pallett, “TIMIT Acoustic-Phonetic Continous Speech Corpus.” Linguistic Data Consortium, Philadelphia, PA, USA, 1993
1993
-
[47]
Do We Need Dereverberation for Hand-Held Telephony?,
M. Jeub, M. Schäfer, H. Krüger, C. M. Nelke, C. Beaugeant, and P. Vary, “Do We Need Dereverberation for Hand-Held Telephony?,” in Proc. of ICA , (Sydney, Australia), pp. 3793– 3799, Aug. 2010
2010
-
[48]
ETSI EG 202 396-1,Speech Processing, Transmission and Qual- ity Aspects (STQ); Speech Quality Performance in the Presence of Background Noise; Part 1: Background Noise Simulation Technique and Background Noise Database , Sept. 2008
2008
-
[49]
Deep Learning for Acoustic Echo Cancellation in Noisy and Double-Talk Scenarios,
H. Zhang and D. Wang, “Deep Learning for Acoustic Echo Cancellation in Noisy and Double-Talk Scenarios,” inProc. of Interspeech, (Hyderabad, India), pp. 3239–3243, Sept. 2018
2018
-
[50]
Evaluation Metrics for Gener- ative Speech Enhancement Methods: Issues and Perspectives,
J. Pirklbauer, M. Sach, K. Fluyt, W. Tirry, W. Wardah, S. Moeller, and T. Fingscheidt, “Evaluation Metrics for Gener- ative Speech Enhancement Methods: Issues and Perspectives,” in Proc. of 15th ITG Conference on Speech Communication , (Aachen, Germany), pp. 265–269, Sep 2023
2023
-
[51]
Double-Talk Detection- Aided Residual Echo Suppression via Spectrogram Masking and Refinement,
E. Shachar, I. Cohen, and B. Berdugo, “Double-Talk Detection- Aided Residual Echo Suppression via Spectrogram Masking and Refinement,” Acoustics, vol. 4, no. 3, pp. 637–655, 2022
2022
-
[52]
Scaling Up Adap- tive Filter Optimizers,
J. Casebeer, N. J. Bryan, and P. Smaragdis, “Scaling Up Adap- tive Filter Optimizers,”arXiv preprint:2403.00977, Mar. 2024. 24 Biographies Ernst Seidel (e.seidel@tu-bs.de) received the M.Sc. degree in electrical engineering in 2021 from Technische Universität Braunschweig, Ger...
2006 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.