Pith. sign in

REVIEW 3 major objections 4 minor 52 references

Neural Kalman Filters for Acoustic Echo Cancellation

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Neural Kalman filters beat classic FDKF on echo and reconvergence.

desk verdict Useful controlled comparison of neural Kalman AEC variants, but the headline claim overreaches what the statistics support. read the letter →

arxiv 2501.16367 v1 pith:CZEENVZN submitted 2025-01-22 eess.AS cs.SD

classification eess.AScs.SD
keywords acousticechocancellationfrequency-domainadaptiveKalmanfilterneuraldouble-talkdeeplearningreturnlossenhancementsystemidentification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that replacing parts of the frequency-domain adaptive Kalman filter (FDKF) with small neural networks yields faster convergence and reconvergence, stronger echo cancellation, and comparable or better near-end speech preservation during double-talk, both for linear and nonlinear loudspeakers. The authors unify four recent hybrid designs in a common state-space notation, retrain them from scratch in one framework on identical data, and evaluate them against the classical FDKF and a fully data-driven baseline. They report that the hybrid methods outperform FDKF in echo-return loss enhancement and reconvergence after room-impulse-response changes, while per-bin designs (DLAC-Kalman and NKF) preserve near-end speech best. If true, the practical takeaway is that a small network estimating the Kalman gain can improve the model-based FDKF without sacrificing its double-talk robustness, and that fully connected hybrids lose to end-to-end networks.

What carries the argument

The central object is the acoustic state-space model, a first-order Markov process $h(n+1)=a\,h(n)+\Delta h(n)$ for the room impulse response combined with the observation $y(n)=s(n)+n(n)+x^{\mathsf T}(n)h(n)$, whose optimal recursive estimator is the frequency-domain adaptive Kalman filter. The FDKF updates a per-bin filter state $\hat{H}_{\ell,k}=A\hat{H}_{\ell-1,k}+A K_{\ell,k} E_{\ell,k}$ with a Kalman gain $K_{\ell,k}$ that balances echo-path tracking against observation noise; its practical weakness is that the process and observation covariances must be supplied a priori. The machinery is to keep this update loop intact and replace individual blocks with DNNs -- the Kalman gain, the transition factor $A_\ell$, the filter-state postprocessing, or the reference signal (distortion model) -- and the comparison is made meaningful by training all variants in one framework with the same log-MSE echo loss and equalized algorithmic delay.

What would settle it

Run the four original released implementations (or author-provided code) with their original filter lengths, DFT sizes, losses, and training procedures on the paper's Dtest and DNL test sets and compare ERLE and (re)convergence against FDKF; if FDKF matches or beats the neural variants on echo reduction or double-talk near-end measures under those settings, the paper's ranking and design recommendations would not transfer to the published methods.

Watch

Extended reading notes

Core claim

The central claim is that the classical FDKF's limitations -- model linearity and the need to hand-set noise covariances -- can be relieved by letting DNNs take over one of four roles: estimating the Kalman gain, estimating the state-transition factor, postprocessing the filter-state update, or modeling loudspeaker distortion as a learned reference signal. Under a fair common framework (same delay, same effective reference length, same data and loss), each of the four neural Kalman filters reaches higher ERLE and faster (re)convergence than the FDKF on linear and nonlinear echo paths, and the gain-estimating variants match or exceed the FDKF's double-talk near-end preservation as measured by PESQ, STOI, and LPS. The authors further observe that methods with a learned distortion model behave like mask-based echo suppressors rather than subtractive cancellers, which explains their aggressive echo reduction but weaker near-end quality; per-bin processing with shared weights gives smaller parameter counts and flexible filter lengths, while fully connected hybrids are outpaced by an end-to-end network in most metrics and resources.

Load-bearing premise

The ranking assumes that the authors' re-implementations of DLAC-Kalman, NKF, NeuralKalman, and DeepAdaptive faithfully represent the original published methods, since every model was retrained from scratch with a common loss, modified filter lengths, DFT sizes, and (for NeuralKalman) a changed kernel, any of which could alter the outcome.

Editorial extensions

If this is right

  • If the central claim is right, a small DNN estimating the Kalman gain (DLAC-Kalman or NKF) can improve FDKF's echo reduction and convergence without harming near-end speech, making it a practical low-parameter upgrade.
  • The learned-distortion hybrids (NeuralKalman, DeepAdaptive) behave as mask-based echo suppressors, so they should be evaluated and deployed with that behavior in mind, not as pure system identifiers.
  • The FDKF remains the minimal-complexity choice when compute or parameters are the binding constraint, while fully connected neural Kalman filters are dominated by end-to-end networks in most metrics and resources.
  • Per-bin architectures decouple frame rate from update rate and can reuse one trained model at different filter lengths, which suits low-latency or variable-latency operation.
  • Residual echo from nonlinearities and long echo tails remains, so a postfilter is still beneficial after neural Kalman filters.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's equalization of algorithmic delay and effective reference length, though fair, may favor OLS-based FDKF and penalize OLA-based hybrids as originally tuned; if original authors' settings change results, the ranking is implementation-sensitive.
  • Because the trained distortion models can absorb the echo path (e.g., learning $\hat{H}=1$), the Kalman loop may decouple from the physical echo path; a testable extension is to constrain the distortion model with an auxiliary loss to keep the filter identifiable.
  • The per-bin weight-sharing designs suggest a path to few-shot or variable-length AEC, where one compact network trained at a fixed DFT size operates at different filter lengths, an extension the paper notes but does not test.
  • The observation that the AECMOS 'other degradation' metric underrates near-end distortion at low SER suggests that objective metrics may mislead hybrid-AEC comparisons; an ASR-based or listening-test evaluation on the paper's test sets would strengthen the near-end preservation ranking.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper revisits the frequency-domain adaptive Kalman filter (FDKF) for acoustic echo cancellation and presents a unified framework for four neural Kalman filter variants (DLAC-Kalman, NKF, NeuralKalman, DeepAdaptive) that replace or augment components such as Kalman gain, state transition, distortion modeling, or filter-state update with DNNs. It then reports a comparative evaluation against FDKF and an end-to-end neural baseline (CGGN16) under linear and nonlinear loudspeaker conditions, including double-talk, RIR switches, and WGN excitation. The central claims are that neural Kalman filters achieve faster (re)convergence, better echo cancellation, and in some cases better near-end speech preservation than FDKF, and that the comparison provides design guidance (e.g., per-bin vs. fully connected processing).

Significance. If the central claims are established, this is a valuable contribution: it provides a common mathematical notation and experimental framework for a recently emerging class of hybrid DNN/Kalman AEC methods, includes a careful attempt to equalize algorithmic delay and effective reference length, and releases code and evaluation scripts. The inclusion of an Oracle-FDKF upper bound, a fully data-driven CGGN16 baseline, complexity and parameter counts, and a language-independent LPS metric are commendable and strengthen the paper's utility as a benchmark. However, the comparative ranking that carries the conclusions is not statistically supported because the authors explicitly omit standard deviations, and the re-implementations of prior methods involve nontrivial modifications, so the significance of the specific rank order remains to be established.

major comments (3)
  1. [§IV-C, Fig. 7] The paper states on p. 14 that 'the broad range of SER values naturally leads to large standard deviations in most metrics. Accordingly, we do not report standard deviations.' This is load-bearing for the central claim because the ranking of methods is based on mean metric differences over 60 files per condition, and Fig. 7 presents bar charts with no error bars or significance tests. For example, the PESQ/STOI/LPS differences between DLAC-Kalman and FDKF in Fig. 7 appear small, and without confidence intervals the observed ordering cannot be distinguished from sampling noise. I request the authors to either report standard deviations or confidence intervals, or perform pairwise significance tests (e.g., matched-pairs bootstrap on the 60 files), and to adjust the conclusions to the strength of the statistical evidence.
  2. [§III-A–D and §IV-A] The manuscript states on p. 9 that the authors 'replicate the authors' methods as closely as possible,' but then lists departures: NeuralKalman is modified to use a 3x3 causal kernel, DLAC-Kalman is implemented in a narrowband variant, all models are retrained from scratch in a common PyTorch framework with the same loss, and DFT sizes and filter lengths are changed (K=512, 1408, 896) to equalize delay and effective reference length. These changes mean the reported ranks and design recommendations apply to the authors' re-implementations rather than necessarily to the original published methods. The paper should either validate the re-implementations against the original code/results or explicitly state in the conclusions that the ranking is specific to the re-implemented versions used here.
  3. [§IV-D, Fig. 6; §V] The conclusion that 'neural Kalman filters reveal better echo performance than FDKF along with better (re)convergence' is contradicted by the paper's own WGN results: in Fig. 6, DeepAdaptive and NeuralKalman are limited to roughly 10 dB ERLE while FDKF reaches a much higher final ERLE, and NKF/DLAC-Kalman also appear close to FDKF. The claim is therefore not uniformly supported across test sets. The conclusions should be qualified to the speech-excitation conditions, or the paper should provide a mechanism (e.g., an interaction analysis of method × excitation) that reconciles this condition with the overall claim.
minor comments (4)
  1. [§IV-F] The 'informal subjective listening' section reports impressions from 106 files but provides no details on the number of listeners, the rating scale, or any statistical treatment; it is anecdotal and should be labeled as such rather than being used to support the comparative ranking.
  2. [Throughout] The name 'DeepAdaptive' is occasionally written as 'Deep Adaptive' (e.g., p. 16, Fig. 5 caption); please standardize the spelling.
  3. [Eq. (13)] The observation noise variance smoothing factor β=0.5 is introduced without comment; a sentence or citation explaining this choice would improve reproducibility.
  4. [§IV-A] The definition of the effective reference input length M = K + (L−1)·R is helpful, but the notation R is used both for frame shift and for RIR length in the introduction (Eq. (1) uses N for RIR length); please ensure the symbols do not clash in the reader's mind.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity; comparative empirical study with self-citations that are not load-bearing, while missing significance testing is a statistical limitation, not circular reasoning.

full rationale

This paper is a comparative benchmark, not a derivation from first principles. The central claim that neural Kalman filters outperform FDKF is an empirical observation over test sets Dtest, DWGN_test, and DNL_test, each with 60 files drawn from TIMIT, the Aachen Impulse Response database, and ETSI noise. The compared methods (DLAC-Kalman [32], NKF [33], NeuralKalman [34], DeepAdaptive [35]) are external methods reimplemented in a common PyTorch framework; training uses the CSTR-VCTK corpus with a log-MSE AEC loss (Eq. 21) and early stopping on Ddev, with no test-set fitting. The FDKF baseline is classical and cited to [5]; although one co-author (Enzner) is an originator of FDKF, the baseline equations are standard and the claimed improvements are measured against independent external implementations, so the self-citation is not load-bearing. There is an explicit limitation: 'Note that the broad range of SER values naturally leads to large standard deviations in most metrics. Accordingly, we do not report standard deviations.' This weakens the statistical support for the headline ranking and belongs under correctness/statistical rigor, not circularity. Likewise, the statements 'we replicate the authors’ methods as closely as possible' and the described modifications to kernel size, DFT size, and filter taps raise external-validity concerns, but they do not make the result equal to its inputs by construction. No equation in the paper defines the predicted metric in terms of a fitted parameter or a self-cited uniqueness theorem; no circular step can be exhibited.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claim is a comparative evaluation, so the ledger lists the assumptions behind the FDKF model and the fidelity of the re-implementations rather than fitted parameters. No new entities are postulated.

assumptions (4)
  • domain assumption First-order Markov model for the time-varying RIR (Eq. 4)
    The FDKF and its neural variants assume the acoustic echo path evolves as h(n+1)=a*h(n)+deltah(n); if the real echo path dynamics violate this, the Kalman gain interpretation is invalid.
  • standard math DFT-domain diagonalization approximation
    FDKF relies on the covariance diagonalization property in the DFT domain to treat frequency bins independently; this is an approximation for finite frames.
  • ad hoc to paper Re-implementations faithfully represent the original published methods
    The comparative claim depends on the authors' implementations of DLAC-Kalman, NKF, NeuralKalman, and DeepAdaptive being accurate, despite modifications such as a 3x3 causal kernel for NeuralKalman and retraining from scratch in a common framework.
  • domain assumption Training and test data represent real-world AEC conditions
    The evaluation uses simulated training data (VCTK, synthetic RIRs, SEF nonlinearity) and test data (TIMIT, Aachen RIRs, sigmoidal nonlinearity); conclusions are limited to these conditions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Neural Kalman Filters for Acoustic Echo Cancellation." pith.science (2026). https://pith.science/paper/CZEENVZN

@misc{pith2026250116367,
  author       = {Pith},
  title        = {Pith review of: Neural Kalman Filters for Acoustic Echo Cancellation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CZEENVZN}},
  note         = {Machine review of arXiv:2501.16367}
}
read the original abstract

Kalman filtering is a powerful approach to adaptive filtering for various problems in signal processing. The frequency-domain adaptive Kalman filter (FDKF), based on the concept of the acoustic state space, provides a unifying solution to the adaptive filter update and the related stepsize control. It was conceived for the problem of acoustic echo cancellation and, as such, is frequently applied in hands-free systems. This article motivates and briefly recapitulates the linear FDKF and investigates how it can be further supported by deep neural networks (DNNs) in various ways, specifically to overcome the challenges and limitations related to the usually required estimation of process and observation noise covariances for the Kalman filter. While the mere FDKF comes with very low computational complexity, its neural Kalman filter variants may deliver faster (re)convergence, better echo cancellation, and even exceed the FDKF in its excellent double-talk near-end speech preservation both under linear and nonlinear loudspeaker conditions. To provide a synopsis of the state of the art, this article contributes a comparison of a range of DNN-based extensions of FDKF in the same training framework and using the same data.

Figures

Figures reproduced from arXiv: 2501.16367 by the authors.

Figure 1
Figure 1. Here, a far-end (FE) speaker’s voice, i.e., the reference signal [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 1
Figure 1. Generalized overview of a hands-free system / speakerphone. The speech signal of the far-end speaker is played [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Overview of a (neural) Kalman filter approach for acoustic echo cancellation. It operates in the discrete Fourier [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figures from the paper (5 more)
Figure 3
Figure 3. Figure 3: Details of the Kalman algorithm block as used in the [PITH_FULL_IMAGE:figures/full_fig_p010_3.png]
Figure 4
Figure 4. Figure 4: Model performance for an example file from test set [PITH_FULL_IMAGE:figures/full_fig_p015_4.png]
Figure 5
Figure 5. Figure 5: Model performance averaged over the STFE sections of all files with far-end speech excitation in test set Dtest. No nonlinearities are employed, but an RIR switch after 4s. strength of NeuralKalman on low-power signals now converts to the opposite, as now a NE signal d…
Figure 6
Figure 6. Figure 6: Model performance averaged over the STFE sections of all files with white Gaussian noise excitation in test set DWGN test . No nonlinearities are employed, but an RIR switch after 4s. nonlinearity, WGN excitation), hybrid models can outperform the classical FDKF in con…
Figure 7
Figure 7. Figure 7: Model performance averaged over the DT sections of our test sets without nonlinear distortions (Dtest, blue bars) and with nonlinear distortions (DNL test, red bars) of the loudspeaker signal. RIR switches take place in the middle of the DT sections. Metrics shown are …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

52 extracted references · 50 canonical work pages

  1. [1]

    Hänsler and G

    E. Hänsler and G. Schmidt,Acoustic Echo and Noise Control: A Practical Approach. Hoboken, NJ, USA: John Wiley & Sons, Ltd, 2004

  2. [2]

    Analysis and Design of Multirate Systems for Cancellation of Acoustical Echoes,

    W. Kellermann, “Analysis and Design of Multirate Systems for Cancellation of Acoustical Echoes,” inProc. of ICASSP, (New York, NY, USA), pp. 2570–2573, Apr. 1988

  3. [3]

    Acoustic Echo Control: An Application of Very-High-Order Adaptive Filters,

    C. Breining, P. Dreiseitel, E. Hänsler, A. Mader, B. Nitsch, H. Puder, T. Schertler, G. Schmidt, and J. Tilp, “Acoustic Echo Control: An Application of Very-High-Order Adaptive Filters,” IEEE Signal Process. Mag. , vol. 16, pp. 42–69, July 1999

  4. [4]

    Benesty, T

    J. Benesty, T. Gänsler, D. Morgan, M. Sondhi, and S. Gay, Advances in Network and Acoustic Echo Cancellation . Berlin, Germany: Springer, 2001

  5. [5]

    Frequency-Domain Adaptive Kalman Filter for Acoustic Echo Control in Hands-Free Telephones,

    G. Enzner and P. Vary, “Frequency-Domain Adaptive Kalman Filter for Acoustic Echo Control in Hands-Free Telephones,” Signal Processing, vol. 86, pp. 1140–1156, June 2006

  6. [6]

    Audio Signal Processing in the 21st Century: The Important Outcomes of the Past 25 Years,

    G. Richard, P. Smaragdis, S. Gannot, P. A. Naylor, S. Makino, W. Kellermann, and A. Sugiyama, “Audio Signal Processing in the 21st Century: The Important Outcomes of the Past 25 Years,” IEEE Signal Process. Mag. , vol. 40, no. 5, pp. 12–26, 2023

  7. [7]

    AEC in A Netshell: On Target and Topology Choices for FCRN Acoustic Echo Cancellation,

    J. Franzen, E. Seidel, and T. Fingscheidt, “AEC in A Netshell: On Target and Topology Choices for FCRN Acoustic Echo Cancellation,” in Proc. of ICASSP , (Toronto, ON, Canada), pp. 156–160, June 2021

  8. [8]

    Y2-Net FCRN for Acoustic Echo and Noise Suppression,

    E. Seidel, J. Franzen, M. Strake, and T. Fingscheidt, “Y2-Net FCRN for Acoustic Echo and Noise Suppression,” inProc. of Interspeech, (Brno, Czech Republic), pp. 4763–4767, Oct. 2021

Show all 52 references
  1. [9]

    Task Splitting for DNN-Based Acoustic Echo and Noise Removal,

    S. Braun and M. L. Valero, “Task Splitting for DNN-Based Acoustic Echo and Noise Removal,” in Proc. of IW AENC , (Bamberg, Germany), pp. 386–390, Sept. 2022

  2. [10]

    Efficient Deep Acoustic Echo Suppression with Condition-Aware Training,

    E. Seidel, P. Mowlaee, and T. Fingscheidt, “Efficient Deep Acoustic Echo Suppression with Condition-Aware Training,” in Proc. of W ASPAA, (New Paltz, NY, USA), pp. 1–5, Oct. 2023

  3. [11]

    Deep Learning for Joint Acoustic Echo and Noise Cancellation with Nonlinear Distor- tions,

    H. Zhang, K. Tan, and D. Wang, “Deep Learning for Joint Acoustic Echo and Noise Cancellation with Nonlinear Distor- tions,” in Proc. of Interspeech, (Graz, Austria), pp. 4255–4259, Sept. 2019

  4. [12]

    Low-Complexity, Real-Time Joint Neural Echo Con- trol and Speech Enhancement Based On PercepNet,

    J.-M. Valin, S. Tenneti, K. Helwani, U. Isik, and A. Krish- naswamy, “Low-Complexity, Real-Time Joint Neural Echo Con- trol and Speech Enhancement Based On PercepNet,” inProc. of ICASSP, (Toronto, ON, Canada), pp. 7133–7137, June 2021

  5. [13]

    Combining Adaptive Filtering and Complex- Valued Deep Postfiltering for Acoustic Echo Cancellation,

    M. Halimeh, T. Haubner, A. Briegleb, A. Schmidt, and W. Kellermann, “Combining Adaptive Filtering and Complex- Valued Deep Postfiltering for Acoustic Echo Cancellation,” in Proc. of ICASSP , (Toronto, ON, Canada), pp. 121–125, June 2021

  6. [14]

    Deep Residual Echo Sup- pression with a Tunable Tradeoff Between Signal Distortion and Echo Suppression,

    A. Ivry, I. Cohen, and B. Berdugo, “Deep Residual Echo Sup- pression with a Tunable Tradeoff Between Signal Distortion and Echo Suppression,” in Proc. of ICASSP , (Toronto, ON, Canada), pp. 126–130, June 2021

  7. [15]

    KalmanNet: Neural Network Aided Kalman Filtering for Partially Known Dynamics,

    G. Revach, N. Shlezinger, X. Ni, A. L. Escoriza, R. J. G. van Sloun, and Y. C. Eldar, “KalmanNet: Neural Network Aided Kalman Filtering for Partially Known Dynamics,”IEEE Trans. Sig. Proc., vol. 70, p. 1532–1547, Jan. 2022

  8. [16]

    Meta-AF: Meta- Learning for Adaptive Filters,

    J. Casebeer, N. J. Bryan, and P. Smaragdis, “Meta-AF: Meta- Learning for Adaptive Filters,”IEEE T-ASLP, vol. 31, pp. 355– 370, 2023

  9. [17]

    Bayesian Inference Model for Applications of Time- Varying Acoustic System Identification,

    G. Enzner, “Bayesian Inference Model for Applications of Time- Varying Acoustic System Identification,” inProc. of EUSIPCO, pp. 2126–2130, Aug. 2010

  10. [18]

    Study of the General Kalman Filter for Echo Cancellation,

    C. Paleologu, J. Benesty, and S. Ciochina, “Study of the General Kalman Filter for Echo Cancellation,”IEEE T-ASLP, vol. 21, pp. 1539–1549, Aug. 2013

  11. [19]

    An Automotive Wideband Stereo Acoustic Echo Canceler Using Frequency- Domain Adaptive Filtering,

    M.-A. Jung, S. Elshamy, and T. Fingscheidt, “An Automotive Wideband Stereo Acoustic Echo Canceler Using Frequency- Domain Adaptive Filtering,” in Proc. of EUSIPCO , (Lisbon, Portugal), pp. 1452–1456, Sept. 2014

  12. [20]

    Frequency-Domain Adap- tive Kalman Filter with Fast Recovery of Abrupt Echo-Path Changes,

    F. Yang, G. Enzner, and J. Yang, “Frequency-Domain Adap- tive Kalman Filter with Fast Recovery of Abrupt Echo-Path Changes,” IEEE SP Letters, vol. 24, pp. 1778–1782, June 2017

  13. [21]

    Efficient Nonlinear Acoustic Echo Cancellation by Dual- Stage Multi-Channel Kalman Filtering,

    M. Schrammen, S. Kühl, S. Markovich-Golan, and P. Jax, “Efficient Nonlinear Acoustic Echo Cancellation by Dual- Stage Multi-Channel Kalman Filtering,” inProc. of ICASSP , (Brighton, UK), pp. 975–979, May 2019

  14. [22]

    ICASSP 2023 Acoustic Echo Cancellation Challenge,

    R. Cutler, A. Saabas, T. Parnamaa, M. Purin, E. Indenbom, N.- C. Ristea, J. Guzhvin, H. Gamper, S. Braun, and R. Aichner, “ICASSP 2023 Acoustic Echo Cancellation Challenge,”arXiv preprint:2309.12553, Sept. 2023

  15. [23]

    An Efficient Residual Echo SupressionforMulti-ChannelAcousticEchoCancellationBased on the Frequency-Domain Adaptive Kalman Filter,

    J. Franzen and T. Fingscheidt, “An Efficient Residual Echo SupressionforMulti-ChannelAcousticEchoCancellationBased on the Frequency-Domain Adaptive Kalman Filter,” inProc. of ICASSP, (Calgary, Canada), pp. 226–230, Apr. 2018

  16. [24]

    Convergence and Performance Analysis of Classical, Hybrid, and Deep Acoustic Echo Control,

    E. Seidel, P. Mowlaee, and T. Fingscheidt, “Convergence and Performance Analysis of Classical, Hybrid, and Deep Acoustic Echo Control,” IEEE T-ASLP, vol. 32, pp. 2857–2870, 2024

  17. [25]

    Haykin, Adaptive Filter Theory

    S. Haykin, Adaptive Filter Theory . Hoboken, NJ, USA: Prentice-Hall, 2002. 23

  18. [26]

    S. L. Gay and J. Benesty, eds.,Acoustic Signal Processing for Telecommunication. USA: Kluwer Academic Publishers, 2000

  19. [27]

    Vary and R

    P. Vary and R. Martin,Digital Speech Transmission. Hoboken, NJ, USA: John Wiley & Sons, Ltd, 2006

  20. [28]

    L. L. Scharf, Statistical Signal Processing . Addison-Wesley Publishing Company, 1991

  21. [29]

    Recursive Bayesian Control of Multi- channel Acoustic Echo Cancellation,

    S. Malik and G. Enzner, “Recursive Bayesian Control of Multi- channel Acoustic Echo Cancellation,”IEEE SP Letters, vol. 18, pp. 619–622, Nov. 2011

  22. [30]

    Fast Implementation of LMS Adaptive Filters,

    E. Ferrara, “Fast Implementation of LMS Adaptive Filters,” vol. 28, pp. 474–475, August 1980

  23. [31]

    State-Space Architec- ture of the Partitioned-Block-Based Acoustic Echo Controller,

    F. Kuech, E. Mabande, and G. Enzner, “State-Space Architec- ture of the Partitioned-Block-Based Acoustic Echo Controller,” in Proc. of ICASSP,(Florence,Italy),pp.1295–1299,May2014

  24. [32]

    End-to-End Deep Learning-Based Adaptation Control for Linear Acoustic Echo Cancellation,

    T. Haubner, A. Brendel, and W. Kellermann, “End-to-End Deep Learning-Based Adaptation Control for Linear Acoustic Echo Cancellation,”IEEE T-ASLP, vol. 32, pp. 227–238, 2024

  25. [33]

    Low- Complexity Acoustic Echo Cancellation with Neural Kalman Filtering,

    D. Yang, F. Jiang, W. Wu, X. Fang, and M. Cao, “Low- Complexity Acoustic Echo Cancellation with Neural Kalman Filtering,” in Proc. of ICASSP , (Rhodes Island, Greece), pp. 7846–7850, June 2023

  26. [34]

    NeuralKal- man: A Learnable Kalman Filter for Acoustic Echo Cancella- tion,

    Y. Zhang, M. Yu, H. Zhang, D. Yu, and D. Wang, “NeuralKal- man: A Learnable Kalman Filter for Acoustic Echo Cancella- tion,” arXiv preprint:2301.12363, Jan. 2023

  27. [35]

    DeepAdaptiveAEC:HybridofDeepLearning and Adaptive Acoustic Echo Cancellation,

    H. Zhang, S. Kandadai, H. Rao, M. Kim, T. Pruthi, and T.Kristjansson,“DeepAdaptiveAEC:HybridofDeepLearning and Adaptive Acoustic Echo Cancellation,” inProc. of ICASSP, (Singapore), pp. 756–760, May 2022

  28. [36]

    Deep Filtering: Signal Ex- traction and Reconstruction Using Complex Time-Frequency Filters,

    W. Mack and E. A. P. Habets, “Deep Filtering: Signal Ex- traction and Reconstruction Using Complex Time-Frequency Filters,” IEEE Signal Processing Letters , vol. 27, pp. 61–65, 2020

  29. [37]

    AECMOS: A Speech Quality Assessment Metric for Echo Impairment,

    M. Purin, S. Sootla, M. Sponza, A. Saabas, and R. Cutler, “AECMOS: A Speech Quality Assessment Metric for Echo Impairment,” in Proc. of ICASSP , (Singapore), pp. 901–905, May 2022

  30. [38]

    P.862.2 Corrigendum 1: Wideband Extension to Rec

    ITU-T Rec. P.862.2 Corrigendum 1: Wideband Extension to Rec. P.862 for the Assessment of Wideband Telephone Networks and Speech Codecs, Oct. 2017

  31. [39]

    An Algorithm for Intelligibility Prediction of Time–Frequency Weighted Noisy Speech,

    C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “An Algorithm for Intelligibility Prediction of Time–Frequency Weighted Noisy Speech,” IEEE T-ASLP , vol. 19, no. 7, pp. 2125–2136, 2011

  32. [40]

    PyTorch: An Imperative Style, High- Performance Deep Learning Library,

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, et al. , “PyTorch: An Imperative Style, High- Performance Deep Learning Library,” in Proc. of NeurIPS , (Vancouver, BC, Canada), pp. 8024–8035, Dec. 2019

  33. [41]

    CSTR VCTK Corpus: English Multi-Speaker Corpus for CSTR Voice Cloning Toolkit

    J. Yamagishi, C. Veaux, and K. MacDonald, “CSTR VCTK Corpus: English Multi-Speaker Corpus for CSTR Voice Cloning Toolkit.” University of Edinburgh. The Centre for Speech Tech- nology Research, 2017

  34. [42]

    Tutorial: Loudspeaker Nonlinearities – Causes, Parameters, Symptoms,

    W. Klippel, “Tutorial: Loudspeaker Nonlinearities – Causes, Parameters, Symptoms,” Journal of the Audio Engineering Society, vol. 54, pp. 907–939, Oct. 2006

  35. [43]

    The Diverse Environ- ments Multi-Channel Acoustic Noise Database: A Database of Multichannel Environmental Noise Recordings,

    J. Thiemann, N. Ito, and E. Vincent, “The Diverse Environ- ments Multi-Channel Acoustic Noise Database: A Database of Multichannel Environmental Noise Recordings,”The Journal of the Acoustical Society of America,vol.133,no.5,pp.3591–3591, 2013

  36. [44]

    The QUT-NOISE-TIMIT Corpus for the Evaluation of Voice Activ- ity Detection Algorithms,

    D. B. Dean, S. Sridharan, R. J. Vogt, and M. W. Mason, “The QUT-NOISE-TIMIT Corpus for the Evaluation of Voice Activ- ity Detection Algorithms,” inProc. of Interspeech, (Makuhari, Japan), p. 3110–3113, Sept. 2010

  37. [45]

    Adam: A Method for Stochastic Optimization,

    D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” inProc. of ICLR, (San Diego, CA, USA), pp. 1– 15, May 2015

  38. [46]

    TIMIT Acoustic-Phonetic Continous Speech Corpus

    J.S.Garofolo,L.F.Lamel,W.M.Fisher,J.G.Fiscus,andD.S. Pallett, “TIMIT Acoustic-Phonetic Continous Speech Corpus.” Linguistic Data Consortium, Philadelphia, PA, USA, 1993

  39. [47]

    Do We Need Dereverberation for Hand-Held Telephony?,

    M. Jeub, M. Schäfer, H. Krüger, C. M. Nelke, C. Beaugeant, and P. Vary, “Do We Need Dereverberation for Hand-Held Telephony?,” in Proc. of ICA , (Sydney, Australia), pp. 3793– 3799, Aug. 2010

  40. [48]

    ETSI EG 202 396-1,Speech Processing, Transmission and Qual- ity Aspects (STQ); Speech Quality Performance in the Presence of Background Noise; Part 1: Background Noise Simulation Technique and Background Noise Database , Sept. 2008

  41. [49]

    Deep Learning for Acoustic Echo Cancellation in Noisy and Double-Talk Scenarios,

    H. Zhang and D. Wang, “Deep Learning for Acoustic Echo Cancellation in Noisy and Double-Talk Scenarios,” inProc. of Interspeech, (Hyderabad, India), pp. 3239–3243, Sept. 2018

  42. [50]

    Evaluation Metrics for Gener- ative Speech Enhancement Methods: Issues and Perspectives,

    J. Pirklbauer, M. Sach, K. Fluyt, W. Tirry, W. Wardah, S. Moeller, and T. Fingscheidt, “Evaluation Metrics for Gener- ative Speech Enhancement Methods: Issues and Perspectives,” in Proc. of 15th ITG Conference on Speech Communication , (Aachen, Germany), pp. 265–269, Sep 2023

  43. [51]

    Double-Talk Detection- Aided Residual Echo Suppression via Spectrogram Masking and Refinement,

    E. Shachar, I. Cohen, and B. Berdugo, “Double-Talk Detection- Aided Residual Echo Suppression via Spectrogram Masking and Refinement,” Acoustics, vol. 4, no. 3, pp. 637–655, 2022

  44. [52]

    Scaling Up Adap- tive Filter Optimizers,

    J. Casebeer, N. J. Bryan, and P. Smaragdis, “Scaling Up Adap- tive Filter Optimizers,”arXiv preprint:2403.00977, Mar. 2024. 24 Biographies Ernst Seidel (e.seidel@tu-bs.de) received the M.Sc. degree in electrical engineering in 2021 from Technische Universität Braunschweig, Ger...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.