REVIEW 3 major objections 2 minor 54 references
Components Loss for Neural Networks in Mask-Based Speech Enhancement
T0 review · 3 major / 2 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A training loss that treats speech preservation, noise suppression, and residual-noise naturalness as separate terms gives mask-based speech-enhancement networks better perceptual quality and stronger noise attenuation than conventional…
desk verdict Clean, honest loss-function paper with a genuinely new third term, but the headline PESQ/SNR gains rest on a single small test set and need statistical backing before they carry the full weight. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The object doing the work is the components loss (CL), a weighted sum of per-component mean-squared errors defined on the filtered speech and filtered noise spectra. For the 2CL variant, the first term $(1-\alpha)\sum_k(|\tilde{S}_\ell(k)|-|S_\ell(k)|)^2$ penalizes attenuation or distortion of the speech component, the second term $\alpha\sum_k|\tilde{D}_\ell(k)|^2$ penalizes residual noise power, and $\alpha\in[0,1]$ sets the trade-off. The 3CL variant adds a third term with weight $\beta$ that compares the normalized filtered noise spectrum $|\tilde{D}_\ell(k)|/\sqrt{\sum_\kappa|\tilde{D}_\ell(\kappa)|^2}$ with the normalized original noise spectrum, penalizing spectral reshaping of the residual noise while remaining zero for a pure fullband attenuation. The loss is naturally differentiable and is used inside the white-box training setup, where the mask is applied inside the network to the noisy magnitude spectrum and both components are available as targets.
What would settle it
Retrain the same mask-based CNN with MSE and with 3CL on a different corpus and a different network topology, then evaluate both on a held-out noise type; if 3CL's PESQ and SNR-improvement advantages over MSE fall below the reported 0.1-point and 0.5-dB thresholds, or reverse, the paper's central claim is not general.
Extended reading notes
Core claim
The central claim is that a mask-estimating CNN for single-channel speech enhancement should not be trained by comparing the enhanced spectrum to the clean spectrum alone, because that leaves the network free to mute low-SNR time–frequency bins, harming both speech detail and residual-noise naturalness. Instead, during training the known clean speech $S_\ell(k)$ and known noise $D_\ell(k)$ can each be multiplied by the estimated mask to form the filtered speech $\tilde{S}_\ell(k)$ and filtered noise $\tilde{D}_\ell(k)$, and the loss can be built from these components. The 2-component loss is $J_\ell^{\text{2CL}} = (1-\alpha)\sum_k (|\tilde{S}_\ell(k)|-|S_\ell(k)|)^2 + \alpha \sum_k |\tilde{D}_\ell(k)|^2$, with $\alpha$ trading speech preservation against noise attenuation. The 3-component loss adds a term comparing the normalized spectra of $\tilde{D}_\ell$ and $D_\ell$, so that a natural-sounding residual noise is preserved; the requirement that the weights satisfy $\alpha+\beta \le 1$ keeps the speech term from being dominated. On a fixed CNN evaluated with PESQ, POLQA, STOI, SSDR, and noise-quality measures, the paper reports that CL-trained networks give the best and most balanced performance, with speech-component quality and total enhanced-speech quality ahead of all three baseline losses.
Load-bearing premise
The load-bearing premise is that results measured on one CNN architecture and one speech/noise corpus are representative enough to support the paper's broad statement that the components loss is not restricted to any specific network topology or application.
Editorial extensions
If this is right
- Switching from MSE to 3CL yields at least 0.1 points higher PESQ on seen noise types and about 0.2 points higher on unseen bus noise, with more than 0.5 dB higher SNR improvement in both cases.
- Speech-component quality improves, with at least 0.5 dB higher SSDR and about 0.1 points higher PESQ on the filtered speech component for seen noises, meaning the enhanced speech retains more of the clean speech's detail.
- Residual noise becomes more natural under 3CL than under MSE or the perceptual weighting filter loss, matching or beating the PESQ-loss baseline in WLAKR, the metric closest to musical-tone annoyance.
- The benefits transfer without retraining the architecture or collecting new data: CL is a drop-in replacement for the loss function and is naturally differentiable.
- Because the weights $\alpha$ and $\beta$ give explicit control over the noise-suppression versus speech-distortion trade-off, a system designer can tune the same network for more aggressive denoising or more conservative speech preservation by changing two scalars.
Reading between the lines
- Because the paper tests only one unseen noise type and one CNN architecture, a natural extension would be to measure whether the 3CL advantage survives across several unseen noise classes and different mask-estimating architectures; the reported margins of 0.1–0.2 PESQ and 0.5 dB SNR give a concrete threshold for such a test.
- The third 3CL term shapes the residual-noise spectrum toward the original noise; an untested corollary is that it may also act as a regularizer that reduces musical-tone artifacts beyond what WLAKR captures, so a listening study or a dedicated tonality metric would be a sharper test.
- Since the loss needs access to clean speech and noise separately during training, it transfers most directly to fully supervised and simulation-based settings; adapting it to self-supervised or real-recording training would require an estimate of the noise component.
- The near-balanced choices $\alpha=0.5$ and $\alpha=1-\alpha-\beta$ in the hyperparameter search suggest that equal weighting between speech preservation and noise suppression may be a robust default for other architectures, a hypothesis the paper does not test.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a components loss (CL) for training mask-based single-channel speech enhancement networks. Two variants are introduced: 2CL, which linearly combines a filtered-speech-preservation term and a residual-noise-power term, and 3CL, which adds a third term that penalizes deviation of the normalized residual-noise spectrum shape from the original noise spectrum. The loss is evaluated with one CNN architecture on Grid speech mixed with CHiME-3 noise, comparing against MSE, a perceptual weighting filter loss (PW-FILT), and a PESQ-based loss (PW-PESQ). The authors report that the CL-trained networks, especially 3CL, achieve higher PESQ, POLQA, SSDR, and ΔSNR on both seen and unseen noise types, with code provided.
Significance. If the empirical claims are reliable, the components loss is a practically useful, differentiable training objective that offers separate control over speech preservation, noise suppression, and residual noise naturalness, and it is not tied to a particular network architecture. Strengths include the coherent and differentiable loss formulation, the use of external metrics (PESQ, POLQA, STOI) for the headline claims, hyperparameter selection on validation rather than test data, and public code. The main significance risk is that the quantitative conclusions rest on a single evaluation with a small number of test speakers and no statistical uncertainty assessment, plus a confounded comparison between 2CL and 3CL.
major comments (3)
- [Section IV.A; Tables IV.a, IV.b, V.a, V.b] The central quantitative claims are based on a single run over a test set of only four Grid speakers (two male, two female) with no confidence intervals, bootstrap estimates, or significance tests. PESQ, POLQA, and STOI are known to be speaker- and utterance-dependent, so the reported differences (for example, roughly 0.25 PESQ for 2CL versus MSE on PED noise in Table IV.a, or about 0.2 PESQ for 3CL on unseen BUS noise in Table V.a) may not be statistically reliable. I request repeated training runs with different seeds and/or bootstrapping across speakers and utterances, together with paired significance tests for the headline metrics, and a statement of variability for every number that supports the abstract and conclusion.
- [Section V.B.1; Tables IV and V] The comparison between 2CL and 3CL confounds the effect of the third loss term with a change in the weighting of the first two terms: 2CL is evaluated at α=0.5 (speech weight 0.5, noise weight 0.5), while 3CL is evaluated at α=0.1, β=0.8 (speech weight 0.1, noise weight 0.1). The observed differences in ΔSNR, WLAKR, and PESQ between 2CL and 3CL therefore cannot be attributed to the third term alone. I recommend an ablation, for example comparing 2CL at α=0.1 with 3CL at α=0.1, β=0.8, or comparing 3CL with β=0 against 3CL with the same α and β>0, to isolate the contribution of the residual-noise-shape term before drawing the mechanistic conclusion that the third term improves balanced performance.
- [Introduction and Section VI] The paper claims that the components loss is "not restricted to any specific network topology or application," but the experiments use exactly one CNN architecture, one corpus, and one mask-estimation framework (magnitude masking with noisy phase). If the authors wish to retain the generality claim, they should either provide evidence with at least one different architecture or task, or explicitly narrow the conclusion to the tested setting; otherwise the broad phrasing in the introduction and conclusion overstates the evidentiary basis.
minor comments (2)
- [Figure 6 caption] The caption says the markers correspond to six SNR conditions "from 20 dB to 5 dB with a step size of −5 dB," which is inconsistent with the captions of Figures 4 and 5; it should read "from 20 dB to −5 dB."
- [Section V.A] The hyperparameter selection procedure states that columns with any measure at or below the baseline MSE are discarded, and then the "best performing" remaining setting is selected, but the multi-metric criterion for this final selection is not formally defined; specifying the exact ordering or scoring rule would make the selection reproducible.
Circularity Check
No significant circularity: the components loss is a training objective validated on held-out external metrics.
full rationale
The proposed components loss is a training objective, not a fitted predictor. The paper optimizes J_2CL and J_3CL on training data (Eqs. 5 and 6) and evaluates on held-out test speakers and noise conditions using external metrics including PESQ, POLQA, STOI, SSDR, and Delta SNR. Hyperparameters alpha, beta, and lambda are selected on a 12.5% validation subset (Section V.A, Tables I-III), not on the test set, so the reported test-table gains do not reduce to fitted values. The loss terms are admittedly aligned with component metrics: Eq. (5) includes filtered-speech MSE and filtered-noise power, and Eq. (6) adds a normalized-noise-spectrum term; the paper openly attributes the Delta SNR and noise-quality improvements to these terms in Section V.B.1. This transparency is a mechanistic explanation of an empirical comparison, not circularity: the central headline claims on PESQ(hat s), POLQA(hat s), and STOI are external and not contained in the loss. Several citations are to the authors' own prior work ([27], [40], [45]), but they are used as architecture, baseline, and white-box references rather than as load-bearing evidence for the new loss. No step reduces by construction to its inputs, and no uniqueness theorem or ansatz is imported from a self-citation. The statistical robustness concern about the small test set is a correctness issue, not a circularity issue.
Assumptions & free parameters
free parameters (5)
- alpha (2CL) =
0.5
- alpha (3CL) =
0.1
- beta (3CL) =
0.8
- lambda1 (PW-PESQ baseline) =
0.2
- lambda2 (PW-PESQ baseline) =
0.8
assumptions (4)
- domain assumption Additive single-channel noise model y(n) = s(n) + d(n).
- domain assumption Clean speech and noise component spectra are available during training.
- domain assumption Minimizing separate errors for speech component, residual noise power, and residual noise spectral shape yields perceptually better enhancement.
- ad hoc to paper Preserving the normalized residual noise spectrum shape improves naturalness of residual noise.
Cite this review
Pith. "Pith review of Components Loss for Neural Networks in Mask-Based Speech Enhancement." pith.science (2026). https://pith.science/paper/ZWXSSPOL
@misc{pith2026190805087,
author = {Pith},
title = {Pith review of: Components Loss for Neural Networks in Mask-Based Speech Enhancement},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZWXSSPOL}},
note = {Machine review of arXiv:1908.05087}
}
read the original abstract
Estimating time-frequency domain masks for single-channel speech enhancement using deep learning methods has recently become a popular research field with promising results. In this paper, we propose a novel components loss (CL) for the training of neural networks for mask-based speech enhancement. During the training process, the proposed CL offers separate control over preservation of the speech component quality, suppression of the residual noise component, and preservation of a naturally sounding residual noise component. We illustrate the potential of the proposed CL by evaluating a standard convolutional neural network (CNN) for mask-based speech enhancement. The new CL obtains a better and more balanced performance in almost all employed instrumental quality metrics over the baseline losses, the latter comprising the conventional mean squared error (MSE) loss and also auditory-related loss functions, such as the perceptual evaluation of speech quality (PESQ) loss and the recently proposed perceptual weighting filter loss. Particularly, applying the CL offers better speech component quality, better overall enhanced speech perceptual quality, as well as a more naturally sounding residual noise. On average, an at least 0.1 points higher PESQ score on the enhanced speech is obtained while also obtaining a higher SNR improvement by more than 0.5 dB, for seen noise types. This improvement is stronger for unseen noise types, where an about 0.2 points higher PESQ score on the enhanced speech is obtained, while also the output SNR is ahead by more than 0.5 dB. The new proposed CL is easy to implement and code is provided at https://github.com/ifnspaml/Components-Loss.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Speech Enhancement Using a Mini mum Mean-Square Error Short-Time Spectral Amplitude Estimato r,
Y . Ephraim and D. Malah, “Speech Enhancement Using a Mini mum Mean-Square Error Short-Time Spectral Amplitude Estimato r,” IEEE T-ASSP, vol. 32, no. 6, pp. 1109–1121, Dec. 1984
work page 1984
-
[2]
Speech Enhancement Using a Minimum Mean-Square Err or Log- Spectral Amplitude Estimator,
——, “Speech Enhancement Using a Minimum Mean-Square Err or Log- Spectral Amplitude Estimator,” IEEE T-ASSP , vol. 33, no. 2, pp. 443– 445, Apr. 1985
work page 1985
-
[3]
Speech Enhancement Based on A Priori Signal to Noise Estimation,
P . Scalart and J. V . Filho, “Speech Enhancement Based on A Priori Signal to Noise Estimation,” in Proc. of ICASSP , Atlanta, GA, USA, May 1996, pp. 629–632
work page 1996
-
[4]
Speech Enhancement by MAP Spectra l Am- plitude Estimation Using a Super-Gaussian Speech Model,
T. Lotter and P . Vary, “Speech Enhancement by MAP Spectra l Am- plitude Estimation Using a Super-Gaussian Speech Model,” EURASIP Journal on Applied Signal Processing , vol. 2005, no. 7, pp. 1110–1126, May 2005
work page 2005
-
[5]
B. Fodor and T. Fingscheidt, “Speech Enhancement Using a Joint Map Estimator with Gaussian Mixture Model for (Non)-Stationar y Noise,” in Proc. of ICASSP , Prague, Czech Republic, May 2011, pp. 4768–4771
work page 2011
-
[6]
Speech Enhancement Using Super-Gaussian Spe ech Models and Noncausal A Priori SNR Estimation,
I. Cohen, “Speech Enhancement Using Super-Gaussian Spe ech Models and Noncausal A Priori SNR Estimation,” Speech Commun. , vol. 47, no. 3, pp. 336–350, Nov. 2005
work page 2005
-
[7]
T. Gerkmann, C. Breithaupt, and R. Martin, “Improved A Po steriori Speech Presence Probability Estimation Based on a Likeliho od Ratio with Fixed Priors,” IEEE T-ASLP , vol. 16, no. 5, pp. 910–919, Jul. 2008
work page 2008
-
[8]
A Data-Driven Ap proach to A Priori SNR Estimation,
S. Suhadi, C. Last, and T. Fingscheidt, “A Data-Driven Ap proach to A Priori SNR Estimation,” IEEE T-ASLP, vol. 19, no. 1, pp. 186–195, Jan. 2011
work page 2011
Show all 54 references
-
[9]
An Iterative Speech Model-Based A Priori SNR Estimator,
S. Elshamy, N. Madhu, W. J. Tirry, and T. Fingscheidt, “An Iterative Speech Model-Based A Priori SNR Estimator,” in Proc. of Interspeech , Dresden, Germany, Sep. 2015, pp. 1740–1744
2015
-
[10]
Ins tantaneous A Priori SNR Estimation by Cepstral Excitation Manipulation ,
S. Elshamy, N. Madhu, W. Tirry, and T. Fingscheidt, “Ins tantaneous A Priori SNR Estimation by Cepstral Excitation Manipulation ,” IEEE/ACM T-ASLP, vol. 25, no. 8, pp. 1592–1605, Aug. 2017
2017
-
[11]
Tracking Speech- presence Uncertainty to Improve Speech Enhancement in Non-stationa ry Noise Environments,
D. Malah, R. V . Cox, and A. J. Accardi, “Tracking Speech- presence Uncertainty to Improve Speech Enhancement in Non-stationa ry Noise Environments,” in Proc. of ICASSP , Phoenix, AZ, USA, Mar. 1999, pp. 789–792
1999
-
[12]
Data-Driven Speech Enha ncement,
T. Fingscheidt and S. Suhadi, “Data-Driven Speech Enha ncement,” in Proc. of ITG Conf. on Speech Communication , Kiel, Germany, Apr. 2006, pp. 1–4
2006
-
[13]
Environment-O ptimized Speech Enhancement,
T. Fingscheidt, S. Suhadi, and S. Stan, “Environment-O ptimized Speech Enhancement,” IEEE T-ASLP, vol. 16, no. 4, pp. 825–834, May 2008
2008
-
[14]
A general Opti mization Procedure for Spectral Speech Enhancement Methods,
J. Erkelens, J. Jensen, and R. Heusdens, “A general Opti mization Procedure for Spectral Speech Enhancement Methods,” in Proc. of EUSIPCO, Florence, Italy, Sep. 2006, pp. 1–5
2006
-
[15]
A Data-Driven Approach to Optimizing Spectral Spe ech En- hancement Methods for V arious Error Criteria,
——, “A Data-Driven Approach to Optimizing Spectral Spe ech En- hancement Methods for V arious Error Criteria,” Speech Communication, vol. 49, no. 7-8, pp. 530–541, Jul. 2007
2007
-
[16]
On Training Targe ts for Supervised Speech Separation,
Y . Wang, A. Narayanan, and D. L. Wang, “On Training Targe ts for Supervised Speech Separation,” IEEE/ACM T-ASLP , vol. 22, no. 12, pp. 1849–1858, Dec. 2014
2014
-
[17]
Discrimina- tively Trained Recurrent Neural Networks for Single-Chann el Speech Separation,
F. Weninger, J. R. Hershey, J. Le Roux, and B. Schuller, “ Discrimina- tively Trained Recurrent Neural Networks for Single-Chann el Speech Separation,” in Proc. of 2nd IEEE GlobalSIP , Atlanta, GA, USA, May 2014, pp. 577–581
2014
-
[18]
Dee p Learning for Monaural Speech Separation,
P . S. Huang, M. Kim, M. H. Johnson, and P . Smaragdis, “Dee p Learning for Monaural Speech Separation,” in Proc. of ICASSP , Florence, Italy, May 2014, pp. 1562–1566
2014
-
[19]
P hase- Sensitive and Recognition-Boosted Speech Separation Usin g Deep Recurrent Neural Networks,
H. Erdogan, J. R. Hershey, S. Watanabe, and J. Le Roux, “P hase- Sensitive and Recognition-Boosted Speech Separation Usin g Deep Recurrent Neural Networks,” in Proc. of ICASSP , Brisbane, QLD, Australia, Aug. 2015, pp. 708–712
2015
-
[20]
A Deep Neural Network for Time-Do main Signal Reconstruction,
Y . Wang and D. L. Wang, “A Deep Neural Network for Time-Do main Signal Reconstruction,” in Proc. of ICASSP , Brisbane, QLD, Australia, Aug. 2015, pp. 4390–4394
2015
-
[21]
Complex Ratio Masking for Monaural Speech Separation,
D. S. Williamson, Y . Wang, and D. L. Wang, “Complex Ratio Masking for Monaural Speech Separation,” IEEE/ACM T-ASLP , vol. 24, no. 3, pp. 483–492, Mar. 2016
2016
-
[22]
A New Ratio Mask Representatio n for CASA-Based Speech Enhancement,
F. Bao and W. H. Abdulla, “A New Ratio Mask Representatio n for CASA-Based Speech Enhancement,” IEEE/ACM TASLP, vol. 27, no. 1, pp. 7–19, Jan. 2019
2019
-
[23]
Supervised Speech Separation Based on Deep Learning: An Overview,
D. L. Wang and J. T. Chen, “Supervised Speech Separation Based on Deep Learning: An Overview,” IEEE/ACM T-ASLP, vol. 26, no. 10, pp. 1702–1726, Oct. 2018
2018
-
[24]
A Regression Approa ch to Single-Channel Speech Separation via High-Resolution Dee p Neural Networks,
J. Du, Y . Tu, L. R. Dai, and C. H. Lee, “A Regression Approa ch to Single-Channel Speech Separation via High-Resolution Dee p Neural Networks,” IEEE/ACM T-ASLP , vol. 24, no. 8, pp. 1424–1437, Apr. 2016
2016
-
[25]
Perception Optimi zed Deep De- noising Autoencoders for Speech Enhancement
P . G. Shivakumar and P . G. Georgiou, “Perception Optimi zed Deep De- noising Autoencoders for Speech Enhancement.” in Proc. of Interspeech, San Francisco, CA, USA, Sep. 2016, pp. 3743–3747
2016
-
[26]
A Percep tually- Weighted Deep Neural Network for Monaural Speech Enhanceme nt in Various Background Noise Conditions,
Q. J. Liu, W. Wang, P . J. B. Jackson, and Y . Tang, “A Percep tually- Weighted Deep Neural Network for Monaural Speech Enhanceme nt in Various Background Noise Conditions,” in Proc. of EUSIPCO , Kos, Greece, Aug. 2017, pp. 1270–1274
2017
-
[27]
A Perceptual W eighting Filter Loss for DNN Training in Speech Enhancement,
Z. Zhao, S. Elshamy, and T. Fingscheidt, “A Perceptual W eighting Filter Loss for DNN Training in Speech Enhancement,” arXiv preprint arXiv:1905.09754, May 2019
1905 arXiv
-
[28]
Lear ning to Dequantize Speech Signals by Primal-Dual Networks: An Appr oach for Acoustic Sensor Networks,
C. Brauer, Z. Zhao, D. Lorenz, and T. Fingscheidt, “Lear ning to Dequantize Speech Signals by Primal-Dual Networks: An Appr oach for Acoustic Sensor Networks,” in Proc. of ICASSP , Brighton, UK, May 2019, pp. 7000–7004
2019
-
[29]
A Deep Learning Loss Function Based on the Perceptual Evalu ation of the Speech Quality,
J. M. Mart´ ın Do˜ nas, A. M. Gomez, J. A. Gonzalez, and A. M . Peinado, “A Deep Learning Loss Function Based on the Perceptual Evalu ation of the Speech Quality,” IEEE SPL, vol. 25, no. 11, pp. 1680–1684, Nov. 2018
2018
-
[30]
DNN- Based Source Enhancement to Increase Objective Sound Quali ty As- sessment Score,
Y . Koizumi, K. Niwa, Y . Hioka, K. Kobayashi, and Y . Haned a, “DNN- Based Source Enhancement to Increase Objective Sound Quali ty As- sessment Score,” IEEE/ACM T-ASLP , vol. 26, no. 10, pp. 1780–1792, Oct. 2018
2018
-
[31]
Monaural Speech En hancement Using Deep Neural Networks by Maximizing a Short-Time Objec tive Intelligibility Measure,
M. Kolbcek, Z. H. Tan, and J. Jensen, “Monaural Speech En hancement Using Deep Neural Networks by Maximizing a Short-Time Objec tive Intelligibility Measure,” in Proc. of ICASSP , Calgary, AB, Canada, Apr. 2018, pp. 5059–5063
2018
-
[32]
Deep Neural Network Based Speech Separation Optimizing an Objective Es timator of Intelligibility for Low Latency Applications,
G. Naithani, J. Nikunen, L. Bramslow, and T. Virtanen, “ Deep Neural Network Based Speech Separation Optimizing an Objective Es timator of Intelligibility for Low Latency Applications,” in Proc. of IWAENC , Tokyo, Japan, Sep. 2018, pp. 386–390. 12
2018
-
[33]
Training Supervise d Speech Separation System to Improve STOI and PESQ Directly,
H. Zhang, X. L. Zhang, and G. L. Gao, “Training Supervise d Speech Separation System to Improve STOI and PESQ Directly,” in Proc. of ICASSP, Calgary, AB, Canada, Apr. 2018, pp. 5374–5378
2018
-
[34]
End-to- End Waveform Utterance Enhancement for Direct Evaluation M etrics Optimization by Fully Convolutional Neural Networks,
S. W. Fu, T. W. Wang, Y . Tsao, X. Lu, and H. Kawai, “End-to- End Waveform Utterance Enhancement for Direct Evaluation M etrics Optimization by Fully Convolutional Neural Networks,” IEEE/ACM T- ASLP, vol. 26, no. 9, pp. 1570–1584, Sep. 2018
2018
-
[35]
Error Concealment by Softb it Speech De- coding,
T. Fingscheidt and P . Vary, “Error Concealment by Softb it Speech De- coding,” in Proc. of ITG-Fachtagung ”Sprachkommunikation” , Frankfurt a.M., Germany, Sep. 1996, pp. 7–10
1996
-
[36]
Softbit Speech Decoding: A New Approach to Error Co nceal- ment,
——, “Softbit Speech Decoding: A New Approach to Error Co nceal- ment,” IEEE T-SAP, vol. 9, no. 3, pp. 240–251, Mar. 2001
2001
-
[37]
A Short-Time Objective Intelligibility Measure for Time-Frequency Wei ghted Noisy Speech,
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “A Short-Time Objective Intelligibility Measure for Time-Frequency Wei ghted Noisy Speech,” in Proc. of ICASSP , Dallas, TX, USA, Jun. 2010, pp. 4214– 4217
2010
-
[38]
ITU, Rec. P.862: Perceptual Evaluation of Speech Quality (PESQ) : An Objective Method for End-To-End Speech Quality Assessme nt of Narrow-Band Telephone Networks and Speech Codecs , International Telecommunication Standardization Sector (ITU-T), Feb. 2 001
-
[39]
On the Optimizat ion of Speech Enhancement Systems Using Instrumental Measures,
S. Gustafsson, R. Martin, and P . Vary, “On the Optimizat ion of Speech Enhancement Systems Using Instrumental Measures,” in Proc. of W orkshop on Qual. Assess. in Speech, Audio, and Image Comm un., Darmstadt, Germany, Mar. 1996, pp. 36–40
1996
-
[40]
Quality Assessment of Sp eech Enhance- ment Systems by Separation of Enhanced Speech, Noise, and Ec ho,
T. Fingscheidt and S. Suhadi, “Quality Assessment of Sp eech Enhance- ment Systems by Separation of Enhanced Speech, Noise, and Ec ho,” in Proc. of Interspeech , Antwerp, Belgium, Aug. 2007, pp. 818–821
2007
-
[41]
A Figure of Merit for Instrume ntal Optimiza- tion of Noise Reduction Algorithms,
H. Yu and T. Fingscheidt, “A Figure of Merit for Instrume ntal Optimiza- tion of Noise Reduction Algorithms,” in Proc. of 5th Biennial W orkshop on DSP for In-V ehicle Systems , Kiel, Germany, Sep. 2011, pp. 1–8
2011
-
[42]
P.1100: Narrowband Hands-Free Communication in Motor Vehicles, International Telecommunication Standardization Secto r (ITU- T), Jan
ITU, Rec. P.1100: Narrowband Hands-Free Communication in Motor Vehicles, International Telecommunication Standardization Secto r (ITU- T), Jan. 2019
2019
-
[43]
P.1110: Wideband Hands-Free Communication in Motor Vehicles, International Telecommunication Standardization Secto r (ITU- T), Jan
——, Rec. P.1110: Wideband Hands-Free Communication in Motor Vehicles, International Telecommunication Standardization Secto r (ITU- T), Jan. 2015
2015
-
[44]
P.1130: Subsystem Requirements for Automotive Speech Services, International Telecommunication Standardization Secto r (ITU- T), Jun
——, Rec. P.1130: Subsystem Requirements for Automotive Speech Services, International Telecommunication Standardization Secto r (ITU- T), Jun. 2015
2015
-
[45]
Convolutional N eural Networks to Enhance Coded Speech,
Z. Zhao, H. J. Liu, and T. Fingscheidt, “Convolutional N eural Networks to Enhance Coded Speech,” IEEE/ACM T-ASLP, vol. 27, no. 4, pp. 663– 678, Apr. 2019
2019
-
[46]
MetricGAN: Ge nerative Adversarial Networks Based Black-Box Metric Scores Optimi zation for Speech Enhancement,
S. Z. Fu, C. F. Liao, Y . Tsao, and S. D. Lin, “MetricGAN: Ge nerative Adversarial Networks Based Black-Box Metric Scores Optimi zation for Speech Enhancement,” arXiv preprint arXiv:1905.04874 , May 2019
1905 arXiv
-
[47]
Residual Networ ks Behave Like Ensembles of Relatively Shallow Networks,
A. Veit, M. J. Wilber, and S. Belongie, “Residual Networ ks Behave Like Ensembles of Relatively Shallow Networks,” in Proc. of NIPS , Barcelona, Spain, Dec. 2016, pp. 550–558
2016
-
[48]
An Audi o-Visual Corpus for Speech Perception and Automatic Speech Recognit ion,
M. Cooke, J. Barker, S. Cunningham, and X. Shao, “An Audi o-Visual Corpus for Speech Perception and Automatic Speech Recognit ion,” The Journal of the Acoustical Society of America , vol. 120, no. 5, pp. 2421– 2424, Jun. 2006
2006
-
[49]
The T hird ‘CHiME’ Speech Separation and Recognition Challenge: Dataset, Tas k and Base- lines,
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, “The T hird ‘CHiME’ Speech Separation and Recognition Challenge: Dataset, Tas k and Base- lines,” in Proc. of ASRU , Scottsdale, AZ, USA, Feb. 2015, pp. 504–511
2015
-
[50]
P.56: Objective Measurement of Active Speech Level , Interna- tional Telecommunication Standardization Sector (ITU-T) , Dec
ITU, Rec. P.56: Objective Measurement of Active Speech Level , Interna- tional Telecommunication Standardization Sector (ITU-T) , Dec. 2011
2011
-
[51]
——, Rec. P .862.2: Corrigendum 1, Wideband Extension to Recomme n- dation P .862 for the Assessment of Wideband Telephone Netwo rks and Speech Codecs, International Telecommunication Standardization Secto r (ITU-T), Oct. 2017
2017
-
[52]
P .863: Perceptual Objective Listening Quality Predic tion (POLQA), International Telecommunication Union, Telecommunicat ion Standardization Sector (ITU-T), Mar
——, Rec. P .863: Perceptual Objective Listening Quality Predic tion (POLQA), International Telecommunication Union, Telecommunicat ion Standardization Sector (ITU-T), Mar. 2018
2018
-
[53]
Black Box Measurement of Musi cal Tones Produced by Noise Reduction Systems,
H. Yu and T. Fingscheidt, “Black Box Measurement of Musi cal Tones Produced by Noise Reduction Systems,” in Proc. of ICASSP , Kyoto, Japan, Aug. 2012, pp. 4573–4576
2012
-
[54]
14) , 3GPP; TSG SA, Mar
3GPP, Mandatory Speech Codec Speech Processing Functions; Adapt ive Multi-Rate (AMR) Speech Codec; Transcoding Functions (3GP P TS 26.090, Rel. 14) , 3GPP; TSG SA, Mar. 2017
2017
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.