REVIEW 1 major objections 4 minor 81 references
30+ Years of Source Separation Research: Achievements and Future Challenges
T0 review · 1 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A 30-year map of source separation: what worked, what is left.
desk verdict Competent, usefully selective review of source separation; the 'major contributions' claim needs a scope caveat, but it's a fair survey for newcomers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing object is the convolutive mixture model $y_m(\tilde{t}) = \sum_n \sum_\tau h_{mn}(\tau) s_n(\tilde{t}-\tau) + v_m(\tilde{t})$, approximated as an instantaneous mixture in the time-frequency domain, $y_{mtf} = \sum_n h_{mnf} s_{ntf} + v_{mtf}$. This formulation generates the field's central technical obstacle, the frequency permutation problem, and the review's narrative machinery is the sequence of mechanisms devised to overcome it: independent vector analysis bundles frequency components statistically; NMF models full-band spectrograms; deep clustering embeds time-frequency units so same-source units are close; permutation-invariant training aligns outputs to targets before the loss; and hybrid systems feed DNN-estimated masks into beamformers. The evaluation culture, standard metrics (SDR, SIR, SAR), SI-SDR, and the challenge infrastructure from SiSEC through the Music Demixing Challenge, is the other machinery, since it is what converts competing algorithms into comparable numbers.
What would settle it
An independent bibliometric or community-wide survey of source-separation publications from 1994 to 2024, ranking papers by citation counts or by adoption in deployed systems, would show whether the milestones highlighted here are in fact the field's pivotal contributions; if the most influential works differ substantially from those featured, the review's claim to cover the major contributions and advancements fails.
Extended reading notes
Core claim
On the authors' account, source separation has moved through three overlapping phases. From the mid-1990s to the 2000s, blind separation of determined mixtures was solved in principle by independent component analysis, extended to convolutive and underdetermined cases by time-frequency masking, independent vector analysis, nonnegative matrix factorization, and full-rank spatial covariance models. The 2010s brought a supervised deep-learning breakthrough: neural networks learned to estimate time-frequency masks, with deep clustering and permutation-invariant training resolving the label-permutation problem for homogeneous sources, and architectures moved from masking to complex-spectrum and waveform estimation. The present phase is hybrid, in which DNNs estimate masks or spatial statistics that drive beamformers, alongside unsupervised and mixture-invariant training to use real recordings without ground truth. The paper's central claim is that this trajectory, supported by shared datasets and metrics (SDR/SIR/SAR, SI-SDR, challenges from SiSEC to the Music Demixing Challenge), has produced a field whose simulated benchmarks have saturated while its remaining difficulties, mismatched real data, unknown and time-varying source counts, limited microphones, and reference-free perceptual evaluation, define the open research agenda.
Load-bearing premise
The review assumes that the works, benchmarks, and challenges it selects give a representative and unbiased history of the field rather than a picture shaped by the authors' own research priorities.
Editorial extensions
If this is right
- If the review's history is right, new methods should be measured against real-recorded, conversational benchmarks rather than fully-overlapped simulated mixtures, where performance has saturated.
- Multichannel separation in realistic conditions will increasingly rely on hybrid designs, DNN-driven mask estimation feeding classical beamformers, since each component handles different signal properties.
- Unsupervised and mixture-invariant training will become standard for leveraging in-the-wild data, because supervised training on simulated mixtures transfers poorly to real rooms.
- Research attention will shift from fixed-source-count separation to joint source activity detection and separation, since real scenes contain unknown, time-varying numbers of sources.
- Evaluation practice will need new reference-free metrics that track human perception for speech, music, and general sounds, replacing alignment-based SDR-style measures.
Reading between the lines
- The authors leave implicit that if simulated benchmarks are saturated and real data is the bottleneck, then progress will depend as much on data collection and simulation-realism engineering as on new network architectures.
- The review's own emphasis on a few state-of-the-art systems (for example, the dual-path/spectrogram-grid architectures and the dereverberation front-end it highlights) suggests that the field's advanced capability is concentrated in a small number of research groups; independent reproductions would test whether the reported gains generalize.
- The brief mention of natural-language-prompted separation hints at a convergence with foundation-model approaches; if that trend holds, source separation may be reframed as a text-conditioned generation problem, which the current benchmark infrastructure does not yet evaluate.
- The stress on low-latency and lightweight models suggests that deployment, not raw separation quality, may be the next differentiator; one could test this by comparing real-time capable systems against offline state-of-the-art on the same real-world data.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a review of source separation (SS) research over the past three decades, written for the occasion of ICASSP's 50th anniversary. It covers the problem formulation, model-based approaches (ICA, IVA, ILRMA, FCA, TF masking), deep-learning approaches (deep clustering, PIT, complex spectral mapping, time-domain and transformer-based systems), hybrid beamforming methods, and the evaluation infrastructure of the field (SiSEC, CHiME, MDX/SDX challenges, metrics, datasets, and open-source tools). It closes with a short discussion of remaining challenges and possible future directions. The manuscript does not present new algorithms or experiments; its contribution is historical and technical synthesis.
Significance. If the coverage issues are addressed, this would be a useful and readable retrospective for the source separation community and for newcomers. The technical descriptions of the core methods (e.g., the convolutive mixture model in Eq. (1), the TF-domain approximation in Eq. (3), ICA, ILRMA, FCA, deep clustering, PIT, and beamforming hybrids) are accurate and appropriately cited. The paper also does a service by documenting the evaluation culture—challenges, metrics, and datasets—that has been crucial to the field. Its significance as a 'major contributions' review, however, is currently weakened by the absence of explicit selection criteria, which makes the representativeness of the chosen topics and references difficult to verify.
major comments (1)
- [I, III-B, IV] The abstract and introduction claim that the paper reviews 'the major contributions and advancements' in source separation over the past three decades, but no inclusion criteria or selection methodology is ever stated. This makes the central claim of the paper difficult to verify. For example, Sec. III-B highlights dual-path architectures and singles out TF-GridNet [48] as the representative modern system without justifying why this particular work is chosen over other recent architectures, and Sec. IV-C lists a few open-source tools without defining the basis for their selection. As written, the reader cannot distinguish a balanced overview from a retrospective centered on the authors' own research priorities. I recommend adding an explicit methodology or scope statement (e.g., criteria for what counts as a 'major contribution', how subtopics were allocated, and any search or selection process), or alternatively reframing the paper as a 'selected overview' and adjusting the title and abstract accordingly.
minor comments (4)
- [IV-C] In Sec. IV-C, the text attributes 'Nara-WPE [60]' as an open implementation of a dereverberation algorithm, but reference [60] is the paper 'Unsupervised training of a deep clustering model for multichannel blind source separation' by Drude et al. This citation appears to be incorrect and should be replaced with the appropriate Nara-WPE or WPE reference.
- [IV-A] The sentence 'This SiSEC initiative was taken over by CHiME 2' is historically imprecise, since SiSEC continued until 2018 and CHiME was launched separately in 2010. Please revise to describe the relationship between these initiatives more accurately.
- [I] The statement that ICASSP has 'consistently received 40–50 submissions to AUD-SEP every year' lacks a citation. Please add a source for this statistic or soften the claim.
- [IV] The opening sentence of Sec. IV says 'in the 2000s, there were no benchmarks in SS research', but the next sentences describe SASSEC in 2007. Please rephrase to 'in the early 2000s' or 'before SASSEC/SiSEC' to avoid the apparent contradiction.
Circularity Check
No circularity: this is a literature survey with no derived predictions or fitted parameters, and its citations are bibliographic rather than load-bearing.
full rationale
The paper is a review of source separation research; it does not claim to derive a new result, fit a parameter, or predict an output from an input. Its stated goal is descriptive: 'we review the major contributions and advancements in the past three decades in the speech, audio, and music SS research field.' A descriptive survey can be incomplete or unrepresentative, but incompleteness is not circularity. None of the enumerated circularity patterns applies: there is no quantity defined in terms of another and then compared with it; no fitted input is relabeled as a prediction; no uniqueness theorem is imported to force a choice; no ansatz is smuggled in by citation; and no known result is merely renamed. The reference list includes several papers by the authors (e.g., [10], [18], [36], [48], [59], [80]), but these are cited as examples of prior contributions within a narrative, not as premises that make the survey's claims true by construction. The absence of explicit inclusion criteria for 'major contributions' is a legitimate methodological criticism of the review's coverage, but it is not a circularity defect. Therefore the honest finding is no significant circularity, score 0.
Assumptions & free parameters
assumptions (2)
- domain assumption The convolutive mixture model (Eq. 1) with additive noise represents the acoustic mixing process.
- domain assumption The time-frequency approximation (Eq. 3) of the convolutive mixture is valid under common conditions.
Cite this review
Pith. "Pith review of 30+ Years of Source Separation Research: Achievements and Future Challenges." pith.science (2026). https://pith.science/paper/G4Y4DFHM
@misc{pith2026250111837,
author = {Pith},
title = {Pith review of: 30+ Years of Source Separation Research: Achievements and Future Challenges},
year = {2026},
howpublished = {\url{https://pith.science/paper/G4Y4DFHM}},
note = {Machine review of arXiv:2501.11837}
}
read the original abstract
Source separation (SS) of acoustic signals is a research field that emerged in the mid-1990s and has flourished ever since. On the occasion of ICASSP's 50th anniversary, we review the major contributions and advancements in the past three decades in the speech, audio, and music SS research field. We will cover both single- and multi-channel SS approaches. We will also look back on key efforts to foster a culture of scientific evaluation in the research field, including challenges, performance metrics, and datasets. We will conclude by discussing current trends and future research directions.
Reference graph
Works this paper leans on
-
[48]
TF-GridNet: Integrating full- and sub-band modeling for speech separation,
Z.-Q. Wang et al., “TF-GridNet: Integrating full- and sub-band modeling for speech separation,” IEEE/ACM TASLP, vol. 31, pp. 3221–3236, 2023
work page 2023
-
[1]
Makino et al
S. Makino et al. , Eds., Blind Speech Separation . Springer, 2007
2007
-
[2]
E. Vincent et al. , Eds., Audio Source Separation and Speech Enhance- ment. Wiley, 2018
work page 2018
-
[3]
Makino, Ed., Audio Source Separation
S. Makino, Ed., Audio Source Separation . Springer, 2018
work page 2018
-
[4]
An information-maximiza tion approach to blind separation and blind deconvolution,
A. J. Bell and T. J. Sejnowski, “An information-maximiza tion approach to blind separation and blind deconvolution,” Neural Computation , vol. 7, pp. 1129–1159, 1995
work page 1995
-
[5]
Independent vector analysis: An extension of ICA to multivariate components,
T. Kim et al. , “Independent vector analysis: An extension of ICA to multivariate components,” in Proc. ICA , 2006, pp. 165–172
work page 2006
-
[6]
A. Hiroe, “Solution of permutation problem in frequency domain ICA, using multivariate probability density functions,” in Proc. ICA, 2006, pp. 601–608
work page 2006
-
[7]
Stable and fast update rules for independent vec tor analysis based on auxiliary function technique,
N. Ono, “Stable and fast update rules for independent vec tor analysis based on auxiliary function technique,” in Proc. WASPAA , 2011, pp. 189–192
work page 2011
Show all 81 references
-
[8]
De- termined blind source separation unifying independent vec tor analysis and nonnegative matrix factorization,
D. Kitamura, N. Ono, H. Sawada, H. Kameoka, and H. Saruwat ari, “De- termined blind source separation unifying independent vec tor analysis and nonnegative matrix factorization,” IEEE/ACM TASLP, vol. 24, no. 9, pp. 1626–1641, 2016
2016
-
[9]
Blind separation of speech mix tures via time- frequency masking,
O. Yilmaz and S. Rickard, “Blind separation of speech mix tures via time- frequency masking,” IEEE Transactions on Signal Processing , vol. 52, no. 7, pp. 1830–1847, 2004
2004
-
[10]
Complex angular central Gaussian mixture model for directional statistics in mask-based microphone array sig nal processing,
N. Ito et al. , “Complex angular central Gaussian mixture model for directional statistics in mask-based microphone array sig nal processing,” in Proc. EUSIPCO , 2016, pp. 1153–1157
2016
-
[11]
A multichannel MMSE-based framework for speech source separation and noise reduction,
M. Souden et al. , “A multichannel MMSE-based framework for speech source separation and noise reduction,” IEEE TASLP, vol. 21, no. 9, pp. 1913–1928, 2013
1913
-
[12]
The NTT CHiME-3 system: Advances in speech enhancement and recognition for mobile multi-microphone d evices,
T. Y oshioka et al. , “The NTT CHiME-3 system: Advances in speech enhancement and recognition for mobile multi-microphone d evices,” in Proc. ASRU, 2015, pp. 436–443
2015
-
[13]
Front-end processing for the CHiME-5 dinner party scenario,
C. Boeddecker et al. , “Front-end processing for the CHiME-5 dinner party scenario,” in Proc. CHiME, 2018, pp. 35–40
2018
-
[14]
Under-determined reverberant audio source separation using a full-rank spatial covariance model,
N. Q. K. Duong et al. , “Under-determined reverberant audio source separation using a full-rank spatial covariance model,” IEEE TASLP , vol. 18, no. 7, pp. 1830–1840, 2010
2010
-
[15]
Nonnegative matrix factorization and spatial covari- ance model for under-determined reverberant audio source s eparation,
S. Arberet et al. , “Nonnegative matrix factorization and spatial covari- ance model for under-determined reverberant audio source s eparation,” in Proc. ISSPA, 2010
2010
-
[16]
Multichannel extensions of non-negative matrix factorization with complex-valued data,
H. Sawada et al. , “Multichannel extensions of non-negative matrix factorization with complex-valued data,” IEEE TASLP , vol. 21, no. 5, pp. 971–982, 2013
2013
-
[17]
Neural full-rank spatial covariance analysis for blind source separation,
Y . Bando et al. , “Neural full-rank spatial covariance analysis for blind source separation,” IEEE Signal Processing Letters , vol. 28, pp. 1670– 1674, 2021
2021
-
[18]
A joint diagonalization based efficient approach to underdetermined blind audio source separation using the mu ltichannel Wiener filter,
N. Ito et al. , “A joint diagonalization based efficient approach to underdetermined blind audio source separation using the mu ltichannel Wiener filter,” IEEE/ACM TASLP, vol. 29, pp. 1950–1965, 2021
1950
-
[19]
Fast multichannel nonnegative matrix factorization with directivity-aware jointly-diagonalizable spatial c ovariance matrices for blind source separation,
K. Sekiguchi et al., “Fast multichannel nonnegative matrix factorization with directivity-aware jointly-diagonalizable spatial c ovariance matrices for blind source separation,” IEEE/ACM TASLP, vol. 28, pp. 2610–2625, 2020
2020
-
[20]
The DUET blind source separation algorith m,
S. Rickard, “The DUET blind source separation algorith m,” in Blind speech separation. Springer, 2007, pp. 217–241
2007
-
[21]
Sound source separat ion: Azimuth discrimination and resynthesis,
D. Barry, B. Lawlor, and E. Coyle, “Sound source separat ion: Azimuth discrimination and resynthesis,” in Proc. DAFX, 2004, pp. 240–244
2004
-
[22]
Projection-based demixing of spatial audio,
D. FitzGerald et al. , “Projection-based demixing of spatial audio,” IEEE/ACM TASLP, vol. 24, no. 9, pp. 1560–1572, 2016
2016
-
[23]
Super-human multi-talker speech recognition: A graphical modeling approach,
J. R. Hershey et al. , “Super-human multi-talker speech recognition: A graphical modeling approach,” Computer Speech and Language , vol. 24, no. 1, pp. 45–66, 2010
2010
-
[24]
Learning the parts of objects by non -negative matrix factorization,
D. Lee and H. Seung, “Learning the parts of objects by non -negative matrix factorization,” Nature, 1999
1999
-
[25]
Monaural sound source separation by nonn egative matrix factorization with temporal continuity and sparseness cri teria,
T. Virtanen, “Monaural sound source separation by nonn egative matrix factorization with temporal continuity and sparseness cri teria,” IEEE TASLP, vol. 15, no. 3, pp. 1066–1074, 2007
2007
-
[26]
Repeating pattern extraction tec hnique (REPET): A simple method for music/voice separation,
Z. Rafii and B. Pardo, “Repeating pattern extraction tec hnique (REPET): A simple method for music/voice separation,” IEEE TASLP , vol. 21, no. 1, pp. 73–84, 2013
2013
-
[27]
Harmonic/percussive separation usin g median filtering,
D. FitzGerald, “Harmonic/percussive separation usin g median filtering,” in Proc. DAFx, 2010
2010
-
[28]
Score-informed source separation for musical audio recordings: An overview,
S. Ewert et al. , “Score-informed source separation for musical audio recordings: An overview,” IEEE Signal Processing Magazine , vol. 31, no. 3, pp. 116–124, 2014
2014
-
[29]
Supervised speech separation base d on deep learning: An overview,
D. Wang and J. Chen, “Supervised speech separation base d on deep learning: An overview,” IEEE/ACM TASLP , vol. 26, pp. 1702–1726, 2017
2017
-
[30]
Towards scaling up classification- based speech separation,
Y . Wang and D. Wang, “Towards scaling up classification- based speech separation,” IEEE TASLP , vol. 21, no. 7, pp. 1381–1390, 2013
2013
-
[31]
Deep clustering: Discriminative embeddings for segmentation and separation,
J. R. Hershey et al. , “Deep clustering: Discriminative embeddings for segmentation and separation,” in Proc. ICASSP , 2016
2016
-
[32]
Permutation invariant training of deep models f or speaker- independent multi-talker speech separation,
D. Y u, “Permutation invariant training of deep models f or speaker- independent multi-talker speech separation,” in Proc. ICASSP, 2017, pp. 241–245
2017
-
[33]
Looking to listen at the cocktail party: A speaker- independent audio-visual model for speech separation,
A. Ephrat et al. , “Looking to listen at the cocktail party: A speaker- independent audio-visual model for speech separation,” ACM TOG , vol. 37, no. 4, apr 2018
2018
-
[34]
Neural target speech extraction: An overview,
K. Zmolikova et al. , “Neural target speech extraction: An overview,” IEEE Signal Processing Magazine , vol. 40, no. 3, pp. 8–29, 2023
2023
-
[35]
Combining spectral and spatial features for deep learning based blind speaker separation,
Z.-Q. Wang and D. Wang, “Combining spectral and spatial features for deep learning based blind speaker separation,” IEEE/ACM TASLP , vol. 27, no. 2, pp. 457–468, 2019
2019
-
[36]
Multi-microphone complex spectral mapping for utterance-wise and continuous speaker separation,
Z.-Q. Wang et al. , “Multi-microphone complex spectral mapping for utterance-wise and continuous speaker separation,” IEEE/ACM TASLP , vol. 29, pp. 2001–2014, 2021
2001
-
[37]
Multichannel speech enhancement by raw waveform- mapping using fully convolutional networks,
C. L. Liu et al. , “Multichannel speech enhancement by raw waveform- mapping using fully convolutional networks,” IEEE/ACM TASLP , vol. 28, pp. 1888–1900, 2020
1900
-
[38]
End-to-end microphone permutation and number invariant multi-channel speech separation,
Y . Luo et al., “End-to-end microphone permutation and number invariant multi-channel speech separation,” in Proc. ICASSP , 2020, pp. 6394– 6398
2020
-
[39]
V arArray: Array-geometry-agnostic continuous speech separation,
T. Y oshioka et al. , “V arArray: Array-geometry-agnostic continuous speech separation,” in Proc. ICASSP , 2022, pp. 6027–6031
2022
-
[40]
Complex ratio masking for monaural speech separation,
D. S. Williamson et al. , “Complex ratio masking for monaural speech separation,” IEEE/ACM TASLP, pp. 483–492, 2016
2016
-
[41]
Learning complex spectral mapping w ith gated convolutional recurrent networks for monaural speech enha ncement,
K. Tan and D. Wang, “Learning complex spectral mapping w ith gated convolutional recurrent networks for monaural speech enha ncement,” IEEE/ACM TASLP, vol. 28, pp. 380–390, 2020
2020
-
[42]
Conv-TasNet: Surpassing idea l time- frequency magnitude masking for speech separation,
Y . Luo and N. Mesgarani, “Conv-TasNet: Surpassing idea l time- frequency magnitude masking for speech separation,” IEEE/ACM TASLP, vol. 27, no. 8, pp. 1256–1266, 2019
2019
-
[43]
Wave-U-Net: A multi-scale neural network for end- to-end audio source separation,
D. Stoller et al. , “Wave-U-Net: A multi-scale neural network for end- to-end audio source separation,” in Proc. ISMIR , 2018, pp. 334–340
2018
-
[44]
Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks,
H. Erdogan et al. , “Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks,” in Proc. ICASSP , 2015, pp. 708–712
2015
-
[45]
Attention is all you need in speech separation,
C. Subakan et al. , “Attention is all you need in speech separation,” in Proc. ICASSP , 2021, pp. 21–25
2021
-
[46]
A convolutional recurrent neural ne twork for real-time speech enhancement,
K. Tan and D. Wang, “A convolutional recurrent neural ne twork for real-time speech enhancement,” in Interspeech, 2018, pp. 3229–3233
2018
-
[47]
Dual-path RNN: Efficient long sequence modeling for time-domain single-channel speech separation,
Y . Luo et al. , “Dual-path RNN: Efficient long sequence modeling for time-domain single-channel speech separation,” in Proc. ICASSP, 2020, pp. 46–50
2020
-
[49]
Open-Unmix - a reference implementation for music source separation,
F.-R. St¨ oter et al., “Open-Unmix - a reference implementation for music source separation,” Journal of Open Source Software , 2019
2019
-
[50]
Improving music source separation based on deep neural networks through data augmentation and network blen ding,
S. Uhlich et al. , “Improving music source separation based on deep neural networks through data augmentation and network blen ding,” in Proc. ICASSP , 2017, pp. 261–265
2017
-
[51]
MMDenseLSTM: An efficient combination of convolutional and recurrent neural networks for audio sour ce separation,
N. Takahashi et al. , “MMDenseLSTM: An efficient combination of convolutional and recurrent neural networks for audio sour ce separation,” in Proc. IWAENC, 2018, pp. 106–110
2018
-
[52]
Hybrid spectrogram and waveform source separation,
A. D´ efossez, “Hybrid spectrogram and waveform source separation,” arXiv:2111.03600, 2022
2022 arXiv
-
[53]
The whole is greater than the sum of its parts: improving music source separation by bridging networks,
R. Sawata et al. , “The whole is greater than the sum of its parts: improving music source separation by bridging networks,” EURASIP J. Audio Speech Music. Process. , vol. 2024, no. 1, p. 39, 2024
2024
-
[54]
Universal sound separation,
I. Kavalerov et al., “Universal sound separation,” in Proc. WASPAA2019, 2019, pp. 175–179
2019
-
[55]
Unsupervised sound separation using mixture invari- ant training,
S. Wisdom et al. , “Unsupervised sound separation using mixture invari- ant training,” in Proc. NeurIPS, vol. 33, 2020
2020
-
[56]
Listen to what you want: Neural network-based universal sound selector,
T. Ochiai et al. , “Listen to what you want: Neural network-based universal sound selector,” in Interspeech, 2020, pp. 1441–1445
2020
-
[57]
Separate anything you describe,
X. Liu et al. , “Separate anything you describe,” arXiv preprint arXiv:2308.05037, 2023
2023 arXiv
-
[58]
BLSTM supported GEV beamformer front-end for the 3rd CHiME challenge,
J. Heymann et al. , “BLSTM supported GEV beamformer front-end for the 3rd CHiME challenge,” in Proc. ASRU, 2015
2015
-
[59]
Integrating DNN-based and spatial clustering-based mask estimation for robust MVDR beamforming,
T. Nakatani et al. , “Integrating DNN-based and spatial clustering-based mask estimation for robust MVDR beamforming,” in Proc. ICASSP , 2017, pp. 286–290
2017
-
[60]
Unsupervised training of a deep clustering model for multichannel blind source separation,
L. Drude et al. , “Unsupervised training of a deep clustering model for multichannel blind source separation,” in Proc. ICASSP , 2019
2019
-
[61]
Unsupervised deep clustering for source separation: Di- rect learning from mixtures using spatial information,
E. Tzinis et al., “Unsupervised deep clustering for source separation: Di- rect learning from mixtures using spatial information,” in Proc. ICASSP, 2019, pp. 81–85
2019
-
[62]
The 2018 signal separation evaluation campaign,
F. R. St¨ oter et al., “The 2018 signal separation evaluation campaign,” in Proc. LVA/ICA, 2018, pp. 293–305
2018
-
[63]
The signal separation evaluation campaign (2007–2010): Achievements and remaining challenges,
E. Vincent et al. , “The signal separation evaluation campaign (2007–2010): Achievements and remaining challenges,” Signal Process- ing, vol. 92, no. 8, pp. 1928–1936, 2012
2007
-
[64]
Music demixing challenge 2021,
Y . Mitsufuji et al., “Music demixing challenge 2021,” Frontiers in Signal Processing, vol. 1, 2022
2021
-
[65]
The sound demixing challenge 2023 - music demixing track,
G. Fabbro et al., “The sound demixing challenge 2023 - music demixing track,” TISMIR, vol. 7, no. 1, pp. 63–84, 2024
2023
-
[66]
The sound demixing challenge 2023 - cinematic demixing track,
S. Uhlich et al. , “The sound demixing challenge 2023 - cinematic demixing track,” TISMIR, vol. 7, no. 1, pp. 44–62, 2024
2023
-
[67]
Performance measurement in blind audio source separation,
E. Vincent et al. , “Performance measurement in blind audio source separation,” IEEE TASLP , vol. 14, no. 4, pp. 1462–1469, 2006
2006
-
[68]
SDR – half-baked or well done?
J. Le Roux et al. , “SDR – half-baked or well done?” in Proc. ICASSP, 2019, pp. 626–630
2019
-
[69]
An algorithm for intelligibility prediction of time- frequency weighted noisy speech,
C. H. Taal et al. , “An algorithm for intelligibility prediction of time- frequency weighted noisy speech,” IEEE TASLP , vol. 19, no. 7, pp. 2125–2136, 2011
2011
-
[70]
Perceptual evaluation of speech quality (PESQ) the new ITU standard for end-to- end speech quality assessment part I–time-delay compensat ion,
A. W. Rix, M. P . Hollier, A. P . Hekstra, and J. G. Beerends , “Perceptual evaluation of speech quality (PESQ) the new ITU standard for end-to- end speech quality assessment part I–time-delay compensat ion,” Journal of the Audio Engineering Society , vol. 50, no. 10, pp. 755–...
2002
-
[71]
WHAMR!: Noisy and reverberant single-channel speech separation,
M. Maciejewski et al., “WHAMR!: Noisy and reverberant single-channel speech separation,” in Proc. ICASSP , 2020
2020
-
[72]
LibriMix: An open-source dataset for generalizable speech separation,
J. Cosentino et al. , “LibriMix: An open-source dataset for generalizable speech separation,” arXiv preprint arXiv:2005.11262 , 2020
2005 arXiv
-
[73]
SMS-WSJ: Database, performance measures, and baseline recipe for multi-channel source separation and re cognition,
L. Drude et al. , “SMS-WSJ: Database, performance measures, and baseline recipe for multi-channel source separation and re cognition,” arXiv preprint arXiv:1910.13934 , 2019
1910 arXiv
-
[74]
EARS: An anechoic fullband speech dataset bench- marked for speech enhancement and dereverberation,
J. Richter et al. , “EARS: An anechoic fullband speech dataset bench- marked for speech enhancement and dereverberation,” in Interspeech, 2024
2024
-
[75]
Pyroomacoustics: A python package for audio room simulation and array processing algorithms,
R. Scheibler et al., “Pyroomacoustics: A python package for audio room simulation and array processing algorithms,” in Proc. ICASSP , 2018
2018
-
[76]
gpuRIR: A python library for room impulse response simulation with GPU acceleration,
D. Diaz-Guerra et al. , “gpuRIR: A python library for room impulse response simulation with GPU acceleration,” Multimedia Tools and Applications, vol. 80, no. 4, pp. 5653–5671, 2021
2021
-
[77]
Asteroid: the PyTorch-based audio source separation toolkit for researchers,
M. Pariente et al., “Asteroid: the PyTorch-based audio source separation toolkit for researchers,” in Interspeech, 2020
2020
-
[78]
ESPnet-SE: End-to-end speech enhancement and separatio n toolkit designed for asr integration,
C. Li et al., “ESPnet-SE: End-to-end speech enhancement and separatio n toolkit designed for asr integration,” in Proc. SLT, 2021, pp. 785–792
2021
-
[79]
Investigating self-supervised learning for speech enhancement and separation,
Z. Huang et al. , “Investigating self-supervised learning for speech enhancement and separation,” in Proc. ICASSP , 2022, pp. 6837–6841
2022
-
[80]
SuperM2M: Supervised and mixture-to-mix ture co- learning for speech enhancement and robust ASR,
Z.-Q. Wang, “SuperM2M: Supervised and mixture-to-mix ture co- learning for speech enhancement and robust ASR,” arxiv preprint arXiv:2403.10271v2, 2024
2024 arXiv
-
[81]
Diffusion-based generative speech source separa- tion,
R. Scheibler et al. , “Diffusion-based generative speech source separa- tion,” in Proc. ICASSP 2023 , 2023, pp. 1–5
2023
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.