{"id":"19947d1e-c87c-4bce-9391-979cf6ac327d","arxiv_id":"2501.11837","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":0.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A comprehensive review of three decades of audio source separation research, presenting no new technical results.","lead":"This paper reviews 30+ years of acoustic source separation research, from classical ICA to modern deep learning. It is a commissioned historical survey for ICASSP's 50th anniversary, useful for newcomers and for tracking the field's evolution.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Coverage claim in '30+ Years of Source Separation Research' is not backed by explicit selection criteria and shows heavy self-citation, making the 'major contributions' claim unverifiable.","rationale":"The reader's weakest assumption—that the selected works, challenges, and directions constitute a representative history—is exactly where the central descriptive claim is most vulnerable. The paper is a useful technical summary and I found no obvious factual errors in the core narrative, but the claim to cover 'the major contributions' is under-supported. A citation-coverage check would settle whether the problem is real. If the check shows high coverage, the paper's claim stands; if it shows significant omissions, the authors should either add a transparent selection methodology or soften the claim. Because the paper is a review with independent historical value, I would not reject it outright; a conditional acceptance requiring methodological transparency or a more balanced reference list is the appropriate adjustment. I therefore align with the reader's identification of representativeness as the weak point and recommend moving from UNVERDICTED to CONDITIONAL rather than leaving the claim unexamined.","tokens_in":10616,"tokens_out":6489,"duration_ms":69447,"concrete_test":"Construct an independent corpus by querying DBLP/Scopus for source separation publications in ICASSP, Interspeech, WASPAA, TASLP, ISMIR, and NeurIPS (1995–2024) with 'source separation', 'speech separation', 'music separation', or 'audio separation' in the title or abstract. Rank the papers by citation count. Compute how many of the top 50 most-cited papers are cited in this review. If more than 15–20% of these top-cited works are absent, the 'major contributions' claim is insufficiently supported. Also compute the fraction of references co-authored by at least one review author and compare this fraction with that of a comparable systematic review; a large discrepancy would confirm self-citation bias.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the paper reviews the major contributions in 30+ years of source separation. This descriptive claim is load-bearing only insofar as the selection of works is demonstrably representative. The paper gives no inclusion criteria for 'major', no systematic search, and no justification for why certain subfields are condensed—sound event separation gets one short paragraph, and music separation is summarized mainly through Open-Unmix and the MDX/SDX challenges. The reference list contains a notable cluster of the authors' own publications, e.g., references [10], [18], [36], [48], [53], [64], and [80], plus challenge reports from the same groups. That is not by itself inappropriate, but it creates a real risk that the narrative over-weights the authors' preferred methods (multi-channel complex mapping, TF-GridNet, Nara-WPE, MDX/SDX) and under-represents other influential lines such as certain end-to-end time-domain architectures, transformer-based separators, and reference-free evaluation metrics. Because the paper never discloses a methodology for selecting 'major contributions', the reader cannot distinguish a balanced overview from a selected retrospective, and the stated goal of reviewing 'the major contributions and advancements' remains unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a review of source separation (SS) research over the past three decades, written for the occasion of ICASSP's 50th anniversary. It covers the problem formulation, model-based approaches (ICA, IVA, ILRMA, FCA, TF masking), deep-learning approaches (deep clustering, PIT, complex spectral mapping, time-domain and transformer-based systems), hybrid beamforming methods, and the evaluation infrastructure of the field (SiSEC, CHiME, MDX/SDX challenges, metrics, datasets, and open-source tools). It closes with a short discussion of remaining challenges and possible future directions. The manuscript does not present new algorithms or experiments; its contribution is historical and technical synthesis.","tokens_in":10826,"tokens_out":6757,"duration_ms":64871,"significance":"If the coverage issues are addressed, this would be a useful and readable retrospective for the source separation community and for newcomers. The technical descriptions of the core methods (e.g., the convolutive mixture model in Eq. (1), the TF-domain approximation in Eq. (3), ICA, ILRMA, FCA, deep clustering, PIT, and beamforming hybrids) are accurate and appropriately cited. The paper also does a service by documenting the evaluation culture—challenges, metrics, and datasets—that has been crucial to the field. Its significance as a 'major contributions' review, however, is currently weakened by the absence of explicit selection criteria, which makes the representativeness of the chosen topics and references difficult to verify.","major_comments":[{"comment":"The abstract and introduction claim that the paper reviews 'the major contributions and advancements' in source separation over the past three decades, but no inclusion criteria or selection methodology is ever stated. This makes the central claim of the paper difficult to verify. For example, Sec. III-B highlights dual-path architectures and singles out TF-GridNet [48] as the representative modern system without justifying why this particular work is chosen over other recent architectures, and Sec. IV-C lists a few open-source tools without defining the basis for their selection. As written, the reader cannot distinguish a balanced overview from a retrospective centered on the authors' own research priorities. I recommend adding an explicit methodology or scope statement (e.g., criteria for what counts as a 'major contribution', how subtopics were allocated, and any search or selection process), or alternatively reframing the paper as a 'selected overview' and adjusting the title and abstract accordingly.","section":"I, III-B, IV"}],"minor_comments":[{"comment":"In Sec. IV-C, the text attributes 'Nara-WPE [60]' as an open implementation of a dereverberation algorithm, but reference [60] is the paper 'Unsupervised training of a deep clustering model for multichannel blind source separation' by Drude et al. This citation appears to be incorrect and should be replaced with the appropriate Nara-WPE or WPE reference.","section":"IV-C"},{"comment":"The sentence 'This SiSEC initiative was taken over by CHiME 2' is historically imprecise, since SiSEC continued until 2018 and CHiME was launched separately in 2010. Please revise to describe the relationship between these initiatives more accurately.","section":"IV-A"},{"comment":"The statement that ICASSP has 'consistently received 40–50 submissions to AUD-SEP every year' lacks a citation. Please add a source for this statistic or soften the claim.","section":"I"},{"comment":"The opening sentence of Sec. IV says 'in the 2000s, there were no benchmarks in SS research', but the next sentences describe SASSEC in 2007. Please rephrase to 'in the early 2000s' or 'before SASSEC/SiSEC' to avoid the apparent contradiction.","section":"IV"}],"recommendation":"major_revision","confidential_remarks":"This appears to be an invited overview paper for an anniversary issue, so the bar for comprehensiveness may be somewhat different from a standard research article. Still, the lack of any stated selection criteria is a legitimate concern for a paper whose central claim is coverage of 'major contributions'. The requested changes are local and should be straightforward to implement: a brief methodology or framing statement and a few citation/history corrections. I therefore see this as a revision rather than a rejection, but the representativeness issue should be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a competent, clearly written review of source separation, not a research contribution. If you want a quick map of the field—model-based methods, deep learning, the CHiME/SiSEC/MDX evaluation ecosystem—it does the job. The technical descriptions are accurate, the historical arc is sensible, and the emphasis on evaluation culture is a real plus. It's the kind of paper you'd hand to a new student.\n\nThe soft spot is exactly the one the stress-test flags: 'major contributions' is a claim without a method. There's no stated selection criteria, no systematic search, and no defense of why certain subfields are compressed. Music separation beyond Open-Unmix and MDX gets thin coverage; sound event separation gets a few lines. And yes, the reference list skews toward the authors' own work (TF-GridNet, Nara-WPE, MDX, their CHiME systems). That's not disqualifying—an invited retrospective by active researchers is naturally going to reflect their view of the field—but it does mean the historical narrative should be read as a selected retrospective, not an objective census. The paper never says that, and a reviewer should ask them to state it.\n\nI wouldn't call this a fatal flaw. The paper is explicit about being a review for ICASSP's anniversary, and for that purpose the selectivity is acceptable. The sections on evaluation metrics, datasets, and open-source tools are genuinely useful and hard to find in one place. The reader's scores (novelty 0, significance 4, soundness 7) seem about right.\n\nWho this is for: anyone entering source separation who wants context, and veterans who want a reminder of where benchmark culture came from. It's not for someone looking for new results or a critical literature analysis.\n\nRecommendation: a serious editor should send this to peer review—not because it's groundbreaking, but because it's an authoritative review from people who helped build the field, and the balance/representativeness issues deserve a careful referee's eye. With a scope caveat and a few added references to other lines of work, it would be a genuinely useful reference.","headline":"Competent, usefully selective review of source separation; the 'major contributions' claim needs a scope caveat, but it's a fair survey for newcomers.","tokens_in":11307,"tokens_out":2862,"would_cite":true,"duration_ms":27374,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 30-year map of source separation: what worked, what is left.","keywords":["source separation","blind source separation","speech separation","music source separation","deep learning","independent component analysis","time-frequency masking","evaluation challenges"],"falsifier":"An independent bibliometric or community-wide survey of source-separation publications from 1994 to 2024, ranking papers by citation counts or by adoption in deployed systems, would show whether the milestones highlighted here are in fact the field's pivotal contributions; if the most influential works differ substantially from those featured, the review's claim to cover the major contributions and advancements fails.","tokens_in":10444,"feed_emoji":"🎧","tokens_out":8341,"duration_ms":79910,"temperature":0.7,"pith_summary":"This review aims to give a reliable, layered account of three decades of acoustic source separation: the problem's mathematical form, the model-based methods (ICA, IVA, NMF, full-rank spatial covariance), the deep-learning turn (masking, permutation-invariant training, time-domain and hybrid beamforming), and the benchmarking culture that made progress measurable. The authors argue that the field is mature enough to have a history and an evaluation infrastructure, yet still far from solved in real conditions. A reader should care because source separation is the technology behind hearing one voice in a crowd, for hearing aids, speech recognition, and music production, and this paper maps both the accumulated toolbox and the concrete bottlenecks (real-world generalization, unknown source counts, perception-based metrics) that now define the agenda.","feed_headline":"30 years of source separation: benchmarks won, real rooms left","feed_subtitle":"A review shows simulated benchmarks have saturated; unknown source counts and real-world noise are the new frontier.","key_machinery":"The organizing object is the convolutive mixture model $y_m(\\tilde{t}) = \\sum_n \\sum_\\tau h_{mn}(\\tau) s_n(\\tilde{t}-\\tau) + v_m(\\tilde{t})$, approximated as an instantaneous mixture in the time-frequency domain, $y_{mtf} = \\sum_n h_{mnf} s_{ntf} + v_{mtf}$. This formulation generates the field's central technical obstacle, the frequency permutation problem, and the review's narrative machinery is the sequence of mechanisms devised to overcome it: independent vector analysis bundles frequency components statistically; NMF models full-band spectrograms; deep clustering embeds time-frequency units so same-source units are close; permutation-invariant training aligns outputs to targets before the loss; and hybrid systems feed DNN-estimated masks into beamformers. The evaluation culture, standard metrics (SDR, SIR, SAR), SI-SDR, and the challenge infrastructure from SiSEC through the Music Demixing Challenge, is the other machinery, since it is what converts competing algorithms into comparable numbers.","core_discovery":"On the authors' account, source separation has moved through three overlapping phases. From the mid-1990s to the 2000s, blind separation of determined mixtures was solved in principle by independent component analysis, extended to convolutive and underdetermined cases by time-frequency masking, independent vector analysis, nonnegative matrix factorization, and full-rank spatial covariance models. The 2010s brought a supervised deep-learning breakthrough: neural networks learned to estimate time-frequency masks, with deep clustering and permutation-invariant training resolving the label-permutation problem for homogeneous sources, and architectures moved from masking to complex-spectrum and waveform estimation. The present phase is hybrid, in which DNNs estimate masks or spatial statistics that drive beamformers, alongside unsupervised and mixture-invariant training to use real recordings without ground truth. The paper's central claim is that this trajectory, supported by shared datasets and metrics (SDR/SIR/SAR, SI-SDR, challenges from SiSEC to the Music Demixing Challenge), has produced a field whose simulated benchmarks have saturated while its remaining difficulties, mismatched real data, unknown and time-varying source counts, limited microphones, and reference-free perceptual evaluation, define the open research agenda.","pith_inferences":["The authors leave implicit that if simulated benchmarks are saturated and real data is the bottleneck, then progress will depend as much on data collection and simulation-realism engineering as on new network architectures.","The review's own emphasis on a few state-of-the-art systems (for example, the dual-path/spectrogram-grid architectures and the dereverberation front-end it highlights) suggests that the field's advanced capability is concentrated in a small number of research groups; independent reproductions would test whether the reported gains generalize.","The brief mention of natural-language-prompted separation hints at a convergence with foundation-model approaches; if that trend holds, source separation may be reframed as a text-conditioned generation problem, which the current benchmark infrastructure does not yet evaluate.","The stress on low-latency and lightweight models suggests that deployment, not raw separation quality, may be the next differentiator; one could test this by comparing real-time capable systems against offline state-of-the-art on the same real-world data."],"forward_implications":["If the review's history is right, new methods should be measured against real-recorded, conversational benchmarks rather than fully-overlapped simulated mixtures, where performance has saturated.","Multichannel separation in realistic conditions will increasingly rely on hybrid designs, DNN-driven mask estimation feeding classical beamformers, since each component handles different signal properties.","Unsupervised and mixture-invariant training will become standard for leveraging in-the-wild data, because supervised training on simulated mixtures transfers poorly to real rooms.","Research attention will shift from fixed-source-count separation to joint source activity detection and separation, since real scenes contain unknown, time-varying numbers of sources.","Evaluation practice will need new reference-free metrics that track human perception for speech, music, and general sounds, replacing alignment-based SDR-style measures."],"supporting_citations":[{"why":"Introduces independent component analysis, the first principled method for determined blind source separation.","marker":"[4]"},{"why":"Unifies independent vector analysis with nonnegative matrix factorization for determined blind separation (ILRMA).","marker":"[8]"},{"why":"Shows that supervised DNNs can learn to estimate ideal binary masks, triggering the deep-learning era in separation.","marker":"[30]"},{"why":"Introduces deep clustering to resolve label permutation and provides the wsj0-2mix dataset creation script.","marker":"[31]"},{"why":"Introduces permutation invariant training (PIT), which aligns estimated sources with true sources before loss computation.","marker":"[32]"},{"why":"Presents Conv-TasNet, showing that end-to-end time-domain separation can surpass ideal time-frequency masking.","marker":"[42]"},{"why":"Presents TF-GridNet, a state-of-the-art architecture that integrates full-band and sub-band modeling across the spectrogram.","marker":"[48]"},{"why":"Documents the 2018 Signal Separation Evaluation Campaign, which the review credits with creating a benchmarking culture.","marker":"[62]"},{"why":"Describes the Music Demixing Challenge 2021, which continued the music separation evaluation line started by SiSEC.","marker":"[64]"},{"why":"Defines the SDR, SIR, and SAR metrics that became the field's standard evaluation toolkit for separated signals.","marker":"[67]"}],"fun_headline_variants":["Source separation: benchmark wins, real-world noise still stumps","From ICA to deep learning: source separation's real-world test remains","Source separation benchmarks saturated; real-world unknowns ahead","30 years in, source separation still can't handle real rooms"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review assumes that the works, benchmarks, and challenges it selects give a representative and unbiased history of the field rather than a picture shaped by the authors' own research priorities.","fun_headline_variants_meta":{"raw":{"variants":["Source separation: benchmark wins, real-world noise still stumps","From ICA to deep learning: source separation's real-world test remains","Source separation benchmarks saturated; real-world unknowns ahead","30 years in, source separation still can't handle real rooms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000973,"raw_usage":{"total_tokens":4098,"prompt_tokens":870,"completion_tokens":3228,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":486,"completion_tokens_details":{"reasoning_tokens":3159}},"tokens_in":486,"tokens_out":3228,"duration_ms":27061,"temperature":1.0,"reasoning_tokens":3159,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:47:13.414341+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent bibliometric or community-wide survey of source-separation publications from 1994 to 2024, ranking papers by citation counts or by adoption in deployed systems, would show whether the milestones highlighted here are in fact the field's pivotal contributions; if the most influential works differ substantially from those featured, the review's claim to cover the major contributions and advancements fails.","supporting_citations":[{"cited_title":"An information-maximiza tion approach to blind separation and blind deconvolution,","cited_arxiv_id":null,"evidence_quote":"Introduces independent component analysis, the first principled method for determined blind source separation."},{"cited_title":"De- termined blind source separation unifying independent vec tor analysis and nonnegative matrix factorization,","cited_arxiv_id":null,"evidence_quote":"Unifies independent vector analysis with nonnegative matrix factorization for determined blind separation (ILRMA)."},{"cited_title":"Towards scaling up classiﬁcation- based speech separation,","cited_arxiv_id":null,"evidence_quote":"Shows that supervised DNNs can learn to estimate ideal binary masks, triggering the deep-learning era in separation."},{"cited_title":"Deep clustering: Discriminative embeddings for segmentation and separation,","cited_arxiv_id":null,"evidence_quote":"Introduces deep clustering to resolve label permutation and provides the wsj0-2mix dataset creation script."},{"cited_title":"Permutation invariant training of deep models f or speaker- independent multi-talker speech separation,","cited_arxiv_id":null,"evidence_quote":"Introduces permutation invariant training (PIT), which aligns estimated sources with true sources before loss computation."},{"cited_title":"Conv-TasNet: Surpassing idea l time- frequency magnitude masking for speech separation,","cited_arxiv_id":null,"evidence_quote":"Presents Conv-TasNet, showing that end-to-end time-domain separation can surpass ideal time-frequency masking."},{"cited_title":"TF-GridNet: Integrating full- and sub-band modeling for speech separation,","cited_arxiv_id":null,"evidence_quote":"Presents TF-GridNet, a state-of-the-art architecture that integrates full-band and sub-band modeling across the spectrogram."},{"cited_title":"The 2018 signal separation evaluation campaign,","cited_arxiv_id":null,"evidence_quote":"Documents the 2018 Signal Separation Evaluation Campaign, which the review credits with creating a benchmarking culture."},{"cited_title":"Music demixing challenge 2021,","cited_arxiv_id":null,"evidence_quote":"Describes the Music Demixing Challenge 2021, which continued the music separation evaluation line started by SiSEC."},{"cited_title":"Performance measurement in blind audio source separation,","cited_arxiv_id":null,"evidence_quote":"Defines the SDR, SIR, and SAR metrics that became the field's standard evaluation toolkit for separated signals."}],"review_version":1}