Pith. sign in

REVIEW 2 cited by

Benchmarks and leaderboards for sound demixing tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.07489 v2 pith:HWZONKMT submitted 2023-05-12 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords demixingmodelsbenchmarksdifferentseparationsoundapproachaudio
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Music demixing is the task of separating different tracks from the given single audio signal into components, such as drums, bass, and vocals from the rest of the accompaniment. Separation of sources is useful for a range of areas, including entertainment and hearing aids. In this paper, we introduce two new benchmarks for the sound source separation tasks and compare popular models for sound demixing, as well as their ensembles, on these benchmarks. For the models' assessments, we provide the leaderboard at https://mvsep.com/quality_checker/, giving a comparison for a range of models. The new benchmark datasets are available for download. We also develop a novel approach for audio separation, based on the ensembling of different models that are suited best for the particular stem. The proposed solution was evaluated in the context of the Music Demixing Challenge 2023 and achieved top results in different tracks of the challenge. The code and the approach are open-sourced on GitHub.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A new public-domain film dataset and a frame-by-frame dialogue-conditioning module improve video-to-music generation on paired-fidelity metrics.

  2. Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models

    cs.SD 2026-07 conditional novelty 3.5 of 10

    MSST unifies training, validation, and inference for many music source-separation architectures and reports small quality gains from TTA, ensembling, and related engineering techniques.

Pith tools