Pith. sign in

REVIEW 5 cited by

Multi-Source Diffusion Models for Simultaneous Music Generation and Separation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.02257 v4 pith:MGAANRMQ submitted 2023-02-04 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords separationgenerationsourcemodelsourcesinferenceintroducemethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we define a diffusion-based generative model capable of both music synthesis and source separation by learning the score of the joint probability density of sources sharing a context. Alongside the classic total inference tasks (i.e., generating a mixture, separating the sources), we also introduce and experiment on the partial generation task of source imputation, where we generate a subset of the sources given the others (e.g., play a piano track that goes well with the drums). Additionally, we introduce a novel inference method for the separation task based on Dirac likelihood functions. We train our model on Slakh2100, a standard dataset for musical source separation, provide qualitative results in the generation settings, and showcase competitive quantitative results in the source separation setting. Our method is the first example of a single model that can handle both generation and separation tasks, thus representing a step toward general audio models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Conditional Flow Matching for Visually-Guided Acoustic Highlighting

    eess.AS 2026-02 conditional novelty 6.0 of 10

    Conditional flow matching with a rollout loss and early audio-visual fusion achieves state-of-the-art results on visually-guided acoustic highlighting.

  2. Unsupervised Single-Channel Speech Separation with Diffusion under Speaker-Embedding Guidance

    eess.AS 2025-09 conditional novelty 6.0 of 10

    Adding speaker-embedding guidance to an unconditional diffusion prior improves unsupervised single-channel speech separation, reaching 9.32 dB SI-SDR on VCTK-2mix, the best among unsupervised baselines.

  3. User-guided Generative Source Separation

    cs.SD 2025-07 conditional novelty 6.0 of 10

    GuideSep separates arbitrary target instruments from a mixture using user-provided waveform mimicry and mel-spectrogram masks, and outperforms a same-architecture mask-prediction baseline in SDR and listening tests.

  4. Training-Free Multi-Step Audio Source Separation

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Iteratively remixing and re-separating the input mixture, with the best blend chosen by a quality metric, improves pretrained one-step audio separation models without any retraining.

  5. A Mixture-Based Framework for Guiding Diffusion Models

    stat.ML 2025-02 conditional novelty 6.0 of 10

    MGDM approximates the intractable guided-diffusion posterior with a weighted mixture of likelihood approximations and samples the mixture using a Gibbs sampler with tunable repetitions.

Pith tools