REVIEW 5 cited by
Multi-Source Diffusion Models for Simultaneous Music Generation and Separation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this work, we define a diffusion-based generative model capable of both music synthesis and source separation by learning the score of the joint probability density of sources sharing a context. Alongside the classic total inference tasks (i.e., generating a mixture, separating the sources), we also introduce and experiment on the partial generation task of source imputation, where we generate a subset of the sources given the others (e.g., play a piano track that goes well with the drums). Additionally, we introduce a novel inference method for the separation task based on Dirac likelihood functions. We train our model on Slakh2100, a standard dataset for musical source separation, provide qualitative results in the generation settings, and showcase competitive quantitative results in the source separation setting. Our method is the first example of a single model that can handle both generation and separation tasks, thus representing a step toward general audio models.
Forward citations
Cited by 5 Pith papers
-
Conditional Flow Matching for Visually-Guided Acoustic Highlighting
Conditional flow matching with a rollout loss and early audio-visual fusion achieves state-of-the-art results on visually-guided acoustic highlighting.
-
Unsupervised Single-Channel Speech Separation with Diffusion under Speaker-Embedding Guidance
Adding speaker-embedding guidance to an unconditional diffusion prior improves unsupervised single-channel speech separation, reaching 9.32 dB SI-SDR on VCTK-2mix, the best among unsupervised baselines.
-
User-guided Generative Source Separation
GuideSep separates arbitrary target instruments from a mixture using user-provided waveform mimicry and mel-spectrogram masks, and outperforms a same-architecture mask-prediction baseline in SDR and listening tests.
-
Training-Free Multi-Step Audio Source Separation
Iteratively remixing and re-separating the input mixture, with the best blend chosen by a quality metric, improves pretrained one-step audio separation models without any retraining.
-
A Mixture-Based Framework for Guiding Diffusion Models
MGDM approximates the intractable guided-diffusion posterior with a weighted mixture of likelihood approximations and samples the mixture using a Gibbs sampler with tunable repetitions.
Discussion (0). Continue with ORCID to comment.