REVIEW 1 cited by
Semi-Supervised Monaural Singing Voice Separation With a Masking Network Trained on Synthetic Mixtures
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We study the problem of semi-supervised singing voice separation, in which the training data contains a set of samples of mixed music (singing and instrumental) and an unmatched set of instrumental music. Our solution employs a single mapping function g, which, applied to a mixed sample, recovers the underlying instrumental music, and, applied to an instrumental sample, returns the same sample. The network g is trained using purely instrumental samples, as well as on synthetic mixed samples that are created by mixing reconstructed singing voices with random instrumental samples. Our results indicate that we are on a par with or better than fully supervised methods, which are also provided with training samples of unmixed singing voices, and are better than other recent semi-supervised methods.
Forward citations
Cited by 1 Pith paper
-
Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed
A waveform-domain convolutional-recurrent network with a remix-silence semi-supervised scheme reaches near state-of-the-art music source separation on MusDB without extra labeled data.
Discussion (0). Continue with ORCID to comment.