Pith. sign in

REVIEW 5 cited by

SingSong: Generating musical accompaniments from singing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.12662 v1 pith:L7R4A3RJ submitted 2023-01-30 cs.SD cs.AIcs.LGcs.MMeess.AS

classification cs.SDcs.AIcs.LGcs.MMeess.AS
keywords singsongaudiogenerationinstrumentalmusicmusicalpairsseparation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present SingSong, a system that generates instrumental music to accompany input vocals, potentially offering musicians and non-musicians alike an intuitive new way to create music featuring their own voice. To accomplish this, we build on recent developments in musical source separation and audio generation. Specifically, we apply a state-of-the-art source separation algorithm to a large corpus of music audio to produce aligned pairs of vocals and instrumental sources. Then, we adapt AudioLM (Borsos et al., 2022) -- a state-of-the-art approach for unconditional audio generation -- to be suitable for conditional "audio-to-audio" generation tasks, and train it on the source-separated (vocal, instrumental) pairs. In a pairwise comparison with the same vocal inputs, listeners expressed a significant preference for instrumentals generated by SingSong compared to those from a strong retrieval baseline. Sound examples at https://g.co/magenta/singsong

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 8 citations worldwide. Full citation record

  1. DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization

    eess.AS 2025-07 conditional novelty 6.0 of 10

    DiffRhythm+ improves full-length lyric-to-song generation via balanced data scaling, MuLan-based multimodal style control, and DPO fine-tuning guided by automated aesthetic scorers.

  2. WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling

    cs.SD 2025-07 conditional novelty 6.0 of 10

    WildFX generates multi-track audio datasets by rendering real DAW effect graphs with commercial plugins inside Docker, and demonstrates the pipeline on blind mixing-graph estimation.

  3. Adaptive Accompaniment with ReaLchords

    cs.SD 2025-06 conditional novelty 6.0 of 10

    An online melody-to-chord accompaniment model, fine-tuned with reinforcement learning and distillation from a future-seeing teacher, recovers from cold starts and mid-song perturbations better than MLE baselines.

  4. Hookpad Aria: A Copilot for Songwriters

    cs.SD 2025-02 conditional novelty 6.0 of 10

    A deployed AI songwriting assistant in Hookpad supports non-sequential generation and collects a live flywheel of accepted and rejected suggestions.

  5. Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation

    cs.SD 2025-08 unverdicted novelty 5.0 of 10

    A bar-level symbolic-score song generator (BACH) is claimed to beat published systems and commercial Suno on human-rated quality, duration, and efficiency, but the supporting full text is corrupted and unverifiable.

Pith tools