Pith. sign in

REVIEW 2 cited by

Accompanied Singing Voice Synthesis with Fully Text-controlled Melody

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.02049 v1 pith:ROE4Z62M submitted 2024-07-02 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords inputmelodylmmidimodelmusicsingingttsongvoice
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-song (TTSong) is a music generation task that synthesizes accompanied singing voices. Current TTSong methods, inherited from singing voice synthesis (SVS), require melody-related information that can sometimes be impractical, such as music scores or MIDI sequences. We present MelodyLM, the first TTSong model that generates high-quality song pieces with fully text-controlled melodies, achieving minimal user requirements and maximum control flexibility. MelodyLM explicitly models MIDI as the intermediate melody-related feature and sequentially generates vocal tracks in a language model manner, conditioned on textual and vocal prompts. The accompaniment music is subsequently synthesized by a latent diffusion model with hybrid conditioning for temporal alignment. With minimal requirements, users only need to input lyrics and a reference voice to synthesize a song sample. For full control, just input textual prompts or even directly input MIDI. Experimental results indicate that MelodyLM achieves superior performance in terms of both objective and subjective metrics. Audio samples are available at https://melodylm666.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation

    cs.SD 2026-08 conditional novelty 6.0 of 10

    MeloCodec quantizes chromagram-derived melody tokens and fuses them with acoustic tokens using a two-stage training scheme, improving pitch consistency and enabling controllable pitch shifting in a singing-voice codec.

  2. DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization

    eess.AS 2025-07 conditional novelty 6.0 of 10

    DiffRhythm+ improves full-length lyric-to-song generation via balanced data scaling, MuLan-based multimodal style control, and DPO fine-tuning guided by automated aesthetic scorers.

Pith tools