Pith. sign in

REVIEW 1 cited by

Drop the beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.15474 v1 pith:3UGXCRPH submitted 2024-08-28 eess.AS cs.SD

classification eess.AScs.SD
keywords freestylergenerationrappingvocalaccompanimentaccompanyingbeatsinputs
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Rap, a prominent genre of vocal performance, remains underexplored in vocal generation. General vocal synthesis depends on precise note and duration inputs, requiring users to have related musical knowledge, which limits flexibility. In contrast, rap typically features simpler melodies, with a core focus on a strong rhythmic sense that harmonizes with accompanying beats. In this paper, we propose Freestyler, the first system that generates rapping vocals directly from lyrics and accompaniment inputs. Freestyler utilizes language model-based token generation, followed by a conditional flow matching model to produce spectrograms and a neural vocoder to restore audio. It allows a 3-second prompt to enable zero-shot timbre control. Due to the scarcity of publicly available rap datasets, we also present RapBank, a rap song dataset collected from the internet, alongside a meticulously designed processing pipeline. Experimental results show that Freestyler produces high-quality rapping voice generation with enhanced naturalness and strong alignment with accompanying beats, both stylistically and rhythmically.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR

    eess.AS 2025-01 conditional novelty 5.0 of 10

    A generative target-speech extractor trained jointly with a Whisper-based transcript-prediction loss achieves better intelligibility and competitive quality on Libri2Mix and WSJ0-2mix than discrete-token and mask-base...

Pith tools