REVIEW 16 cited by
DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion
read the original abstract
Recent advancements in music generation have garnered significant attention, yet existing approaches face critical limitations. Some current generative models can only synthesize either the vocal track or the accompaniment track. While some models can generate combined vocal and accompaniment, they typically rely on meticulously designed multi-stage cascading architectures and intricate data pipelines, hindering scalability. Additionally, most systems are restricted to generating short musical segments rather than full-length songs. Furthermore, widely used language model-based methods suffer from slow inference speeds. To address these challenges, we propose DiffRhythm, the first latent diffusion-based song generation model capable of synthesizing complete songs with both vocal and accompaniment for durations of up to 4m45s in only ten seconds, maintaining high musicality and intelligibility. Despite its remarkable capabilities, DiffRhythm is designed to be simple and elegant: it eliminates the need for complex data preparation, employs a straightforward model structure, and requires only lyrics and a style prompt during inference. Additionally, its non-autoregressive structure ensures fast inference speeds. This simplicity guarantees the scalability of DiffRhythm. Moreover, we release the complete training code along with the pre-trained model on large-scale data to promote reproducibility and further research.
Forward citations
Cited by 16 Pith papers
-
Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation
UniSinger unifies speaker-cloned song generation and accompaniment co-generation SVC in one multimodal diffusion transformer model trained with curriculum learning via task-specific modality masking.
-
YingMusic-Singer: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance
YingMusic-Singer-Plus is a diffusion model for singing voice synthesis that preserves melody from a reference clip while allowing flexible lyric changes without manual alignment, outperforming Vevo2 and introducing th...
-
MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline
MIDI-SAG generates consistent long-form singing accompaniments by feeding symbolic MIDI timing, chords, and structure labels into a compositional pipeline built from pre-trained modules.
-
MusicMark: A Robust Generative Watermarking Framework for Music Generation
Embedding watermark bits into diffusion semantic latents via a frozen-backbone adapter yields far more robust music provenance than post-hoc watermarking under codecs and cover-song attacks.
-
MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation
Adding SVS-style phoneme conditioning and a length regulator to melody-guided cover generation cuts phoneme error rate from 45.6% to 18.7% while holding melody metrics.
-
Qwen-Music Technical Report
Qwen-Music generates high-fidelity vocal songs via 25 Hz semantic tokens, Melody-CoT planning, and DiT rendering, claiming SOTA on 13/16 metrics and expert preference over proprietary systems.
-
MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres
Current singing voice synthesis models fail to differentiate musical genres, defaulting to pop-like output regardless of input genre, unless given genre-specific fine-tuning data.
-
SingFox: A Multi-Lingual Singfake Detection Corpus
SingFox is a large-scale dataset of 113802 audio clips totaling 126.32 hours from 1150 singers across 20 languages, organized into six tracks for singing deepfake detection and source verification benchmarks.
-
An Empirical Analysis of AI Slop in Music Streaming
Empirical study finds 93% of AI music on Spotify gets negligible plays, distributors have inconsistent unenforced AI policies, and detection methods are unreliable, suggesting slop may become self-sustaining.
-
Probing Token Spaces under Generator Shift in AI-Generated Music Detection
Experiments on an open dataset show X-Codec tokens perform best under Udio shift while MERT tokens perform best under Suno-v3.5 shift, indicating token space choice is a key variable for generator-robust detection.
-
S2Accompanist: A Semantic-Aware and Structure-Guided Diffusion Model for Music Accompaniment Generation
S2Accompanist is a 402M-parameter semantic-aware diffusion model that achieves SOTA on the ATTM Grand Challenge benchmark for music accompaniment generation via automated data processing and structure-guided VAE fine-tuning.
-
SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment
SongBench is a new fine-grained benchmark for song quality assessment with seven dimensions and an expert-annotated dataset of 11,717 samples showing high correlation with professional ratings.
-
YingMusic-Singer: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance
A diffusion SVS model with curriculum training and GRPO edits lyrics while preserving melody without manual alignment, outperforming Vevo2 on LyricEditBench.
-
LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training
LeVo 2 presents a hierarchical LLM-Diffusion model with progressive post-training stages to generate full-length songs that balance semantic planning, track-specific acoustics, and musicality.
-
SketchSong: Hierarchical Song Generation with Sketch Planning and Fine-Grained Multi-Track Modeling
SketchSong uses temporal sketch planning with high-level tokens and explicit modeling of four tracks (vocals, bass, drums, other) to generate more coherent songs than baselines.
-
SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision
SongFormer achieves state-of-the-art strict boundary detection and functional label accuracy in music structure analysis by fusing SSL representations and using learned source embeddings on a new 14k-song corpus and e...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.