REVIEW 10 cited by
Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent years have seen the rapid development of large generative models for text; however, much less research has explored the connection between text and another "language" of communication -- music. Music, much like text, can convey emotions, stories, and ideas, and has its own unique structure and syntax. In our work, we bridge text and music via a text-to-music generation model that is highly efficient, expressive, and can handle long-term structure. Specifically, we develop Mo\^usai, a cascading two-stage latent diffusion model that can generate multiple minutes of high-quality stereo music at 48kHz from textual descriptions. Moreover, our model features high efficiency, which enables real-time inference on a single consumer GPU with a reasonable speed. Through experiments and property analyses, we show our model's competence over a variety of criteria compared with existing music generation models. Lastly, to promote the open-source culture, we provide a collection of open-source libraries with the hope of facilitating future work in the field. We open-source the following: Codes: https://github.com/archinetai/audio-diffusion-pytorch; music samples for this paper: http://bit.ly/44ozWDH; all music samples for all models: https://bit.ly/audio-diffusion.
Forward citations
Cited by 10 Pith papers
-
Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model
Diff-Symbo generates long, text-controlled symbolic music by autoregressively extending 8-bar latent diffusion segments conditioned on the previous segment's latent.
-
Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment
Beat-guided contrastive alignment of pretrained MotionBERT/MERT features plus ControlNet conditioning of AudioLDM improves dance–music alignment on AIST++ over a MusicGen textual-inversion baseline while remaining com...
-
Next Tokens Denoising for Speech Synthesis
Dragon-FM generates speech autoregressively over two-second chunks while using flow matching inside each chunk, achieving fast synthesis at 12.5 discrete audio tokens per second.
-
MusGO: A Community-Driven Framework For Assessing Openness in Music-Generative AI
MusGO is a community-refined framework with 13 openness categories, applied to 16 music-generative models to produce a public openness leaderboard.
-
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
OSSL is the first self-hosted, mood-annotated video-music dataset, and a video adapter on MusicGen-Medium improves film music generation over text-only baselines.
-
A Mixture-Based Framework for Guiding Diffusion Models
MGDM approximates the intractable guided-diffusion posterior with a weighted mixture of likelihood approximations and samples the mixture using a Gibbs sampler with tunable repetitions.
-
CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio
CoDiCodec unifies continuous and discrete audio compression in one consistency-trained autoencoder, using FSQ-dropout to serve both continuous ~11 Hz embeddings and 2.38 kbps discrete tokens.
-
Workflow-Based Evaluation of Music Generation Systems
A single-producer workflow evaluation of eight music AI tools finds they work as idea and sound generators but not as complete composers, and proposes a reusable framework.
-
FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration
FlowSonic combines deterministic rectified-flow inversion, cached cross-attention injection, and a 'seeded' third-order Adams-Bashforth solver to report better timbre and genre edits on small datasets.
-
ASAudio: A Survey of Advanced Spatial Audio Research
A comprehensive survey that systematically categorizes spatial audio research by representation, task, dataset, and evaluation.
Discussion (0). Continue with ORCID to comment.