Pith. sign in

REVIEW 10 cited by

Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.11757 v3 pith:U6L7SYBO submitted 2023-01-27 cs.CL cs.LGcs.SDeess.AS

classification cs.CLcs.LGcs.SDeess.AS
keywords musicmodeltextgenerationmodelsopen-sourcediffusionhttps
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent years have seen the rapid development of large generative models for text; however, much less research has explored the connection between text and another "language" of communication -- music. Music, much like text, can convey emotions, stories, and ideas, and has its own unique structure and syntax. In our work, we bridge text and music via a text-to-music generation model that is highly efficient, expressive, and can handle long-term structure. Specifically, we develop Mo\^usai, a cascading two-stage latent diffusion model that can generate multiple minutes of high-quality stereo music at 48kHz from textual descriptions. Moreover, our model features high efficiency, which enables real-time inference on a single consumer GPU with a reasonable speed. Through experiments and property analyses, we show our model's competence over a variety of criteria compared with existing music generation models. Lastly, to promote the open-source culture, we provide a collection of open-source libraries with the hope of facilitating future work in the field. We open-source the following: Codes: https://github.com/archinetai/audio-diffusion-pytorch; music samples for this paper: http://bit.ly/44ozWDH; all music samples for all models: https://bit.ly/audio-diffusion.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model

    cs.SD 2026-08 conditional novelty 6.0 of 10

    Diff-Symbo generates long, text-controlled symbolic music by autoregressively extending 8-bar latent diffusion segments conditioned on the previous segment's latent.

  2. Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Beat-guided contrastive alignment of pretrained MotionBERT/MERT features plus ControlNet conditioning of AudioLDM improves dance–music alignment on AIST++ over a MusicGen textual-inversion baseline while remaining com...

  3. Next Tokens Denoising for Speech Synthesis

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Dragon-FM generates speech autoregressively over two-second chunks while using flow matching inside each chunk, achieving fast synthesis at 12.5 discrete audio tokens per second.

  4. MusGO: A Community-Driven Framework For Assessing Openness in Music-Generative AI

    cs.SD 2025-07 conditional novelty 6.0 of 10

    MusGO is a community-refined framework with 13 openness categories, applied to 16 music-generative models to produce a public openness leaderboard.

  5. Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

    cs.SD 2025-06 conditional novelty 6.0 of 10

    OSSL is the first self-hosted, mood-annotated video-music dataset, and a video adapter on MusicGen-Medium improves film music generation over text-only baselines.

  6. A Mixture-Based Framework for Guiding Diffusion Models

    stat.ML 2025-02 conditional novelty 6.0 of 10

    MGDM approximates the intractable guided-diffusion posterior with a weighted mixture of likelihood approximations and samples the mixture using a Gibbs sampler with tunable repetitions.

  7. CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio

    cs.SD 2025-09 conditional novelty 5.0 of 10

    CoDiCodec unifies continuous and discrete audio compression in one consistency-trained autoencoder, using FSQ-dropout to serve both continuous ~11 Hz embeddings and 2.38 kbps discrete tokens.

  8. Workflow-Based Evaluation of Music Generation Systems

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A single-producer workflow evaluation of eight music AI tools finds they work as idea and sound generators but not as complete composers, and proposes a reusable framework.

  9. FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration

    cs.SD 2026-07 reject novelty 4.0 of 10

    FlowSonic combines deterministic rectified-flow inversion, cached cross-attention injection, and a 'seeded' third-order Adams-Bashforth solver to report better timbre and genre edits on small datasets.

  10. ASAudio: A Survey of Advanced Spatial Audio Research

    eess.AS 2025-08 unverdicted novelty 3.0 of 10

    A comprehensive survey that systematically categorizes spatial audio research by representation, task, dataset, and evaluation.

Pith tools