Pith. sign in

REVIEW 5 cited by

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.06660 v1 pith:UD2LZBYG submitted 2024-12-09 cs.SD cs.MMeess.AS

classification cs.SDcs.MMeess.AS
keywords musicmulti-modalgenerationimagesmodelsmumu-llamaunderstandingvideos
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Research on large language models has advanced significantly across text, speech, images, and videos. However, multi-modal music understanding and generation remain underexplored due to the lack of well-annotated datasets. To address this, we introduce a dataset with 167.69 hours of multi-modal data, including text, images, videos, and music annotations. Based on this dataset, we propose MuMu-LLaMA, a model that leverages pre-trained encoders for music, images, and videos. For music generation, we integrate AudioLDM 2 and MusicGen. Our evaluation across four tasks--music understanding, text-to-music generation, prompt-based music editing, and multi-modal music generation--demonstrates that MuMu-LLaMA outperforms state-of-the-art models, showing its potential for multi-modal music applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A new audio benchmark, TORUS, shows unified audio models are not self-coherent: the best model answers 50.5% of questions about its own generations, below a 63.2% cascaded specialist baseline.

  2. Assessing Factual Music Comprehension in Large Audio Language Models

    cs.SD 2025-11 conditional novelty 6.0 of 10

    Standard NLP metrics fail to capture factual music understanding in audio-language models; a CLAP-based metric and an LLM-parsed factual QA protocol measure it more directly.

  3. Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

    cs.SD 2025-06 conditional novelty 6.0 of 10

    OSSL is the first self-hosted, mood-annotated video-music dataset, and a video adapter on MusicGen-Medium improves film music generation over text-only baselines.

  4. WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation

    cs.SD 2025-09 reject novelty 4.0 of 10

    An open multi-agent system that orchestrates specialized music models for understanding, composition, and synthesis, with local or hosted deployment.

  5. CoComposer: LLM Multi-agent Collaborative Music Composition

    cs.SD 2025-08 conditional novelty 4.0 of 10

    A five-agent LLM system for ABC-notation composition scores modestly higher than ComposerX and a single LLM on an automated aesthetic model, but no error bars or significance tests are reported.

Pith tools