REVIEW 5 cited by
MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Research on large language models has advanced significantly across text, speech, images, and videos. However, multi-modal music understanding and generation remain underexplored due to the lack of well-annotated datasets. To address this, we introduce a dataset with 167.69 hours of multi-modal data, including text, images, videos, and music annotations. Based on this dataset, we propose MuMu-LLaMA, a model that leverages pre-trained encoders for music, images, and videos. For music generation, we integrate AudioLDM 2 and MusicGen. Our evaluation across four tasks--music understanding, text-to-music generation, prompt-based music editing, and multi-modal music generation--demonstrates that MuMu-LLaMA outperforms state-of-the-art models, showing its potential for multi-modal music applications.
Forward citations
Cited by 5 Pith papers
-
TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models
A new audio benchmark, TORUS, shows unified audio models are not self-coherent: the best model answers 50.5% of questions about its own generations, below a 63.2% cascaded specialist baseline.
-
Assessing Factual Music Comprehension in Large Audio Language Models
Standard NLP metrics fail to capture factual music understanding in audio-language models; a CLAP-based metric and an LLM-parsed factual QA protocol measure it more directly.
-
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
OSSL is the first self-hosted, mood-annotated video-music dataset, and a video adapter on MusicGen-Medium improves film music generation over text-only baselines.
-
WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation
An open multi-agent system that orchestrates specialized music models for understanding, composition, and synthesis, with local or hosted deployment.
-
CoComposer: LLM Multi-agent Collaborative Music Composition
A five-agent LLM system for ABC-notation composition scores modestly higher than ComposerX and a single LLM on an automated aesthetic model, but no error bars or significance tests are reported.
Discussion (0). Continue with ORCID to comment.