Pith. sign in

REVIEW 4 cited by

Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.11276 v1 pith:K6W7MAQY submitted 2023-08-22 cs.SD cs.AIcs.CLcs.MMeess.AS

classification cs.SDcs.AIcs.CLcs.MMeess.AS
keywords musicansweringmodelquestionaudiodatasetdatasetsgeneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-music generation (T2M-Gen) faces a major obstacle due to the scarcity of large-scale publicly available music datasets with natural language captions. To address this, we propose the Music Understanding LLaMA (MU-LLaMA), capable of answering music-related questions and generating captions for music files. Our model utilizes audio representations from a pretrained MERT model to extract music features. However, obtaining a suitable dataset for training the MU-LLaMA model remains challenging, as existing publicly accessible audio question answering datasets lack the necessary depth for open-ended music question answering. To fill this gap, we present a methodology for generating question-answer pairs from existing audio captioning datasets and introduce the MusicQA Dataset designed for answering open-ended music-related questions. The experiments demonstrate that the proposed MU-LLaMA model, trained on our designed MusicQA dataset, achieves outstanding performance in both music question answering and music caption generation across various metrics, outperforming current state-of-the-art (SOTA) models in both fields and offering a promising advancement in the T2M-Gen research field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Assessing Factual Music Comprehension in Large Audio Language Models

    cs.SD 2025-11 conditional novelty 6.0 of 10

    Standard NLP metrics fail to capture factual music understanding in audio-language models; a CLAP-based metric and an LLM-parsed factual QA protocol measure it more directly.

  2. MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation

    cs.AI 2025-07 reject novelty 5.0 of 10

    Fine-tuning MU-LLaMA on 3,371 pseudo-labeled video-music examples lets it produce scene captions that yield small subjective gains in video background music generation over music-only captions.

  3. Selective Invocation for Multilingual ASR: A Cost-effective Approach Adapting to Speech Recognition Difficulty

    cs.SD 2025-05 conditional novelty 5.0 of 10

    A spoken large language model learns to judge its own transcription difficulty and routes only hard speech to a stronger ASR model, cutting cost and improving word error rate.

  4. The TEA-ASLP System for Multilingual Conversational Speech Recognition and Speech Diarization in MLC-SLM 2025 Challenge

    cs.SD 2025-07 conditional novelty 4.0 of 10

    Combining dual encoders, LID-routed MoE LoRA, and CTC prompts yields top challenge results for multilingual conversational ASR and speech diarization.

Pith tools