REVIEW 12 cited by
Foundation Models for Music: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In recent years, foundation models (FMs) such as large language models (LLMs) and latent diffusion models (LDMs) have profoundly impacted diverse sectors, including music. This comprehensive review examines state-of-the-art (SOTA) pre-trained models and foundation models in music, spanning from representation learning, generative learning and multimodal learning. We first contextualise the significance of music in various industries and trace the evolution of AI in music. By delineating the modalities targeted by foundation models, we discover many of the music representations are underexplored in FM development. Then, emphasis is placed on the lack of versatility of previous methods on diverse music applications, along with the potential of FMs in music understanding, generation and medical application. By comprehensively exploring the details of the model pre-training paradigm, architectural choices, tokenisation, finetuning methodologies and controllability, we emphasise the important topics that should have been well explored, like instruction tuning and in-context learning, scaling law and emergent ability, as well as long-sequence modelling etc. A dedicated section presents insights into music agents, accompanied by a thorough analysis of datasets and evaluations essential for pre-training and downstream tasks. Finally, by underscoring the vital importance of ethical considerations, we advocate that following research on FM for music should focus more on such issues as interpretability, transparency, human responsibility, and copyright issues. The paper offers insights into future challenges and trends on FMs for music, aiming to shape the trajectory of human-AI collaboration in the music realm.
Forward citations
Cited by 12 Pith papers
-
Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset
The paper builds the first popular-music audio-to-score dataset and a model that improves classical A2S error rates by 67.5% while establishing a 20.92% benchmark on the new data.
-
RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling
RPPNet generates melodies by planning variable-length perceptually grouped rhythm-pitch primitives first and then decoding them into notes, beating bar-level baselines in subjective structure and musicality ratings.
-
Unified Audio Intelligence Without Regressing on Text Intelligence
A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.
-
Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets
A new benchmark using expert musicological annotations shows that no single music AI encoder wins across all tagging tasks.
-
LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model
LLaSO releases a 3.8B speech-language model, 25.5M training instances, and an evaluation benchmark, claiming a normalized score of 0.72.
-
Emergent musical properties of a transformer under contrastive self-supervised learning
Sequence tokens of a class-token-only contrastively trained ViT-1D carry usable beat, chord, and onset information that emerges during training without any frame-level supervision.
-
Universal Music Representations? Evaluating Foundation Models on World Music Corpora
Five audio foundation models are evaluated across six Western and non-Western music corpora, showing a consistent Western-centric bias and only limited generalization to culturally distant traditions.
-
Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models
Diff-TONE selects the prompt-swap timestep using a distilled latent instrument classifier, improving content preservation over fixed timestep baselines in text-to-music diffusion editing.
-
CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
CMI-Bench converts standard MIR annotations into instruction-following tasks and shows current audio-text LLMs underperform supervised MIR systems across nearly all 14 tasks.
-
Temporal Adaptation of Pre-trained Foundation Models for Music Structure Analysis
Temporal adaptation, extending the audio window and downsampling the time axis during fine-tuning, improves music structure labeling accuracy on Harmonix and RWC-Pop.
-
Workflow-Based Evaluation of Music Generation Systems
A single-producer workflow evaluation of eight music AI tools finds they work as idea and sound generators but not as complete composers, and proposes a reusable framework.
-
Semantic-Aware Interpretable Multimodal Music Auto-Tagging
A music auto-tagging system using grouped perceptual audio and lyric features with EM-BANDED gives competitive results on MTG-Jamendo and group-level explanations that loosely match amateur listeners.
Discussion (0). Continue with ORCID to comment.