Pith. sign in

REVIEW 12 cited by

Foundation Models for Music: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.14340 v3 pith:DMWAKR3B submitted 2024-08-26 cs.SD cs.AIcs.CLcs.LGeess.AS

classification cs.SDcs.AIcs.CLcs.LGeess.AS
keywords musicmodelsfoundationlearningdiverseinsightsissuespre-training
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, foundation models (FMs) such as large language models (LLMs) and latent diffusion models (LDMs) have profoundly impacted diverse sectors, including music. This comprehensive review examines state-of-the-art (SOTA) pre-trained models and foundation models in music, spanning from representation learning, generative learning and multimodal learning. We first contextualise the significance of music in various industries and trace the evolution of AI in music. By delineating the modalities targeted by foundation models, we discover many of the music representations are underexplored in FM development. Then, emphasis is placed on the lack of versatility of previous methods on diverse music applications, along with the potential of FMs in music understanding, generation and medical application. By comprehensively exploring the details of the model pre-training paradigm, architectural choices, tokenisation, finetuning methodologies and controllability, we emphasise the important topics that should have been well explored, like instruction tuning and in-context learning, scaling law and emergent ability, as well as long-sequence modelling etc. A dedicated section presents insights into music agents, accompanied by a thorough analysis of datasets and evaluations essential for pre-training and downstream tasks. Finally, by underscoring the vital importance of ethical considerations, we advocate that following research on FM for music should focus more on such issues as interpretability, transparency, human responsibility, and copyright issues. The paper offers insights into future challenges and trends on FMs for music, aiming to shape the trajectory of human-AI collaboration in the music realm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset

    cs.SD 2026-08 conditional novelty 7.0 of 10

    The paper builds the first popular-music audio-to-score dataset and a model that improves classical A2S error rates by 67.5% while establishing a 20.92% benchmark on the new data.

  2. RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling

    cs.SD 2026-07 conditional novelty 6.0 of 10

    RPPNet generates melodies by planning variable-length perceptually grouped rhythm-pitch primitives first and then decoding them into notes, beating bar-level baselines in subjective structure and musicality ratings.

  3. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  4. Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets

    cs.SD 2025-09 conditional novelty 6.0 of 10

    A new benchmark using expert musicological annotations shows that no single music AI encoder wins across all tagging tasks.

  5. LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    LLaSO releases a 3.8B speech-language model, 25.5M training instances, and an evaluation benchmark, claiming a normalized score of 0.72.

  6. Emergent musical properties of a transformer under contrastive self-supervised learning

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Sequence tokens of a class-token-only contrastively trained ViT-1D carry usable beat, chord, and onset information that emerges during training without any frame-level supervision.

  7. Universal Music Representations? Evaluating Foundation Models on World Music Corpora

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Five audio foundation models are evaluated across six Western and non-Western music corpora, showing a consistent Western-centric bias and only limited generalization to culturally distant traditions.

  8. Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Diff-TONE selects the prompt-swap timestep using a distilled latent instrument classifier, improving content preservation over fixed timestep baselines in text-to-music diffusion editing.

  9. CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following

    eess.AS 2025-06 conditional novelty 6.0 of 10

    CMI-Bench converts standard MIR annotations into instruction-following tasks and shows current audio-text LLMs underperform supervised MIR systems across nearly all 14 tasks.

  10. Temporal Adaptation of Pre-trained Foundation Models for Music Structure Analysis

    cs.SD 2025-07 conditional novelty 5.0 of 10

    Temporal adaptation, extending the audio window and downsampling the time axis during fine-tuning, improves music structure labeling accuracy on Harmonix and RWC-Pop.

  11. Workflow-Based Evaluation of Music Generation Systems

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A single-producer workflow evaluation of eight music AI tools finds they work as idea and sound generators but not as complete composers, and proposes a reusable framework.

  12. Semantic-Aware Interpretable Multimodal Music Auto-Tagging

    cs.LG 2025-05 conditional novelty 4.0 of 10

    A music auto-tagging system using grouped perceptual audio and lyric features with EM-BANDED gives competitive results on MTG-Jamendo and group-level explanations that loosely match amateur listeners.

Pith tools