REVIEW 20 cited by
MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Self-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech. Although SSL has been proven effective in speech and audio, its application to music audio has yet to be thoroughly explored. This is partially due to the distinctive challenges associated with modelling musical knowledge, particularly tonal and pitched characteristics of music. To address this research gap, we propose an acoustic Music undERstanding model with large-scale self-supervised Training (MERT), which incorporates teacher models to provide pseudo labels in the masked language modelling (MLM) style acoustic pre-training. In our exploration, we identified an effective combination of teacher models, which outperforms conventional speech and audio approaches in terms of performance. This combination includes an acoustic teacher based on Residual Vector Quantisation - Variational AutoEncoder (RVQ-VAE) and a musical teacher based on the Constant-Q Transform (CQT). Furthermore, we explore a wide range of settings to overcome the instability in acoustic language model pre-training, which allows our designed paradigm to scale from 95M to 330M parameters. Experimental results indicate that our model can generalise and perform well on 14 music understanding tasks and attain state-of-the-art (SOTA) overall scores.
Forward citations
Cited by 20 Pith papers
-
MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing
MixAssist is the first audio-grounded, multi-turn conversational dataset for co-creative music mixing instruction, and fine-tuning Qwen-Audio on it yields human-comparable mixing advice.
-
Do Music Foundation Models Embed Pitch in Helical Structure?
Intermediate layers of Jukebox and MusicGen represent pitch as a conical helix whose clarity depends on octave-equivalent harmonics in the input.
-
StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems
StemFX predicts tokenized per-stem audio-effect chains with a jointly-trained Transformer encoder-decoder, beating contrastive and prior FX-encoding methods on effect-chain retrieval and real-mix style transfer.
-
Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment
Beat-guided contrastive alignment of pretrained MotionBERT/MERT features plus ControlNet conditioning of AudioLDM improves dance–music alignment on AIST++ over a MusicGen textual-inversion baseline while remaining com...
-
MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations
MADB is a 9,999-track music aesthetics benchmark with multi-dimensional professional annotations revealing that current pretrained audio models capture only partial aesthetic information.
-
Assessing Factual Music Comprehension in Large Audio Language Models
Standard NLP metrics fail to capture factual music understanding in audio-language models; a CLAP-based metric and an LLM-parsed factual QA protocol measure it more directly.
-
UniVerse-1: Unified Audio-Video Generation via Stitching of Experts
A unified audio-video generator built by stitching pre-trained video and music diffusion models, trained on 7,600 hours of data, with a new evaluation benchmark.
-
Training a Perceptual Model for Evaluating Auditory Similarity in Music Adversarial Attack
PAMT, a psychoacoustically conditioned transformer on frozen MERT embeddings, matches human similarity judgments better than prior metrics and improves music-model robustness to perceptually constrained adversarial attacks.
-
Scaling Self-Supervised Representation Learning for Symbolic Piano Performance
Self-supervised pretraining on 60,000 hours of symbolic piano music produces a generative model and contrastive embeddings that beat leading baselines on continuation quality and several MIR classification benchmarks.
-
Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction
Speaker-recognition pre-trained models (x-vector, ECAPA) outperform other speech and music models for singing voice MOS prediction, and their fusion via a Bhattacharyya-distance loss sets a new reported state of the a...
-
Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis
Off-the-shelf deep audio embeddings generally improve unsupervised music boundary detection over spectrogram features, with CBM the best segmenter, but not all models help and standard scores are inflated by edge boundaries.
-
Affect-aware Cross-Domain Recommendation for Art Therapy via Music Preference Elicitation
A 200-person study of music-driven cross-domain recommendation for art therapy shows music-based and visual-based engines perform equally, contradicting the paper's 'outperforming' claim.
-
MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation
Fine-tuning MU-LLaMA on 3,371 pseudo-labeled video-music examples lets it produce scene captions that yield small subjective gains in video background music generation over music-only captions.
-
OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction
OMAR-RQ, an open 580M-parameter music audio model trained with multi-codebook, multi-feature masked token prediction on 330k hours, reports leading open-model results on several MIR benchmarks.
-
DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding
DEL is a new audio-visual transformer framework that reports state-of-the-art temporal action localization on UnAV-100, THUMOS14, ActivityNet 1.3, and EPIC-Kitchens-100.
-
Workflow-Based Evaluation of Music Generation Systems
A single-producer workflow evaluation of eight music AI tools finds they work as idea and sound generators but not as complete composers, and proposes a reusable framework.
-
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
Multimodal foundation models outperform speech and music models for closed-set source attribution of singing voice deepfakes on CtrSVDD, with Chernoff-distance fusion of LanguageBind and ImageBind reaching 91.2% accuracy.
-
Towards Unified Music Emotion Recognition across Dimensional and Categorical Models
A multitask learning plus knowledge distillation framework with MERT, chord, and key features unifies categorical and dimensional music emotion labels and reports improved MTG-Jamendo performance over listed baselines.
-
Enhancing In-Domain and Out-Domain EmoFake Detection via Cooperative Multilingual Speech Foundation Models
Multilingual speech foundation models, fused with a Tucker-Hadamard module, are reported to reach about 1% equal error rate on emotion fake audio detection, a large drop from prior benchmarks.
-
Semantic-Aware Interpretable Multimodal Music Auto-Tagging
A music auto-tagging system using grouped perceptual audio and lyric features with EM-BANDED gives competitive results on MTG-Jamendo and group-level explanations that loosely match amateur listeners.
Discussion (0). Continue with ORCID to comment.