Pith. sign in

REVIEW 20 cited by

MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.00107 v5 pith:H5DLRUIQ submitted 2023-05-31 cs.SD cs.AIcs.CLcs.LGeess.AS

classification cs.SDcs.AIcs.CLcs.LGeess.AS
keywords acousticmusicmodelteacheraudiolarge-scalemodelsself-supervised
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech. Although SSL has been proven effective in speech and audio, its application to music audio has yet to be thoroughly explored. This is partially due to the distinctive challenges associated with modelling musical knowledge, particularly tonal and pitched characteristics of music. To address this research gap, we propose an acoustic Music undERstanding model with large-scale self-supervised Training (MERT), which incorporates teacher models to provide pseudo labels in the masked language modelling (MLM) style acoustic pre-training. In our exploration, we identified an effective combination of teacher models, which outperforms conventional speech and audio approaches in terms of performance. This combination includes an acoustic teacher based on Residual Vector Quantisation - Variational AutoEncoder (RVQ-VAE) and a musical teacher based on the Constant-Q Transform (CQT). Furthermore, we explore a wide range of settings to overcome the instability in acoustic language model pre-training, which allows our designed paradigm to scale from 95M to 330M parameters. Experimental results indicate that our model can generalise and perform well on 14 music understanding tasks and attain state-of-the-art (SOTA) overall scores.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 27 citations worldwide. Full citation record

  1. MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing

    cs.SD 2025-07 conditional novelty 7.0 of 10

    MixAssist is the first audio-grounded, multi-turn conversational dataset for co-creative music mixing instruction, and fine-tuning Qwen-Audio on it yields human-comparable mixing advice.

  2. Do Music Foundation Models Embed Pitch in Helical Structure?

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Intermediate layers of Jukebox and MusicGen represent pitch as a conical helix whose clarity depends on octave-equivalent harmonics in the input.

  3. StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems

    cs.SD 2026-07 conditional novelty 6.0 of 10

    StemFX predicts tokenized per-stem audio-effect chains with a jointly-trained Transformer encoder-decoder, beating contrastive and prior FX-encoding methods on effect-chain retrieval and real-mix style transfer.

  4. Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Beat-guided contrastive alignment of pretrained MotionBERT/MERT features plus ControlNet conditioning of AudioLDM improves dance–music alignment on AIST++ over a MusicGen textual-inversion baseline while remaining com...

  5. MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations

    cs.SD 2026-07 accept novelty 6.0 of 10

    MADB is a 9,999-track music aesthetics benchmark with multi-dimensional professional annotations revealing that current pretrained audio models capture only partial aesthetic information.

  6. Assessing Factual Music Comprehension in Large Audio Language Models

    cs.SD 2025-11 conditional novelty 6.0 of 10

    Standard NLP metrics fail to capture factual music understanding in audio-language models; a CLAP-based metric and an LLM-parsed factual QA protocol measure it more directly.

  7. UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A unified audio-video generator built by stitching pre-trained video and music diffusion models, trained on 7,600 hours of data, with a new evaluation benchmark.

  8. Training a Perceptual Model for Evaluating Auditory Similarity in Music Adversarial Attack

    cs.SD 2025-09 conditional novelty 6.0 of 10

    PAMT, a psychoacoustically conditioned transformer on frozen MERT embeddings, matches human similarity judgments better than prior metrics and improves music-model robustness to perceptually constrained adversarial attacks.

  9. Scaling Self-Supervised Representation Learning for Symbolic Piano Performance

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Self-supervised pretraining on 60,000 hours of symbolic piano music produces a generative model and contrastive embeddings that beat leading baselines on continuation quality and several MIR classification benchmarks.

  10. Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction

    eess.AS 2025-06 reject novelty 6.0 of 10

    Speaker-recognition pre-trained models (x-vector, ECAPA) outperform other speech and music models for singing voice MOS prediction, and their fusion via a Bhattacharyya-distance loss sets a new reported state of the a...

  11. Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis

    cs.SD 2026-03 conditional novelty 5.0 of 10

    Off-the-shelf deep audio embeddings generally improve unsupervised music boundary detection over spectrogram features, with CBM the best segmenter, but not all models help and standard scores are inflated by edge boundaries.

  12. Affect-aware Cross-Domain Recommendation for Art Therapy via Music Preference Elicitation

    cs.IR 2025-07 reject novelty 5.0 of 10

    A 200-person study of music-driven cross-domain recommendation for art therapy shows music-based and visual-based engines perform equally, contradicting the paper's 'outperforming' claim.

  13. MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation

    cs.AI 2025-07 reject novelty 5.0 of 10

    Fine-tuning MU-LLaMA on 3,371 pseudo-labeled video-music examples lets it produce scene captions that yield small subjective gains in video background music generation over music-only captions.

  14. OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction

    cs.SD 2025-07 conditional novelty 5.0 of 10

    OMAR-RQ, an open 580M-parameter music audio model trained with multi-codebook, multi-feature masked token prediction on 330k hours, reports leading open-model results on several MIR benchmarks.

  15. DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    DEL is a new audio-visual transformer framework that reports state-of-the-art temporal action localization on UnAV-100, THUMOS14, ActivityNet 1.3, and EPIC-Kitchens-100.

  16. Workflow-Based Evaluation of Music Generation Systems

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A single-producer workflow evaluation of eight music AI tools finds they work as idea and sound generators but not as complete composers, and proposes a reusable framework.

  17. Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models

    eess.AS 2025-06 conditional novelty 5.0 of 10

    Multimodal foundation models outperform speech and music models for closed-set source attribution of singing voice deepfakes on CtrSVDD, with Chernoff-distance fusion of LanguageBind and ImageBind reaching 91.2% accuracy.

  18. Towards Unified Music Emotion Recognition across Dimensional and Categorical Models

    cs.SD 2025-02 conditional novelty 5.0 of 10

    A multitask learning plus knowledge distillation framework with MERT, chord, and key features unifies categorical and dimensional music emotion labels and reports improved MTG-Jamendo performance over listed baselines.

  19. Enhancing In-Domain and Out-Domain EmoFake Detection via Cooperative Multilingual Speech Foundation Models

    eess.AS 2025-07 conditional novelty 4.0 of 10

    Multilingual speech foundation models, fused with a Tucker-Hadamard module, are reported to reach about 1% equal error rate on emotion fake audio detection, a large drop from prior benchmarks.

  20. Semantic-Aware Interpretable Multimodal Music Auto-Tagging

    cs.LG 2025-05 conditional novelty 4.0 of 10

    A music auto-tagging system using grouped perceptual audio and lyric features with EM-BANDED gives competitive results on MTG-Jamendo and group-level explanations that loosely match amateur listeners.

Pith tools