Pith. sign in

REVIEW 5 cited by

Contrastive Learning of Musical Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.09410 v2 pith:CFSNBJ25 submitted 2021-03-17 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords musiclearningrepresentationsdatasetdatasetsclmrlabeledmagnatagatune
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While deep learning has enabled great advances in many areas of music, labeled music datasets remain especially hard, expensive, and time-consuming to create. In this work, we introduce SimCLR to the music domain and contribute a large chain of audio data augmentations to form a simple framework for self-supervised, contrastive learning of musical representations: CLMR. This approach works on raw time-domain music data and requires no labels to learn useful representations. We evaluate CLMR in the downstream task of music classification on the MagnaTagATune and Million Song datasets and present an ablation study to test which of our music-related innovations over SimCLR are most effective. A linear classifier trained on the proposed representations achieves a higher average precision than supervised models on the MagnaTagATune dataset, and performs comparably on the Million Song dataset. Moreover, we show that CLMR's representations are transferable using out-of-domain datasets, indicating that our method has strong generalisability in music classification. Lastly, we show that the proposed method allows data-efficient learning on smaller labeled datasets: we achieve an average precision of 33.1% despite using only 259 labeled songs in the MagnaTagATune dataset (1% of the full dataset) during linear evaluation. To foster reproducibility and future research on self-supervised learning in music, we publicly release the pre-trained models and the source code of all experiments of this paper.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems

    cs.SD 2026-07 conditional novelty 6.0 of 10

    StemFX predicts tokenized per-stem audio-effect chains with a jointly-trained Transformer encoder-decoder, beating contrastive and prior FX-encoding methods on effect-chain retrieval and real-mix style transfer.

  2. Scaling Self-Supervised Representation Learning for Symbolic Piano Performance

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Self-supervised pretraining on 60,000 hours of symbolic piano music produces a generative model and contrastive embeddings that beat leading baselines on continuation quality and several MIR classification benchmarks.

  3. Bridging the Gap Between Semantic and User Preference Spaces for Multi-modal Music Representation Learning

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A two-stage contrastive training method aligns song audio with text semantics and then with user-favored song pairs, improving music classification and recommendation over prior models.

  4. Yambda-5B -- A Large-Scale Multi-modal Dataset for Ranking And Retrieval

    cs.IR 2025-05 conditional novelty 6.0 of 10

    A new open 4.79B-interaction music dataset from Yandex Music with an is_organic flag, audio embeddings, and a Global Temporal Split benchmark protocol.

  5. Multi-Distillation from Speech and Music Representation Models

    eess.AS 2025-06 conditional novelty 4.0 of 10

    A 23M-parameter student distilled from HuBERT/WavLM and MERT gets close to teacher-level average accuracy on speech and music benchmarks and outperforms its teachers in few-shot classification.

Pith tools