Pith. sign in

REVIEW 5 cited by

MAP-Music2Vec: A Simple and Effective Baseline for Self-Supervised Music Audio Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.02508 v1 pith:Y7V6V5BT submitted 2022-12-05 cs.SD cs.AIcs.LGcs.MMeess.AS

classification cs.SDcs.AIcs.LGcs.MMeess.AS
keywords learningmusicmodelself-supervisedaudioframeworkhuggingfaceachieves
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The deep learning community has witnessed an exponentially growing interest in self-supervised learning (SSL). However, it still remains unexplored how to build a framework for learning useful representations of raw music waveforms in a self-supervised manner. In this work, we design Music2Vec, a framework exploring different SSL algorithmic components and tricks for music audio recordings. Our model achieves comparable results to the state-of-the-art (SOTA) music SSL model Jukebox, despite being significantly smaller with less than 2% of parameters of the latter. The model will be released on Huggingface(Please refer to: https://huggingface.co/m-a-p/music2vec-v1)

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Universal Music Representations? Evaluating Foundation Models on World Music Corpora

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Five audio foundation models are evaluated across six Western and non-Western music corpora, showing a consistent Western-centric bias and only limited generalization to culturally distant traditions.

  2. Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction

    eess.AS 2025-06 reject novelty 6.0 of 10

    Speaker-recognition pre-trained models (x-vector, ECAPA) outperform other speech and music models for singing voice MOS prediction, and their fusion via a Bhattacharyya-distance loss sets a new reported state of the a...

  3. Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models

    eess.AS 2025-06 conditional novelty 5.0 of 10

    Multimodal foundation models outperform speech and music models for closed-set source attribution of singing voice deepfakes on CtrSVDD, with Chernoff-distance fusion of LanguageBind and ImageBind reaching 91.2% accuracy.

  4. BeatFM: Improving Beat Tracking with Pre-trained Music Foundation Model

    cs.SD 2025-08 reject novelty 4.0 of 10

    The abstract claims BeatFM achieves state-of-the-art beat tracking, but the body describes a different model, HingeNet, so the BeatFM claim is unsupported.

  5. Enhancing In-Domain and Out-Domain EmoFake Detection via Cooperative Multilingual Speech Foundation Models

    eess.AS 2025-07 conditional novelty 4.0 of 10

    Multilingual speech foundation models, fused with a Tucker-Hadamard module, are reported to reach about 1% equal error rate on emotion fake audio detection, a large drop from prior benchmarks.

Pith tools