REVIEW 5 cited by
MAP-Music2Vec: A Simple and Effective Baseline for Self-Supervised Music Audio Representation Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The deep learning community has witnessed an exponentially growing interest in self-supervised learning (SSL). However, it still remains unexplored how to build a framework for learning useful representations of raw music waveforms in a self-supervised manner. In this work, we design Music2Vec, a framework exploring different SSL algorithmic components and tricks for music audio recordings. Our model achieves comparable results to the state-of-the-art (SOTA) music SSL model Jukebox, despite being significantly smaller with less than 2% of parameters of the latter. The model will be released on Huggingface(Please refer to: https://huggingface.co/m-a-p/music2vec-v1)
Forward citations
Cited by 5 Pith papers
-
Universal Music Representations? Evaluating Foundation Models on World Music Corpora
Five audio foundation models are evaluated across six Western and non-Western music corpora, showing a consistent Western-centric bias and only limited generalization to culturally distant traditions.
-
Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction
Speaker-recognition pre-trained models (x-vector, ECAPA) outperform other speech and music models for singing voice MOS prediction, and their fusion via a Bhattacharyya-distance loss sets a new reported state of the a...
-
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
Multimodal foundation models outperform speech and music models for closed-set source attribution of singing voice deepfakes on CtrSVDD, with Chernoff-distance fusion of LanguageBind and ImageBind reaching 91.2% accuracy.
-
BeatFM: Improving Beat Tracking with Pre-trained Music Foundation Model
The abstract claims BeatFM achieves state-of-the-art beat tracking, but the body describes a different model, HingeNet, so the BeatFM claim is unsupported.
-
Enhancing In-Domain and Out-Domain EmoFake Detection via Cooperative Multilingual Speech Foundation Models
Multilingual speech foundation models, fused with a Tucker-Hadamard module, are reported to reach about 1% equal error rate on emotion fake audio detection, a large drop from prior benchmarks.
Discussion (0). Continue with ORCID to comment.