REVIEW 4 cited by
The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce the Song Describer dataset (SDD), a new crowdsourced corpus of high-quality audio-caption pairs, designed for the evaluation of music-and-language models. The dataset consists of 1.1k human-written natural language descriptions of 706 music recordings, all publicly accessible and released under Creative Common licenses. To showcase the use of our dataset, we benchmark popular models on three key music-and-language tasks (music captioning, text-to-music generation and music-language retrieval). Our experiments highlight the importance of cross-dataset evaluation and offer insights into how researchers can use SDD to gain a broader understanding of model performance.
Forward citations
Cited by 4 Pith papers
-
CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
CMI-Bench converts standard MIR annotations into instruction-following tasks and shows current audio-text LLMs underperform supervised MIR systems across nearly all 14 tasks.
-
Qwen-Audio-3.0-Gen-Preview Technical Report
One diffusion-transformer system with a shared audio codec generates standalone audio, multi-speaker dialogue, and time-structured mixed scenes, with strongest measured advantages in speaker similarity, cross-turn con...
-
CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages
A contrastive learning framework (CLaMP 3) aligns three music modalities with multilingual text, enabling text-to-music retrieval, cross-lingual retrieval for unseen languages, and emergent cross-modal retrieval.
-
JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata
A dataset of 362,000 Jamendo instrumental tracks pairs each song with a generated caption and imputed metadata fields, created with a retrieval-based local-LLM pipeline.
Discussion (0). Continue with ORCID to comment.