Pith. sign in

REVIEW 4 cited by

The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.10057 v3 pith:CCQNRIXX submitted 2023-11-16 cs.SD cs.AIcs.CLeess.AS

classification cs.SDcs.AIcs.CLeess.AS
keywords datasetevaluationmusic-and-languagecorpusdescribermodelsmusicsong
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce the Song Describer dataset (SDD), a new crowdsourced corpus of high-quality audio-caption pairs, designed for the evaluation of music-and-language models. The dataset consists of 1.1k human-written natural language descriptions of 706 music recordings, all publicly accessible and released under Creative Common licenses. To showcase the use of our dataset, we benchmark popular models on three key music-and-language tasks (music captioning, text-to-music generation and music-language retrieval). Our experiments highlight the importance of cross-dataset evaluation and offer insights into how researchers can use SDD to gain a broader understanding of model performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following

    eess.AS 2025-06 conditional novelty 6.0 of 10

    CMI-Bench converts standard MIR annotations into instruction-following tasks and shows current audio-text LLMs underperform supervised MIR systems across nearly all 14 tasks.

  2. Qwen-Audio-3.0-Gen-Preview Technical Report

    eess.AS 2026-07 conditional novelty 5.0 of 10

    One diffusion-transformer system with a shared audio codec generates standalone audio, multi-speaker dialogue, and time-structured mixed scenes, with strongest measured advantages in speaker similarity, cross-turn con...

  3. CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages

    cs.SD 2025-02 conditional novelty 5.0 of 10

    A contrastive learning framework (CLaMP 3) aligns three music modalities with multilingual text, enabling text-to-music retrieval, cross-lingual retrieval for unseen languages, and emergent cross-modal retrieval.

  4. JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata

    cs.SD 2025-02 conditional novelty 5.0 of 10

    A dataset of 362,000 Jamendo instrumental tracks pairs each song with a generated caption and imputed metadata fields, created with a retrieval-based local-LLM pipeline.

Pith tools