Pith. sign in

REVIEW 2 cited by

Codified audio language modeling learns useful representations for music information retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.05677 v1 pith:GAKD27IB submitted 2021-07-12 cs.SD cs.IRcs.LGcs.MMeess.AS

classification cs.SDcs.IRcs.LGcs.MMeess.AS
keywords representationsaudiojukeboxcodifiedlanguagemodelsmodelingmusic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We demonstrate that language models pre-trained on codified (discretely-encoded) music audio learn representations that are useful for downstream MIR tasks. Specifically, we explore representations from Jukebox (Dhariwal et al. 2020): a music generation system containing a language model trained on codified audio from 1M songs. To determine if Jukebox's representations contain useful information for MIR, we use them as input features to train shallow models on several MIR tasks. Relative to representations from conventional MIR models which are pre-trained on tagging, we find that using representations from Jukebox as input features yields 30% stronger performance on average across four MIR tasks: tagging, genre classification, emotion recognition, and key detection. For key detection, we observe that representations from Jukebox are considerably stronger than those from models pre-trained on tagging, suggesting that pre-training via codified audio language modeling may address blind spots in conventional approaches. We interpret the strength of Jukebox's representations as evidence that modeling audio instead of tags provides richer representations for MIR.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following

    eess.AS 2025-06 conditional novelty 6.0 of 10

    CMI-Bench converts standard MIR annotations into instruction-following tasks and shows current audio-text LLMs underperform supervised MIR systems across nearly all 14 tasks.

  2. Segment Transformer: AI-Generated Music Detection via Music Structural Analysis

    cs.SD 2025-09 conditional novelty 4.0 of 10

    A two-stage transformer framework classifies AI-generated music from short clips and beat-segmented full tracks, reporting 99.9% accuracy on SONICS without releasing code or ablations.

Pith tools