Pith. sign in

REVIEW 3 cited by

MT3: Multi-Task Multitrack Music Transcription

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.03017 v4 pith:DPWIHBYX submitted 2021-11-04 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords instrumentstranscriptionmusicdatasetslow-resourcemulti-taskacrossautomatic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automatic Music Transcription (AMT), inferring musical notes from raw audio, is a challenging task at the core of music understanding. Unlike Automatic Speech Recognition (ASR), which typically focuses on the words of a single speaker, AMT often requires transcribing multiple instruments simultaneously, all while preserving fine-scale pitch and timing information. Further, many AMT datasets are "low-resource", as even expert musicians find music transcription difficult and time-consuming. Thus, prior work has focused on task-specific architectures, tailored to the individual instruments of each task. In this work, motivated by the promising results of sequence-to-sequence transfer learning for low-resource Natural Language Processing (NLP), we demonstrate that a general-purpose Transformer model can perform multi-task AMT, jointly transcribing arbitrary combinations of musical instruments across several transcription datasets. We show this unified training framework achieves high-quality transcription results across a range of datasets, dramatically improving performance for low-resource instruments (such as guitar), while preserving strong performance for abundant instruments (such as piano). Finally, by expanding the scope of AMT, we expose the need for more consistent evaluation metrics and better dataset alignment, and provide a strong baseline for this new direction of multi-task AMT.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 4 citations worldwide. Full citation record

  1. MulTTiPop: A Multitrack Transcription Dataset for Pop Music

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A new 572-segment dataset pairs commercial pop audio with multitrack MIDI, revealing state-of-the-art transcription models achieve only 38% Onset F1.

  2. VioPTT: Violin Technique-Aware Transcription from Synthetic Data Augmentation

    cs.SD 2025-09 conditional novelty 6.0 of 10

    A cascade model transcribes violin pitch, onset, offset, and playing technique, trained on 76 hours of synthetic VST-rendered audio.

  3. PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music

    cs.SD 2025-09 conditional novelty 4.0 of 10

    PianoBind, a trimodal audio-MIDI-text embedding model trained on piano data, beats general-purpose music embedding models on pop-piano text-to-music retrieval benchmarks.

Pith tools