Pith. sign in

REVIEW 1 cited by

NOTA: Multimodal Music Notation Understanding for Visual Large Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.14893 v1 pith:X7HO3CBI submitted 2025-02-17 cs.CV cs.AIcs.LGcs.SDeess.AS

NOTA: Multimodal Music Notation Understanding for Visual Large Language Model

classification cs.CV cs.AIcs.LGcs.SDeess.AS
keywords musicnotationlanguagetraininglargenotaunderstandingvisual
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Symbolic music is represented in two distinct forms: two-dimensional, visually intuitive score images, and one-dimensional, standardized text annotation sequences. While large language models have shown extraordinary potential in music, current research has primarily focused on unimodal symbol sequence text. Existing general-domain visual language models still lack the ability of music notation understanding. Recognizing this gap, we propose NOTA, the first large-scale comprehensive multimodal music notation dataset. It consists of 1,019,237 records, from 3 regions of the world, and contains 3 tasks. Based on the dataset, we trained NotaGPT, a music notation visual large language model. Specifically, we involve a pre-alignment training phase for cross-modal alignment between the musical notes depicted in music score images and their textual representation in ABC notation. Subsequent training phases focus on foundational music information extraction, followed by training on music notation analysis. Experimental results demonstrate that our NotaGPT-7B achieves significant improvement on music understanding, showcasing the effectiveness of NOTA and the training pipeline. Our datasets are open-sourced at https://huggingface.co/datasets/MYTH-Lab/NOTA-dataset.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores

    cs.SD 2025-11 accept novelty 7.0

    MSU-Bench is a new benchmark with 1,800 QA pairs from classical scores that exposes modality gaps and multilevel consistency problems in LLMs and VLMs for musical understanding.