MuSViT is the first foundation vision model for sheet music, pre-trained on 9.7M IMSLP pages, that outperforms general encoders on recognition, detection, and classification tasks while encoding symbolic structure in its embeddings.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
MuSViT: A Foundation Vision Model for Sheet Music Representation
MuSViT is the first foundation vision model for sheet music, pre-trained on 9.7M IMSLP pages, that outperforms general encoders on recognition, detection, and classification tasks while encoding symbolic structure in its embeddings.