Pith. sign in

Paper Citation Record · LEDGER

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models

As of 20 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 5 inbound Pith citation observations for arXiv:2506.23009.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23009 v3

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:57:08.122917Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T02:12:58.120295Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T02:17:46.305033Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact2
  • verified fuzzy28
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f5cc43c0-2f24-47c8-9aa3-c5a96aeb3049 · outbound

This paper cites Acrobat AI Assistant, 2024.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Acrobat AI Assistant, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:14.300132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:03.029002Z digest=sha256:0c61cdc3b3d1670cecfb164cbe85986f63cd491d1341ebfded036ca94aeeb470

Observation 940c5739-7079-4665-a70c-b0ff2e4571b7 · outbound

This paper cites Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:03.094544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:03.094544Z digest=sha256:8d6e4b9c27a06f39e6920f861afd9d41c01a71ccc382ea05d42e3effbcbca974

Observation d0e24319-069e-4c41-8abe-6d1fc94fce4e · outbound

This paper cites The basics of reading music.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models The basics of reading music

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:14.081606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:03.170985Z digest=sha256:490edf7434e2095f1eaef67cd34d44c4126f04b6d4ef7a70af45e4b78b7d89a7

Observation ccb2d247-30ec-4ed1-99b3-f219450ede58 · outbound

This paper cites Reading sheet music facilitates sensorimotor mu- desynchronization in musicians.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Reading sheet music facilitates sensorimotor mu- desynchronization in musicians

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:13.876029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:03.243569Z digest=sha256:dd196a64b9b8a9cca0f663ee6039801f861fa896646d6a43f10119f9be21a5f2

Observation 444834fa-8cde-4875-89f3-9e5004298c0b · outbound

This paper cites Optical Music Recognition: State of the Art and Major Challenges.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Optical Music Recognition: State of the Art and Major Challenges

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:57:08.894312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:03.307574Z digest=sha256:78cbf3b273dc94d92ff43234a22ce31985eed1ef8edc83461213d59d070f384f

Observation b20ed1a5-0a3d-49c0-87c4-c2fcf0ba7def · outbound

This paper cites Understanding optical music recognition.ACM Computing Surveys (CSUR), 53(4):1–35, 2020.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Understanding optical music recognition.ACM Computing Surveys (CSUR), 53(4):1–35, 2020

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:13.647645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:03.381957Z digest=sha256:f2fbeb089b47251d94b74ee8a80569262d25dfe990cb7f744055c60db0c0fafd

Observation 0ad181d0-58c5-4a4d-a5c6-6942d6910ebe · outbound

This paper cites Optical music recognition: state-of-the-art and open issues.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Optical music recognition: state-of-the-art and open issues

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:13.437024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:03.441555Z digest=sha256:a3b534012b9df8438345b63e4c8b8b5e0e358ed1a4c4c4daded0a3bcd96e781a

Observation 3cedc974-803e-4108-9133-15be13d1487a · outbound

This paper cites The challenge of opti- cal music recognition.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models The challenge of opti- cal music recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:13.262422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:03.530806Z digest=sha256:43381cbde83357be888c09633bf044144afaf33ade323e994f59223d4352ad14

Observation a453cfc9-af66-446f-9777-0952ee80cedc · outbound

This paper cites Optical music recognition using pro- jections.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Optical music recognition using pro- jections

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:13.096993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:03.612404Z digest=sha256:23b649a84527df6ec1f3c67e650c5b662755c4e8663d7b87e246791aeeeea6a7

Observation 60d3e49a-7c68-4e09-9e4d-418e773af6dc · outbound

This paper cites Gui agents: A survey.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Gui agents: A survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:03.704090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:03.704090Z digest=sha256:849f0a937cba9df6328157bbbdd697fde0ea1a1f4f233fe0189b5e39515fc3f2

Observation ea595109-5701-468c-b1e6-3ea06fd2ff0f · outbound

This paper cites Natural language understand- ing and inference with mllm in visual question answer- ing: A survey.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Natural language understand- ing and inference with mllm in visual question answer- ing: A survey

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:12.868582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:03.796792Z digest=sha256:0a9e20f1f297b246d9c4ec161606c7d765dffb599129d50834f040005361e381

Observation 5a9369b0-1cf4-4a8c-be2d-cce557455e5b · outbound

This paper cites Internvl: Scal- ing up vision foundation models and aligning for generic visual-linguistic tasks.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Internvl: Scal- ing up vision foundation models and aligning for generic visual-linguistic tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:12.665173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:03.801270Z digest=sha256:a75f2acfd726d81a1079cd2e0bdc3c42e935186a323ab799255071670da53549

Observation 320ef79e-a083-4b28-85bf-72cc75c4383a · outbound

This paper cites MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:03.931871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:03.931871Z digest=sha256:d122a93a2377521b18934d815a271033943e9a87c6a06f6704183f2d98c70550

Observation aa2550e8-38f8-4dd0-8235-07e20b3ea904 · outbound

This paper cites Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision- language benchmark.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision- language benchmark

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:12.448383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:04.036358Z digest=sha256:f491aa3508ce93300b7bbbb87501da2b999b6f0506987f5eed5a084b3a243b5f

Observation c8f6ce39-1320-4708-a18d-31ee29f69ce0 · outbound

This paper cites PP-OCR: A Practical Ultra Lightweight OCR System.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models PP-OCR: A Practical Ultra Lightweight OCR System

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:04.136009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:04.136009Z digest=sha256:56fcaeb325b321d3f9c104fc9672fc9e785e8234bddc616782726be46e61f465

Observation 8a612c9c-cc42-4bb0-a4a4-3444743860b3 · outbound

This paper cites Tex- tocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Tex- tocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:12.295573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:04.213329Z digest=sha256:44ec7e74afea25075985b826ecd1dbc9fe6ea3a1629bf640b7a9d257644f9b72

Observation 74b8b8f5-9138-4277-8ef6-e870f4d84a8b · outbound

This paper cites CVC-MUSCIMA: A ground-truth of hand- written music score images for writer identification and staff removal.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models CVC-MUSCIMA: A ground-truth of hand- written music score images for writer identification and staff removal

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:12.108776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:04.318571Z digest=sha256:bc44593d19e97ee4046561b8a7b696f7a8cd3f4af230081217df0f35a7baabf0

Observation bf1091cd-1469-437c-9c43-c6694acdb18c · outbound

This paper cites Knowledge Discovery in Optical Music Recognition: Enhancing Information Retrieval with Instance Segmentation.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Knowledge Discovery in Optical Music Recognition: Enhancing Information Retrieval with Instance Segmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:04.402673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:04.402673Z digest=sha256:5b4612b8ea8644874713280d43604f753b479eb78b0868da2bf075814742e18f

Observation 25e45c23-f6c6-4e85-8d72-e701d2888fbe · outbound

This paper cites Deepscores-a dataset for segmentation, detection and classification of tiny objects.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Deepscores-a dataset for segmentation, detection and classification of tiny objects

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:11.927933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:04.520127Z digest=sha256:e6f117845ef5f70acfd28b47dedd7f0bea9427448da4afc844f1264bc03bbe41

Observation c8cf186a-0bc7-471d-a910-eb80ebd255d3 · outbound

This paper cites End-to- end neural optical music recognition of monophonic scores.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models End-to- end neural optical music recognition of monophonic scores

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:11.748658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:04.652406Z digest=sha256:0d77240f89d21ee2782c19223c29df6ccc2988e62b501c11dd841431ab4732f6

Observation dde69ff4-d1c0-4e62-bb9f-2b8729f92b64 · outbound

This paper cites DoReMi: First glance at a universal OMR dataset.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models DoReMi: First glance at a universal OMR dataset

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:57:08.584722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:04.841644Z digest=sha256:efc0f74337e1ac8a126dd58d5b5399ea769535cd27a8215f1ff1352670ee2eb8

Observation 58138aea-ac0a-4327-bed4-ac749362497e · outbound

This paper cites A uni- fied representation framework for the evaluation of optical music recognition systems.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models A uni- fied representation framework for the evaluation of optical music recognition systems

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:11.509970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:04.953302Z digest=sha256:c164910a4caa189e5e511cd039dcb605a322ff38dfbaa855e577fee119eb2d77

Observation b9ccfb9a-77e1-4966-b55f-2c25af092de3 · outbound

This paper cites Practical end-to-end optical music recognition for pianoform music.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Practical end-to-end optical music recognition for pianoform music

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:11.238511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:05.038931Z digest=sha256:6a70f4c05a48f0ef15294bee5e1b35e43ca03420d5ab3edd9debd3881d0e5f82

Observation 0099d064-ce21-4194-ad6f-d8db0da76193 · outbound

This paper cites Breezewhite/oemer: v0.1.7, October 2023.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Breezewhite/oemer: v0.1.7, October 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:11.038775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:05.155593Z digest=sha256:f96a5e88302028a010cb38a520b8b63292a8c07c2e06afa6ff75732e969f5f2c

Observation acb50a2c-7375-4bd9-9de1-923c41045971 · outbound

This paper cites Optical music recognition in manuscripts from the ricordi archive.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Optical music recognition in manuscripts from the ricordi archive

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:10.853442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:05.249120Z digest=sha256:eb6589cba5456e57d4574b86ba6a739bd825be0675fcc518087f1c16fc7f38f3

Observation 5d5ad310-4581-4137-910f-76be0515c404 · outbound

This paper cites Optical music recognition with convolutional sequence-to-sequence models.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Optical music recognition with convolutional sequence-to-sequence models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:10.606216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:05.356627Z digest=sha256:141916f52599b4b8a57451c78d81d7901d2bc09dc00cd9997249ae1ea7a4e3e9

Observation 3f4661b1-32d6-469b-a1fd-69aa8172e667 · outbound

This paper cites Tromr:transformer-based polyphonic optical music recognition.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Tromr:transformer-based polyphonic optical music recognition

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:10.417118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:05.523218Z digest=sha256:a4c32ee7a64690ef33d9729c065d4d0bbbcb962b1442ae069c01e4517fd64159

Observation 1f102df9-df64-4f5e-98b4-c40740b6ed78 · outbound

This paper cites Sheet music transformer: End-to-end optical music recognition beyond monophonic transcription, 2024.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Sheet music transformer: End-to-end optical music recognition beyond monophonic transcription, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:10.244977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:05.682538Z digest=sha256:3914d2797b6d42aa8aa22976d090c9ef9e3a6540ae63f7215650e9b057e766f4

Observation e6bbf8a2-3f79-4115-9454-0853be89d164 · outbound

This paper cites End-to-End Full-Page Optical Music Recognition for Pianoform Sheet Music.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models End-to-End Full-Page Optical Music Recognition for Pianoform Sheet Music

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:05.787561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:05.787561Z digest=sha256:997006ce1b9ac9c3f6c31ba98cdeaa8e7c64245484f8af3ad8d1871f31ec11bb

Observation b25e32c3-7c57-4510-a441-9fc3db13dafe · outbound

This paper cites ChatMusician: Understanding and Generating Music Intrinsically with LLM.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models ChatMusician: Understanding and Generating Music Intrinsically with LLM

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:05.952693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:05.952693Z digest=sha256:ae5f03285050b497aa1cb481b7e8d0b437121ecab4e1c924bff39a5b96d42ab2

Observation 14ff5879-dfb9-45d9-8db0-5ea7db605ac8 · outbound

This paper cites MusicAgent: An AI Agent for Music Understanding and Generation with Large Language Models.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models MusicAgent: An AI Agent for Music Understanding and Generation with Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:06.012244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:06.012244Z digest=sha256:a28640f053f9f1b0c68d87af7a2ebf7d514632c8848ff03165a71688a433ded5

Observation d66af441-d8e1-4675-89bf-06959580841b · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:06.098344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:06.098344Z digest=sha256:25d5c7ba2cec32a21a28302e8a22d1183dc1bdfe7cfbd04a43517a8ecab214f6

Observation df89a664-d190-445b-81c5-1049c1ca726c · outbound

This paper cites GPT-4o System Card.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models GPT-4o System Card

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:06.179397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:06.179397Z digest=sha256:7a3b60a47c3e39d9404c02b506edb194bd4346f11cec58603f68b9ed7f7981d0

Observation a7908338-7a3c-4c00-b9bc-dd65d225b05f · outbound

This paper cites DeepSeek-V3 Technical Report.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models DeepSeek-V3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:06.325770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:06.325770Z digest=sha256:30c9897dfa0786300f4419433667b4c84a1aafea4f8afd7156d5f8e2d60d8007

Observation 369ae79e-0a88-43c2-9875-83f009103400 · outbound

This paper cites Musical scales and the generalized circle of fifths.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Musical scales and the generalized circle of fifths

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:10.090242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:06.425346Z digest=sha256:24d1f88856f7f810c9b9a42e951066828faddc373231f9d9bb97e59b4916f02e

Observation 0c53bfd0-4e4c-4c08-995b-4f015c15d54b · outbound

This paper cites MusiXTEX.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models MusiXTEX

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:09.932628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:06.520548Z digest=sha256:c8a7788d1f44761b1ea43881f64b523cf1305938b97aa8cec27e27e44b24de18

Observation 606c7818-3166-49b8-b16d-772e748a5a56 · outbound

This paper cites Harmonic experience: Tonal harmony from its natural origins to its modern expression.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Harmonic experience: Tonal harmony from its natural origins to its modern expression

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:09.751371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:06.597284Z digest=sha256:76a4b6c8568d082f6b2c5261064b60c791be7deb29d143b725a82d3f42c59348

Observation 4676f4b5-d193-46bb-9439-90bd09a41d67 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:06.742697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:06.742697Z digest=sha256:b179be974ae59e441d6054389600fc3c48d4dd68a2704b8824f0ec5bcb677764

Observation caa80afb-5dd9-4a46-8e4a-9b2ebfd4c561 · outbound

This paper cites Trins: Towards multimodal language models that can read.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Trins: Towards multimodal language models that can read

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:09.577261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:06.908892Z digest=sha256:cc8587c05c161995d08f471d838a38850fa64c946ea56d103386443151e9032b

Observation 97993c1c-9a6d-4154-9aff-ae7d5edad3c3 · outbound

This paper cites LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:07.053168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:07.053168Z digest=sha256:b565fd2d70e58b1f7d97e9a46c236b70c3b7081cc0c61a8d50acce523e6caa95

Observation 4e90a859-a7cb-491c-bffa-8b713cd65997 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:07.157500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:07.157500Z digest=sha256:88e0907fb18bdf29eaba03bb46bc6ada5c0025de891b17d89cd05304c30d2e3d

Observation c63f5167-f5c7-4aaa-978c-3c7e8c33e36c · outbound

This paper cites Music information processing using the humdrum toolkit: Concepts, examples, and lessons.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Music information processing using the humdrum toolkit: Concepts, examples, and lessons

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:09.407200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:07.301593Z digest=sha256:ea863bec4aa63e4126ec4437f2d36fe608dd42f6b019cea6b58da5dbe3b0f0df

Observation 1bdd0ad0-d2cc-4fb0-9f03-9aee76356e82 · outbound

This paper cites MMR: Evaluating Reading Ability of Large Multimodal Models.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models MMR: Evaluating Reading Ability of Large Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:07.432546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:07.432546Z digest=sha256:cc49d86c4db0411a91ffc43dab3e0d2eaa54a99629ea465ef9c4727e2c8deb59

Observation 7e0f840c-504c-4ca6-b0f3-13ac051aa37e · outbound

This paper cites Decoupled Weight Decay Regularization.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Decoupled Weight Decay Regularization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:07.577643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:07.577643Z digest=sha256:a413e8927c5c6068705b42993ac3bebe2c524b9c6bf32d639f4d41372b281e63

Observation 8b1be671-60c9-4af2-9463-b46312da531f · outbound

This paper cites Retrieval-augmented generation for knowledge- intensive nlp tasks.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Retrieval-augmented generation for knowledge- intensive nlp tasks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:07.694170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:07.694170Z digest=sha256:78bb937eb3921a63cecc6ab3305b55a59904acd6fba7d3426539ca65e9eb1355

Observation ffefab5a-ed6e-485d-8181-41742bf42140 · outbound

This paper cites Layoutgpt: Compositional visual planning and generation with large language models.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Layoutgpt: Compositional visual planning and generation with large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:09.227377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:07.851431Z digest=sha256:42c2ee6b360ff3a66be67f969ea320373bcd5a38d88daf603778e14217924160

Observation a1c2e789-ffe0-480d-9692-972da2b8b98d · outbound

This paper cites TextLap: Customizing Language Models for Text-to-Layout Planning.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models TextLap: Customizing Language Models for Text-to-Layout Planning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:07.934980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:07.934980Z digest=sha256:5928ce4e8c007eb380991695983fe337cb951728f646a6b5c3986ec504cd8e28

Observation 29ef0fcd-2bc9-4672-916c-740cfec8368e · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:08.059391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:08.059391Z digest=sha256:75061f6f2f5f39bef1c7fed0720ea70ba3d2587a673d34482625b6bfbefa9078

Observation b87514ec-34f8-4519-91d0-5228c6da73f1 · outbound

This paper cites Information not found.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Information not found

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:09.066912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T21:57:08.122917Z digest=sha256:4d9e8cb4cdd385527ba1dfb8adcc21ef6d0b4af57f19e16bc99094bd41ddbc3c

Pith citing papers

Observation d96761d1-6795-4eab-9920-7f5e5c1453f1 · inbound

ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence cites this paper.

ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:04:17.568533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T23:03:01.356594Z digest=sha256:a012db9d57403ff3456dabe57e1dd8a94a2bc5aa2b9c9d3aa8e75d115e9eface

Observation d3580fe4-5cda-4565-a885-f9ca87d51454 · inbound

Direct content-based retrieval from music scores images cites this paper.

Direct content-based retrieval from music scores images MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:24:43.335938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T07:21:33.390581Z digest=sha256:efc3f33dcc4b068304368ab6dc2dcb7174eb89bec6cf305826bd27e12aeee431

Observation eb9799a3-f455-45e9-8103-3613ff422f8a · inbound

Direct content-based retrieval from music scores images cites this paper.

Direct content-based retrieval from music scores images MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:56.953564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T17:23:39.285332Z digest=sha256:7dc496516301dc37fe287422dffc494d04c6c1ab959718ac089c21ce0348aeb8

Observation 9c192de3-7c3e-4328-ba8d-4758b25ca7c8 · inbound

LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding cites this paper.

LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T02:17:46.326748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-11T02:12:58.120295Z digest=sha256:2ee24e13dec1b0d3dbd2f1c0f6479b386d5aa0d38fd7e204648d9bad650e30dc

Observation 297aee09-8d2f-435a-a6c9-577e60120dae · inbound

Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music cites this paper.

Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T19:35:32.938812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T19:28:58.523932Z digest=sha256:65b2cb741fe890f8ba259f05510ecc8dde3e8be673e4e413f231d32db5feb058