Pith. sign in

Paper Citation Record · LEDGER

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2412.06660.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06660 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:49:44.099128Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 29b928b4-861d-4832-8757-7750994f7a18 · inbound

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections cites this paper.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.099128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.099128Z digest=sha256:9f556a69a8e7985afc80723c9a18d03957429e9fd12d74c20ff0070a46f5fd1f

Observation e0a56444-c904-473d-bc8c-293cddf235e7 · inbound

CoComposer: LLM Multi-agent Collaborative Music Composition cites this paper.

CoComposer: LLM Multi-agent Collaborative Music Composition MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T14:08:49.450762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:08:49.450762Z digest=sha256:21ce625c678ddeacead3fab4f3256e91c0f618d08c153cd52ba29e4820a457fd

Observation e22d0eea-1baa-4009-9178-c7d91576540d · inbound

WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation cites this paper.

WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:26.252375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:26.252375Z digest=sha256:a8809dc5d7bb7f3404b924dbc313163c7b26011a263633e602389827b025462b

Observation 15dfa6d9-63a3-407b-b1a1-ef073fbfabcc · inbound

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach cites this paper.

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:41:22.691207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T12:39:56.396231Z digest=sha256:c7b054d481d3fb77707a2f39d2df76a3633ba7484627b0e90060974f4b718267

Observation 7c7d4301-72d4-4c73-8a61-293e4a56bae5 · inbound

Assessing Factual Music Comprehension in Large Audio Language Models cites this paper.

Assessing Factual Music Comprehension in Large Audio Language Models MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T00:29:46.845901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:29:46.845901Z digest=sha256:1667f8423309113109c29beb7767c918e18b221d4f99e12030328a62185af0be

Observation cb1a608a-ac30-449e-8508-c744c0187524 · inbound

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs cites this paper.

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:03:14.327192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T07:58:51.272971Z digest=sha256:fb00b57999312dae62837ed724a2d2115528b9696b55591939529bfb25b28b91

Observation 6f86704a-5c1a-4950-9d1b-daabb4d47bc4 · inbound

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions cites this paper.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:24.675619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:64a2eac5a8efc881b3136ad8c74c26af5570857b454aab298960bb8ce3ab1647

Observation fb1dc0a5-c76d-40de-b689-6b9ebc421660 · inbound

EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement cites this paper.

EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-28T12:42:08.960299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T12:34:06.024192Z digest=sha256:46d6cf4c5c36b678ddc5755ace959a4cacbd14f6d34f5d20819795d0e49b24db

Observation f9ce5774-f600-457c-93c2-fc45393e91ee · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.767747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:42eb4f63617162389209cdc2116b9afe2b377a44a146d4a86a3aae82de092471

Observation 4bb19c6a-ee30-4b9a-9400-4353f2b0f2dd · inbound

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models cites this paper.

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T01:28:27.544162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:28:27.544162Z digest=sha256:5b42d11a458bb817b8f9c50358de27896aaae5e199c9fa249e871b307a33124f