Pith. sign in

Paper Citation Record · LEDGER

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

As of 14 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 10 inbound Pith citation observations for arXiv:2412.06660.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06660 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:31:44.023128Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:49:44.099128Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 52747d69-ecaa-42ae-944b-b6c5af526adc · outbound

This paper cites , " * write output.state after.block = add.period write newline.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.705432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.705432Z digest=sha256:2508470047db3121f2ce34de2400245f3113107e583daac4ec6b8cdf6bc74a5b

Observation 0ba7e298-946c-4c5b-aee2-046c69434faf · outbound

This paper cites write newline.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.711036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.711036Z digest=sha256:dbbdad60d5e980b82afe1ccf4bee44522f4c4beec97a8800ef0c7594cf9ea0d9

Observation d359a3da-9657-47eb-ba09-8ad9fedf5f58 · outbound

This paper cites MusicLM: Generating Music From Text.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models MusicLM: Generating Music From Text

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.716458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.716458Z digest=sha256:858407c99360588ab7b8e736ab580e0d68e900e6b747e1f8a9c38af6d41f0e7a

Observation 2aab8bef-3821-4dfa-9b42-4a40c0798e34 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:45.069968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.721855Z digest=sha256:021e4110286fcb507937356be40ec0c31d57bd8a67b043a4789f2f09e83608f2

Observation 88ba6a51-11ed-41f5-8d46-10efce83322a · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:45.056091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.726511Z digest=sha256:ad297537e986e18fe1ab6990b9edf5860496ac56da0a42483fe96b5df00064e1

Observation 8928c1fc-95d9-4b4e-aba5-526d748473b0 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:45.042995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.730971Z digest=sha256:b091a813df08ed4e978b8edf02b06257aea3227f7db51d6cc656b900b84928cf

Observation e0d86532-eb3e-412c-9e36-7a1f0beff2c6 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:45.029690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.735376Z digest=sha256:2edbba42c2823d878c4726753beb6ab84d93833f652490bfeda9abd599be3d09

Observation 80f68da0-f51d-40a7-ac0c-b12f9ec1b657 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.739675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.739675Z digest=sha256:e1df626545489ef5f01911a64f539dbb2b2d355bc7486ac4d69d172cf6623d96

Observation b1f525db-bb51-4be1-96d3-848765969541 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:45.016332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.744803Z digest=sha256:aad79af738739853d21edff3f46e2f8d3a6f3dd48b4f51717d2ead4a215902bc

Observation 83817e5c-7fe5-459b-aad0-537c6c10f49e · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:45.001971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.749175Z digest=sha256:5700199925f5c1284d19326da4e9c74e51a18e800f930d266b2365ca43141f53

Observation 3feb965a-9521-4145-99a7-74a4b43be759 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.987889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.753827Z digest=sha256:bf145ac706cc16f42e38f4eab418115dfba66efff093b1bf5f69dbf83bf0ff8b

Observation c7093b46-d3e5-406a-a6fc-ec4ee3e93aed · outbound

This paper cites Simple and Controllable Music Generation.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Simple and Controllable Music Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.758323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.758323Z digest=sha256:daac191a4906b1506b288cbb6e4b46edc212ca746defc3092ae32a7b2fd073ed

Observation 5b0b685d-f88a-4002-89ce-00f4d21d3269 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.974799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.763005Z digest=sha256:58294d36b0d654976c4a45796de2727bffe22eb807564a18917d0d9f469ade49

Observation 81f03d1d-2b10-43df-a194-58202a90c243 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.961749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.767022Z digest=sha256:f60cbc58db8b1bef3dc4705cee57b2e8b967ebfb6bb650e623a5beaafd4a96b1

Observation 09192207-9f26-44b1-b4c2-c6c38798cefe · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.949835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.770611Z digest=sha256:5234dfe5835e23c8de45b88da3d2e5f76451a43f4ebf4cdd999934ec99e340a8

Observation 8a88e08a-9bfa-43b6-9806-b4a8e16df373 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.774088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.774088Z digest=sha256:fc0f393c0e677259e8c76d5d389c13cbf90dc011e4a30df3696a8f8c4b40095e

Observation f6097280-c2f2-47f8-b9d0-f5c224f2f2cd · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.937143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.778088Z digest=sha256:23e16b38764bf6691d5c8149de8dfbef2d4e808692cb6a6743014fbbcb470dd7

Observation 42cb127e-1952-4855-87e1-c32ff338a4cf · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.781732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.781732Z digest=sha256:79e03b3a67c6d5fc043c22e45b3cd828df7a603927a1d178fa7f1286453fb9a9

Observation c83293ac-b572-44e4-a03e-2a2e4eec0f1a · outbound

This paper cites Planting a SEED of Vision in Large Language Model.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Planting a SEED of Vision in Large Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.785719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.785719Z digest=sha256:3558935aeb63e62b021b52cecef3cfa7e1ce976d912517f3d474dac46278674b

Observation 78898fc8-66cf-4b12-b871-78e0c9f0cdf3 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Making LLaMA SEE and Draw with SEED Tokenizer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.789931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.789931Z digest=sha256:6daafa4c46ce530947fd5837dfcc685210a15ac70211b471564b2c8889e6d22b

Observation a4d8a4db-3c20-419f-8fc8-2ea0cc59f4af · outbound

This paper cites F.; Ellis, D.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models F.; Ellis, D

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:31:44.924567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.793698Z digest=sha256:1023cad3d0f90f96d62af56605b966ede29005b57fb89e92881c0790140b6a70

Observation cb8d17b4-f7d8-46cf-8426-0e4ef07eaa56 · outbound

This paper cites V.; Joulin, A.; and Misra, I.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models V.; Joulin, A.; and Misra, I

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:31:44.911841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.797283Z digest=sha256:620e3cba53cadcecae07f209d3d7732ada4e2715d7cf34ffbe5a64fe433ee6e8

Observation e0842bcb-6311-4948-8dc5-4b55942112be · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.898439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.800922Z digest=sha256:119d7259485a654d11babac224cbec2c5e5f8bfce4a2b8d495386c509ae91c4f

Observation 7e376fd6-1a0e-4692-82fe-5598a7617139 · outbound

This paper cites Listen, Think, and Understand.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Listen, Think, and Understand

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.804426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.804426Z digest=sha256:0a1a363896ff2646bc3091f9cf1085cf935c0dc29748a2ad5dfd460ef34520c2

Observation baeb32e2-5344-40fc-922a-7c98d4eaf7cb · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.885307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.808007Z digest=sha256:e5a9a7fbb2987ba84789a4988b5602a7c14a1708211a22e52a6e61269be10550

Observation a6bf126f-83b2-464b-8675-01c2cbbe3763 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.872926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.811397Z digest=sha256:37e1e4756166bc6c3c807b73fd0180f0d15997e41619ce0775e5bebf38e139cf

Observation c601d7a0-ab0d-4afa-92ed-e20d9e70784c · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.815043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.815043Z digest=sha256:151f0dc51a1041aa6b74fa7d3920eb5831259b9a73e6e224067609e5dcb11c1d

Observation 06b2e3f3-6fe6-4a5e-b8af-917de5d3ce31 · outbound

This paper cites InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.818818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.818818Z digest=sha256:07b0f530bee1cbdcbbac120d02084e64229e0d1de48dbdbe68fceeaff24b316e

Observation 1f62d293-4a55-4a37-bd16-b44ea017a713 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.859884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.822501Z digest=sha256:63cb82e4852edefec375cfa026f55b3b305ae2460fb3260e4d109c4566537588

Observation 69052478-f215-4616-8153-e1276cff4bd0 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.847372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.826558Z digest=sha256:0282f3144718e5b34b40d0664ec68825a401a665f19b9cb6f60144e5cad335c5

Observation 84b756ee-1ead-4558-93c0-3770b665280a · outbound

This paper cites J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:31:44.834131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.830782Z digest=sha256:dc831c94fdd24cc4e572d0f1284e2955e44feb56ba21112b1b5b2d7f3f2cff6f

Observation 15a81bfd-452f-4c7a-8bcd-67b378a9de1a · outbound

This paper cites AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.834996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.834996Z digest=sha256:7c33ed00acd66a7eb4ddba3bd52ecc8831046cf701ce7d3e2d3e4c2481bae7a4

Observation e4791fb4-4bc1-42b8-8642-add112dd2ead · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.820518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.839551Z digest=sha256:722d5a25b7f7ca4a6c822d5ef981ecd6a7e5d0f0e9bc62642aa118d32cb31215

Observation dc8013f5-4ab4-46a3-994c-6c29a2339efc · outbound

This paper cites Mistral 7B.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Mistral 7B

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.843669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.843669Z digest=sha256:4880ff1b9febda1c90369fd393642acf78c2ac803253872a192b7071bb4f4994

Observation fb776e67-9be2-4663-aec3-6aca20e9d67e · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.807986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.848004Z digest=sha256:be9fb0aa14632b1d61d00ce8d573699e6469adb1855eb190d8841a738cc64cf1

Observation 689e732f-3887-489d-a2eb-001b8d873cde · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.795267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.852203Z digest=sha256:d03a0bdb2ff22c220def35822933a37cc779808adad613c9bd510d5838e5e5af

Observation 5ebf0617-3da7-46c6-8701-c814db27bf03 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.783318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.856476Z digest=sha256:6fa02069c500a0d16cac7156bec3caa01f000653eb3dfd0c4da599732bc608d0

Observation 4849aeb2-9e27-45af-a0a9-683d9500c0d7 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.860736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.860736Z digest=sha256:66d7c6236760992389d6b37570ace6e09a1425c1c410f0e71aef84dc9e41eba8

Observation b5be3d1a-7846-4d7f-a791-0cf5174197af · outbound

This paper cites MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.864877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.864877Z digest=sha256:cde16feb3325703bdb5d0ae29e240619840d28670b2b8746b0cf02969d8b8a6e

Observation c725aae5-cdcb-44f2-9cce-da35d0be786b · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.869161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.869161Z digest=sha256:2166943c22d581c095489751b67aeb97f8f95fedce29ecdebfc40157d6ec367d

Observation adbe5ba0-3e8d-4b71-8940-b63ba98a2d1a · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.771539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.873566Z digest=sha256:31421359f94df1d71fc98d807ab78134eaa4a74cc1afdbba1f7bb521e78b1fa8

Observation 7dec7a92-355f-4f77-a4f7-50b6ca51471e · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.759888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.877655Z digest=sha256:b350b0c0205fc5f5c287b6312f38bae80161e642460a48890281833dc5c950da

Observation 55b9226c-4cb1-4cf6-b331-4103cb85e500 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.747874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.881898Z digest=sha256:7e25372e2dad69999f11e314080a68a736918c9c644ab71f916affb6debfce50

Observation 06d17157-3251-42da-8043-e289a7c4f7b1 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.735330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.885967Z digest=sha256:c29bfe5759e6b4ab583bd551fa4690abe78fc9fa9f9c517bc683880193b1a051

Observation d415168c-88d2-4d38-bbdc-e30ffb24f81e · outbound

This paper cites Visual Instruction Tuning.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Visual Instruction Tuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.890048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.890048Z digest=sha256:df78aad218266d2ba7e699415d0d8f554a522a4e046e573e1b6fc86c6bb6b601

Observation b236c06c-d809-4d78-85e8-82618aae45c2 · outbound

This paper cites AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.893881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.893881Z digest=sha256:4235bc9b1989a023a5790d53bb73945fb38b0495fbe4f9fe3025ff5adb0fe672

Observation 4edca8fd-f310-411c-aeb6-b66e85c338ce · outbound

This paper cites Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.897874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.897874Z digest=sha256:240695cf57a83294040dd789f2f934cdc9f120c2cb1b8d59f0136cf71693a19e

Observation 98d8bb8f-6348-421a-afb2-7b281e9f1bfc · outbound

This paper cites WavJourney: Compositional Audio Creation with Large Language Models.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models WavJourney: Compositional Audio Creation with Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.901563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.901563Z digest=sha256:4e91860da5ef9316f3a86a2d3b0765807bf84594a31d36484bdc7c7c8c96321b

Observation ca5a84a2-40c7-4805-9831-6311674156ae · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.722677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.905562Z digest=sha256:f461dfb740b995fddde3c1787a7a5f2ed58aeef513b414bc92a07d6c85d237db

Observation 919c5fa8-8e2c-4e6f-bb69-deb37b4425b7 · outbound

This paper cites Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.909252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.909252Z digest=sha256:b2be3d872a51b84b9727653ceaf72cd9274148d196bbc978f75e2caf8f686682

Observation 4bb54209-c6a5-412d-bf0a-4d788672e223 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.708046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.913490Z digest=sha256:c4869940d215e1b239e203b75d53cfc0899da57eb6a9d4f1dd694f664e93c500

Observation da59dc1a-1487-4b6a-9753-c1a332670814 · outbound

This paper cites D.; and Wang, W.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models D.; and Wang, W

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:31:44.693653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.917114Z digest=sha256:5fe0913ae0b3982d368cbf8f90751d1136a3d6c359afd6960049972288cedcda

Observation 6f88dc4b-655b-4c6d-8a65-a96c73c210bd · outbound

This paper cites H.; Lee, H.; Shin, W.; Kim, Y.-H.; and Choi, E.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models H.; Lee, H.; Shin, W.; Kim, Y.-H.; and Choi, E

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:31:44.680404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.920668Z digest=sha256:c4c3f0cc1a3e71462d63f2069be8209d10764ce4695cf60ea8e7da2a7f35f016

Observation 39898fb8-8f9c-4568-a9d2-4a67acd28b4f · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.924005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.924005Z digest=sha256:5937528718b34479505bf7cba199f2ceb86775a24b201c594de279932f0f0dd5

Observation c0a99d80-a4bd-4ebc-bd64-0cb2eb7e69b5 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.667148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.927630Z digest=sha256:eac6b5cfd5c2a8cb49bf4ef89ee02fe6f7cb08b8a0f5c1d3d76dd46090c6a7b2

Observation 08a4e1f3-be52-4eda-ba5b-462b371474f0 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.653987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.930885Z digest=sha256:edd8630da63b941f26d7209931b101b7a8869f9c6874ef122e3f0cf013e56909

Observation 24979a9e-48f5-458c-b7a9-0ceb7918634d · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.642068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.934500Z digest=sha256:6ab77bd503ee57c443ba30d89224163024587e075ed9f28c8580a43d11a0817e

Observation 6a6e6337-2662-435a-a33c-a7994f71352c · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models High-Resolution Image Synthesis with Latent Diffusion Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.937901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.937901Z digest=sha256:066a62828afd5905f050e8c9ed3f1645e2f05ebd7c21c437c38d555f442f3784

Observation 09797383-3f95-4f8a-8e03-7331d5477f42 · outbound

This paper cites 3D-GPT: Procedural 3D Modeling with Large Language Models.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models 3D-GPT: Procedural 3D Modeling with Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.942283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.942283Z digest=sha256:6be70e525c24b8472e96dd4ca54cb6348934dabf868b8f08cd1207500411fe5d

Observation 3f8a90fa-b2eb-4a95-90ca-708a44c1ea6f · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.946807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.946807Z digest=sha256:5c48f29ca36143cf07a876297f07d123df3b85e64a3ccd65dbb557ab95e88158

Observation 4c53b14c-d68c-4132-894e-658527a1e44a · outbound

This paper cites Any-to-Any Generation via Composable Diffusion.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Any-to-Any Generation via Composable Diffusion

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.951230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.951230Z digest=sha256:954066b09e2b6cb28a13769292f80b6ccd8c4f7e89bfe2857f74ffdbcd4e031d

Observation 18dc28b7-52c9-47de-8221-1caa30c9d611 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.629925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.955474Z digest=sha256:5d1d53570856bb8278e8c42272dfaccaf7c375327e065c31e2aa8c98173c86bd

Observation f26c9e4d-38dd-459d-91ca-0183f3514fe2 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.617172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.959510Z digest=sha256:4f6dc14b6a34dd7e735cae3ad8d002442ded9f2e50d3dcbfe90839429e62fefd

Observation 47cfa9fa-e367-42fe-80ff-2c5cd8f8067e · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.963415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.963415Z digest=sha256:4809a1842686c890054cf81af23800463ff52b9d7871da7213a6513b4ebf3765

Observation 82ddab59-89c4-4e42-9fb9-70384bb357aa · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.967483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.967483Z digest=sha256:fc2510d20d5a444e0c76dc0cf0250c19ebcde866f40e3139a22688651635ab94

Observation 8bdfeabf-4b3b-45b9-bee5-f3c996a5441b · outbound

This paper cites N.; Kaiser, .; and Polosukhin, I.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models N.; Kaiser, .; and Polosukhin, I

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:31:44.603891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.971624Z digest=sha256:ac735316209094915c6ffe8f86267abb3651df6ddcd1ef96e91c7c105da9983a

Observation b5b553fa-f43a-4502-a6c4-debe01289e51 · outbound

This paper cites AUDIT: Audio Editing by Following Instructions with Latent Diffusion Models.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models AUDIT: Audio Editing by Following Instructions with Latent Diffusion Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.975732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.975732Z digest=sha256:b603e7c481bc263933f0cd8827f437f66d088f2153571796ff0d97bfd8463a6d

Observation 326ea139-20b5-4057-994c-1c66eaeec16e · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.980072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.980072Z digest=sha256:b8ad7973a8e70678d9c2db92586bb0dc28a30d228efb00650f2abc1c2d60e819

Observation 23ca5c1a-42a3-4b83-8e68-b5d7ab8094b9 · outbound

This paper cites an unresolved cited work.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:31:44.589181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.984307Z digest=sha256:25e7101f415ed0e56d6d46a7520be96e1178055b13fb965caad5697af2684a61

Observation fdbf86e9-be1b-4249-ac9a-e77a0b9907c4 · outbound

This paper cites H.; Fan, L.; and Wu, X.-M.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models H.; Fan, L.; and Wu, X.-M

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:31:44.576398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:43.988171Z digest=sha256:1bc9d080a0d8177a64d969d74f99a2043704c937910808cfb1c1381daa6f190a

Observation 87ca801d-b520-4f0f-88bc-98fafc8c6841 · outbound

This paper cites PointLLM: Empowering Large Language Models to Understand Point Clouds.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models PointLLM: Empowering Large Language Models to Understand Point Clouds

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.992134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.992134Z digest=sha256:d206cbeb1614c69dad7b17c31a4ddb779e03aa5feb326e47083e3dcb6cfb3790

Observation 43600cc5-b3bf-4b74-81bd-f52404f90fc6 · outbound

This paper cites TEAL: Tokenize and Embed ALL for Multi-modal Large Language Models.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models TEAL: Tokenize and Embed ALL for Multi-modal Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.996573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.996573Z digest=sha256:1cc90aad5279b9986fd5c9b4e618cf23eb715ac3464f336d38903dd537e0836c

Observation cc1d4927-f74e-49f2-9f80-230c6843b3f4 · outbound

This paper cites A Survey on Multimodal Large Language Models.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models A Survey on Multimodal Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:44.000451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:44.000451Z digest=sha256:c1ddb674505d3cf65f18cb7126e98d62aeff47e905fab864a03bd9781bddce99

Observation db60093f-5b2c-4ecd-bb64-70d7f5f447ab · outbound

This paper cites Vis2Mus: Exploring Multimodal Representation Mapping for Controllable Music Generation.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Vis2Mus: Exploring Multimodal Representation Mapping for Controllable Music Generation

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:31:44.092583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:44.004385Z digest=sha256:50ee9d1499c0761c075c1858f0ea686aa9404c622710c73162c7bac4078b3afa

Observation 7cc0f05a-2758-4178-a122-9745b03256b9 · outbound

This paper cites Q.; and Artzi, Y.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Q.; and Artzi, Y

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:31:44.562649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:44.008431Z digest=sha256:b35256512ecb44afc6ff168e258eed2c6134e98e16cad96dcaa6191088dfa998

Observation d38993f5-6468-44a0-bfa4-316f18c18cff · outbound

This paper cites Q.; and Artzi, Y.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Q.; and Artzi, Y

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:31:44.548941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:44.012106Z digest=sha256:90c14e772b441e113dbd3fe1a8a6d521761517e81e7cbeacfc97971900df6c6a

Observation 20c9109e-81e2-4bf3-a976-c96c84423cd9 · outbound

This paper cites Loop Copilot: Conducting AI Ensembles for Music Generation and Iterative Editing.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Loop Copilot: Conducting AI Ensembles for Music Generation and Iterative Editing

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:44.015610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:44.015610Z digest=sha256:dba90c60c9c82be54813004343efb9b53e40b9127ed743bf203edb18f1f52c92

Observation 06aeaa55-97f4-41d7-b525-65f38280117e · outbound

This paper cites a henb \.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models a henb \

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:31:44.534909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T19:31:44.019472Z digest=sha256:9bbcc5b13e01223cba3afaaf5da7587ff705d6f9ab9ae39828e67b6c8593ce30

Observation 959a04ba-52b2-48f6-8ac3-de7c8176ae1a · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:44.023128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:44.023128Z digest=sha256:7606970b3647c81250cfc67124be838570a6b402cef687d9ea61c5fa63e7ce14

Pith citing papers

Observation 29b928b4-861d-4832-8757-7750994f7a18 · inbound

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections cites this paper.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.099128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.099128Z digest=sha256:3fec74fe4b8000cfa952ca4ce78f63766cf247cb9f0f098d33438884090210a0

Observation e0a56444-c904-473d-bc8c-293cddf235e7 · inbound

CoComposer: LLM Multi-agent Collaborative Music Composition cites this paper.

CoComposer: LLM Multi-agent Collaborative Music Composition MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T14:08:49.450762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:08:49.450762Z digest=sha256:c4f84d143473c61d9aa20632237641031a4abe4cd02c2b1a14445792ca5153d7

Observation e22d0eea-1baa-4009-9178-c7d91576540d · inbound

WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation cites this paper.

WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T17:01:26.252375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:01:26.252375Z digest=sha256:76327e66b2d4292ff241a68b91151ab1ab4892e5d552148775d6b475834cb8e9

Observation 15dfa6d9-63a3-407b-b1a1-ef073fbfabcc · inbound

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach cites this paper.

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:41:22.691207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T12:39:56.396231Z digest=sha256:eef4dc267be03495fa20ed80321b748e6350701fb9bb9eaf02173d7a5d51d09d

Observation 7c7d4301-72d4-4c73-8a61-293e4a56bae5 · inbound

Assessing Factual Music Comprehension in Large Audio Language Models cites this paper.

Assessing Factual Music Comprehension in Large Audio Language Models MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T00:29:46.845901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:29:46.845901Z digest=sha256:9240af617946f047393c74ad7562b6ea4f90a5d813d290630c3a6b3c2b9c78a1

Observation cb1a608a-ac30-449e-8508-c744c0187524 · inbound

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs cites this paper.

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:03:14.327192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T07:58:51.272971Z digest=sha256:126bc9e595db64ba4c758606d4f51f384c026929d941f1ff0757b3c9118851ce

Observation 6f86704a-5c1a-4950-9d1b-daabb4d47bc4 · inbound

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions cites this paper.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:24.675619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:5f5304279d6d342122787e847f0f735c94dd16126d3bfb5cc1323fa75de631d2

Observation fb1dc0a5-c76d-40de-b689-6b9ebc421660 · inbound

EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement cites this paper.

EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-28T12:42:08.960299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T12:34:06.024192Z digest=sha256:823a6bd32adbadd8b7d1644317182f721771fd945c5e14d6f13c5451d9b65b4d

Observation f9ce5774-f600-457c-93c2-fc45393e91ee · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.767747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:05dfd2537d3e329e0257a04fadb1a6a93ff7a1aa4a3307038c8d0528abdd97f3

Observation 4bb19c6a-ee30-4b9a-9400-4353f2b0f2dd · inbound

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models cites this paper.

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T01:28:27.544162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:28:27.544162Z digest=sha256:256fc6ff67b6f761cfd6439a14314e8d8127d1a449de58841473fb8e581d23d3