Pith. sign in

Paper Citation Record · LEDGER

Mustango: Toward Controllable Text-to-Music Generation

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2311.08355.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.08355 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:45:19.847605Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.335915Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bc78afad-bff9-43ca-be72-4027c2db333b · inbound

ETTA: Elucidating the Design Space of Text-to-Audio Models cites this paper.

ETTA: Elucidating the Design Space of Text-to-Audio Models Mustango: Toward Controllable Text-to-Music Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.847605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.847605Z digest=sha256:1ec60b895cfce7f348fe55578212e25bdd0d80ad096f8b3f3266a60e4e02521e

Observation abcdffa6-0467-4f04-94e8-a8f31291579b · inbound

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation cites this paper.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Mustango: Toward Controllable Text-to-Music Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:56.333453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:56.333453Z digest=sha256:96e20c66f33500f3d6d15564089ad1ad8546716fd8f69d4dd2d364d00b226a31

Observation 3158d74e-844f-4719-8b6e-4ec31bbdd231 · inbound

Genre Controlled Music Generation via Activation Steering cites this paper.

Genre Controlled Music Generation via Activation Steering Mustango: Toward Controllable Text-to-Music Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:36:41.140747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:36:41.140747Z digest=sha256:e2a79bd0b927d25bfeaece6195bdb6d5349fb6b110da0195c38710a7fbc3d0a6

Observation 4a2b1cb3-759e-4290-96a2-5d7f6423fc9b · inbound

DanceChat: Large Language Model-Guided Music-to-Dance Generation cites this paper.

DanceChat: Large Language Model-Guided Music-to-Dance Generation Mustango: Toward Controllable Text-to-Music Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:26:42.195662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:26:42.195662Z digest=sha256:614ee142bdd3b8c5a39baa10b0e76c3f65f75ced71a17bc7d2ebbf108ed5c611

Observation 712b6267-1f37-4ea1-94e7-9cfb31d71bfd · inbound

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections cites this paper.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Mustango: Toward Controllable Text-to-Music Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.437966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.437966Z digest=sha256:8f0011cad21d037e2253e078f8c1158d38effb50e32c5fd6561c16332e0be0c1

Observation 7b451098-c377-4e48-ac5e-9248f2331f46 · inbound

MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners cites this paper.

MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners Mustango: Toward Controllable Text-to-Music Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:19:51.505249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:19:51.505249Z digest=sha256:1a27039346ecb35d875c20f5e1caefefdd76c41dec20f5fddea81edfdf239cc9

Observation 2570ed29-32d7-4832-a1d2-7ab9448d92ce · inbound

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models cites this paper.

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models Mustango: Toward Controllable Text-to-Music Generation

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:42:44.834686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T03:42:44.523919Z digest=sha256:e5c679ca1eaa5b40d6a925c7b057dd649851b16587df62f6979b062f5bf5496b

Observation f2198e82-a9f1-4308-994b-d4b265a5eaf5 · inbound

DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization cites this paper.

DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization Mustango: Toward Controllable Text-to-Music Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:40:49.274001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:40:49.274001Z digest=sha256:a5fe4c0a784aefbc8b6dffc973ec0ccc979c936c27884f228843dab96d1c6ce3

Observation 96a2debe-1a19-49b5-b9b0-f38a10d7574e · inbound

Controllable Video-to-Music Generation with Multiple Time-Varying Conditions cites this paper.

Controllable Video-to-Music Generation with Multiple Time-Varying Conditions Mustango: Toward Controllable Text-to-Music Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T13:31:12.444751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:31:12.444751Z digest=sha256:8f3741291c3fd0a37871e90f0f4f79f32a3967abecf8a2a8ed6e58a6cd8b21bd

Observation 4d39e694-9084-43a8-bad0-ca9069f4858c · inbound

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment cites this paper.

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment Mustango: Toward Controllable Text-to-Music Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:14:37.561225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:14:37.561225Z digest=sha256:276d98187f106a05d6af1e26c2c7265108c404393e0f4a69466f428002db6b42

Observation c5536b7b-197b-472b-b36a-929ede8c8c6d · inbound

A Survey on Evaluation Metrics for Music Generation cites this paper.

A Survey on Evaluation Metrics for Music Generation Mustango: Toward Controllable Text-to-Music Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:27.397749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:27.397749Z digest=sha256:dd716b329f9889160817882f3526c1029f13e685c16eb1258ed876cb52f8b7f2

Observation afb0178d-35fc-47ae-b92f-6849f77c9448 · inbound

AImoclips: A Benchmark for Evaluating Emotion Conveyance in Text-to-Music Generation cites this paper.

AImoclips: A Benchmark for Evaluating Emotion Conveyance in Text-to-Music Generation Mustango: Toward Controllable Text-to-Music Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T13:16:08.841052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:16:08.841052Z digest=sha256:f17f3fb73ce3d2e8c7caec8f09307065e866d729584fc287b7a940337aac1779

Observation 1f5ac5ce-20ee-479a-98d8-f6bcb284dddb · inbound

TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization cites this paper.

TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization Mustango: Toward Controllable Text-to-Music Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T13:12:14.509681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:12:14.509681Z digest=sha256:af14f1355994a93bdf5f1b0c0ebb54dbffe42d3edd2fa511628741c68894787b

Observation 24ef680d-7125-4f95-ab00-09148a105fa6 · inbound

Segment Transformer: AI-Generated Music Detection via Music Structural Analysis cites this paper.

Segment Transformer: AI-Generated Music Detection via Music Structural Analysis Mustango: Toward Controllable Text-to-Music Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T20:54:14.265474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:54:14.265474Z digest=sha256:be86d07bce3bbd2f8258d6d4a9c8de7914dd3407994556781149bf2b6d8a5af4

Observation 394b1895-d2f0-4adf-a7a0-5b1e62f1066c · inbound

Steering Autoregressive Music Generation with Recursive Feature Machines cites this paper.

Steering Autoregressive Music Generation with Recursive Feature Machines Mustango: Toward Controllable Text-to-Music Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:10:54.487022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T05:06:22.119543Z digest=sha256:7811fcd432a6b687f603229199d2a684483816501bf790fed817ca8894878317

Observation 8b75afb5-e8bf-4f16-b92d-4f948b9b7f8d · inbound

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions cites this paper.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Mustango: Toward Controllable Text-to-Music Generation

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:56:24.665422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:3d2a019cc83c2ca6bb878ace76b4ac0e1c50f6df1ee87c29ec9481737ddba96b

Observation dc665989-1d9d-4ef0-8f1b-81cc125c072d · inbound

FIGMA: Towards FIne-Grained Music retrievAl cites this paper.

FIGMA: Towards FIne-Grained Music retrievAl Mustango: Toward Controllable Text-to-Music Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:47:06.213111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T23:32:09.401023Z digest=sha256:95876af9d6ccca85e5fb4d20c0a70c7a413bc500b73adfc67e9603d76c41712c

Observation 11dd6db8-79d2-4a3b-bea3-e6693339f5e8 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Mustango: Toward Controllable Text-to-Music Generation

Reference 80

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.337322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:308bc4bc824325a37b2968f1066c0e089901323f9cf7d1e9ae626cc7dbfe0eba

Observation 59ea047e-8644-4b8f-93a2-04a424b2b992 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Mustango: Toward Controllable Text-to-Music Generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:2c00a1968744927350e2d83c27883e14b645e8c1f1f73b8643ad812c1d780e3e

Observation 8fc956f7-7bdc-42f8-9c5b-bb93f9cf7fbc · inbound

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model cites this paper.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Mustango: Toward Controllable Text-to-Music Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.979881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.979881Z digest=sha256:89519ad30aa02d255f67ad480d16960f1b2152f0d79de68f938dc7d80d91adc7