Pith. sign in

Paper Citation Record · LEDGER

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models

As of 17 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2411.12641.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12641 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:22:18.146028Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact2
  • verified fuzzy17
  • unresolved49
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 990d35bb-ab15-4a86-b710-5824cc4d88dc · outbound

This paper cites MusicLM: Generating Music From Text.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models MusicLM: Generating Music From Text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.824039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.824039Z digest=sha256:d7919b7d8d1bed8b49bf3b4349cc34664a8a430a4d0bf656256bc203c32a43a2

Observation d34cade4-29f9-4155-a4e3-af00fe1ce649 · outbound

This paper cites Look, listen, and learn more: Design choices for deep audio embeddings.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Look, listen, and learn more: Design choices for deep audio embeddings

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.319985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:17.846659Z digest=sha256:03f1f7fbbee1cbeaab4ec36899b401b7f04c899fac1d80f52853b0eb7d85f293

Observation c1dbbc6d-6d7c-4958-b17f-ec351a2cf76e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.850829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.850829Z digest=sha256:51c4202215c17bbb45559d7d5f4fa19d256494c5ed21f6b21b738c533c71519d

Observation e601f507-1b25-4895-bf77-6822b18f1367 · outbound

This paper cites LP-MusicCaps: LLM-Based Pseudo Music Captioning.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models LP-MusicCaps: LLM-Based Pseudo Music Captioning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.865512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.865512Z digest=sha256:152435008d894e4855e2b8047a129cde0312fe7c51d25dfa6340c3e2782219f0

Observation 86c8d089-5c58-4df5-955c-7d032858bf5d · outbound

This paper cites SingSong: Generating musical accompaniments from singing.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models SingSong: Generating musical accompaniments from singing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.869956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.869956Z digest=sha256:f9f2607028dbe603dcc3f9cdb3840c931c04186a8cf9e103c80bb4f6fad92908

Observation 5b17c951-e43c-4d6e-ba22-54cbf1d17685 · outbound

This paper cites Joint music and language attention models for zero-shot music tagging.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Joint music and language attention models for zero-shot music tagging

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.306287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:17.874701Z digest=sha256:664b3514265c9c2e539319c8cafae8ff1b68bcee52d6b7a6fccf3c5b99b80263

Observation c37ee982-b157-47eb-8672-99a29fa3df47 · outbound

This paper cites Long-form music generation with latent diffusion.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Long-form music generation with latent diffusion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.879005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.879005Z digest=sha256:64def27133672983e290b1ee371a1eae22adece7d549844b9357435283157151

Observation 214f2509-3365-4596-b4c2-900f898a8b6c · outbound

This paper cites VampNet: Music Generation via Masked Acoustic Token Modeling.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models VampNet: Music Generation via Masked Acoustic Token Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.883184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.883184Z digest=sha256:aa93697c2770c6a57f06e33e23f9033e55bd3ceb802b2a27926c88a3e18606c6

Observation 506b6021-d676-4bda-b7e1-a3572e26369f · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Audio set: An ontology and human-labeled dataset for audio events

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.294067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:17.887501Z digest=sha256:6f11b3780fc119da32db9c7963684bfe82efbd7b9daa152c9c8f474c2e4845a3

Observation 40bc01b0-0948-4894-8ce8-254c2ef195ae · outbound

This paper cites CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.891727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.891727Z digest=sha256:521c874679483aaad5916542e52d3dbc7673d7d97bc7569bed6f1856ddaa75d4

Observation eca96272-26cb-4398-8d2b-0d765900055d · outbound

This paper cites InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.896554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.896554Z digest=sha256:9261a0bdede2f8a5fd8f61be961c9b5bdf826ef9732271d55cdd89235ca7b396

Observation de35cce0-8fe5-440a-9b0a-fdbdb3d63b6b · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.905866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.905866Z digest=sha256:48ca02acce9f801d05610db33f9c4758d1bda974b4fa9daabc886474668c7d0d

Observation 524d5f25-d188-48a5-8ee9-4bfb8847f6c7 · outbound

This paper cites M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.914502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.914502Z digest=sha256:bc9f84244fbbe8d174285a9025b5b91b7eead884304fb54a0183c5d167a8a3f3

Observation b5579418-771a-4633-a14e-002d6cbfc947 · outbound

This paper cites an unresolved cited work.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:22:19.280963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:17.919166Z digest=sha256:c591d91ed80e4b5903f6b9e7fd32762bdeefba6f917c852ade4ca386f797dc45

Observation b0111806-a9f6-4f14-af0e-a808c2b3e499 · outbound

This paper cites Retrieval Augmented Generation of Symbolic Music with LLMs.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Retrieval Augmented Generation of Symbolic Music with LLMs

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-12T17:22:18.874918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:17.928087Z digest=sha256:dbe034125608a9c5fcb623b3de207da1089ea950faf518340269fbc87b5458c3

Observation 7959ab72-2e7a-45c1-94b1-1a48bcaa3460 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.932668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.932668Z digest=sha256:b231e8f327f2978bfd45f36ad93a523e4c95e51901ded001cca8e2c762f3039f

Observation 3bb143b2-f6e4-474b-86cc-4b67d6bef4cc · outbound

This paper cites Auto-Encoding Variational Bayes.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Auto-Encoding Variational Bayes

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.941042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.941042Z digest=sha256:ebae89a712b4239037ee84e1c64e7842779a4a67b66d78e80e763b68353a6b51

Observation c664983b-0f43-43a8-8a62-8aaabf079cfa · outbound

This paper cites SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.949611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.949611Z digest=sha256:e24ec97735c919b62bf5b05c373c0f5ef00df28daae06f88eceed2a6686f9adf

Observation a7f48f30-1d70-4f13-99b9-16df11f82b7d · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Efficient Training of Audio Transformers with Patchout

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.953788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.953788Z digest=sha256:93bd0011b0012af8a36cb4ccb49e69fa3297df1b47408919ae64bd538f2398b3

Observation e13e77da-de29-4894-a1f3-33bcdf016491 · outbound

This paper cites PerTok: Expressive Encoding and Modeling of Symbolic Musical Ideas and Variations.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models PerTok: Expressive Encoding and Modeling of Symbolic Musical Ideas and Variations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.957983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.957983Z digest=sha256:6096eff1ff7fc19b3f7fb5d52a8fa83a2699ec941fda1cf66b14f6ff6f1050c8

Observation 1f7e1e4f-e84d-4d2f-83ce-0fc66c659fb9 · outbound

This paper cites Content-based Controls For Music Large Language Modeling.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Content-based Controls For Music Large Language Modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.962109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.962109Z digest=sha256:cbab5708fcc5f8f611aa105a56775e43eef607252303e73ece75b7a53e518e7c

Observation 1d9d8692-1128-4c04-a4e4-833456fff121 · outbound

This paper cites Arrange, Inpaint, and Refine: Steerable Long-term Music Audio Generation and Editing via Content-based Controls.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Arrange, Inpaint, and Refine: Steerable Long-term Music Audio Generation and Editing via Content-based Controls

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.966629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.966629Z digest=sha256:eeca740c3c441a6cda5075afdd172b45e673fa614f4c02a61644934a9be8293e

Observation 6e8c9f1b-e2b0-4c2d-9d5a-ed0418f56f1d · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.970735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.970735Z digest=sha256:a893e46d2020311c233495ecf7eb13ee3ab0d3c7708da9b1f0e3e1bef336d8f0

Observation 7f74cdce-4863-49dc-8d77-68473f76c2d0 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.974840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.974840Z digest=sha256:da7d89d9363872db5bc6afb043eb089f96d820288cfa79ef5a44fc81dc55bd66

Observation d54e9a4e-c430-4d5f-bba0-18693ff386d3 · outbound

This paper cites Novice-AI music co-creation via AI-steering tools for deep generative models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Novice-AI music co-creation via AI-steering tools for deep generative models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.252792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:17.979035Z digest=sha256:971012528ac51f45154de65182a2f4f4327c6ea60543866a6ecc60ce8137984c

Observation 3ae3d660-10ae-42dd-be2d-40be5d851534 · outbound

This paper cites MuseCoco: Generating Symbolic Music from Text.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models MuseCoco: Generating Symbolic Music from Text

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.982833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.982833Z digest=sha256:cc38b721ddbeef8fa8f7060afc6352a9a9dffecb5a1fb9eb8bd33dfd97daebdf

Observation 45537720-6e91-4814-9d5e-00371d492646 · outbound

This paper cites Learning Disentangled Representations of Timbre and Pitch for Musical Instrument Sounds Using Gaussian Mixture Variational Autoencoders.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Learning Disentangled Representations of Timbre and Pitch for Musical Instrument Sounds Using Gaussian Mixture Variational Autoencoders

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.986909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.986909Z digest=sha256:99f2a7da24097a5d4e84272d40a8d7b1e1ed7f860443b2efc1c23449777d58b2

Observation 1f262fcf-3f34-4c01-86fd-2ff4870db5d6 · outbound

This paper cites The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.991759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.991759Z digest=sha256:403e2f0b2aa7e97590387f8d8d1cb2bd0c246945d180c2336bd47412b6aac463

Observation 6bec592a-7b1c-4b21-ad97-ddd546f84ca9 · outbound

This paper cites Cutting music source separation some Slakh: A dataset to study the impact of training data quality and quantity.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Cutting music source separation some Slakh: A dataset to study the impact of training data quality and quantity

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.240664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:17.996053Z digest=sha256:41d7fafde0eedc7432f30b5afdcb217f52451e4c5ac179113a39b658f5759998

Observation e2c73553-2e83-4041-9037-efc495714b3e · outbound

This paper cites 2019.8937170.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models 2019.8937170

Reference 41

Resolution
malformed identifier
no resolver link, observed 2026-08-12T17:22:18.000506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.000506Z digest=sha256:056cdcc24c20dc9311f333cd4a36467ac40d6f6438c5bbad3476c2b15a5e67fd

Observation 2b7d0675-2c8f-4c2f-9d98-1e41803c0269 · outbound

This paper cites Stem- gen: A music generation model that listens.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Stem- gen: A music generation model that listens

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.228074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:18.012749Z digest=sha256:ce3a82a83fd55b79462dad0ae8bf68743e3773397570fe374e45ff82c2bc8336

Observation 94ce5295-f9f1-49ab-9bfe-e085ba0f4556 · outbound

This paper cites Zero-shot image-to-image translation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Zero-shot image-to-image translation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.216388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:18.016660Z digest=sha256:04c37b7d7debb2e665c5b88102d76e943f306037bcea2722b0219efa38609341

Observation 1e186346-0f81-47ff-b771-832e2a9a2092 · outbound

This paper cites Glove: Global vectors for word representation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Glove: Global vectors for word representation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.203529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:18.020738Z digest=sha256:7f412a114a859f5a0e73f2cc7dd19465b8e67d5c71c87cbc00e5a439f07872e2

Observation 6505c8e3-76a5-42cf-97d2-7044012f2901 · outbound

This paper cites Deep contextualized word representations.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Deep contextualized word representations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.033393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.033393Z digest=sha256:cec589e9caff2613e08acd263760e01efbcf64938c36a60fd9d8cfd648af34e8

Observation da217417-315c-4b64-89c0-cda45a3910bb · outbound

This paper cites Generalized multi-source inference for text conditioned music diffusion models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Generalized multi-source inference for text conditioned music diffusion models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.171499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:18.037046Z digest=sha256:16a3ff3c1069a0f7fc8ee4f9ffb519bddeac7e916e4a542bd81dab4202924c18

Observation 72792de0-2abd-4dfd-ba04-aea455bc04d8 · outbound

This paper cites Paguri: a user experience study of creative interaction with text-to-music models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Paguri: a user experience study of creative interaction with text-to-music models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.041327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.041327Z digest=sha256:b78520e6da18a60915e4274b724e17ec4e396570bf2fce2b10f3d7cb9a0ea9e6

Observation a8511da5-b9d9-409a-846c-139d67be53a5 · outbound

This paper cites Audio Conditioning for Music Generation via Discrete Bottleneck Features.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Audio Conditioning for Music Generation via Discrete Bottleneck Features

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.045976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.045976Z digest=sha256:f7d77ec120f4558868502eb74192139abee2c194906ddd4709bf6b45863b5e13

Observation c3572c32-6a17-4956-85e2-3fb128887240 · outbound

This paper cites an unresolved cited work.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:22:19.157232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:18.050497Z digest=sha256:a026b10e9b9dc83bee585f796ae344a59da6852b1705f3512b6b8f0525080886

Observation b6032a99-19c9-43ad-a077-09b2585925e6 · outbound

This paper cites URL https://doi.org/10.1109/ICASSP.2019.8683855.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models URL https://doi.org/10.1109/ICASSP.2019.8683855

Reference 54

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T17:22:18.502761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:18.062013Z digest=sha256:fb184de52d3accec86d2ebe285d42e3806192cf98c9dd3ecf0d172595febd156

Observation 2f28cfb2-25c1-4eb9-8f49-b09ede1b1942 · outbound

This paper cites Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.066712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.066712Z digest=sha256:2b8feaf324e8da15268a5efff576bd36fc0dbe1375f1db78180d1334b261193c

Observation bee62ef0-c927-4586-b0fa-f3c5a1a85f37 · outbound

This paper cites Music FaderNets: Controllable Music Generation Based On High-Level Features via Low-Level Feature Modelling.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Music FaderNets: Controllable Music Generation Based On High-Level Features via Low-Level Feature Modelling

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.071838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.071838Z digest=sha256:e0bde6395d574c1318a8c42353d7139da44d045d0ea500c43886b9df2e3b692b

Observation 4ae40fbd-b2da-456a-8d2d-f086b3d3373c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models LLaMA: Open and Efficient Foundation Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.077072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.077072Z digest=sha256:061ffbdb546f285d26e74382e47c33b7a7a35cec88559180e5902a85acb2a389

Observation 36568fc4-9e0d-42c1-8a3e-651dd21d23c3 · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Plug-and-play diffusion features for text-driven image-to-image translation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.143365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:18.081609Z digest=sha256:69d3bea8d5720328831d139c13ca6c6ddd5e25e7075bce3d76d7010692c37d64

Observation 5f8e728e-061b-4156-9bd6-00b3c5458970 · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models WaveNet: A Generative Model for Raw Audio

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.085750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.085750Z digest=sha256:74ddb274ad3e331cd4ec8eed9f31bff16f833cb03984c6b91c4138e3368ece93

Observation e141c705-cb91-474c-a455-2ec8c01301d8 · outbound

This paper cites Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.094014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.094014Z digest=sha256:8d63c91a69aa94e74bdd5d4bf7ab102cbb759346094945bb9f6bb09caa01ea1c

Observation 7fa745de-7c50-45fe-b3b1-1506c629e7b0 · outbound

This paper cites Exploring the Efficacy of Pre-trained Checkpoints in Text-to-Music Generation Task.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Exploring the Efficacy of Pre-trained Checkpoints in Text-to-Music Generation Task

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.098324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.098324Z digest=sha256:282ad05a625434ad9623abb251135147a9763fccc0c92f03083e1b046f35822b

Observation 5b21ed84-4989-4f8d-8f77-4858ca1e080e · outbound

This paper cites Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.102619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.102619Z digest=sha256:4781720971244fcb329f19ddec01cbf3d0feee4df77a6783105159a1da53d2e4

Observation deaafc50-5a0a-4bd1-8afd-eccaff0c3403 · outbound

This paper cites Iteratta: An interface for exploring both text prompts and audio priors in generating music with text-to-audio models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Iteratta: An interface for exploring both text prompts and audio priors in generating music with text-to-audio models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.129595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:18.106545Z digest=sha256:6bc825c0739eed000d7fa13785ac570b6777483f58c01e2bf6f6293b9e2ba787

Observation 2703ed64-5d4f-4d4a-a51c-7e9cbb6c0e02 · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.110641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.110641Z digest=sha256:7f38efc793551daaef800f541b74d4aca2ec03eb2301853565c6ce524a23e60d

Observation d7d3a452-f80f-4be2-879c-54b11508117a · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.114817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.114817Z digest=sha256:43b681e86eaedac7f2a74289dab12fba107ff67d687a8403965b4d54b690febe

Observation 12264b63-e1f6-41f7-bfb4-7b46197d5f43 · outbound

This paper cites JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-12T17:22:18.262901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:18.119324Z digest=sha256:0ef4f0426c70887d45c5ee67d402eb0f281cca31b68bb653724639b24900cacf

Observation 7aaae18c-fb86-4487-a150-aefcaa3649a6 · outbound

This paper cites ChatMusician: Understanding and Generating Music Intrinsically with LLM.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models ChatMusician: Understanding and Generating Music Intrinsically with LLM

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.123882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.123882Z digest=sha256:e34636bfa843f31f51c835b5f4d8571d6640d4886836ead30368ad9971260246

Observation 3dedc977-28b4-44cd-a5de-2858df3024b6 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.128284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.128284Z digest=sha256:f44481a63585bb0529527aa66257ecd2ddfc0f1230662da7d9498192c764ed4f

Observation 7d503132-40d5-45a0-8a85-75d23867d85b · outbound

This paper cites Cosmic: A conversational interface for human-ai music co-creation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Cosmic: A conversational interface for human-ai music co-creation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.115111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:18.132747Z digest=sha256:ac8c8e062a3b4dddc16c7d439b59fec3e6dda72a3e75d73d65b9c7600c1dab45

Observation 28009de3-3aab-4a83-b2f9-07d17531f7d8 · outbound

This paper cites Interpreting Song Lyrics with an Audio-Informed Pre-trained Language Model.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Interpreting Song Lyrics with an Audio-Informed Pre-trained Language Model

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.136706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.136706Z digest=sha256:4085b11f3d5ab06c94e07a125e3c9712569963500cdc526b4dcaa02a23951c4c

Observation 4c418d10-b949-4159-add6-a74d84190332 · outbound

This paper cites A Survey of Large Language Models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models A Survey of Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.141425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.141425Z digest=sha256:491288338fd8e514545274bab557afbc980c8455e3a156bc803657f52efd076a

Observation b00db70a-dfd9-42df-9df6-e87647de670c · outbound

This paper cites Masked Audio Generation using a Single Non-Autoregressive Transformer.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Masked Audio Generation using a Single Non-Autoregressive Transformer

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.146028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.146028Z digest=sha256:e8101f3ee311e9a5304443132e5af4b6bed0f08d0f1f0473b234577fd1629271

Observation f06bdcae-e553-4106-812e-acbba9ecca99 · outbound

This paper cites Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models

Reference 1986

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.008497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.008497Z digest=sha256:9571a82117a33f7562438854318abff492c8cd7bb0a704a7e237ac09b1807da6

Observation a57bb2bc-47fd-4c40-afeb-66e9746f6412 · outbound

This paper cites Exploring XAI for the Arts: Explaining Latent Space in Generative Music.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Exploring XAI for the Arts: Explaining Latent Space in Generative Music

Reference 1996

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.829212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.829212Z digest=sha256:f7aab5addcab9d0b345a57100fa2fb196c45b067940bfa1398ca2269210492fc

Observation 6bbe475b-5897-4987-86d2-2ab251e89d31 · outbound

This paper cites 2003.819861.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models 2003.819861

Reference 2004

Resolution
malformed identifier
no resolver link, observed 2026-08-12T17:22:18.090093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.090093Z digest=sha256:e425b84a649cdb97b4293a36ef7b72652a078ba357decbb9f4cc07e56420fa62

Observation 370d21ad-03b8-494f-a600-fbf4f2ef17dc · outbound

This paper cites Improving Text-To-Audio Models with Synthetic Captions.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Improving Text-To-Audio Models with Synthetic Captions

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.945138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.945138Z digest=sha256:c9767209460211eb6c15bf2fa67b772215ff572a63b2634c9fa2e3b998a47aac

Observation d803a4f9-532f-4a98-844a-4347d9fe9899 · outbound

This paper cites MoisesDB: A dataset for source separation beyond 4-stems.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models MoisesDB: A dataset for source separation beyond 4-stems

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.189399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:18.024699Z digest=sha256:f3007a0866bd21e011da0cb7adbfef7eeeda185a49265a7748153c240bf71d23

Observation b6bd2d2e-5f02-4445-ac86-2279b1482464 · outbound

This paper cites MidiCaps: A large-scale MIDI dataset with text captions.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models MidiCaps: A large-scale MIDI dataset with text captions

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.004318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.004318Z digest=sha256:e8a80298068d4d3e0e07d76066dc07343b9506d75075bd89b5e1930b6b2cc59b

Observation 91e7ac71-eccd-4b9e-a718-16e7989b14e1 · outbound

This paper cites A Comprehensive Survey on Deep Music Generation: Multi-level Representations, Algorithms, Evaluations, and Future Directions.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models A Comprehensive Survey on Deep Music Generation: Multi-level Representations, Algorithms, Evaluations, and Future Directions

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.923559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.923559Z digest=sha256:040c4247f086c9be89b3e3902942266f4c0f029067ffe5d4682cca5dff0cb09d

Observation bf5cb6ad-9aa4-430d-ae8a-95a11b661c7a · outbound

This paper cites Au- diocaps: Generating captions for audios in the wild.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Au- diocaps: Generating captions for audios in the wild

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.265481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:17.936966Z digest=sha256:2181009ea448c2962c4e04b1609f1fce1ec932f3c78a0675439f9ea7c053aa3c

Observation 0035bc81-1603-4cf1-8b24-b65ea1e86213 · outbound

This paper cites MuLan: A Joint Embedding of Music Audio and Natural Language.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.910125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.910125Z digest=sha256:22aa461eda87ee5849caf8cdb220a57a777491e0ba65ad5969d006c49b6afbe3

Observation ca295f9c-16c9-4a67-9d7e-c3ddf8763784 · outbound

This paper cites SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.861124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.861124Z digest=sha256:efc21fe1f0d1d9d2145bdbfc1c4be2ce7b4f7439284c13a8f7dc55dc91520f01

Observation 08a68021-b2fe-4f82-9568-cfe8712a4d44 · outbound

This paper cites Jukebox: A Generative Model for Music.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Jukebox: A Generative Model for Music

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.855404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.855404Z digest=sha256:7ca99b1fa21e5eca17fd67aa562823e3d4d0fa0bf8bd5771f8a07b4e46454b10

Observation 5decd465-1463-4b20-b74a-8ff02ef74acb · outbound

This paper cites Musicldm: Enhancing novelty in text- to-music generation using beat-synchronous mixup strategies.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Musicldm: Enhancing novelty in text- to-music generation using beat-synchronous mixup strategies

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.343974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:17.838339Z digest=sha256:e062abbac9fdf1c3cd37508409a72505aacaff934acad53916da28a8e329db80

Observation 56994326-27b9-4dee-bca5-b81e86bbbaf9 · outbound

This paper cites Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.355936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:17.833825Z digest=sha256:ade03a932546ba65ff86a6efe07f10ba167f9cdc6320c4b2fd9e594d0c247c6d

Observation 0a50e756-deba-4490-bf58-6fdf0d6ea5b5 · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.332606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T17:22:17.842502Z digest=sha256:d9a97145d8c52b377dc51e686265272a5f185600b982b59425b22e711205458f

Pith citing papers

No inbound Pith citation observations are available.