Pith. sign in

Paper Citation Record · LEDGER

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning

As of 11 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2501.14680.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.14680 v2

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:59:13.413269Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 566a6805-e497-41c4-86bf-0f81baf37c77 · outbound

This paper cites Denoising diffusion probabilistic models,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Denoising diffusion probabilistic models,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:13.226462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:13.226462Z digest=sha256:34c94e4d794363e4281f11e9786bfb6d495868490253a37313fc3a307994104a

Observation ee1aee23-d93d-4c61-b40e-210b61353d5a · outbound

This paper cites Attention is all you need,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Attention is all you need,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:13.231462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:13.231462Z digest=sha256:1f27818c404d2f7bed0e123409c6a5fbd48c494ec20f1c8821cb6de466a6c990

Observation 966344f7-3780-459f-b958-da3755165e1f · outbound

This paper cites High-resolution image synthesis with latent diffusion models,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning High-resolution image synthesis with latent diffusion models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.949050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.237567Z digest=sha256:b3ff64face9b2d8f70a2bc849a3b805f7887bd3126d74f8b83973789f251eadf

Observation ffb23d7d-3ea3-4714-b2d0-77e5973346b2 · outbound

This paper cites AudioLDM: Text-to-audio generation with latent diffusion models,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning AudioLDM: Text-to-audio generation with latent diffusion models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.937191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.243347Z digest=sha256:bc53a08f9c2150220544ab4f7e88f4e33cb6749f917fe9fbb3e0b66dea8d2006

Observation 1ea920f5-f649-4d73-bd31-7294ce13535f · outbound

This paper cites AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.924064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.247862Z digest=sha256:47dedd5742ccd3480a84918b04db4aa603e0a4244fe03009848c8075a518414f

Observation f61aed68-8741-4be0-bd98-592b7aa181bc · outbound

This paper cites Fast Timing-Conditioned Latent Audio Diffusion.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Fast Timing-Conditioned Latent Audio Diffusion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:13.252059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:13.252059Z digest=sha256:fd3dff0bd4c219f920a60591fab44c0742fb840084b5bd242ef0a3be363e24f0

Observation 47cf95b6-e3a7-4277-abf1-539d7ff7656c · outbound

This paper cites Grad- TTS: A diffusion probabilistic model for text-to-speech,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Grad- TTS: A diffusion probabilistic model for text-to-speech,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.910987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.257841Z digest=sha256:91fb2ecf6a2d69ae253c7134bcda7d2b1183bee3bfc8bba36c36b6484c5ee159

Observation 283bd3a9-de8c-4763-b624-c8fe25a07278 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Learning transferable visual models from natural language supervision,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.898448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.261630Z digest=sha256:d6bf4216a239900b9abdde216ecf2ed5afd8dfddcb801737049d6f0a3b64ba48

Observation 5e8729c8-8429-4fc4-99da-92bfcde07589 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.885498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.265355Z digest=sha256:413885d08bea44133b427be162387208f2e6a2def6122eac876acf4f15808dcf

Observation 3160c0eb-c05e-459a-b0fa-8ec633876133 · outbound

This paper cites MuLan: A joint embedding of music audio and natural language,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning MuLan: A joint embedding of music audio and natural language,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.873273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.268900Z digest=sha256:0b755eaa2d9eb36673014e4ad330a5eb254a946ddd9218e344517050ddfd712d

Observation 902b2cc3-17b9-4d6d-9cde-b8918855bf36 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.860511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.272842Z digest=sha256:4b9e5d653ce65f10f05238d5d88ba94220513c58df0c7ea4728fe2e3a063116a

Observation 0776d894-663a-481a-8e79-878116b01567 · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning BERT: Pre-training of deep bidirectional transformers for language understanding,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:13.276656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:13.276656Z digest=sha256:beed8b15f50c96396df49e0615b036a66f78ad8ff3cc8da5f0a008196370db6c

Observation f5aad878-1222-40dc-8dce-9ece449697d7 · outbound

This paper cites MusicLM: Generating Music From Text.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning MusicLM: Generating Music From Text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:13.280766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:13.280766Z digest=sha256:139193aa7d84d47c12ffda38878522a3a7f5e160073c1c5e1284761dd2edc403

Observation 929e3b7e-73e3-4564-82bb-5f0724e54828 · outbound

This paper cites Audio-Text models do not yet leverage natural language,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Audio-Text models do not yet leverage natural language,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.841488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.285632Z digest=sha256:03d0997264df2e0fdd6e5c0b348cd44a3960811f79041138490ee12a8baf8e23

Observation f197e251-1377-44db-9ba4-6107f6aa9036 · outbound

This paper cites Simple and controllable music generation,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Simple and controllable music generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.828233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.290069Z digest=sha256:ad9d76acd923f31813dced589c9a500a2001ea263beaf517f491411341858ae7

Observation 3fcc17df-b228-41c0-95d1-6a3e309169d7 · outbound

This paper cites T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:13.293802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:13.293802Z digest=sha256:366aefa8dccfe57b7cd97edcdd49c063852a45002b9c248fe9872702b973f7f5

Observation acd9d5b1-9e68-4c48-9885-16227857af87 · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:13.299055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:13.299055Z digest=sha256:12dae841582a2caa1c873e36b4b433140a859eeaa45212a99763c22a9c218265

Observation 7f63d4d4-9f04-49f0-a456-e3488eb8d161 · outbound

This paper cites Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:13.303276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:13.303276Z digest=sha256:de780cf57cde46e66e28a4625d495e60b6285a0097de6405f479f8275bbeaccd

Observation 35c20cb4-a0b8-49c4-906f-ee4c6615b9af · outbound

This paper cites Scaling instruction-finetuned language models,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Scaling instruction-finetuned language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.815112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.308605Z digest=sha256:3396481482223d2bf81287852dc508a69142830514214f5157756996342e9cf4

Observation cb11ae3a-d7c1-49ff-8c19-59fd83b599cf · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.803876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.312383Z digest=sha256:c09cd79f7aa0716142e3ec3f9d8b051a18c223d415dc62f2fe181327bc68b739

Observation 74a009b1-1e61-42d2-be28-81a8aa34ae0e · outbound

This paper cites Sentence-T5: Scalable sentence encoders from pre-trained text-to-text models,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Sentence-T5: Scalable sentence encoders from pre-trained text-to-text models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.791721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.316545Z digest=sha256:85a33b0854a3156c5e7740adc17639f53ddba3ea4a7b480a5aa1458029283356

Observation b914bd54-6e1c-478d-b1ce-7810777e15b3 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:13.321116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:13.321116Z digest=sha256:16eab062262e440b36ce3219f593c1b701895a5860b1bb2e407d52c71269382b

Observation 0f89a955-997d-46cd-bcd9-db00163ef125 · outbound

This paper cites Whitening Sentence Representations for Better Semantics and Faster Retrieval.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Whitening Sentence Representations for Better Semantics and Faster Retrieval

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:13.325103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:13.325103Z digest=sha256:8f5137c7c2d4563573d7da4a6a86daa0ca75fc65ce1e605f93825cb67b89edbf

Observation 2ae17606-0a6b-4b87-96a7-2d04485d714e · outbound

This paper cites On the sentence embeddings from pre-trained language models,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning On the sentence embeddings from pre-trained language models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.780308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.329613Z digest=sha256:72187a678f478e4df79ffa1f9bc9587291e19e5d5158e5705873c1955bfdd4f8

Observation a221d86c-6487-4f1a-832e-9caeea2f7ffa · outbound

This paper cites SimCSE: Simple contrastive learning of sentence embeddings,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning SimCSE: Simple contrastive learning of sentence embeddings,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.768717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.333615Z digest=sha256:c650fddb6b8ad228f2b7d2859b5b1c31909cc8e98e9640949f7195d62252ae61

Observation 860d36a4-6d5d-4059-815f-5ddf185f3a97 · outbound

This paper cites FiLM: Visual reasoning with a general conditioning layer,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning FiLM: Visual reasoning with a general conditioning layer,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.757072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.338536Z digest=sha256:6d377b9f8a854cd997e1b337512433255009f1020fe753edfdc2763bde1cf0e1

Observation 758ceddd-7816-45e9-81a8-96cb144ecea3 · outbound

This paper cites Self-attention encoding and pooling for speaker recognition,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Self-attention encoding and pooling for speaker recognition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.746245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.342375Z digest=sha256:270dcc15c27bfd2627ff893f85168e44754507dfe9533a06b041a836f8146e22

Observation 3a3d6c11-7b22-4857-afc5-00bc06521ecc · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Progressive Distillation for Fast Sampling of Diffusion Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:13.346143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:13.346143Z digest=sha256:b36f0b11d10cc230cbb7ddd6b21da0ef428e1d1cedb200adb553a575662a8973

Observation 11db5f13-3cbd-4732-b402-6d7f233250eb · outbound

This paper cites Variational diffusion models,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Variational diffusion models,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.733767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.350325Z digest=sha256:1c1897d6133bfc3ad46e7eba37e742ec10a573087c9a6b955c32e753d5f5a012

Observation 9af5f36e-2048-4843-9127-d4deb2529023 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Classifier-Free Diffusion Guidance

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:13.354010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:13.354010Z digest=sha256:b5b479d4f0d30c57095eb8b1ec3d32f4cafba8301f7d10eca999174d2250c7d7

Observation e3d21dee-0e18-4b06-a7ba-4f6e64ec2b5c · outbound

This paper cites The MTG-Jamendo dataset for automatic music tagging,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning The MTG-Jamendo dataset for automatic music tagging,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.721866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.357952Z digest=sha256:bcad8497604d73c6426805ff7b25dddff9cb548cfc825d644b01e6d6e474832e

Observation 5a6eb880-7b8b-4f2d-8ef6-148e5ad6c50c · outbound

This paper cites FMA: A dataset for music analysis,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning FMA: A dataset for music analysis,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.710277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.361577Z digest=sha256:217d88b7fa74c29c4fb7db42ac9eb587afb362ed95a67a2fe2fff5707394726c

Observation d3f9f4c4-3504-4c0a-93fa-6e0bed2dce2f · outbound

This paper cites an unresolved cited work.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:59:13.696574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.365285Z digest=sha256:11970f55dad3874748481cb1374ec80bd1402d4046dd573c7cdbaf18428899a0

Observation 0e4b77eb-f779-4c16-bbab-3acac67ab291 · outbound

This paper cites Hybrid transformers for music source separation,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Hybrid transformers for music source separation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.684095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.368689Z digest=sha256:cd29a1b31126bcc5ca22320810f51a102b79680847db7de80c1c4de7896261b2

Observation ae4fa1a8-10e2-45e4-818b-2323d5a6c4e8 · outbound

This paper cites Mustango: Toward controllable text-to-music generation,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Mustango: Toward controllable text-to-music generation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.671617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.372334Z digest=sha256:eaccfc4471d0654b4ea3f1dcd2f835310a3ce6625c7907e0586c371c3dc32f9a

Observation 02fd6909-8ce7-4ff8-9d79-a7315916f4e4 · outbound

This paper cites AudioLDM training, finetuning, inference and evaluation.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning AudioLDM training, finetuning, inference and evaluation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.660452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.376244Z digest=sha256:e5eb5b60635530f4194f5179bddf2d9b07e0231301774bf37566c86ef267e2da

Observation 0b17c31f-fd4c-4ddc-abc1-1d2f4979e6de · outbound

This paper cites LAION-AI/CLAP.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning LAION-AI/CLAP

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.648992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.379968Z digest=sha256:3b264d05a6e549d43093b6fd0bc633e8a185499a0203db63bcae1569b93b825a

Observation 2dc6ea61-dc3e-47dd-9904-901eb5c6d0d8 · outbound

This paper cites Decoupled weight decay regularization,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Decoupled weight decay regularization,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:13.383522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:13.383522Z digest=sha256:38775ae3fc64cf734487dce90e4ce3d9907b12c25cb53ea0d27de34526b7f24d

Observation f46d9230-614a-4a67-887b-75909a96cc60 · outbound

This paper cites Denoising diffusion implicit models,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Denoising diffusion implicit models,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:13.387181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:13.387181Z digest=sha256:c5a689a0e745d4141242907265c11fbb72e6c7cd88539b7c358519204e55860d

Observation 255e0e09-02aa-4f52-b371-69376eddd4de · outbound

This paper cites MusicLDM: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning MusicLDM: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.622726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.390565Z digest=sha256:437e9c73bd712ee91d190a503b17a84ded2e6245da536920a81a2060784aa295

Observation 16834dba-2422-4e04-a209-04fab56c8001 · outbound

This paper cites Stable Audio Open.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Stable Audio Open

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:13.394342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:13.394342Z digest=sha256:f38617ba88f8c4bbf65f3b861ee8f15b6090fae420693718c226fa1a6b3a62a4

Observation b2deda01-6e87-4709-9ca8-374ec85eb6cc · outbound

This paper cites Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.609315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.398451Z digest=sha256:ad13d552031f711e5b558994220e0a64a1592115fdb224121bb7d297e71d84ff

Observation ce8ed845-9eb0-4e2a-a345-f59323737667 · outbound

This paper cites Audio generation evaluation.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Audio generation evaluation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.597342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.402113Z digest=sha256:34fb2ac4aa7dfebaac33d8f6f6625ed6b6f776b2e33cada0e73eb7bba8916612

Observation 727c1c7e-0988-4d59-8e07-eedfc6b08e08 · outbound

This paper cites CNN architectures for large-scale audio classification,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning CNN architectures for large-scale audio classification,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.584621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.405811Z digest=sha256:b11e1a40c1095d290ac6aad843fa38760c481fba078acafef00fa1a8b616cf09

Observation 95c438e2-aca3-4bc5-aecc-dbf2e809f8c4 · outbound

This paper cites PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.572898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.409490Z digest=sha256:344cab5aae3b11423caa7600f82e9b25cbbd7ee1e5a78760394a6fddd1b8730a

Observation e4f311e3-c33e-4515-bc6c-705ae7dd259a · outbound

This paper cites CoLLAT: On adding fine-grained audio understanding to language models using token-level locked-language tuning,.

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning CoLLAT: On adding fine-grained audio understanding to language models using token-level locked-language tuning,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:13.560633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:59:13.413269Z digest=sha256:df28598f2d532490878b614c719c645267296d31a57dac6a58f0a6918fa78188

Pith citing papers

No inbound Pith citation observations are available.