Pith. sign in

Paper Citation Record · LEDGER

AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 67 inbound Pith citation observations for arXiv:2301.12503.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.12503 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 67 of 67 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:26:46.757326Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:59:50.469153Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4244a785-d3d1-4482-8f44-0db941ed88c7 · inbound

DGSNA: Dynamic Generative Scene-based Noise Addition method cites this paper.

DGSNA: Dynamic Generative Scene-based Noise Addition method AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:45:46.270816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T17:43:47.086524Z digest=sha256:240d11a9ca50ac7a65afd2440bd0d67de3155d52a7e90ef22cff38fea9002157

Observation 2f0e8708-68cd-4cad-9cde-f14bd224c3b7 · inbound

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT cites this paper.

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T14:26:46.757326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:26:46.757326Z digest=sha256:aabb08cedc70ddbf5329fa65856f0ca0b9b02f7b8a5dc0ac4cd1a1e9271777bd

Observation ead1e612-6841-4a2b-817b-ab9faebc6d7d · inbound

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet cites this paper.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:57.138047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:57.138047Z digest=sha256:c25ddf8f114fae8b16ba5b7ed6857a3f914d1fb3f62c0f34071570269546eb67

Observation 9d153604-3a49-4782-899a-23fbe1ff70e2 · inbound

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation cites this paper.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:55.918085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:55.918085Z digest=sha256:1a1712a70164e984e30c6b1ca597545681b874d495e1118d24c2c3f2fef1ca73

Observation 5cd6f166-9bb0-45a3-93d8-cf787e908f48 · inbound

In-the-wild Audio Spatialization with Flexible Text-guided Localization cites this paper.

In-the-wild Audio Spatialization with Flexible Text-guided Localization AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:48.144450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:59:48.144450Z digest=sha256:90d73bd96fbb663e683f61ca5fe95d724222934582a741b8a048cee2d941c690

Observation 833355d6-6bc7-45e7-b768-abac05b45178 · inbound

How Far Are We from Generating Missing Modalities with Foundation Models? cites this paper.

How Far Are We from Generating Missing Modalities with Foundation Models? AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:15:33.620840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T08:15:12.947854Z digest=sha256:8d32852ab0ed2f33c26997c0147c050315a6ea931ee987ac756d306d2249cd67

Observation ad315d80-4b72-4627-83d5-86978052c1e7 · inbound

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction cites this paper.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.194385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.194385Z digest=sha256:35e9d33411c5141db2b3d305152236e4d3532a5da5262944d92ab70002bc501f

Observation 87fd9334-847c-4c91-abcb-77ad8bf938a9 · inbound

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching cites this paper.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.973851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.973851Z digest=sha256:9a06e81066cc0f562929a49a768b3fa0932e204bb3db028fd7ed142430bd2921

Observation 401ae98b-4024-41b3-ad2e-c5b3acd1d977 · inbound

Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models cites this paper.

Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:58:34.577052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:58:34.577052Z digest=sha256:d2edb3bce0a56141ec782f050897f1f8083a29450fff36082af991cd2dda80f7

Observation 3384e3c5-011a-4867-b15e-0451f2ba050a · inbound

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation cites this paper.

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:39:54.589254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:39:54.589254Z digest=sha256:2094dd1db4ab3edb6a52ed7f56b0a34ee25c67c99faa8e0d77743cf2fb454381

Observation 8afa710d-257f-449e-a116-60b3e028de26 · inbound

ADMC: Attention-based Diffusion Model for Missing Modalities Feature Completion cites this paper.

ADMC: Attention-based Diffusion Model for Missing Modalities Feature Completion AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:42.399897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:42.399897Z digest=sha256:31274dd398f2637c117ed331da451cfba5890fccccecc3902a5677f91fe2c19d

Observation c1a90ca5-8730-466f-96ed-44324b7d35cd · inbound

Diffusion Models for Time Series Forecasting: A Survey cites this paper.

Diffusion Models for Time Series Forecasting: A Survey AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T15:58:50.494374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:58:50.494374Z digest=sha256:b6241d017aba3bab7dea3ebf480eccc7998665023914d30862c3d643b16a887a

Observation 9139e409-ca8c-43d2-9ce7-0a57256d1a80 · inbound

CHORDS: Diffusion Sampling Accelerator with Multi-core Hierarchical ODE Solvers cites this paper.

CHORDS: Diffusion Sampling Accelerator with Multi-core Hierarchical ODE Solvers AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:47:02.687551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:47:02.687551Z digest=sha256:95d8e43675989e523bda196b6d3e099b9e5413c3d347a0fec99a6a4b5ad36046

Observation 035f79e7-9d79-4d62-83f6-e690c2962950 · inbound

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations cites this paper.

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:04:00.776212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:04:00.776212Z digest=sha256:ad1a227dfce5893c3751757fbab8455a365ce73f6e50f99dff00079c1e2d7cfe

Observation 511d80a0-aa90-46af-93dc-954e0950af30 · inbound

Flow Matching Policy Gradients cites this paper.

Flow Matching Policy Gradients AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.495265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.495265Z digest=sha256:e0082c7011eff5a981061f005ccb3ceadd91053dcd655c3c808358b08302d337

Observation d212bfcf-8c24-4765-b564-a318947e4b00 · inbound

Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs cites this paper.

Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:19:01.944976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:19:01.944976Z digest=sha256:3abb5cb755e89f6bf646230562d30024bb09283442612606e6c709e89825dc16

Observation eae0bf53-89bc-4d9b-adb5-4745a5dce62d · inbound

Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning cites this paper.

Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T11:14:46.151866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:14:46.151866Z digest=sha256:7c00a1d8138f2e8441b926cb94137985a78dd7a61f2b2eec4fa460322b419397

Observation 1d75771a-4415-4798-b175-2d463d71af6d · inbound

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation cites this paper.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.913997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.913997Z digest=sha256:f99f9838561d34f4dea01b2525977224e0283509cc7a7b8c0380fd496955f089

Observation 64148a22-c14a-43ff-a71a-acb113314683 · inbound

Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation cites this paper.

Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:56.915651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:40:56.915651Z digest=sha256:5fa0a656a67c5265c1f168ff1c8d71489986d66a32ebcafa96ae389c7a71dd69

Observation e0fee5c5-370f-447c-94d1-65406a061189 · inbound

Inference-time Scaling for Diffusion-based Audio Super-resolution cites this paper.

Inference-time Scaling for Diffusion-based Audio Super-resolution AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.134039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.134039Z digest=sha256:4fddb1074d2b7ddc22b59a375ca1aec2d8d2a43ead6b921983017c781c169060

Observation 986d2272-7240-4034-913a-4c095f3bd0ca · inbound

ASAudio: A Survey of Advanced Spatial Audio Research cites this paper.

ASAudio: A Survey of Advanced Spatial Audio Research AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-05T22:54:55.090118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:54:55.090118Z digest=sha256:f167fe635b0dc15a235fe54e48348ff937150e7077f78c3f7e26add957e0f1dc

Observation b40285ca-fcc7-4c76-86fb-d4cc7f16590e · inbound

A Sharp KL-Convergence Analysis for Diffusion Models under Minimal Assumptions cites this paper.

A Sharp KL-Convergence Analysis for Diffusion Models under Minimal Assumptions AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:10.648281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:10.648281Z digest=sha256:43801ede7a0b742bf82f1ce21381ca33106d19181118148b77bd8f4055eb6cc2

Observation b0ef9092-7b6a-45a5-a957-900a8070585c · inbound

Audio-Guided Visual Editing with Complex Multi-Modal Prompts cites this paper.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.753877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.753877Z digest=sha256:2b273ca11d0813ce217817fd58101b2a24f17fd2f9fd179bdc2ff58808cbf848

Observation 01204637-e03e-46bd-86aa-ef69bf55b79c · inbound

WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration cites this paper.

WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T14:36:37.617270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:36:37.617270Z digest=sha256:8d53de6f34288dcc3d0c85dc33b1c7cd313813e5e1e8ca28a65977a4904ab79a

Observation 29112a42-28a7-48ca-a089-aee4158f9105 · inbound

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance cites this paper.

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.790250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T13:10:18.700497Z digest=sha256:1e5c376215e8aece45a42bc641abdb3def8c725fde7552554f390ecb3099fb19

Observation cf197551-189e-4413-bfbb-75800a2a5d56 · inbound

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction cites this paper.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.198477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.198477Z digest=sha256:f05399bd0a62602a46e72551080b5d7df2fdb7543d5586d2e28de7e8524de3fd

Observation 9ba08cfd-6e72-4426-9c96-74149d070b56 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:33.883026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:33.883026Z digest=sha256:ee22a0794d3709c77f274e7eea22911d6edf7ce275355891c9ef40335c820e0c

Observation 5cab15e9-428c-48d5-8f1c-1590be20ebf8 · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:46:03.658946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T07:42:43.077644Z digest=sha256:aa7846fb56226a92d4e10f0fa818fc09cddc58156a9b21b517deaa5951dde567

Observation 7f33acef-db98-43a1-926c-d821ecfb8475 · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:00:39.130192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T20:56:37.533183Z digest=sha256:d3dce2c4d807889463f09fd16cfa17ac95e571a68f5a43e0a38f7833f12e2668

Observation 34d78cc4-a91d-4143-a1b0-bab32f6be7fa · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T09:49:43.654091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:49:43.654091Z digest=sha256:267f38b55f218a5eee65e7d96713e8d157e3bd4874968b8b65beafa8d5d33761

Observation ad9ddeab-612d-4c4d-806d-6591e5de0714 · inbound

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion cites this paper.

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:57:42.900012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:55:32.044506Z digest=sha256:7f4cc0366e0f39345d2b5dc0d3bb355f36edb0be28c7d998f0bac54a611555ab

Observation ad26e427-935b-4ab3-bf78-e31e4327eb8f · inbound

Dual-End Consistency Model cites this paper.

Dual-End Consistency Model AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:40:39.814660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:40:36.406150Z digest=sha256:03e0bc6096165e38b2225a44db63e840cd85870c5104f478228fcd4b183077c9

Observation a5fe170f-ff47-4c70-9b9d-e78aba20a207 · inbound

Dual-End Consistency Model cites this paper.

Dual-End Consistency Model AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T01:04:20.684172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:04:20.684172Z digest=sha256:5ce5cf4a111e1009086f3064d21beaa27575bc2629433f851265b564bf950424

Observation d8e4dd18-5751-40fe-bcf0-56855f426025 · inbound

Diffusion Models Memorize in Training -- and Generalize in Inference cites this paper.

Diffusion Models Memorize in Training -- and Generalize in Inference AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:54:07.907159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T10:52:31.849094Z digest=sha256:98277ec51be3270ed9b219f5682acc4b3549ff59026e94d238dc25b91be55ada

Observation 6f6fbe12-1056-481f-a394-17d714e80274 · inbound

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training cites this paper.

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T18:03:40.092528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:03:40.092528Z digest=sha256:ff2e14e147b5fbcbf25f6c5df236b57339595dc7e96ca750e13216f114565831

Observation 7d8d74b6-99e3-412b-a589-d61ee8b439db · inbound

Woosh: A Sound Effects Foundation Model cites this paper.

Woosh: A Sound Effects Foundation Model AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:15.805338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T20:51:08.144573Z digest=sha256:c99e875fd36734834e80ff9ed3f40979783536dd76f7d96a15dc92a11b59b555

Observation 01351671-3543-4f01-a0f1-0ac538ec27e2 · inbound

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips cites this paper.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.725801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:3dc304f900f7642572471669d408e4db3c40f4711919e9442494ebc811fb0823

Observation 7a567944-8d5b-41e9-8911-b1b29291eb45 · inbound

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan cites this paper.

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:21:00.633260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:40:10.657590Z digest=sha256:0bb5dad2b0fa9457dd70f35734a3c4082efd9cf2c67c59812cfd7bd5dd0fa3f6

Observation 29ea7c1a-3e59-4755-a1a7-e76c6cf6c59b · inbound

Latent Fourier Transform cites this paper.

Latent Fourier Transform AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:21:07.061004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T03:45:07.892234Z digest=sha256:33e45aa968e626e6d6495ef353a6c51e6fddc0150308cd5d1956bb4d023b5113

Observation 01797019-b68d-4ae6-ac3f-1015173d9bd7 · inbound

ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis cites this paper.

ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:11:22.315786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T02:11:12.460695Z digest=sha256:16402c53db6821da846dd24722394f297136558c2764744f7e4cd05da41402fd

Observation 39d88a97-f0b0-48b8-9ac7-7861846c0866 · inbound

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation cites this paper.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:44.540336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:fe632b458105928ce865d496a95a98c135b1ad1e8878f96c5c250b091219a479

Observation 2d7908dd-aed0-4ad3-87ba-adc684434783 · inbound

Stage-adaptive audio diffusion modeling cites this paper.

Stage-adaptive audio diffusion modeling AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:41:06.288613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T17:21:34.140699Z digest=sha256:dd05ebf2f457964e1e9ff3a0e10b9b4abadbf7e635b8e129ab3f409dc7d16e5c

Observation 9213e690-cbd2-494a-b509-1434c2bb22a9 · inbound

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems cites this paper.

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:28.002815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:58:26.634355Z digest=sha256:74ef5fe3016f1673f40305dc9698e8bd261e15545e88960772d8c7b8a17b3dd0

Observation a8c7fa61-6d59-4926-9f7a-96b99a4fd859 · inbound

DiffATS: Diffusion in Aligned Tensor Space cites this paper.

DiffATS: Diffusion in Aligned Tensor Space AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:26.589603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:38:12.633086Z digest=sha256:eec32f390e2a7dddb57864e5989e433cbb035c04c796d29f8814813bdb244fd1

Observation a5609b5b-8f1c-4de5-8d38-02ceb2028256 · inbound

HapticLDM: A Diffusion Model for Text-to-Vibrotactile Generation cites this paper.

HapticLDM: A Diffusion Model for Text-to-Vibrotactile Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:28.188916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:11:29.782569Z digest=sha256:2d314e2aaafbecf8464ae874911ec62f27e86fdab753932706e1f1c289cc4bad

Observation 1e51d76d-72e3-48e3-9143-9a44fa42d66a · inbound

PoDAR: Power-Disentangled Audio Representation for Generative Modeling cites this paper.

PoDAR: Power-Disentangled Audio Representation for Generative Modeling AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:41:35.176549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:04:08.420600Z digest=sha256:883299b2bbad4a52c26dc4f8576340233dea11d64ab4b6123a647aa5823e3677

Observation 875fd0db-094b-4cac-a577-9f4e223b7d81 · inbound

WavFlow: Audio Generation in Waveform Space cites this paper.

WavFlow: Audio Generation in Waveform Space AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:38:09.626883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T07:33:35.243337Z digest=sha256:ef9bf99988e5c94f14cea1ff5b914c54b38c24129542b675a454fa9949608ca8

Observation 9e9be8ae-4fc9-4935-9355-e05862a46376 · inbound

EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction cites this paper.

EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:28.537734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T10:25:14.873276Z digest=sha256:86afed1a30ffd9bd7721c019ffabd85957685162e3317e7e14bcb869a18bc149

Observation 917ed80e-e9b2-4838-8169-7486d16d7403 · inbound

Mental Damage: Caption Poisoning Attacks on Retrieval-Augmented Text-to-Music Generation cites this paper.

Mental Damage: Caption Poisoning Attacks on Retrieval-Augmented Text-to-Music Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:15:01.184596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T18:57:35.126894Z digest=sha256:9386ff2631a6cd919cc6f557effae458b0d0837f7b3879a9a2eadfc375fdee0b

Observation d8b7e1ac-56f6-48d5-9137-35997613f8c2 · inbound

dMoE: dLLMs with Learnable Block Experts cites this paper.

dMoE: dLLMs with Learnable Block Experts AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:45.089912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T22:50:51.900169Z digest=sha256:e6e0baae0955856c8740a8074d5c65a4bedbcf8ead908fb3c044e2cfb7667563

Observation b957d6e4-935a-412e-a3f5-2f0b3e30d1e4 · inbound

Entropy as a Structural Prior: How a Log-Barrier on DiT Belief Space Drives Musical Diversity and Development cites this paper.

Entropy as a Structural Prior: How a Log-Barrier on DiT Belief Space Drives Musical Diversity and Development AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.253964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T21:06:45.886467Z digest=sha256:6d608fa3e0651b2afa2b41d038cb75c27b1b8efe67577c952204a64012f8d1b9

Observation 2987213b-e846-46ff-9bad-cb73d2f87e50 · inbound

Flow Matching with In-Context Priors for Out-of-Distribution Brain Dynamics cites this paper.

Flow Matching with In-Context Priors for Out-of-Distribution Brain Dynamics AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:27:46.755613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:44:39.810293Z digest=sha256:b999cbb79d7958ab56acbb0393e0464329466c98dc9cb01f5d5ab6931a82e9da

Observation b2be330d-b942-43c6-81c4-adebd15d638f · inbound

Net-Ev$^2$: A Generative Simulator for Network Event Evolution cites this paper.

Net-Ev$^2$: A Generative Simulator for Network Event Evolution AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:57.609498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:09:20.657051Z digest=sha256:5ce39f041f8a60b26db8932ebe0ec715af9783617773c4d7133bb821931c8c05

Observation 6553f382-ad31-49dd-8943-a4627ad387a8 · inbound

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation cites this paper.

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.471182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T07:26:28.527338Z digest=sha256:d663cda737393c7ec86333cc94b9c5bd4b1ec671c3b97d2ae95d532bc9b42c94

Observation 201afcc3-58ad-49f4-b226-f6f680c2fc2d · inbound

ARIA: Adaptive Region-Based Importance Allocation for Conditional Diffusion Distillation cites this paper.

ARIA: Adaptive Region-Based Importance Allocation for Conditional Diffusion Distillation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:44.954421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T08:41:11.951712Z digest=sha256:daaa3ac54022151612f84e0f1018c4e23695bd17dba9f4b7f4124a1228fb60ad

Observation 1f58b6b3-7635-4d6a-b9b4-127dc290554b · inbound

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control cites this paper.

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:24:21.554371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T07:18:20.501369Z digest=sha256:896f91831d81efc814519819407ed02c5a413cbc6c56d5f5b611567b83d852be

Observation d870af0f-cf64-45bc-add2-1507338a977b · inbound

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control cites this paper.

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T17:03:01.432145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:03:01.432145Z digest=sha256:034f93c08765df087e9de93fbb87442246eb15c4c9e0c841e150f3887728615b

Observation 09db36e7-d0b1-4508-9387-bb85347cb181 · inbound

ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models cites this paper.

ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:05:37.221480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T06:47:53.308909Z digest=sha256:ca5ee8d394c0521a49339bfb9eeb84eb08aa660476a290dbf5c92501f0ecf781

Observation eda6be9b-5c73-4779-a2c7-3b4ba7304da7 · inbound

An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation cites this paper.

An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.292439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T05:04:23.295592Z digest=sha256:8111f4af839a8bbaf2eb8708b39e42314fdfb76c9aad519a27ea520fa5a3686a

Observation 6b618a80-5245-450a-b915-ba7b6712dcc8 · inbound

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation cites this paper.

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T12:34:20.057072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T12:34:20.057072Z digest=sha256:1e80ab8e058effdcb9b74bdc49bc3d442e813ca1a13122b4d750c0127dbf8f41

Observation 46dabe77-610a-4901-8c3c-bfea8c1bbebd · inbound

Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment cites this paper.

Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T10:59:22.914974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:59:22.914974Z digest=sha256:be873cca6c4d42ae1d273d475819998124a3a64eb5113aecffa3354a0af4cfb2

Observation c0eb786b-3987-46b9-8017-92c2c3547882 · inbound

FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving cites this paper.

FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:40.274605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:40.274605Z digest=sha256:b16c033d826cfa016f11a59ffebdb597feaf8d1c9b06cb918a576eff7503e3d9

Observation 961127b7-5d90-47d2-a72e-207eddba16ef · inbound

Analytic Distribution of Classifier-Free Guidance for Schedule Design cites this paper.

Analytic Distribution of Classifier-Free Guidance for Schedule Design AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T11:59:45.840064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:59:45.840064Z digest=sha256:64e0b5d552ec084a8469f7c1363fc21bdcabe8654797fa377142eefc1c5982fd

Observation 09755735-db25-461a-a381-ef03fee5233a · inbound

RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling cites this paper.

RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T11:48:09.922353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:48:09.922353Z digest=sha256:fec0cec8f17249eafe7b009f4583d1a211bd73ac745972501824d610b1225fcf

Observation 9bbb7541-3030-47ac-9cc4-64cb3ef81e9f · inbound

Amortized Moment Matching for Visual Generation cites this paper.

Amortized Moment Matching for Visual Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-30T18:58:28.076405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:58:28.076405Z digest=sha256:7d702124b8c73038920b2712be2fe8dc9c3f38b397e9d24b2d7a7f2474435b97

Observation 8e6509e6-dac5-48c8-9391-0c7876ebc211 · inbound

Exploring Efficient Waveform Diffusion Models for Foley Sound Generation cites this paper.

Exploring Efficient Waveform Diffusion Models for Foley Sound Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T12:46:58.947785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:46:58.947785Z digest=sha256:dc195398931b1ffa361a21366926e807ed779686569dc0cce96c7b0ea2d41527

Observation 5c16e57b-094a-4a42-a1dd-9b50a8f604f0 · inbound

AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities cites this paper.

AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:43:18.242065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:43:18.242065Z digest=sha256:8a225d23cf0b4e58df6f75ee813518a3cb2775f06044837a3622edc9a9361594