Pith. sign in

REVIEW 35 cited by

AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.12503 v3 pith:KSUBW4F6 submitted 2023-01-29 cs.SD cs.AIcs.MMeess.ASeess.SP

classification cs.SDcs.AIcs.MMeess.ASeess.SP
keywords audioldmaudiogenerationlatentsystemclapcomputationalembedding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-audio (TTA) system has recently gained attention for its ability to synthesize general audio based on text descriptions. However, previous studies in TTA have limited generation quality with high computational costs. In this study, we propose AudioLDM, a TTA system that is built on a latent space to learn the continuous audio representations from contrastive language-audio pretraining (CLAP) latents. The pretrained CLAP models enable us to train LDMs with audio embedding while providing text embedding as a condition during sampling. By learning the latent representations of audio signals and their compositions without modeling the cross-modal relationship, AudioLDM is advantageous in both generation quality and computational efficiency. Trained on AudioCaps with a single GPU, AudioLDM achieves state-of-the-art TTA performance measured by both objective and subjective metrics (e.g., frechet distance). Moreover, AudioLDM is the first TTA system that enables various text-guided audio manipulations (e.g., style transfer) in a zero-shot fashion. Our implementation and demos are available at https://audioldm.github.io.

Discussion (0). Sign in to comment.

Forward citations

Cited by 35 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Analytic Distribution of Classifier-Free Guidance for Schedule Design

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Deterministic CFG samples from p_t0 times an exponential path integral of the score discrepancy, and the resulting schedule DG-CFG reduces sampling steps at high guidance.

  2. MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    MAVIN proposes boundary-aware attention, ID-aware propagation, a multi-agent scripting pipeline, and the MAVINSet dataset as the first framework for multi-shot audio-visual generation with narrative control, claiming ...

  3. A Sharp KL-Convergence Analysis for Diffusion Models under Minimal Assumptions

    stat.ML 2025-08 conditional novelty 7.0 of 10

    A new analysis shows O~(d/epsilon) steps suffice for KL-close diffusion sampling under only L2 score error and finite second moment assumptions, improving the known O~(d/epsilon^2).

  4. Flow Matching Policy Gradients

    cs.LG 2025-07 conditional novelty 7.0 of 10

    FPO trains flow-based policies with PPO by replacing the likelihood ratio with an exponentiated flow matching loss difference.

  5. Amortized Moment Matching for Visual Generation

    cs.LG 2026-07 accept novelty 6.0 of 10

    Amortized Fréchet Distance uses neural nets to match conditional means and covariances, yielding stronger one-step visual generators than explicit FD-loss or multi-step teachers.

  6. RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling

    cs.SD 2026-07 conditional novelty 6.0 of 10

    RPPNet generates melodies by planning variable-length perceptually grouped rhythm-pitch primitives first and then decoding them into notes, beating bar-level baselines in subjective structure and musicality ratings.

  7. FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving

    cs.DC 2026-07 conditional novelty 6.0 of 10

    FlashDiff reduces diffusion serving latency by 30–97% and raises throughput 1.2–2.2× by adaptively skipping refinement of latent regions that no longer need it.

  8. Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Beat-guided contrastive alignment of pretrained MotionBERT/MERT features plus ControlNet conditioning of AudioLDM improves dance–music alignment on AIST++ over a MusicGen textual-inversion baseline while remaining com...

  9. SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation

    cs.SD 2026-07 conditional novelty 6.0 of 10

    SynSFX provides a multi-generator sound-effect deepfake corpus showing speech detectors fail, joint training mitigates forgetting, but generalization to unseen generators remains poor due to artifact overfitting.

  10. Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training

    cs.SD 2026-03 conditional novelty 6.0 of 10

    Jointly training the watermark embedder/detector with the source separator enables ~1% bit-error-rate recovery of per-stem watermarks after mixing and separation, where independent training yields 15–35%.

  11. Dual-End Consistency Model

    cs.CV 2026-02 unverdicted novelty 6.0 of 10

    DE-CM trains a flow-map consistency model on three sub-trajectories (coupling, instantaneous, noise-to-noisy) and reports 1.70 FID one-step on ImageNet 256.

  12. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

    eess.AS 2025-10 conditional novelty 6.0 of 10

    A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.

  13. Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Separating the text condition into a video caption and a visually grounded audio caption, and fusing the diffusion towers with dual cross-attention, gives the reported-best text-to-sounding-video quality and synchroni...

  14. Audio-Guided Visual Editing with Complex Multi-Modal Prompts

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A training-free framework that maps audio embeddings into Stable Diffusion's text space and fuses multiple audio/text prompts via per-patch residual noise selection, outperforming text-only editors on new audio-visual...

  15. Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new interactive method lets users click on an object in a video and generates audio for just that object, using mask-conditioned contrastive fine-tuning and latent diffusion.

  16. Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Diff-TONE selects the prompt-swap timestep using a distilled latent instrument classifier, improving content preservation over fixed timestep baselines in text-to-music diffusion editing.

  17. In-the-wild Audio Spatialization with Flexible Text-guided Localization

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A text-guided latent diffusion model converts monaural audio into binaural audio whose perceived directions and distances follow user-specified text prompts.

  18. AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A training-free multi-agent framework that decomposes multimodal inputs into audio events, selects specialized generators, and self-corrects outputs to produce multiple audio types.

  19. SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A ControlNet branch plus a frequency-aware feature aligner lets a pretrained masked generative TTA model produce video-synchronized foley, beating several from-scratch models on VGGSound.

  20. Exploring Efficient Waveform Diffusion Models for Foley Sound Generation

    eess.AS 2026-07 conditional novelty 5.0 of 10

    Dual-path attention over time-frequency representations lets a 3.26M-parameter waveform diffusion model match the quality of 50M+ parameter Foley generators.

  21. Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models

    cs.CL 2025-10 unverdicted novelty 5.0 of 10

    Introduces CCG-CFG with inconsistency-based dynamic scales and hard-sample mining distillation to boost emotional alignment in auto-regressive TTS, reporting up to 12% absolute gains in emotion recognition accuracy.

  22. Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation

    cs.SD 2025-08 unverdicted novelty 5.0 of 10

    A bar-level symbolic-score song generator (BACH) is claimed to beat published systems and commercial Suno on human-rated quality, duration, and efficiency, but the supporting full text is corrupted and unverifiable.

  23. AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

    cs.SD 2025-08 conditional novelty 5.0 of 10

    A single multimodal diffusion transformer generates video-synchronized general audio, speech, and song from flexible combinations of video, text, and lyrics inputs.

  24. Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning

    cs.SD 2025-07 conditional novelty 5.0 of 10

    A multi-view SSL framework with combined reconstruction and separation-based contrastive losses obtains disentangled pitch and instrument subspaces without the accuracy loss seen with contrastive-only training on NSynth.

  25. Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs

    cs.CV 2025-07 conditional novelty 5.0 of 10

    An integrated storytelling framework that generates text, scene graphs, images, and sound together reports higher expert-rated coherence than a sequential baseline, but the evaluation is small and qualitative.

  26. SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations

    cs.SD 2025-07 conditional novelty 5.0 of 10

    A three-stage diffusion pipeline maps 3D Gaussian Splatting object representations to position-dependent impact sounds, trained first on text captions and then on real recordings.

  27. CHORDS: Diffusion Sampling Accelerator with Multi-core Hierarchical ODE Solvers

    cs.LG 2025-07 conditional novelty 5.0 of 10

    CHORDS accelerates diffusion sampling by running hierarchical ODE solvers on multiple cores, with slower solvers rectifying faster ones, achieving up to 2.9x speedup without retraining.

  28. ADMC: Attention-based Diffusion Model for Missing Modalities Feature Completion

    cs.AI 2025-07 conditional novelty 5.0 of 10

    An attention-based diffusion model that generates missing modality features, combined with independently trained extractors, achieves state-of-the-art emotion and intent recognition on IEMOCAP and MIntRec.

  29. UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

    cs.SD 2025-06 conditional novelty 5.0 of 10

    UmbraTTS jointly synthesizes speech and environmental audio via conditional flow matching, conditioned on text and acoustic context, with controllable background volume.

  30. Inference-time Scaling for Diffusion-based Audio Super-resolution

    cs.SD 2025-08 conditional novelty 4.0 of 10

    Generating 120 candidate super-resolved audios and choosing the best by task-specific verifiers improves speech, music, and sound effects over single-sample diffusion output, at 120x compute.

  31. Diffusion Models for Time Series Forecasting: A Survey

    stat.ML 2025-07 conditional novelty 4.0 of 10

    A survey classifies diffusion-based time series forecasting models into a two-axis taxonomy by conditioning source and integration method.

  32. WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction

    cs.SD 2025-06 conditional novelty 4.0 of 10

    WhisQ uses Whisper and Qwen with co-attention and optimal transport to predict music quality and text-alignment scores, but its reported improvements do not match its own data.

  33. AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities

    cs.SD 2026-08 conditional novelty 3.0 of 10

    A narrative review of 30 recent papers classifies AI sound-effect generators by input modality and summarizes reported progress and remaining limitations in temporal sync, evaluation, and controllability.

  34. WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration

    cs.SD 2025-08 conditional novelty 3.0 of 10

    WaveLLDM, a lightweight latent diffusion model with a neural codec, achieves low spectral distortion (LSD 0.48-0.60) on speech restoration but scores far below SOTA on PESQ and STOI.

  35. ASAudio: A Survey of Advanced Spatial Audio Research

    eess.AS 2025-08 unverdicted novelty 3.0 of 10

    A comprehensive survey that systematically categorizes spatial audio research by representation, task, dataset, and evaluation.

Pith tools