Pith. sign in

REVIEW 45 cited by

Audiobox: Unified Audio Generation with Natural Language Prompts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.15821 v1 pith:6A3RYLST submitted 2023-12-25 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords audiogenerationmodelsspeechaudioboxsoundgeneratingstyles
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Audio is an essential part of our life, but creating it often requires expertise and is time-consuming. Research communities have made great progress over the past year advancing the performance of large scale audio generative models for a single modality (speech, sound, or music) through adopting more powerful generative models and scaling data. However, these models lack controllability in several aspects: speech generation models cannot synthesize novel styles based on text description and are limited on domain coverage such as outdoor environments; sound generation models only provide coarse-grained control based on descriptions like "a person speaking" and would only generate mumbling human voices. This paper presents Audiobox, a unified model based on flow-matching that is capable of generating various audio modalities. We design description-based and example-based prompting to enhance controllability and unify speech and sound generation paradigms. We allow transcript, vocal, and other audio styles to be controlled independently when generating speech. To improve model generalization with limited labels, we adapt a self-supervised infilling objective to pre-train on large quantities of unlabeled audio. Audiobox sets new benchmarks on speech and sound generation (0.745 similarity on Librispeech for zero-shot TTS; 0.77 FAD on AudioCaps for text-to-sound) and unlocks new methods for generating audio with novel vocal and acoustic styles. We further integrate Bespoke Solvers, which speeds up generation by over 25 times compared to the default ODE solver for flow-matching, without loss of performance on several tasks. Our demo is available at https://audiobox.metademolab.com/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 45 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation

    cs.SD 2026-08 conditional novelty 7.0 of 10

    SonicWeave routes chunks of audio through specialized experts, using a learned gate between text-derived prior and local evidence, improving compositional fidelity in unified text-to-audio scene generation.

  2. AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation

    cs.CV 2024-12 conditional novelty 7.0 of 10

    AV-Link unifies video-to-audio and audio-to-video generation by aligning frozen diffusion-model activations with temporally matched rotary position embeddings in a shared Fusion Block.

  3. VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching

    cs.SD 2026-08 conditional novelty 6.0 of 10

    VoxAudio generates audio scenes with intelligible, temporally placed quoted speech by combining chunk-wise causal flow matching with multi-reward fine-tuning and a large transcript-annotated corpus.

  4. VIOLET: High-Fidelity Violin Synthesis with Techniques and Dynamics

    eess.AS 2026-08 conditional novelty 6.0 of 10

    VIOLET, a latent diffusion model, synthesizes violin audio that follows MIDI, playing technique, and dynamics controls, outperforming the prior neural baseline and approaching a commercial virtual instrument.

  5. CustomDance: Customized 3D Dance Generation with Coarse-to-Fine Human-Centered Interactive Control

    cs.HC 2026-08 conditional novelty 6.0 of 10

    CustomDance combines an MLLM-based choreographic planner, multimodal dance-phrase retrieval, and diffusion inpainting into one three-stage interactive system for user-customized 3D dance generation.

  6. AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A structured soundscape benchmark with 25,707 binary semantic rubrics shows that rubric-based, audio-grounded evaluation tracks human semantic judgments better than CLAP-style global similarity for text-to-audio models.

  7. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  8. SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation

    cs.SD 2026-07 conditional novelty 6.0 of 10

    SynSFX provides a multi-generator sound-effect deepfake corpus showing speech detectors fail, joint training mitigates forgetting, but generalization to unseen generators remains poor due to artifact overfitting.

  9. Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability

    cs.SD 2026-01 conditional novelty 6.0 of 10

    A new metric, MCLP, uses a frozen audio-LLM's continuation likelihood to grade speaking-style consistency and doubles as a reward that improves role-play TTS on a new 1,435-hour drama dataset.

  10. iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation

    cs.CV 2025-11 conditional novelty 6.0 of 10

    iMontage repurposes a pretrained video diffusion model to generate coherent yet highly dynamic image sets from arbitrary numbers of input images.

  11. UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement

    cs.SD 2025-10 conditional novelty 6.0 of 10

    A 63M-parameter decoder-only LM, UniSE, unifies speech restoration, target speaker extraction, and speech separation by generating BiCodec discrete tokens under task-specific prompts.

  12. Testing chatbots on the creation of encoders for audio conditioned image generation

    cs.SD 2025-09 conditional novelty 6.0 of 10

    All chatbot-designed audio encoders failed to align with CLIP text embeddings and produced incoherent images, while showing a surprising architectural similarity across chatbots.

  13. DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization

    eess.AS 2025-07 conditional novelty 6.0 of 10

    DiffRhythm+ improves full-length lyric-to-song generation via balanced data scaling, MuLan-based multimodal style control, and DPO fine-tuning guided by automated aesthetic scorers.

  14. Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction

    eess.AS 2025-07 conditional novelty 6.0 of 10

    A causal audio language model with continuous-valued tokens and masked next-token prediction matches diffusion-based text-to-audio quality with smaller, streamable models.

  15. Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Under matched training conditions, auto-regressive models slightly outperform flow-matching on music quality and temporal control, while flow-matching offers faster inference and better inpainting flexibility.

  16. In-the-wild Audio Spatialization with Flexible Text-guided Localization

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A text-guided latent diffusion model converts monaural audio into binaural audio whose perceived directions and distances follow user-specified text prompts.

  17. FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A 1.5B unified multimodal model trained with discrete flow matching and metric-induced probability paths matches autoregressive baselines of similar size on generation and understanding benchmarks.

  18. VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation

    eess.AS 2025-05 conditional novelty 6.0 of 10

    VoiceStar uses a progress-based rotary position embedding and mixed prompt training to give zero-shot voice cloning precise duration control and much longer output than training clips.

  19. RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new 13,000-hour dataset with 24 million LLM-generated text descriptions enables the first open-source text-driven TTS for 24 Indian languages, with reported high speaker, emotion, and cross-lingual control.

  20. Learning to Highlight Audio by Watching Movies

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A transformer model learns to transform poorly mixed audio into well-balanced audio guided by video content, trained on a new pseudo-data set derived from movies.

  21. OmniAudio: Generating Spatial Audio from 360-Degree Video

    eess.AS 2025-04 conditional novelty 6.0 of 10

    OmniAudio generates First-order Ambisonics audio directly from 360-degree video using dual-branch video encoding and flow-matching pre-training, and it introduces the Sphere360 dataset and benchmark.

  22. RenderBox: Expressive Performance Rendering with Text Control

    eess.AS 2025-02 conditional novelty 6.0 of 10

    RenderBox is a text-and-score conditioned diffusion model that renders expressive, controllable audio performances across piano, guitar, saxophone, violin, and orchestral instruments.

  23. LoRP-TTS: Low-Rank Personalized Text-To-Speech

    cs.SD 2025-02 conditional novelty 6.0 of 10

    LoRP-TTS shows that per-prompt LoRA fine-tuning with one short recording improves speaker similarity in Voicebox-based zero-shot TTS, at some cost in inference time.

  24. Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction

    cs.SD 2025-02 conditional novelty 6.0 of 10

    A zero-shot TTS model using controllable masked speech prediction and a dual speaker encoder that can both remove and preserve acoustic background from the prompt, selected by a binary control signal.

  25. Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

    cs.SD 2025-02 conditional novelty 6.0 of 10

    Vevo achieves zero-shot timbre, accent, and emotion imitation by using VQ-VAE codebook size on HuBERT features to create content and content-style tokens.

  26. CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions

    eess.AS 2025-01 conditional novelty 6.0 of 10

    CosyAudio improves text-to-audio generation by generating synthetic captions with confidence scores and iteratively refining training data through a self-evolving loop.

  27. FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching

    eess.AS 2025-01 conditional novelty 6.0 of 10

    A single-step flow matching model with a data-dependent prior matches or beats diffusion-based audio super-resolution models on VCTK while using one function evaluation.

  28. TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

    cs.SD 2024-12 conditional novelty 6.0 of 10

    A fast flow-matching text-to-audio model aligned via CLAP-ranked self-generated preference pairs reports state-of-the-art AudioCaps and human-evaluation scores.

  29. ETTA: Elucidating the Design Space of Text-to-Audio Models

    cs.SD 2024-12 conditional novelty 6.0 of 10

    ETTA, a text-to-audio model trained on a large synthetic caption dataset, outperforms public-data baselines on AudioCaps and MusicCaps and rivals proprietary-data systems.

  30. SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor

    eess.AS 2024-12 conditional novelty 6.0 of 10

    A language-model song generator adapted for multi-task editing, supporting segment-wise and track-wise modifications alongside full-song generation.

  31. SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

    cs.MM 2024-12 conditional novelty 6.0 of 10

    SyncFlow jointly generates temporally aligned 16 FPS video and 48kHz audio from text using a dual-diffusion-transformer with modality adaptors, and reports better audio-video alignment than cascaded and contrastive baselines.

  32. Qwen-Audio-3.0-Gen-Preview Technical Report

    eess.AS 2026-07 conditional novelty 5.0 of 10

    One diffusion-transformer system with a shared audio codec generates standalone audio, multi-speaker dialogue, and time-structured mixed scenes, with strongest measured advantages in speaker similarity, cross-turn con...

  33. FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation

    cs.CV 2025-11 conditional novelty 5.0 of 10

    FlowerDance pairs MeanFlow few-step flow matching with a bidirectional Mamba backbone and physical-consistency losses, reporting state-of-the-art dance quality at 2008 FPS on FineDance and AIST++.

  34. Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation

    cs.SD 2025-07 conditional novelty 5.0 of 10

    Segment-level EER overstates deployment readiness for partial fake speech localizers, which drop from 7.6% to above 40% EER on out-of-domain test sets.

  35. UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

    cs.SD 2025-06 conditional novelty 5.0 of 10

    UmbraTTS jointly synthesizes speech and environmental audio via conditional flow matching, conditioned on text and acoustic context, with controllable background volume.

  36. InfiniteAudio: Infinite-Length Audio Generation with Consistency

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A training-free inference strategy, FIFO sampling plus attention-guided curved denoising, enables diffusion text-to-audio models to produce long, temporally consistent audio with constant memory.

  37. MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers

    cs.SD 2025-05 conditional novelty 5.0 of 10

    A zero-shot pipeline that creates character voices from AI-generated faces and LLM-written prosody instructions can produce expressive audiobooks without extra training or manual annotation, though human quality score...

  38. OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching

    cs.SD 2025-05 conditional novelty 5.0 of 10

    OZSpeech is a one-step zero-shot TTS system that starts from learned content and mean-style codes and uses flow matching to refine them, achieving very low word error rates with a small model.

  39. Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

    cs.SD 2025-02 conditional novelty 5.0 of 10

    A masked generative model pre-trained on unlabeled speech then fine-tuned per task matches or beats task-specific systems across TTS, voice conversion, speaker extraction, enhancement, and lip-to-speech.

  40. VisualSpeech: Enhancing Prosody Modeling in TTS Using Video

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Adding video-derived visual features to FastSpeech2 improves pitch, energy, and duration prediction on a 33-hour movie TTS dataset, with UTMOS rising from 2.91 to 3.13.

  41. FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

    eess.AS 2025-01 conditional novelty 5.0 of 10

    A multi-stage TTS framework that accepts arbitrary combinations of text, audio, and face prompts to control both style and timbre of generated speech.

  42. VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis

    eess.AS 2024-12 conditional novelty 5.0 of 10

    VoiceDiT generates speech and matching environmental sounds from text, audio, or image prompts, and reports better speech intelligibility than the VoiceLDM baseline.

  43. Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation

    cs.CV 2025-02 reject novelty 4.0 of 10

    VLFM models video latent patches as a HiPPO-LegS polynomial flow and trains a flow matching model to generate frames, claiming bounded interpolation and extrapolation error.

  44. Overview of the Amphion Toolkit (v0.2)

    cs.SD 2025-01 conditional novelty 4.0 of 10

    Amphion v0.2 is an open-source toolkit for audio, music, and speech generation, adding a 101K-hour multilingual dataset, processing pipelines, and pretrained models.

  45. YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

    cs.SD 2024-12 reject novelty 4.0 of 10

    A video-guided sound effects model with a learnable audio-visual aggregator and multi-modal chain-of-thought module reports strong VGGSound benchmark scores, but its few-shot claim rests on three qualitative samples.

Pith tools