REVIEW 45 cited by
Audiobox: Unified Audio Generation with Natural Language Prompts
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Audio is an essential part of our life, but creating it often requires expertise and is time-consuming. Research communities have made great progress over the past year advancing the performance of large scale audio generative models for a single modality (speech, sound, or music) through adopting more powerful generative models and scaling data. However, these models lack controllability in several aspects: speech generation models cannot synthesize novel styles based on text description and are limited on domain coverage such as outdoor environments; sound generation models only provide coarse-grained control based on descriptions like "a person speaking" and would only generate mumbling human voices. This paper presents Audiobox, a unified model based on flow-matching that is capable of generating various audio modalities. We design description-based and example-based prompting to enhance controllability and unify speech and sound generation paradigms. We allow transcript, vocal, and other audio styles to be controlled independently when generating speech. To improve model generalization with limited labels, we adapt a self-supervised infilling objective to pre-train on large quantities of unlabeled audio. Audiobox sets new benchmarks on speech and sound generation (0.745 similarity on Librispeech for zero-shot TTS; 0.77 FAD on AudioCaps for text-to-sound) and unlocks new methods for generating audio with novel vocal and acoustic styles. We further integrate Bespoke Solvers, which speeds up generation by over 25 times compared to the default ODE solver for flow-matching, without loss of performance on several tasks. Our demo is available at https://audiobox.metademolab.com/
Forward citations
Cited by 45 Pith papers
-
SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation
SonicWeave routes chunks of audio through specialized experts, using a learned gate between text-derived prior and local evidence, improving compositional fidelity in unified text-to-audio scene generation.
-
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation
AV-Link unifies video-to-audio and audio-to-video generation by aligning frozen diffusion-model activations with temporally matched rotary position embeddings in a shared Fusion Block.
-
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching
VoxAudio generates audio scenes with intelligible, temporally placed quoted speech by combining chunk-wise causal flow matching with multi-reward fine-tuning and a large transcript-annotated corpus.
-
VIOLET: High-Fidelity Violin Synthesis with Techniques and Dynamics
VIOLET, a latent diffusion model, synthesizes violin audio that follows MIDI, playing technique, and dynamics controls, outperforming the prior neural baseline and approaching a commercial virtual instrument.
-
CustomDance: Customized 3D Dance Generation with Coarse-to-Fine Human-Centered Interactive Control
CustomDance combines an MLLM-based choreographic planner, multimodal dance-phrase retrieval, and diffusion inpainting into one three-stage interactive system for user-customized 3D dance generation.
-
AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation
A structured soundscape benchmark with 25,707 binary semantic rubrics shows that rubric-based, audio-grounded evaluation tracks human semantic judgments better than CLAP-style global similarity for text-to-audio models.
-
Unified Audio Intelligence Without Regressing on Text Intelligence
A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.
-
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation
SynSFX provides a multi-generator sound-effect deepfake corpus showing speech detectors fail, joint training mitigates forgetting, but generalization to unseen generators remains poor due to artifact overfitting.
-
Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability
A new metric, MCLP, uses a frozen audio-LLM's continuation likelihood to grade speaking-style consistency and doubles as a reward that improves role-play TTS on a new 1,435-hour drama dataset.
-
iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation
iMontage repurposes a pretrained video diffusion model to generate coherent yet highly dynamic image sets from arbitrary numbers of input images.
-
UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement
A 63M-parameter decoder-only LM, UniSE, unifies speech restoration, target speaker extraction, and speech separation by generating BiCodec discrete tokens under task-specific prompts.
-
Testing chatbots on the creation of encoders for audio conditioned image generation
All chatbot-designed audio encoders failed to align with CLIP text embeddings and produced incoherent images, while showing a surprising architectural similarity across chatbots.
-
DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization
DiffRhythm+ improves full-length lyric-to-song generation via balanced data scaling, MuLan-based multimodal style control, and DPO fine-tuning guided by automated aesthetic scorers.
-
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
A causal audio language model with continuous-valued tokens and masked next-token prediction matches diffusion-based text-to-audio quality with smaller, streamable models.
-
Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation
Under matched training conditions, auto-regressive models slightly outperform flow-matching on music quality and temporal control, while flow-matching offers faster inference and better inpainting flexibility.
-
In-the-wild Audio Spatialization with Flexible Text-guided Localization
A text-guided latent diffusion model converts monaural audio into binaural audio whose perceived directions and distances follow user-specified text prompts.
-
FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
A 1.5B unified multimodal model trained with discrete flow matching and metric-induced probability paths matches autoregressive baselines of similar size on generation and understanding benchmarks.
-
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
VoiceStar uses a progress-based rotary position embedding and mixed prompt training to give zero-shot voice cloning precise duration control and much longer output than training clips.
-
RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations
A new 13,000-hour dataset with 24 million LLM-generated text descriptions enables the first open-source text-driven TTS for 24 Indian languages, with reported high speaker, emotion, and cross-lingual control.
-
Learning to Highlight Audio by Watching Movies
A transformer model learns to transform poorly mixed audio into well-balanced audio guided by video content, trained on a new pseudo-data set derived from movies.
-
OmniAudio: Generating Spatial Audio from 360-Degree Video
OmniAudio generates First-order Ambisonics audio directly from 360-degree video using dual-branch video encoding and flow-matching pre-training, and it introduces the Sphere360 dataset and benchmark.
-
RenderBox: Expressive Performance Rendering with Text Control
RenderBox is a text-and-score conditioned diffusion model that renders expressive, controllable audio performances across piano, guitar, saxophone, violin, and orchestral instruments.
-
LoRP-TTS: Low-Rank Personalized Text-To-Speech
LoRP-TTS shows that per-prompt LoRA fine-tuning with one short recording improves speaker similarity in Voicebox-based zero-shot TTS, at some cost in inference time.
-
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
A zero-shot TTS model using controllable masked speech prediction and a dual speaker encoder that can both remove and preserve acoustic background from the prompt, selected by a binary control signal.
-
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
Vevo achieves zero-shot timbre, accent, and emotion imitation by using VQ-VAE codebook size on HuBERT features to create content and content-style tokens.
-
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
CosyAudio improves text-to-audio generation by generating synthetic captions with confidence scores and iteratively refining training data through a self-evolving loop.
-
FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching
A single-step flow matching model with a data-dependent prior matches or beats diffusion-based audio super-resolution models on VCTK while using one function evaluation.
-
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
A fast flow-matching text-to-audio model aligned via CLAP-ranked self-generated preference pairs reports state-of-the-art AudioCaps and human-evaluation scores.
-
ETTA: Elucidating the Design Space of Text-to-Audio Models
ETTA, a text-to-audio model trained on a large synthetic caption dataset, outperforms public-data baselines on AudioCaps and MusicCaps and rivals proprietary-data systems.
-
SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor
A language-model song generator adapted for multi-task editing, supporting segment-wise and track-wise modifications alongside full-song generation.
-
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
SyncFlow jointly generates temporally aligned 16 FPS video and 48kHz audio from text using a dual-diffusion-transformer with modality adaptors, and reports better audio-video alignment than cascaded and contrastive baselines.
-
Qwen-Audio-3.0-Gen-Preview Technical Report
One diffusion-transformer system with a shared audio codec generates standalone audio, multi-speaker dialogue, and time-structured mixed scenes, with strongest measured advantages in speaker similarity, cross-turn con...
-
FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation
FlowerDance pairs MeanFlow few-step flow matching with a bidirectional Mamba backbone and physical-consistency losses, reporting state-of-the-art dance quality at 2008 FPS on FineDance and AIST++.
-
Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation
Segment-level EER overstates deployment readiness for partial fake speech localizers, which drop from 7.6% to above 40% EER on out-of-domain test sets.
-
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
UmbraTTS jointly synthesizes speech and environmental audio via conditional flow matching, conditioned on text and acoustic context, with controllable background volume.
-
InfiniteAudio: Infinite-Length Audio Generation with Consistency
A training-free inference strategy, FIFO sampling plus attention-guided curved denoising, enables diffusion text-to-audio models to produce long, temporally consistent audio with constant memory.
-
MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers
A zero-shot pipeline that creates character voices from AI-generated faces and LLM-written prosody instructions can produce expressive audiobooks without extra training or manual annotation, though human quality score...
-
OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
OZSpeech is a one-step zero-shot TTS system that starts from learned content and mean-style codes and uses flow matching to refine them, achieving very low word error rates with a small model.
-
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
A masked generative model pre-trained on unlabeled speech then fine-tuned per task matches or beats task-specific systems across TTS, voice conversion, speaker extraction, enhancement, and lip-to-speech.
-
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video
Adding video-derived visual features to FastSpeech2 improves pitch, energy, and duration prediction on a 33-hour movie TTS dataset, with UTMOS rising from 2.91 to 3.13.
-
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
A multi-stage TTS framework that accepts arbitrary combinations of text, audio, and face prompts to control both style and timbre of generated speech.
-
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
VoiceDiT generates speech and matching environmental sounds from text, audio, or image prompts, and reports better speech intelligibility than the VoiceLDM baseline.
-
Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation
VLFM models video latent patches as a HiPPO-LegS polynomial flow and trains a flow matching model to generate frames, claiming bounded interpolation and extrapolation error.
-
Overview of the Amphion Toolkit (v0.2)
Amphion v0.2 is an open-source toolkit for audio, music, and speech generation, adding a 101K-hour multilingual dataset, processing pipelines, and pretrained models.
-
YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls
A video-guided sound effects model with a learnable audio-visual aggregator and multi-modal chain-of-thought module reports strong VGGSound benchmark scores, but its few-shot claim rests on three qualitative samples.
Discussion (0). Continue with ORCID to comment.