REVIEW 6 cited by
Separate Anything You Describe
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA). LASS aims to separate a target sound from an audio mixture given a natural language query, which provides a natural and scalable interface for digital audio applications. Recent works on LASS, despite attaining promising separation performance on specific sources (e.g., musical instruments, limited classes of audio events), are unable to separate audio concepts in the open domain. In this work, we introduce AudioSep, a foundation model for open-domain audio source separation with natural language queries. We train AudioSep on large-scale multimodal datasets and extensively evaluate its capabilities on numerous tasks including audio event separation, musical instrument separation, and speech enhancement. AudioSep demonstrates strong separation performance and impressive zero-shot generalization ability using audio captions or text labels as queries, substantially outperforming previous audio-queried and language-queried sound separation models. For reproducibility of this work, we will release the source code, evaluation benchmark and pre-trained model at: https://github.com/Audio-AGI/AudioSep.
Forward citations
Cited by 6 Pith papers
-
CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents
CodecSep performs prompt-driven universal sound separation directly in neural audio codec latents by combining a frozen DAC backbone with a lightweight FiLM-conditioned Transformer masker driven by CLAP embeddings, yi...
-
ZeroSep: Separate Anything in Audio with Zero Training
Latent inversion of a mixed audio into a pretrained text-guided diffusion model, followed by denoising with classifier-free guidance weight 1, performs zero-training source separation.
-
Beyond Speaker Identity: Text Guided Target Speech Extraction
StyleTSE extracts target speech from mixtures using natural-language speaking style descriptions, optionally combined with reference audio, trained on the new TextrolMix dataset.
-
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation
A text-and-video conditioned flow transformer that generates onscreen plus offscreen audio, evaluated on a new curated benchmark and on VGGSound.
-
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
A systematic survey that categorizes audio-language models by architecture, training objective, and application, covering speech, music, and general audio.
-
30+ Years of Source Separation Research: Achievements and Future Challenges
A comprehensive review of three decades of audio source separation research, presenting no new technical results.
Discussion (0). Continue with ORCID to comment.