Pith. sign in

REVIEW 6 cited by

Separate Anything You Describe

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.05037 v3 pith:PAW5ZN2D submitted 2023-08-09 eess.AS cs.AIcs.MMcs.SD

classification eess.AScs.AIcs.MMcs.SD
keywords audioseparationaudioseplassnaturalseparatesourcelanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA). LASS aims to separate a target sound from an audio mixture given a natural language query, which provides a natural and scalable interface for digital audio applications. Recent works on LASS, despite attaining promising separation performance on specific sources (e.g., musical instruments, limited classes of audio events), are unable to separate audio concepts in the open domain. In this work, we introduce AudioSep, a foundation model for open-domain audio source separation with natural language queries. We train AudioSep on large-scale multimodal datasets and extensively evaluate its capabilities on numerous tasks including audio event separation, musical instrument separation, and speech enhancement. AudioSep demonstrates strong separation performance and impressive zero-shot generalization ability using audio captions or text labels as queries, substantially outperforming previous audio-queried and language-queried sound separation models. For reproducibility of this work, we will release the source code, evaluation benchmark and pre-trained model at: https://github.com/Audio-AGI/AudioSep.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents

    cs.SD 2025-09 unverdicted novelty 6.0 of 10

    CodecSep performs prompt-driven universal sound separation directly in neural audio codec latents by combining a frozen DAC backbone with a lightweight FiLM-conditioned Transformer masker driven by CLAP embeddings, yi...

  2. ZeroSep: Separate Anything in Audio with Zero Training

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Latent inversion of a mixed audio into a pretrained text-guided diffusion model, followed by denoising with classifier-free guidance weight 1, performs zero-training source separation.

  3. Beyond Speaker Identity: Text Guided Target Speech Extraction

    eess.AS 2025-01 conditional novelty 6.0 of 10

    StyleTSE extracts target speech from mixtures using natural-language speaking style descriptions, optionally combined with reference audio, trained on the new TextrolMix dataset.

  4. VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A text-and-video conditioned flow transformer that generates onscreen plus offscreen audio, evaluated on a new curated benchmark and on VGGSound.

  5. Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

    cs.SD 2025-01 conditional novelty 5.0 of 10

    A systematic survey that categorizes audio-language models by architecture, training objective, and application, covering speech, music, and general audio.

  6. 30+ Years of Source Separation Research: Achievements and Future Challenges

    eess.AS 2025-01 unverdicted

    A comprehensive review of three decades of audio source separation research, presenting no new technical results.

Pith tools