Pith. sign in

REVIEW 2 cited by

Bridging Language Gaps in Audio-Text Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.07012 v2 pith:4766DFEO submitted 2024-06-11 cs.SD cs.CLeess.AS

classification cs.SDcs.CLeess.AS
keywords audio-textretrievaldatatextaudiocontentencoderenglish
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio-text retrieval is a challenging task, requiring the search for an audio clip or a text caption within a database. The predominant focus of existing research on English descriptions poses a limitation on the applicability of such models, given the abundance of non-English content in real-world data. To address these linguistic disparities, we propose a language enhancement (LE), using a multilingual text encoder (SONAR) to encode the text data with language-specific information. Additionally, we optimize the audio encoder through the application of consistent ensemble distillation (CED), enhancing support for variable-length audio-text retrieval. Our methodology excels in English audio-text retrieval, demonstrating state-of-the-art (SOTA) performance on commonly used datasets such as AudioCaps and Clotho. Simultaneously, the approach exhibits proficiency in retrieving content in seven other languages with only 10% of additional language-enhanced training data, yielding promising results. The source code is publicly available https://github.com/zyyan4/ml-clap.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

    cs.SD 2025-01 conditional novelty 5.0 of 10

    A systematic survey that categorizes audio-language models by architecture, training objective, and application, covering speech, music, and general audio.

  2. AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models

    eess.AS 2024-11 conditional novelty 5.0 of 10

    An automated pipeline produced AudioSetCaps, a 1.9 million pair audio-caption dataset, and training on it improved audio-text retrieval and audio captioning benchmarks.

Pith tools