Pith. sign in

REVIEW 3 cited by

Adapting Language-Audio Models as Few-Shot Audio Learners

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.17719 v1 pith:LVQLV2BZ submitted 2023-05-28 eess.AS cs.SD

classification eess.AScs.SD
keywords adaptertreffaudioclassificationcalmclapdesignedfew-shot
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We presented the Treff adapter, a training-efficient adapter for CLAP, to boost zero-shot classification performance by making use of a small set of labelled data. Specifically, we designed CALM to retrieve the probability distribution of text-audio clips over classes using a set of audio-label pairs and combined it with CLAP's zero-shot classification results. Furthermore, we designed a training-free version of the Treff adapter by using CALM as a cosine similarity measure. Experiments showed that the proposed Treff adapter is comparable and even better than fully-supervised methods and adaptation methods in low-shot and data-abundant scenarios. While the Treff adapter shows that combining large-scale pretraining and rapid learning of domain-specific knowledge is non-trivial for obtaining generic representations for few-shot learning, it is still limited to audio classification tasks. In the future, we will explore how to use audio-language models in diverse audio domains.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CLAP-S: Support Set Based Adaptation for Downstream Fiber-optic Acoustic Recognition

    eess.AS 2025-01 conditional novelty 4.0 of 10

    CLAP-S and CLAP-S+ adapt CLAP models to fiber-optic acoustic recognition by combining support-set retrieval with a fine-tuned adapter, reporting improved few-shot classification accuracy over existing methods.

  2. TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification

    cs.SD 2024-12 reject novelty 4.0 of 10

    Task-specific prompt ensembling with GPT-4-generated attributes and sources improves some zero-shot audio classification datasets while degrading others.

  3. Multiple Consistency-guided Test-Time Adaptation for Contrastive Audio-Language Models with Unlabeled Audio

    cs.SD 2024-12 conditional novelty 4.0 of 10

    A consistency-guided test-time prompt adaptation method improves CLAP zero-shot audio classification by 4.41% relative on average over DA CLAP across 12 datasets.

Pith tools