Pith. sign in

REVIEW 1 cited by

PALM: Few-Shot Prompt Learning for Audio Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.19806 v1 pith:2SHNJJ6N submitted 2024-09-29 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords audiolearningmodelspromptpalmtextalmsapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio-Language Models (ALMs) have recently achieved remarkable success in zero-shot audio recognition tasks, which match features of audio waveforms with class-specific text prompt features, inspired by advancements in Vision-Language Models (VLMs). Given the sensitivity of zero-shot performance to the choice of hand-crafted text prompts, many prompt learning techniques have been developed for VLMs. We explore the efficacy of these approaches in ALMs and propose a novel method, Prompt Learning in Audio Language Models (PALM), which optimizes the feature space of the text encoder branch. Unlike existing methods that work in the input space, our approach results in greater training efficiency. We demonstrate the effectiveness of our approach on 11 audio recognition datasets, encompassing a variety of speech-processing tasks, and compare the results with three baselines in a few-shot learning setup. Our method is either on par with or outperforms other approaches while being computationally less demanding. Code is available at https://asif-hanif.github.io/palm/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Domain Adaptation Method and Modality Gap Impact in Audio-Text Models for Prototypical Sound Classification

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A background-profile subtraction method improves zero-shot sound classification accuracy under noisy conditions, and narrowing the audio-text modality gap further boosts performance.

Pith tools