Pith. sign in

REVIEW 2 cited by

Investigating the Emergent Audio Classification Ability of ASR Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.09363 v2 pith:3A6OVADA submitted 2023-11-15 cs.CL

classification cs.CL
keywords zero-shotfoundationmodelsperformanceabilityclassificationaudiodata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text and vision foundation models can perform many tasks in a zero-shot setting, a desirable property that enables these systems to be applied in general and low-resource settings. There has been far less work, however, on the zero-shot abilities of ASR foundation models, with these systems typically fine-tuned to specific tasks or constrained to applications that match their training criterion and data annotation. In this work we investigate the ability of Whisper and MMS, ASR foundation models trained primarily for speech recognition, to perform zero-shot audio classification. We use simple template-based text prompts at the decoder and use the resulting decoding probabilities to generate zero-shot predictions. Without training the model on extra data or adding any new parameters, we demonstrate that Whisper shows promising zero-shot classification performance on a range of 8 audio-classification datasets, outperforming the accuracy of existing state-of-the-art zero-shot baselines by an average of 9%. One important step to unlock the emergent ability is debiasing, where a simple unsupervised reweighting method of the class probabilities yields consistent significant performance gains. We further show that performance increases with model size, implying that as ASR foundation models scale up, they may exhibit improved zero-shot performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples

    eess.AS 2025-05 conditional novelty 6.0 of 10

    A contrastive-style adapter trained on LLM-generated positive and negative audio descriptions improves audio hallucination accuracy to 77.5 percent and audio question answering to 84.3 percent, without changing the fr...

  2. Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect

    cs.SD 2025-05 conditional novelty 5.0 of 10

    Linear probes on frozen speech embeddings predict seven voice quality ratings and transfer zero-shot across languages, tasks, and affect.

Pith tools