Pith. sign in

Clotho: an audio captioning dataset

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.SD 1

years

2024 1

verdicts

CONDITIONAL 1

representative citing papers

Vision Language Models Are Few-Shot Audio Spectrogram Classifiers

cs.SD · 2024-11-18 · conditional · novelty 6.0

GPT-4o classifies environmental sounds from spectrogram images in a few-shot setting with 59% cross-validated accuracy on ESC-10, beating a commercial audio language model and roughly matching human experts.

citing papers explorer

Showing 1 of 1 citing paper.

  • Vision Language Models Are Few-Shot Audio Spectrogram Classifiers cs.SD · 2024-11-18 · conditional · none · ref 7

    GPT-4o classifies environmental sounds from spectrogram images in a few-shot setting with 59% cross-validated accuracy on ESC-10, beating a commercial audio language model and roughly matching human experts.