Pith. sign in

REVIEW 1 cited by

Benchmarking Representations for Speech, Music, and Acoustic Events

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.00934 v1 pith:HWPX2ASW submitted 2024-05-02 eess.AS cs.LGcs.SD

classification eess.AScs.LGcs.SD
keywords methodsmodelsarchaudiodatasetspre-trainedacousticbenchmarking
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Limited diversity in standardized benchmarks for evaluating audio representation learning (ARL) methods may hinder systematic comparison of current methods' capabilities. We present ARCH, a comprehensive benchmark for evaluating ARL methods on diverse audio classification domains, covering acoustic events, music, and speech. ARCH comprises 12 datasets, that allow us to thoroughly assess pre-trained SSL models of different sizes. ARCH streamlines benchmarking of ARL techniques through its unified access to a wide range of domains and its ability to readily incorporate new datasets and models. To address the current lack of open-source, pre-trained models for non-speech audio, we also release new pre-trained models that demonstrate strong performance on non-speech datasets. We argue that the presented wide-ranging evaluation provides valuable insights into state-of-the-art ARL methods, and is useful to pinpoint promising research directions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluation of Deep Audio Representations for Hearables

    cs.SD 2025-02 conditional novelty 6.0 of 10

    DEAR, a new hearable-focused benchmark of 1,158 audio tracks, shows that the BEATs audio foundation model outperforms Wav2Vec2, HuBERT, and WavLM across context, source, and acoustic-property tasks.

Pith tools