REVIEW 5 cited by
EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Audio self-supervised learning (SSL) pre-training, which aims to learn good representations from unlabeled audio, has made remarkable progress. However, the extensive computational demands during pre-training pose a significant barrier to the potential application and optimization of audio SSL models. In this paper, inspired by the success of data2vec 2.0 in image modality and Audio-MAE in audio modality, we introduce Efficient Audio Transformer (EAT) to further improve the effectiveness and efficiency in audio SSL. The proposed EAT adopts the bootstrap self-supervised training paradigm to the audio domain. A novel Utterance-Frame Objective (UFO) is designed to enhance the modeling capability of acoustic events. Furthermore, we reveal that the masking strategy is critical in audio SSL pre-training, and superior audio representations can be obtained with large inverse block masks. Experiment results demonstrate that EAT achieves state-of-the-art (SOTA) performance on a range of audio-related tasks, including AudioSet (AS-2M, AS-20K), ESC-50, and SPC-2, along with a significant pre-training speedup up to ~15x compared to existing audio SSL models.
Forward citations
Cited by 5 Pith papers
-
Quantitative Analysis of Proxy Tasks for Anomalous Sound Detection
Better proxy-task performance does not generally improve anomalous sound detection; only source separation showed a strong, consistent positive correlation.
-
SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes
SSLAM pre-trains audio transformers on partially mixed audio clips with a source retention loss, improving polyphonic sound tagging while keeping monophonic benchmark scores.
-
Hidden-Domain Routing for All-Type Audio Deepfake Detection
A router-then-specialist audio deepfake detector, which classifies audio type first and then applies type-specific models and thresholds, achieved 96.10% Macro-F1 and first place on AT-ADD Track2.
-
IMPACT: Industrial Machine Perception via Acoustic Cognitive Transformer
With DINOS, a new 1,093-hour industrial-sound dataset, the authors show that pretraining a transformer (IMPACT, an EAT adaptation) on machine audio beats general audio models on 24 of 30 self-built monitoring tasks.
-
RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection
A reconstruction-perception-reinforcement-attention framework for audio deepfake detection reports state-of-the-art EERs on ASVspoof 2019/2021 and competitive results on sound and singing benchmarks.
Discussion (0). Continue with ORCID to comment.