Pith. sign in

REVIEW 5 cited by

EAT: Self-Supervised Pre-Training with Efficient Audio Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.03497 v1 pith:SHJZZE5Q submitted 2024-01-07 eess.AS cs.AIcs.CLcs.LGcs.SD

classification eess.AScs.AIcs.CLcs.LGcs.SD
keywords audiopre-trainingself-supervisedefficientmodalitymodelsrepresentationssignificant
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio self-supervised learning (SSL) pre-training, which aims to learn good representations from unlabeled audio, has made remarkable progress. However, the extensive computational demands during pre-training pose a significant barrier to the potential application and optimization of audio SSL models. In this paper, inspired by the success of data2vec 2.0 in image modality and Audio-MAE in audio modality, we introduce Efficient Audio Transformer (EAT) to further improve the effectiveness and efficiency in audio SSL. The proposed EAT adopts the bootstrap self-supervised training paradigm to the audio domain. A novel Utterance-Frame Objective (UFO) is designed to enhance the modeling capability of acoustic events. Furthermore, we reveal that the masking strategy is critical in audio SSL pre-training, and superior audio representations can be obtained with large inverse block masks. Experiment results demonstrate that EAT achieves state-of-the-art (SOTA) performance on a range of audio-related tasks, including AudioSet (AS-2M, AS-20K), ESC-50, and SPC-2, along with a significant pre-training speedup up to ~15x compared to existing audio SSL models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Quantitative Analysis of Proxy Tasks for Anomalous Sound Detection

    eess.AS 2026-01 conditional novelty 6.0 of 10

    Better proxy-task performance does not generally improve anomalous sound detection; only source separation showed a strong, consistent positive correlation.

  2. SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

    cs.SD 2025-06 conditional novelty 6.0 of 10

    SSLAM pre-trains audio transformers on partially mixed audio clips with a source retention loss, improving polyphonic sound tagging while keeping monophonic benchmark scores.

  3. Hidden-Domain Routing for All-Type Audio Deepfake Detection

    cs.SD 2026-08 accept novelty 5.0 of 10

    A router-then-specialist audio deepfake detector, which classifies audio type first and then applies type-specific models and thresholds, achieved 96.10% Macro-F1 and first place on AT-ADD Track2.

  4. IMPACT: Industrial Machine Perception via Acoustic Cognitive Transformer

    cs.SD 2025-07 conditional novelty 5.0 of 10

    With DINOS, a new 1,093-hour industrial-sound dataset, the authors show that pretraining a transformer (IMPACT, an EAT adaptation) on machine audio beats general audio models on 24 of 30 self-built monitoring tasks.

  5. RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection

    cs.SD 2025-05 conditional novelty 5.0 of 10

    A reconstruction-perception-reinforcement-attention framework for audio deepfake detection reports state-of-the-art EERs on ASVspoof 2019/2021 and competitive results on sound and singing benchmarks.

Pith tools