Pith. sign in

REVIEW 37 cited by

AST: Audio Spectrogram Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.01778 v3 pith:UTUOOLIP submitted 2021-04-05 cs.SD cs.AI

classification cs.SDcs.AI
keywords audioclassificationaccuracymodelnetworksneuralpurelyspectrogram
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the past decade, convolutional neural networks (CNNs) have been widely adopted as the main building block for end-to-end audio classification models, which aim to learn a direct mapping from audio spectrograms to corresponding labels. To better capture long-range global context, a recent trend is to add a self-attention mechanism on top of the CNN, forming a CNN-attention hybrid model. However, it is unclear whether the reliance on a CNN is necessary, and if neural networks purely based on attention are sufficient to obtain good performance in audio classification. In this paper, we answer the question by introducing the Audio Spectrogram Transformer (AST), the first convolution-free, purely attention-based model for audio classification. We evaluate AST on various audio classification benchmarks, where it achieves new state-of-the-art results of 0.485 mAP on AudioSet, 95.6% accuracy on ESC-50, and 98.1% accuracy on Speech Commands V2.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 37 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing

    cs.SD 2025-07 conditional novelty 7.0 of 10

    MixAssist is the first audio-grounded, multi-turn conversational dataset for co-creative music mixing instruction, and fine-tuning Qwen-Audio on it yields human-comparable mixing advice.

  2. Device Invariance using Domain Adaptation on Acoustic Scene Classification

    eess.AS 2026-07 conditional novelty 6.0 of 10

    On the DCASE 2020 device-shift benchmark, DANN improves acoustic scene classification across CNN and transformer features, but CDAN fails to converge with the PaSST transformer.

  3. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  4. Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Automatically constructed synthetic exact-GT data plus multi-model pseudo-labels with interval-aware GRPO rewards improve LALM open-vocabulary audio event grounding on AEGBench and DESED.

  5. Missingness as Signal: Channel-Independent Spectrogram Learning for Clinical Time Series Prediction

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A channel-independent spectrogram model with an aligned missingness stream improves ICU mortality prediction by treating which measurements were taken as structured 2D signal.

  6. VAN-AD: Visual Masked Autoencoder with Normalizing Flow For Time Series Anomaly Detection

    cs.LG 2026-03 unverdicted novelty 6.0 of 10

    VAN-AD adapts a pretrained visual MAE with distribution mapping and normalizing flow modules to detect anomalies in time series data more effectively across different datasets.

  7. EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

    cs.CV 2025-12 conditional novelty 6.0 of 10

    EchoingPixels prunes audio-visual LLM input tokens jointly across modalities and re-tunes RoPE frequencies so that 5–20% of tokens retain roughly full-model performance.

  8. DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation

    cs.SD 2025-11 conditional novelty 6.0 of 10

    DHAuDS is a new audio benchmark that corrupts four existing datasets with dynamically varying and diverse acoustic noise, and evaluates three classifiers under test-time adaptation.

  9. Region-Specific Audio Tagging for Spatial Sound

    eess.AS 2025-09 conditional novelty 6.0 of 10

    Region-specific audio tagging: a model trained to tag sound events within a specified angular or distance region, tested on a new simulated benchmark and STARSS23.

  10. AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition

    cs.SD 2025-09 conditional novelty 6.0 of 10

    AudioRWKV, an RWKV7-based audio backbone with 2D convolution and bidirectional WKV, beats AST and AuM baselines on five benchmarks at linear complexity.

  11. Full-Frequency Temporal Patching and Structured Masking for Enhanced Audio Classification

    cs.SD 2025-08 conditional novelty 6.0 of 10

    Replacing square spectrogram patches with full-frequency temporal patches plus patch-aligned masking improves audio classification accuracy and reduces compute for Transformer and Mamba models.

  12. Tactile Beyond Pixels: Multisensory Touch Representations for Robot Manipulation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    Sparsh-X is a transformer trained on about one million unlabeled touch interactions that fuses image, audio, motion, and pressure into representations that boost downstream robot manipulation performance over tactile-...

  13. SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

    cs.SD 2025-06 conditional novelty 6.0 of 10

    SSLAM pre-trains audio transformers on partially mixed audio clips with a source retention loss, improving polyphonic sound tagging while keeping monophonic benchmark scores.

  14. PromptTSS: A Prompting-Based Approach for Interactive Multi-Granularity Time Series Segmentation

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A prompt-conditioned transformer segments time series at coarse and fine granularities in one model, reporting 24.49% and 17.88% accuracy gains over baselines and up to 599.24% in transfer settings.

  15. Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A contrastive audio-text framework detects hate speech in synthesized speech across six languages and outperforms baselines, with a new 127k-sample dataset.

  16. Acoustic Classification of Maritime Vessels using Learnable Filterbanks

    cs.SD 2025-05 conditional novelty 6.0 of 10

    CATFISH, a learnable Gabor filterbank model with attention pooling and optional environmental data, achieves 96.63% test accuracy on the multi-scenario VTUAD underwater vessel classification benchmark.

  17. Bridging the Gap Between Semantic and User Preference Spaces for Multi-modal Music Representation Learning

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A two-stage contrastive training method aligns song audio with text semantics and then with user-favored song pairs, improving music classification and recommendation over prior models.

  18. Time to Embed: Unlocking Foundation Models for Time Series with Channel Descriptions

    cs.LG 2025-05 reject novelty 6.0 of 10

    CHARM is a 7M-parameter self-supervised embedding model for multivariate time series that uses channel descriptions to beat specialized baselines on forecasting, classification, and anomaly detection.

  19. Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026

    cs.SD 2026-07 conditional novelty 5.0 of 10

    For BirdCLEF+ 2026, a frozen Perch-v2 probe plus a trained HGNetV2-B0 SED net and non-bird prototype heads reach private LB 0.936, while WavTokenizer codec tokens collapse and four general audio transformers lag under...

  20. Semantic Audio-driven Understanding for Dynamic Humanoid Whole Body Control

    cs.RO 2026-07 conditional novelty 5.0 of 10

    A multi-modal audio router maps streaming music and speech to imitation-learned whole-body policies for a Unitree G1 humanoid, achieving 84.8% chunk-level retrieval accuracy in simulation.

  21. Distance-aware Soft Prompt Guidance for Multimodal Valence-Arousal Estimation

    cs.CV 2026-03 reject novelty 5.0 of 10

    Distance-aware soft prompts over a 3×3 emotion grid with CLIP text prototypes and audio-visual GRU fusion achieve CCC_mean 0.5361 on Aff-Wild2, beating only the paper's self-defined baselines.

  22. Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models

    cs.SD 2026-03 conditional novelty 5.0 of 10

    Using CMA-ES to jointly optimize activation quantization scales keeps speech-model accuracy near full precision under full INT8 and INT4 quantization.

  23. Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment

    cs.CV 2025-07 reject novelty 5.0 of 10

    LMAC-Net reports state-of-the-art Spearman correlations on the RG and Fis-V benchmarks by aligning attention centers across RGB, optical flow, and audio branches.

  24. Your Attention Matters: to Improve Model Robustness to Noise and Spurious Correlations

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Doubly stochastic attention is the most robust of five ViT attention mechanisms to fog corruption in relative accuracy, based on single-seed experiments on CIFAR-10, CIFAR-100, and Imagenette.

  25. Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition

    cs.CV 2025-07 reject novelty 5.0 of 10

    DICCAE dynamically weights a confusion loss using measured inter-class overlap and reports 65.5% audio-visual top-1 on VGGSound, but the evaluation protocol uses test data during training.

  26. IMPACT: Industrial Machine Perception via Acoustic Cognitive Transformer

    cs.SD 2025-07 conditional novelty 5.0 of 10

    With DINOS, a new 1,093-hour industrial-sound dataset, the authors show that pretraining a transformer (IMPACT, an EAT adaptation) on machine audio beats general audio models on 24 of 30 self-built monitoring tasks.

  27. OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A shared-backbone transformer with pairwise modality training reports top results across 25 datasets spanning 12 modalities.

  28. Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types

    cs.SD 2026-07 conditional novelty 4.5 of 10

    On WiWa, factorised multi-task bird species and call-type classification with adaptive loss balancing improves call-type recognition most consistently, while preferred weighting and adaptation depth depend on backbone...

  29. LGFNet: A CTC-Guided Local-Global Fusion Framework for Single-Channel Sleep Staging

    cs.CV 2026-07 conditional novelty 4.0 of 10

    A CTC-guided local-global fusion network achieves state-of-the-art single-channel sleep staging, with notable gains on the N1 stage and across datasets with different sampling rates and montages.

  30. CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning

    cs.SD 2025-08 conditional novelty 4.0 of 10

    A self-supervised masked-spectrogram pretraining method for cough audio produces representations that match or exceed AudioSet-pretrained AST on three cough classification tasks.

  31. A Two-Step Learning Framework for Enhancing Sound Event Localization and Detection

    cs.SD 2025-07 conditional novelty 4.0 of 10

    A two-step SELD framework with separate DoA and SED training, trackwise label reordering, and beamformed feature fusion achieves a 0.3891 SELD score on the 2023 DCASE Task 3 development test set.

  32. Angle-distance decomposition based on deep learning for active sonar detection

    eess.SP 2025-07 conditional novelty 4.0 of 10

    A two-network deep learning system for active sonar, separating angle estimation from distance estimation, beats classical beamforming and spectrogram classifiers on simulated data.

  33. Deciphering GunType Hierarchy through Acoustic Analysis of Gunshot Recordings

    cs.SD 2025-06 conditional novelty 4.0 of 10

    A CNN on log-mel spectrograms classifies five gun types from curated audio with mAP 0.58, beating an SVM baseline (0.39), but falls to 0.35 on noisy web data.

  34. Comparison of spectrogram scaling in multi-label Music Genre Recognition

    cs.SD 2025-06 conditional novelty 4.0 of 10

    On a custom 18k-song multi-label dataset, Mel-scaled spectrograms outperform standard spectrograms for music genre classification with transfer-learned ResNets.

  35. Patient Domain Supervised Contrastive Learning for Lung Sound Classification Using Mobile Phone

    cs.SD 2025-05 conditional novelty 4.0 of 10

    Patient Domain Supervised Contrastive Learning (PD-SCL) improves lung sound classification on mobile-phone recordings by 2.4 points over an AST baseline, but the result relies on a small private dataset with no error bars.

  36. Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A superpoint-guided, scale-normalized tokenizer lets a frozen CLIP model perform 3D segmentation and classification without fine-tuning.

  37. 15,500 Seconds: Lean UAV Classification Using EfficientNet and Lightweight Fine-Tuning

    cs.LG 2025-05 reject novelty 4.0 of 10

    On a private 3,100-clip, 31-class drone audio dataset, full fine-tuning of EfficientNet-B0 with three augmentations reached 95.95% validation accuracy, the best of all compared models and PEFT methods.

Pith tools