REVIEW 5 cited by
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Transformers have rapidly become the preferred choice for audio classification, surpassing methods based on CNNs. However, Audio Spectrogram Transformers (ASTs) exhibit quadratic scaling due to self-attention. The removal of this quadratic self-attention cost presents an appealing direction. Recently, state space models (SSMs), such as Mamba, have demonstrated potential in language and vision tasks in this regard. In this study, we explore whether reliance on self-attention is necessary for audio classification tasks. By introducing Audio Mamba (AuM), the first self-attention-free, purely SSM-based model for audio classification, we aim to address this question. We evaluate AuM on various audio datasets - comprising six different benchmarks - where it achieves comparable or better performance compared to well-established AST model.
Forward citations
Cited by 5 Pith papers
-
Deep Active Speech Cancellation with Mamba-Masking Network
DeepASC, a Mamba-masking multi-band network with a per-example optimal-target loss, reports up to 7.2 dB NMSE improvement in simulated active noise and speech cancellation.
-
Latent Mamba Operator for Partial Differential Equations
LaMO replaces attention in latent-token neural operators with bidirectional state-space models and reports consistent accuracy gains on six PDE benchmarks.
-
TAME: Temporal Audio-based Mamba for Enhanced Drone Trajectory Estimation and Classification
TAME applies parallel Mamba state-space models to audio spectrograms and reports state-of-the-art drone trajectory estimation and classification on MMAUD, with unresolved evaluation concerns.
-
BadScan: An Architectural Backdoor Attack on Visual State Space Models
BadScan is a trigger-activated architectural backdoor for VMamba that replaces the standard 2D selective scan with malformed scans at inference time.
-
Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation
A Transformer-Mamba model that adds a learned correction signal to degraded speech beats adapted active-noise-control baselines on denoising, dereverberation, and declipping in simulation.
Discussion (0). Continue with ORCID to comment.