Pith. sign in

REVIEW 4 cited by

SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.11831 v2 pith:W4PCLO7K submitted 2024-05-20 eess.AS cs.LG

classification eess.AScs.LG
keywords audiossambaself-supervisedlearningmambamodeltasksrepresentation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformers have revolutionized deep learning across various tasks, including audio representation learning, due to their powerful modeling capabilities. However, they often suffer from quadratic complexity in both GPU memory usage and computational inference time, affecting their efficiency. Recently, state space models (SSMs) like Mamba have emerged as a promising alternative, offering a more efficient approach by avoiding these complexities. Given these advantages, we explore the potential of SSM-based models in audio tasks. In this paper, we introduce Self-Supervised Audio Mamba (SSAMBA), the first self-supervised, attention-free, and SSM-based model for audio representation learning. SSAMBA leverages the bidirectional Mamba to capture complex audio patterns effectively. We incorporate a self-supervised pretraining framework that optimizes both discriminative and generative objectives, enabling the model to learn robust audio representations from large-scale, unlabeled datasets. We evaluated SSAMBA on various tasks such as audio classification, keyword spotting, and speaker identification. Our results demonstrate that SSAMBA outperforms the Self-Supervised Audio Spectrogram Transformer (SSAST) in most tasks. Notably, SSAMBA is approximately 92.7% faster in batch inference speed and 95.4% more memory-efficient than SSAST for the tiny model size with an input token size of 22k. These efficiency gains, combined with superior performance, underscore the effectiveness of SSAMBA's architectural innovation, making it a compelling choice for a wide range of audio processing applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deep Active Speech Cancellation with Mamba-Masking Network

    cs.SD 2025-02 conditional novelty 6.0 of 10

    DeepASC, a Mamba-masking multi-band network with a per-example optimal-target loss, reports up to 7.2 dB NMSE improvement in simulated active noise and speech cancellation.

  2. Recognizing Dementia from Neuropsychological Tests with State Space Models

    cs.LG 2025-07 reject novelty 5.0 of 10

    Demenba applies Mamba/VMamba state space models to full-length neuropsychological interview audio and reports improved dementia classification AUC over an EfficientNet baseline, but the evaluation has significant sele...

  3. PARROT: Synergizing Mamba and Attention-based SSL Pre-Trained Models via Parallel Branch Hadamard Optimal Transport for Speech Emotion Recognition

    eess.AS 2025-06 conditional novelty 5.0 of 10

    PARROT reports state-of-the-art speech emotion recognition by fusing Audio-MAMBA with attention-based SSL models using parallel Hadamard product and optimal transport branches.

  4. Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation

    eess.AS 2025-05 conditional novelty 3.0 of 10

    A Transformer-Mamba model that adds a learned correction signal to degraded speech beats adapted active-noise-control baselines on denoising, dereverberation, and declipping in simulation.

Pith tools