Pith. sign in

REVIEW 7 cited by

Mamba in Speech: Towards an Alternative to Self-Attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.12609 v6 pith:24I6DOP7 submitted 2024-05-21 eess.AS cs.SD

classification eess.AScs.SD
keywords speechmambaprocessingtasksalternativeself-attentiontransformerbidirectional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformer and its derivatives have achieved success in diverse tasks across computer vision, natural language processing, and speech processing. To reduce the complexity of computations within the multi-head self-attention mechanism in Transformer, Selective State Space Models (i.e., Mamba) were proposed as an alternative. Mamba exhibited its effectiveness in natural language processing and computer vision tasks, but its superiority has rarely been investigated in speech signal processing. This paper explores solutions for applying Mamba to speech processing by discussing two typical speech processing tasks: speech recognition, which requires semantic and sequential information, and speech enhancement, which focuses primarily on sequential patterns. The experimental results confirm that bidirectional Mamba (BiMamba) consistently outperforms vanilla Mamba, highlighting the advantages of a bidirectional design for speech processing. Moreover, experiments demonstrate the effectiveness of BiMamba as an alternative to the self-attention module in the Transformer model and its derivates, particularly for the semantic-aware task. The crucial technologies for transferring Mamba to speech are then summarized in ablation studies and the discussion section, offering insights for extending this research to a broader scope of tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deep Active Speech Cancellation with Mamba-Masking Network

    cs.SD 2025-02 conditional novelty 6.0 of 10

    DeepASC, a Mamba-masking multi-band network with a per-example optimal-target loss, reports up to 7.2 dB NMSE improvement in simulated active noise and speech cancellation.

  2. Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A controlled 120-configuration study of EEND-VC speaker diarization finds finetuned WavLM encoders, Conformer/Mamba decoders, and longer chunks give the biggest gains, with the best system reaching state-of-the-art on...

  3. Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss

    eess.AS 2025-02 conditional novelty 5.0 of 10

    HMamba, a hierarchical Mamba-based model with a decoupled cross-entropy loss, jointly performs pronunciation scoring and mispronunciation detection, reaching an MDD F1 of 63.85% on speechocean762.

  4. The Alpha-Alternator: Dynamic Adaptation To Varying Noise Levels In Sequences Using The Vendi Score For Improved Robustness and Performance

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A Vendi Score-based adaptive gating mechanism, trained with input masking, improves neural decoding and time-series forecasting accuracy relative to several baselines.

  5. Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and Maliseet

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Multilingual training improves objective quality for low-resource Ojibwe, Mi'kmaq, and Maliseet TTS, and attention-free architectures match self-attention with lower memory, but the improvement may be due mostly to la...

  6. ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

    cs.SD 2025-06 conditional novelty 4.0 of 10

    A unified open-source toolkit with pretrained models for speech enhancement, separation, super-resolution, and multimodal target extraction, plus a new face-conditioned extraction model.

  7. Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation

    eess.AS 2025-05 conditional novelty 3.0 of 10

    A Transformer-Mamba model that adds a learned correction signal to degraded speech beats adapted active-noise-control baselines on denoising, dereverberation, and declipping in simulation.

Pith tools