REVIEW 7 cited by
Mamba in Speech: Towards an Alternative to Self-Attention
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Transformer and its derivatives have achieved success in diverse tasks across computer vision, natural language processing, and speech processing. To reduce the complexity of computations within the multi-head self-attention mechanism in Transformer, Selective State Space Models (i.e., Mamba) were proposed as an alternative. Mamba exhibited its effectiveness in natural language processing and computer vision tasks, but its superiority has rarely been investigated in speech signal processing. This paper explores solutions for applying Mamba to speech processing by discussing two typical speech processing tasks: speech recognition, which requires semantic and sequential information, and speech enhancement, which focuses primarily on sequential patterns. The experimental results confirm that bidirectional Mamba (BiMamba) consistently outperforms vanilla Mamba, highlighting the advantages of a bidirectional design for speech processing. Moreover, experiments demonstrate the effectiveness of BiMamba as an alternative to the self-attention module in the Transformer model and its derivates, particularly for the semantic-aware task. The crucial technologies for transferring Mamba to speech are then summarized in ablation studies and the discussion section, offering insights for extending this research to a broader scope of tasks.
Forward citations
Cited by 7 Pith papers
-
Deep Active Speech Cancellation with Mamba-Masking Network
DeepASC, a Mamba-masking multi-band network with a per-example optimal-target loss, reports up to 7.2 dB NMSE improvement in simulated active noise and speech cancellation.
-
Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering
A controlled 120-configuration study of EEND-VC speaker diarization finds finetuned WavLM encoders, Conformer/Mamba decoders, and longer chunks give the biggest gains, with the best system reaching state-of-the-art on...
-
Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss
HMamba, a hierarchical Mamba-based model with a decoupled cross-entropy loss, jointly performs pronunciation scoring and mispronunciation detection, reaching an MDD F1 of 63.85% on speechocean762.
-
The Alpha-Alternator: Dynamic Adaptation To Varying Noise Levels In Sequences Using The Vendi Score For Improved Robustness and Performance
A Vendi Score-based adaptive gating mechanism, trained with input masking, improves neural decoding and time-series forecasting accuracy relative to several baselines.
-
Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and Maliseet
Multilingual training improves objective quality for low-resource Ojibwe, Mi'kmaq, and Maliseet TTS, and attention-free architectures match self-attention with lower memory, but the improvement may be due mostly to la...
-
ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment
A unified open-source toolkit with pretrained models for speech enhancement, separation, super-resolution, and multimodal target extraction, plus a new face-conditioned extraction model.
-
Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation
A Transformer-Mamba model that adds a learned correction signal to degraded speech beats adapted active-noise-control baselines on denoising, dereverberation, and declipping in simulation.
Discussion (0). Continue with ORCID to comment.