REVIEW 4 major objections 1 minor 68 references
Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's Alternative
T0 review · 4 major / 1 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Fake-Mamba claims bidirectional Mamba with an XLSR front-end detects synthetic speech at 0.97% EER on ASVspoof 21 LA, 1.74% on 21 DF, and 5.85% on In-The-Wild, outperforming XLSR-Conformer and XLSR-Mamba while preserving real-time inference
desk verdict The submitted full text is an unrelated medical-imaging paper, so the abstract's EER claims for Fake-Mamba are unverifiable and the manuscript cannot be seriously reviewed as submitted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the bidirectional Mamba encoder, a state-space sequence model that processes speech forward and backward, used as a drop-in alternative to self-attention. The paper proposes three variants; the named best performer is PN-BiMamba. Paired with the XLSR front-end, the encoder extracts both local and global artifacts that distinguish synthetic from natural speech.
What would settle it
Retrain XLSR-Conformer and XLSR-Mamba under identical training folds, optimizer settings, and evaluation scripts on ASVspoof 21 LA, then compare equal error rates with Fake-Mamba's reported 0.97%. If the gap closes or reverses, the central superiority claim fails.
Extended reading notes
Core claim
On the paper's own terms, bidirectional Mamba is competitive with or superior to self-attention for detecting synthetic speech. The system, Fake-Mamba, uses an XLSR pre-trained front-end for linguistic representations, then one of three proposed Mamba encoders—TransBiMamba, ConBiMamba, or PN-BiMamba—to model local and global artifacts. PN-BiMamba is the best at capturing subtle cues of synthetic speech. Evaluated on standard benchmarks, it reports state-of-the-art equal error rates and real-time inference across utterance lengths.
Load-bearing premise
The reported equal error rates beat the baselines only if XLSR-Conformer and XLSR-Mamba were retrained under the exact same data splits, evaluation protocol, and hyperparameter budget; the abstract does not spell out those controls.
Editorial extensions
If this is right
- If the EER numbers hold under matched conditions, Mamba-based encoders become a strong candidate for real-time synthetic speech detection on long utterances.
- The success of bidirectional Mamba suggests attention is not necessary for capturing both local and global artifacts in speech forensics; other audio classification tasks may adopt similar encoders.
- The three encoder variants form a lightweight design space, and the best one, PN-BiMamba, could be plugged into other front-ends for forensic or anti-spoofing systems.
- Real-time inference across utterance lengths implies deployability in streaming or interactive settings where latency constraints previously favored lighter architectures.
Reading between the lines
- The reported gains are relative to specific baselines; if XLSR-Conformer and XLSR-Mamba were not retrained with identical data splits, evaluation protocol, and compute budgets, the margins may shrink or reverse.
- The architecture may generalize beyond the three tested benchmarks to other spoofing types or unseen generators, but that is not demonstrated here.
- The XLSR front-end likely dominates parameter count and inference cost; the real-time claim may hold only for the encoder component, not the full pipeline.
- If the Mamba encoder works on raw features rather than fine-tuned front-end outputs, it could transfer to other audio forensics tasks such as voice conversion or replay detection.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submission, as identified by its abstract (arXiv:2508.09294), claims to present Fake-Mamba, a speech deepfake detector combining an XLSR front-end with bidirectional Mamba encoders (TransBiMamba, ConBiMamba, PN-BiMamba). The abstract reports EERs of 0.97%, 1.74%, and 5.85% on ASVspoof 21 LA, ASVspoof 21 DF, and In-The-Wild, respectively, with 'substantial relative gains' over XLSR-Conformer and XLSR-Mamba, and real-time inference. However, the full text supplied is an entirely different manuscript: 'Ethical Medical Image Synthesis' by Jin et al. (arXiv:2508.09293). The body contains no description of Fake-Mamba, no architectural details, no experimental setup, no EER tables, no baseline comparisons, and no runtime measurements. The central claims of the abstract are therefore completely unsupported by the accompanying document.
Significance. If the reported results were accompanied by a proper method description, experimental protocol, and reproducible evaluation, the work would be significant for the speech deepfake detection community: replacing self-attention with bidirectional Mamba while maintaining real-time performance on three standard benchmarks would be a useful contribution. As submitted, however, the significance cannot be assessed. The manuscript supplies no derivations, no ablation, no statistical tests, and no falsifiable details. The only verifiable content is a medical-image-synthesis ethics paper that is unrelated to the claimed topic. Thus, the potential significance of the underlying idea does not translate into significance of this submission.
major comments (4)
- [Full text (all sections)] The body of the manuscript is 'Ethical Medical Image Synthesis' by Jin, Sinha, Abhishek, and Hamarneh (arXiv:2508.09293), which contains no mention of Fake-Mamba, TransBiMamba, ConBiMamba, PN-BiMamba, XLSR, ASVspoof, In-The-Wild, EER, or speech deepfake detection. The abstract's quantitative claims — 0.97%, 1.74%, and 5.85% EER — are asserted with no supporting derivation or evaluation in the document. This is a load-bearing evidentiary failure: the central claim of the paper is unverifiable from the submitted text.
- [Abstract (comparison claims)] The abstract claims 'substantial relative gains' over XLSR-Conformer and XLSR-Mamba. No table, figure, or textual description reports how these baselines were configured, whether they were retrained under identical data splits and hyperparameter budgets, or which scoring metric was used. Without this information, the claimed improvements cannot be checked and may reflect protocol differences rather than architectural merit.
- [Abstract (real-time claim)] The abstract states that the framework 'maintains real-time inference across utterance lengths.' No latency measurements, hardware specifications, batch sizes, or utterance-length sweeps are provided anywhere in the manuscript. This claim is therefore unsupported and not reproducible from the submitted document.
- [Abstract (method description)] The core innovation is described only as three named encoders: TransBiMamba, ConBiMamba, and PN-BiMamba. The manuscript contains no equations, no architectural diagrams, no pseudocode, and no ablations isolating the contribution of each encoder. The central research question — whether bidirectional Mamba can serve as a competitive alternative to self-attention — cannot be evaluated because the proposed architecture is never defined.
minor comments (1)
- [General] The GitHub repository link in the abstract cannot be assessed from the manuscript; no code snapshot, license, or documentation is included. If this submission is the result of a metadata/upload error, the authors should be asked to provide the corrected full text corresponding to the abstract.
Circularity Check
No circular derivation observable: Fake-Mamba's claimed results have no supporting text in the supplied manuscript, so there is no derivation chain to be circular.
full rationale
The supplied full text is arXiv:2508.09293, 'Ethical Medical Image Synthesis' by Jin et al., which shares no content with the Fake-Mamba abstract: there is no description of TransBiMamba, ConBiMamba, PN-BiMamba, the XLSR front-end, the ASVspoof/In-The-Wild protocols, EER tables, or real-time benchmarks. The abstract of arXiv:2508.09294 asserts the central results (0.97%, 1.74%, 5.85% EER) and the SOTA comparison, but no method, equations, fitted parameters, or evaluation details are present to walk. An absent derivation cannot be circular: there is no equation that reduces to an input, no fitted value renamed as a prediction, and no load-bearing self-citation. The manuscript-text mismatch is a serious verifiability/evidentiary failure, and the quantitative claims should be treated as unsupported by this document, but it is not a circularity defect under the stated criteria. Hence the circularity score is 0.
Assumptions & free parameters
assumptions (2)
- domain assumption The XLSR front-end provides rich linguistic representations that encode subtle cues of synthetic speech.
- domain assumption Bidirectional Mamba can capture both local and global artifacts in speech better than self-attention.
Cite this review
Pith. "Pith review of Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's Alternative." pith.science (2026). https://pith.science/paper/7GWET7AC
@misc{pith2026250809294,
author = {Pith},
title = {Pith review of: Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's Alternative},
year = {2026},
howpublished = {\url{https://pith.science/paper/7GWET7AC}},
note = {Machine review of arXiv:2508.09294}
}
read the original abstract
Advances in speech synthesis intensify security threats, motivating real-time deepfake detection research. We investigate whether bidirectional Mamba can serve as a competitive alternative to Self-Attention in detecting synthetic speech. Our solution, Fake-Mamba, integrates an XLSR front-end with bidirectional Mamba to capture both local and global artifacts. Our core innovation introduces three efficient encoders: TransBiMamba, ConBiMamba, and PN-BiMamba. Leveraging XLSR's rich linguistic representations, PN-BiMamba can effectively capture the subtle cues of synthetic speech. Evaluated on ASVspoof 21 LA, 21 DF, and In-The-Wild benchmarks, Fake-Mamba achieves 0.97%, 1.74%, and 5.85% EER, respectively, representing substantial relative gains over SOTA models XLSR-Conformer and XLSR-Mamba. The framework maintains real-time inference across utterance lengths, demonstrating strong generalization and practical viability. The code is available at https://github.com/xuanxixi/Fake-Mamba.
Reference graph
Works this paper leans on
-
[1]
Y. Jeon, Y. Kim, and G. G. Lee, ``Enhancing zero-shot multi-speaker tts with negated speaker representations,'' in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 16, 2024, pp. 18\,336--18\,344
work page 2024
-
[2]
C.-J. Hsu, Y.-C. Lin, C.-C. Lin, W.-C. Chen, H. L. Chung, C.-A. Li, Y.-C. Chen, C.-Y. Yu, M.-J. Lee, C.-C. Chen, R.-H. Huang, H. yi Lee, and D.-S. Shiu, ``Breezyvoice: Adapting tts for taiwanese mandarin with enhanced polyphone disambiguation -- challenges and insights,'' 2025. [Online]. Available: https://arxiv.org/abs/2501.17790
arXiv 2025
-
[3]
Y. K. Kan, K. Xu, H. Li et al., ``Voicedefense: Protecting automatic speaker verification models against black-box adversarial attacks,'' in Proc. Interspeech 2024, 2024, pp. 517--521
work page 2024
-
[5]
X. Xuan, K. kui Sin, Y. Zhou, and C. Kit, ``Translaw: Benchmarking large language models in multi-agent simulation of the collaborative translation,'' 2025. [Online]. Available: https://arxiv.org/abs/2507.00875
work page Pith review arXiv 2025
-
[6]
W. Zhang and C. Luo, ``Ge-gnn: Gated edge-augmented graph neural network for fraud detection,'' IEEE Transactions on Big Data, vol. 11, no. 4, pp. 1664--1676, 2025
work page 2025
-
[7]
B. Ding, R. Han, Z. Ma, and X. Xuan, ``Crowd density estimation based on multi-level attention maps,'' in 2021 IEEE 5th Information Technology,Networking,Electronic and Automation Control Conference (ITNEC), vol. 5, 2021, pp. 1759--1765
work page 2021
- [8]
-
[9]
J. Du, X. Chen, H. Wu, L. Zhang, I.-M. Lin, I.-H. Chiu, W. Ren, Y. Tseng, Y. Tsao, J.-S. R. Jang, and H. yi Lee, ``Codecfake-omni: A large-scale codec-based deepfake speech dataset,'' CoRR, vol. abs/2501.08238, January 2025. [Online]. Available: https://doi.org/10.48550/arXiv.2501.08238
Show all 68 references
-
[10]
Gulati, J
A. Gulati, J. Qin, C. C. Chiu et al., ``Conformer: Convolution-augmented transformer for speech recognition,'' pp. 5036--5040, 2020
2020
-
[11]
X. Xuan, R. Han, and J. Gao, ``Conformer-based speaker recognition model for real-time multi-scenarios,'' Computer Engineering and Applications, vol. 60, no. 7, pp. 147--156, 2024
2024
-
[12]
H. Shin, J. Heo, J. Kim et al., ``Hm-conformer: A conformer-based audio deepfake detection system with hierarchical pooling and multi-level classification token aggregation methods,'' in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing ...
2024
-
[13]
D. T. Truong, R. Tao, T. Nguyen et al., ``Temporal-channel modeling in multi-head self-attention for synthetic speech detection,'' in Proceedings of Interspeech 2024, 2024, pp. 537--541
2024
-
[14]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, L. Kaiser, and I. Polosukhin, ``Attention is all you need,'' in Advances in Neural Information Processing Systems, vol. 30. 1em plus 0.5em minus 0.4em NeurIPS, 2017. [Online]. Available: https://arxiv.org/abs...
2017 arXiv
-
[15]
J. Yang, R. K. Das, and H. Li, ``Significance of subband features for synthetic speech detection,'' IEEE Transactions on Information Forensics and Security, vol. 15, pp. 2160--2170, 2019
2019
-
[16]
Sriskandaraja, V
K. Sriskandaraja, V. Sethu, P. N. Le et al., ``Investigation of sub-band discriminative information between spoofed and genuine speech,'' in Interspeech, 2016, pp. 1710--1714
2016
-
[17]
Zhang, D
W. Zhang, D. Xu, X. Xuan, L. Jiang, G. Yao, R. Han, X. Lang, and C. Luo, ``Addressing noise and stochasticity in fraud detection for service networks,'' 2025. [Online]. Available: https://arxiv.org/abs/2505.00946
2025 arXiv
-
[18]
Zhang, D
W. Zhang, D. Xu, G. Yao, X. Lin, R. Guan, C. Du, R. Han, X. Xuan, and C. Luo, ``Frect: Frequency-augmented convolutional transformer for robust time series anomaly detection,'' in Advanced Intelligent Computing Technology and Applications, D.-S. Huang, W. Chen, Y. Pan, and H. ...
2025
-
[19]
Zhang and C
W. Zhang and C. Luo, ``Decomposition-based multi-scale transformer framework for time series anomaly detection,'' Neural Networks, vol. 187, p. 107399, 2025. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0893608025002783
2025
-
[20]
Z. Lin, J. Wang, R. Li, F. Shen, and X. Xuan, ``Primek-net: Multi-scale spectral learning via group prime-kernel convolutional neural networks for single channel speech enhancement,'' in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processin...
2025
-
[21]
Wu, H.-L
H. Wu, H.-L. Chung, Y.-C. Lin, Y.-K. Wu, X. Chen, Y.-C. Pai, H.-H. Wang, K.-W. Chang, A. Liu, and H.-y. Lee, ``Codec- SUPERB : An in-depth analysis of sound codec models,'' in Findings of the Association for Computational Linguistics: ACL 2024, L.-W. Ku, A. Martins, and V. Sri...
2024
-
[22]
Ren, Y.-C
W. Ren, Y.-C. Lin, H.-C. Chou, H. Wu, Y.-C. Wu, C.-C. Lee, H.-Y. Lee, H.-M. Wang, and Y. Tsao, ``Emo-codec: An in-depth look at emotion preservation capacity of legacy and neural codec models with subjective and objective evaluations,'' in 2024 Asia Pacific Signal and Informat...
2024
-
[23]
Gu and T
A. Gu and T. Dao, ``Mamba: Linear-time sequence modeling with selective state spaces,'' in First Conference on Language Modeling, 2024. [Online]. Available: https://openreview.net/forum?id=tEYskw1VY2
2024
-
[24]
H. Zhao, M. Zhang, W. Zhao, and et al., `` Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference ,'' in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 10, 2025, pp. 10\,421--10\,429
2025
-
[25]
B. Lenz, O. Lieber, A. Arazi, and et al., `` Jamba: Hybrid Transformer-Mamba Language Models ,'' in The Thirteenth International Conference on Learning Representations, 2025
2025
-
[26]
Waleffe, W
R. Waleffe, W. Byeon, D. Riach et al., ``An empirical study of mamba-based language models,'' arXiv preprint arXiv:2406.07887, 2024
2024 arXiv
-
[27]
L. Zhu, B. Liao, Q. Zhang et al., ``Vision mamba: Efficient visual representation learning with bidirectional state space model,'' in Forty-first International Conference on Machine Learning, 2024
2024
-
[28]
Z. Wang, F. Kong, S. Feng, and et al., `` Is Mamba Effective for Time Series Forecasting? '' Neurocomputing, vol. 619, p. 129178, 2025
2025
-
[29]
Q. Li, J. Qin, D. Cui, and et al., `` CMMamba: Channel Mixing Mamba for Time Series Forecasting ,'' Journal of Big Data, vol. 11, no. 1, p. 153, 2024
2024
-
[30]
Yamagishi, X
J. Yamagishi, X. Wang, M. Todisco et al., ``Asvspoof 2021: Accelerating progress in spoofed and deepfake speech detection,'' in ASVspoof 2021 Workshop - Automatic Speaker Verification and Spoofing Countermeasures Challenge, 2021
2021
-
[31]
2783--2787
Nicolas Müller and Pavel Czempin and Franziska Diekmann and Adam Froghyar and Konstantin Böttinger , `` Does Audio Deepfake Detection Generalize? '' in Interspeech 2022 , 2022 , pp. 2783--2787
2022
-
[32]
M. H. Erol, A. Senocak, J. Feng, and et al., `` Audio Mamba: Bidirectional State Space Model for Audio Representation Learning ,'' IEEE Signal Processing Letters, 2024
2024
-
[33]
Yadav and Z.-H
S. Yadav and Z.-H. Tan, `` Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations ,'' in Proceedings of the 25th International Conference on Speech Communication and Technology (Interspeech 2024), 2024, pp. 552--556
2024
-
[34]
Shams, S
S. Shams, S. S. Dindar, X. Jiang, and et al., `` SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model ,'' in 2024 IEEE Spoken Language Technology Workshop (SLT). 1em plus 0.5em minus 0.4em IEEE, 2024, pp. 1053--1059
2024
-
[35]
Lee and J
D. Lee and J. W. Choi, `` DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification ,'' in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1em plus 0.5em minus 0.4em IEEE, 2025, pp. 1--5
2025
-
[36]
Zhang and S
T. Zhang and S. Ruan, `` VM-ASR: A Lightweight Dual-Stream U-Net Model for Efficient Audio Super-Resolution ,'' IEEE Transactions on Audio, Speech and Language Processing, 2025
2025
-
[37]
Gao and N
X. Gao and N. F. Chen, `` Speech-Mamba: Long-Context Speech Recognition with Selective State Space Models ,'' in 2024 IEEE Spoken Language Technology Workshop (SLT). 1em plus 0.5em minus 0.4em IEEE, 2024, pp. 1--8
2024
-
[38]
Chao, W.-H
R. Chao, W.-H. Cheng, M. La Quatra, and et al., ``An investigation of incorporating mamba for speech enhancement,'' in 2024 IEEE Spoken Language Technology Workshop (SLT). 1em plus 0.5em minus 0.4em IEEE, 2024, pp. 302--308
2024
-
[39]
W. Ren, H. Wu et al., ``Leveraging joint spectral and spatial learning with mamba for multichannel speech enhancement,'' in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025, pp. 1--5
2025
-
[40]
Jiang, C
X. Jiang, C. Han, and N. Mesgarani, `` Dual-Path Mamba: Short and Long-Term Bidirectional Selective Structured State Space Models for Speech Separation ,'' in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1em plus 0.5em m...
2025
-
[41]
T. H. Avenstrup, B. Elek, I. L. M \'a di, and et al., `` SepMamba: State-Space Models for Speaker Separation Using Mamba ,'' in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1em plus 0.5em minus 0.4em IEEE, 2025, pp. 1--5
2025
-
[42]
Plaquet, N
A. Plaquet, N. Tawara, M. Delcroix, and et al., `` Mamba-Based Segmentation Model for Speaker Diarization ,'' in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1em plus 0.5em minus 0.4em IEEE, 2025, pp. 1--5
2025
-
[43]
C. Fan, Y. Gao, Z. Pan, J. Zhang, H. Zhang, J. Zhang, and Z. Lv, ``Improved feature extraction network for neuro-oriented target speaker extraction,'' in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025, pp. 1--5
2025
-
[44]
X. Xuan, J. Dong, and T. Xuan, ``Research on front-end of asv system based on mel spectrum in noise scenario,'' in 2022 IEEE 10th Joint International Information Technology and Artificial Intelligence Conference (ITAIC), vol. 10, 2022, pp. 2638--2642
2022
-
[45]
Xuan and R
X. Xuan and R. Han, ``Research on acoustic feature extractor for automatic speaker verification systerm,'' in 2022 IEEE 10th Joint International Information Technology and Artificial Intelligence Conference (ITAIC), vol. 10, 2022, pp. 2628--2633
2022
-
[46]
X. Xuan, R. Jin, T. Xuan, G. Du, and K. Xuan, ``Multi-scene robust speaker verification system built on improved ecapa-tdnn,'' in 2022 IEEE 6th Advanced Information Technology, Electronic and Automation Control Conference (IAEAC ), 2022, pp. 1689--1693
2022
-
[47]
X. Xuan, R. Han, and B. Ding, ``Research on speaker identification models based on cnn and additive angular margin loss,'' in 2021 2nd International Conference on Electronics, Communications and Information Technology (CECIT), 2021, pp. 1046--1050
2021
-
[48]
X. Xuan, Z. Zhu, and C. Kit, ``Efficient real-time multi-scenario speaker recognition with mel-spectrogram-based hybrid tdnn for edge system,'' in INTERSPEECH 2024-Young Female* Researchers in Speech Workshop (YFRSW 2024), 2024
2024
-
[49]
Y. Chen, J. Yi, J. Xue et al., ``Rawbmamba: End-to-end bidirectional state space model for audio deepfake detection,'' pp. 2720--2724, 2024
2024
-
[50]
Xiao and R
Y. Xiao and R. K. Das, ``Xlsr-mamba: A dual-column bidirectional state space model for spoofing attack detection,'' IEEE Signal Processing Letters, vol. 32, pp. 1276--1280, 2025
2025
-
[51]
J. D. Hamilton, ``State-space models,'' Handbook of Econometrics, vol. 4, pp. 3039--3080, 1986
1986
-
[52]
2278--2282
Arun Babu and Changhan Wang and Andros Tjandra and Kushal Lakhotia and Qiantong Xu and Naman Goyal and Kritika Singh and Patrick von Platen and Yatharth Saraf and Juan Pino and Alexei Baevski and Alexis Conneau and Michael Auli , `` XLS-R: Self-supervised Cross-lingual Speech ...
2022
-
[53]
Baevski, Y
A. Baevski, Y. Zhou, A. Mohamed, and et al., ``wav2vec 2.0: A framework for self-supervised learning of speech representations,'' Advances in Neural Information Processing Systems, vol. 33, pp. 12\,449--12\,460, 2020
2020
-
[54]
X. Xuan, Y. Xiao, R. K. Das, and T. Kinnunen, ``Multilingual source tracing of speech deepfakes: A first benchmark,'' arXiv preprint arXiv:2508.04143, 2025
2025 arXiv
-
[55]
Wang, Z.-C
S.-H. Wang, Z.-C. Chen, J. Shi, M.-T. Chuang, G.-T. Lin, K.-P. Huang, D. Harwath, S.-W. Li, and H. yi Lee, ``How to learn a new language? an efficient solution for self-supervised learning models unseen languages adaption in low-resource scenario,'' 2025. [Online]. Available: ...
2025 arXiv
-
[56]
Lin, Y.-C
H.-C. Lin, Y.-C. Lin et al., ``Improving speech emotion recognition in under-resourced languages via speech-to-speech translation with bootstrapping data selection,'' in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025,...
2025
-
[57]
W., ``Investigating self-supervised front ends for speech spoofing countermeasures,'' in The Speaker and Language Recognition Workshop (Odyssey 2022), 2022, pp
X. W., ``Investigating self-supervised front ends for speech spoofing countermeasures,'' in The Speaker and Language Recognition Workshop (Odyssey 2022), 2022, pp. 112--119
2022
-
[58]
Zhang, Q
X. Zhang, Q. Zhang, H. Liu, T. Xiao, X. Qian, B. Ahmed, E. Ambikairajah, H. Li, and J. Epps, ``Mamba in speech: Towards an alternative to self-attention,'' IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 1933--1948, 2025
1933
-
[59]
Rosello, A
E. Rosello, A. Gomez-Alanis, A. M. Gomez, and A. Peinado, ``A conformer-based classifier for variable-length utterance processing in anti-spoofing,'' in Interspeech 2023, 2023, pp. 5281--5285
2023
-
[60]
Zhang, S
Q. Zhang, S. Wen, and T. Hu, `` Audio deepfake detection with self-supervised XLS-R and SLS classifier ,'' in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 6765--6773
2024
-
[61]
Todisco, X
M. Todisco, X. Wang et al., `` ASVspoof 2019: Future horizons in spoofed and fake audio detection ,'' pp. 1008--1012, 2019
2019
-
[62]
H. Tak, M. Kamble, J. Patino et al., ``Rawboost: A raw data boosting and augmentation method applied to automatic speaker verification anti-spoofing,'' in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1em plus 0.5em minus 0...
2022
-
[63]
Wang, W.-N
C. Wang, W.-N. Hsu, Y. Adi, A. Polyak, A. Lee, P.-J. Chen, J. Gu, and J. Pino, ``fairseq s 2: A scalable and integrable speech synthesis toolkit,'' in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, H. Adel and S. ...
2021
-
[64]
Y. Gao, C. Herold, Z. Yang, and H. Ney, ``Revisiting checkpoint averaging for neural machine translation,'' in Findings of the Association for Computational Linguistics: AACL-IJCNLP 2022, Y. He, H. Ji, S. Li, Y. Liu, and C.-H. Chang, Eds. 1em plus 0.5em minus 0.4em Association...
2022
-
[65]
X. Xuan, R. Han, S. Ji, and B. Ding, ``Research on clothing image classification models based on cnn and transfer learning,'' in 2021 IEEE 5th Advanced Information Technology, Electronic and Automation Control Conference (IAEAC), vol. 5, 2021, pp. 1461--1466
2021
-
[66]
Sholokhov, M
A. Sholokhov, M. Sahidullah, and T. Kinnunen, ``Semi-supervised speech activity detection with an application to automatic speaker verification,'' Computer Speech & Language, vol. 47, pp. 132--156, 2018
2018
-
[67]
Arora, W
S. Arora, W. Hu, and P. K. Kothari, ``An analysis of the t-sne algorithm for data visualization,'' in Conference on Learning Theory. 1em plus 0.5em minus 0.4em PMLR, 2018, pp. 1455--1462
2018
-
[68]
A. Cui, C. Zhao, X. Deng, G. Jiang, Y. Yang, G. Yao, R. Han, W. Zhang, and X. Xuan, ``Unlocking the full potential of separable convolutions on tensor cores,'' in International Conference on Intelligent Computing. 1em plus 0.5em minus 0.4em Springer, 2025, pp. 39--50
2025
-
[69]
fairseq S 2: A Scalable and Integrable Speech Synthesis Toolkit
11em plus .33em minus .07em 4000 4000 100 4000 4000 500 `\.=1000 = #1 \@IEEEnotcompsoconly \@IEEEcompsoconly #1 * [1] 0pt [0pt][0pt] #1 * [1] 0pt [0pt][0pt] #1 * \| ** #1 \@IEEEauthorblockNstyle \@IEEEcompsocnotconfonly \@IEEEauthorblockAstyle \@IEEEcompsocnotconfonly \@IEEEco...
2024 arXiv
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.