REVIEW 2 major objections 5 minor 3 cited by
ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech
T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read ASVspoof 5 is a crowdsourced, 2,000-speaker benchmark for speech deepfake and spoofing detection.
desk verdict The ASVspoof 5 database paper is the definitive description of the next generation benchmark for speech deepfake detection; it deserves a serious referee and will be widely cited. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the ASVspoof 5 database itself: a collection of genuine and generated speech split into speaker-disjoint training, development, and evaluation partitions, plus a separate auxiliary set from about 30,000 additional speakers for training speaker encoders. Three mechanisms carry the argument. First, the MLS English source provides acoustic diversity because each volunteer recorded in their own home or office, which is what makes the benchmark more representative than the earlier studio-quality corpus of about 100 speakers. Second, surrogate ASV and detector models are trained and held out so that attack contributors can tune their algorithms against open-source stand-ins, simulating a realistic black-box adversary. Third, a post-processing pipeline rescales peak amplitude, randomly trims leading and trailing non-speech segments, caps utterance duration, and removes metadata, and the evaluation set adds twelve encoding and compression conditions, all to prevent detectors from exploiting accidental cues rather than genuine signs of synthesis.
What would settle it
Train a spoofing detector on the ASVspoof 5 training set and test it on an independent corpus of conversational or telephone speech containing both genuine and synthetic utterances; if the error-rate ranking of attacks on that corpus diverges sharply from the ranking reported on ASVspoof 5, or if the post-processed shortcut-artefact distributions still separate genuine from synthetic audio for a newly contributed attack, the claim that the benchmark reflects in-the-wild conditions would be weakened. A more direct check is to compute the equal error rate of a classifier that uses only the five identified shortcut cues on the released evaluation set, since any cue that still yields an equal error rate well below 50% would show that the cleaning pipeline has not fully removed the artefact.
Extended reading notes
Core claim
The central discovery the paper seeks to establish is that the ASVspoof 5 database is a valid, more realistic replacement for earlier ASVspoof benchmarks. Genuine (bona fide) utterances are drawn from the MLS English corpus of volunteer audiobook recordings, giving thousands of speakers in varied, non-studio acoustic environments, while spoofed and deepfake utterances are contributed by independent TTS and voice-conversion experts using 32 attack algorithms, seven of them adversarial filtering attacks applied on top of TTS/VC outputs. The protocol splits the MLS source into seven speaker-disjoint partitions so that attack models, surrogate detectors, and the final training, development, and evaluation sets never share speakers. Validation with two spoofing-detector baselines, one automatic speaker verification baseline, and two spoofing-aware verification baselines shows pooled detector equal error rates of 29.49% and 17.83% on development data and 36.04% and 29.12% on evaluation data, with higher error rates under encoding and compression, and shows that legacy unit-selection TTS attacks remain among the most dangerous. The paper also reports a post-processing pipeline that visibly reduces shortcut artefacts such as peak amplitude and utterance-duration cues, so that performance differences are less likely to come from dataset quirks.
Load-bearing premise
The load-bearing premise is that home-recorded volunteer audiobook speech is representative enough of real-world speech that detector performance on ASVspoof 5 will generalize to operational settings such as telephone calls and online media; if that representativeness fails, the main motivation for replacing the studio-quality corpus weakens.
Editorial extensions
If this is right
- Detectors trained and evaluated on ASVspoof 5 should yield performance estimates that transfer better to non-studio, codec-degraded speech than estimates from studio-quality datasets.
- Legacy unit-selection TTS attacks such as A12 and A19 remain among the most effective at defeating both speaker verification and spoofing detectors, so countermeasure research cannot concentrate exclusively on modern neural TTS and voice conversion.
- Adversarial filtering attacks tuned against surrogate ASV and detector models transfer to raise error rates of unseen baselines, particularly Malacopula on verification and Malafide on detectors, and their combination is especially potent.
- Encoding and compression, including neural codecs and telephone-bandwidth pipelines, consistently degrade all baseline systems and should be part of realistic evaluation protocols.
- Because the database and baselines are publicly released, future attack algorithms can be benchmarked under the same protocols, making detector progress measurable across time.
Reading between the lines
- If the volunteer audiobook acoustics turn out to match operational conditions such as telephone or videoconference speech, the codec conditions in the evaluation set make ASVspoof 5 a natural testbed for encoding-robust detectors, a capability earlier studio benchmarks lacked.
- The 30,000-speaker auxiliary partition could be used to measure how the scale of speaker-encoder training data changes attack success, a question the paper's protocol enables but does not itself answer.
- Because adversarial attacks were tuned against open-source surrogate models, real-world adversaries who target a specific deployed system may achieve even higher error rates; the reported threat levels are plausibly a lower bound rather than an upper bound.
- The shortcut-reduction pipeline could be adopted as a standard cleaning step for future synthetic-speech corpora, since it targets cues that are unrelated to whether speech is genuine.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports the design, collection, and validation of the ASVspoof 5 database, the fifth edition of a well-known benchmark for speech spoofing and deepfake detection. The database is built from the MLS/LibriVox source corpus, increasing the number of speakers to roughly 2,000 and introducing crowdsourced acoustic variability, in contrast to earlier editions derived from studio-quality VCTK data. It contains 32 attack algorithms, including TTS, VC, and adversarial attacks, organized into seven speaker-disjoint partitions that separate data for attack training, surrogate model development/evaluation, and the main train/dev/eval sets. The paper also describes a post-processing pipeline that reduces shortcut artefacts (peak amplitude, non-speech durations, total duration, energy) and a set of codec/compression conditions applied to the evaluation set. Validation experiments with one ASV baseline, two CM baselines, and two SASV baselines show that the evaluation set is challenging, especially for unit-selection TTS and adversarial attacks, and that encoding/compression degrades detection performance.
Significance. If the resource is delivered as described, it will be a significant community asset: it is substantially larger in speaker coverage, includes the first adversarial attacks in the ASVspoof series, adopts disjoint protocols designed for attack and surrogate-model tuning, and is made publicly available with open baseline implementations. The shortcut-artifact analysis in Figure 5 provides a concrete, quantitative check that the post-processing reduces known cues, and the A17 'seen' speaker check in Section 3.4 is a transparent treatment of an unavoidable data-overlap risk. The main weakness is that the paper's motivating claim of improved in-the-wild reliability is not directly validated by any cross-dataset experiment; the current experiments only measure performance on data from the same MLS source distribution.
major comments (2)
- [Section 1 and Section 7] The paper motivates the new database by arguing that MLS/LibriVox recordings are more representative of in-the-wild conditions than VCTK studio data and 'should help provide more reliable estimates' of attack impact and CM performance (Section 1, third paragraph). However, all validation experiments in Section 7 are performed on the ASVspoof 5 evaluation set, which is drawn from the same MLS source distribution, so the design hypothesis is never tested directly. No cross-dataset experiment (e.g., training a CM on ASVspoof 5 and testing on ASVspoof 2019 LA, or on an independent noisy/conversational corpus) is reported. I recommend either adding such an experiment or explicitly restating the in-the-wild generalization as an untested design goal rather than a validated property.
- [Section 2.1 and Section 2.5] The paper states in Section 2.1 that MLS recordings were made in diverse acoustic conditions, and Section 2.5 describes speaker/utterance selection criteria, but no quantitative description of the acoustic variability across the seven partitions is provided (e.g., SNR, background noise type, device distribution, number of recording sessions per speaker). Without such measurements, the central 'more representative than VCTK' claim rests on an appeal to the source corpus's nature rather than on evidence about the actual partitions used in the database. Please add at least a summary of basic acoustic statistics for the train, development, and evaluation partitions.
minor comments (5)
- [Section 7.5 and Section 7.1] The MOS-based analysis in Section 7.5 uses a pre-trained MOS predictor without human validation, and Section 7.1 notes that no human listeners were recruited. The comparative statements about perceptual quality (e.g., that A12 and A19 are easily distinguishable by human listeners) therefore rest on an unvalidated proxy; please add a caveat or, if feasible, a small human listening test for a subset of attacks.
- [Abstract and Section 2.6] The abstract mentions an auxiliary set from 'an additional 30k speakers', but Section 2.6 describes two subsets of approximately 13k speakers each. Please clarify the total unique speaker count to resolve the apparent discrepancy.
- [Section 8] The conclusion calls the validation 'comprehensive', but only two CM baselines and one ASV baseline are used. I suggest softening this wording or adding a third, feature-based CM baseline to make the validation more broadly representative.
- [Section 3.4, Section 7.1, Appendix D] There are several small typos: 'MarryTTS' should be 'MaryTTS' in Section 3.4; the baseline B04 is referred to as 'B4' in Section 7.1; and the title of Appendix D reads 'encoding/compressiong'. Please correct these.
- [Figure 8 and Section 6] The surrogate optimization results in Figure 8 cover only three of the five data contributors. A sentence explaining why the other two are absent (e.g., no optimization rounds were needed) would make the figure self-contained.
Circularity Check
No significant circularity: ASVspoof 5 is a resource/dataset paper whose validation baselines are empirical measurements, not derivations from the database design.
full rationale
The paper's central contribution is the construction and public release of a benchmark database, not a mathematical derivation whose conclusion is equivalent to an input. The motivating claim that MLS/LibriVox speech is more acoustically diverse than VCTK is an empirical assumption about the source corpus, not a result derived from the paper's own outputs. The baseline ASV/CM/SASV experiments in Section 7 are standard supervised evaluations: baselines are trained on the ASVspoof 5 training partition and scored on disjoint development and evaluation partitions, so the reported EERs and a-DCF values are external to the protocol design rather than fitted predictions. The surrogate-model attack tuning described in Sections 2.4 and 6 is openly disclosed as optimization of attack algorithms against open-source stand-in detectors, and the paper does not present the resulting attack EERs as evidence that the attacks generalize to unseen detectors; the architectural overlap between surrogate AASIST/RawNet2 and baseline B01/B02 is a realistic black-box attack scenario, not a hidden reuse of test labels. The shortcut-artifact post-processing in Section 4.2 removes the specific cues it targets, and Figure 5 confirms the distributions overlap after processing; this is a stated design objective, and the paper does not claim the EER reductions as independent predictions. The acknowledged lack of human MOS listening tests is a limitation for perceptual quality assessment, not a circularity. Self-citations, including the Malafide and Malacopula attack implementations, describe externally published tools borrowed for attack generation; they are descriptive and not load-bearing justifications of the database's validity. No circular step can be exhibited, so the appropriate score is 0.
Assumptions & free parameters
assumptions (2)
- domain assumption MLS/LibriVox recordings represent diverse, non-studio acoustic conditions similar to real-world deployments
- domain assumption The chosen baselines (ECAPA-TDNN ASV, RawNet2 and AASIST CMs, SASV systems) are representative enough to validate database difficulty
Cite this review
Pith. "Pith review of ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech." pith.science (2026). https://pith.science/paper/UDUTZOYK
@misc{pith2026250208857,
author = {Pith},
title = {Pith review of: ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech},
year = {2026},
howpublished = {\url{https://pith.science/paper/UDUTZOYK}},
note = {Machine review of arXiv:2502.08857}
}
read the original abstract
ASVspoof 5 is the fifth edition in a series of challenges which promote the study of speech spoofing and deepfake attacks as well as the design of detection solutions. We introduce the ASVspoof 5 database which is generated in a crowdsourced fashion from data collected in diverse acoustic conditions (cf. studio-quality data for earlier ASVspoof databases) and from ~2,000 speakers (cf. ~100 earlier). The database contains attacks generated with 32 different algorithms, also crowdsourced, and optimised to varying degrees using new surrogate detection models. Among them are attacks generated with a mix of legacy and contemporary text-to-speech synthesis and voice conversion models, in addition to adversarial attacks which are incorporated for the first time. ASVspoof 5 protocols comprise seven speaker-disjoint partitions. They include two distinct partitions for the training of different sets of attack models, two more for the development and evaluation of surrogate detection models, and then three additional partitions which comprise the ASVspoof 5 training, development and evaluation sets. An auxiliary set of data collected from an additional 30k speakers can also be used to train speaker encoders for the implementation of attack algorithms. Also described herein is an experimental validation of the new ASVspoof 5 database using a set of automatic speaker verification and spoof/deepfake baseline detectors. With the exception of protocols and tools for the generation of spoofed/deepfake speech, the resources described in this paper, already used by participants of the ASVspoof 5 challenge in 2024, are now all freely available to the community.
Figures
Figures from the paper (6 more)
Forward citations
Cited by 3 Pith papers
-
Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models
Speech DF Arena standardizes audio deepfake detection benchmarking across 14 datasets and 15 systems, showing that most open-source detectors have high error rates on out-of-domain attacks.
-
Open-Set Source Tracing of Audio Deepfake Systems
Softmax energy, a modified out-of-distribution score, improves open-set source tracing of audio deepfake systems, achieving a 31% relative FPR95 reduction and best FPR95 of 8.3% with augmentation.
-
Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis
Rectified flows with a tunable sampling temperature offer the best naturalness-diversity trade-off among stochastic prosody predictors for text-to-speech.
Reference graph
Works this paper leans on
-
[1]
Evans, T
N. Evans, T. Kinnunen, J. Yamagishi, Spoofing and countermeasures for automatic speaker verification, in: Proc. Interspeech, 2013, pp. 925–929
2013
-
[2]
Z. Wu, T. Kinnunen, N. Evans, J. Yamagishi, C. Hanil¸ ci, M. Sahidullah, A. Sizov, ASVspoof 2015: the first automatic speaker verification spoofing and countermeasures challenge, in: Proc. Interspeech, 2015, pp. 2037–2041
2015
-
[3]
Z. Wu, J. Yamagishi, T. Kinnunen, C. Hanil¸ ci, M. Sahidullah, A. Sizov, N. Evans, M. Todisco, H. Delgado, ASVspoof: The automatic speaker verification spoofing and countermeasures challenge, IEEE Journal of Selected Topics in Signal Processing 11 (4) (2017) 588–604
2017
-
[4]
Kinnunen, M
T. Kinnunen, M. Sahidullah, H. Delgado, M. Todisco, N. Evans, J. Yamagishi, K. Lee, The ASVspoof 2017 challenge: Assessing the limits of replay spoofing attack detection, in: Proc. Interspeech, 2017, pp. 2–6
2017
-
[5]
Todisco, X
M. Todisco, X. Wang, V. Vestman, M. Sahidullah, H. Delgado, A. Nautsch, J. Yamagishi, N. Evans, T. Kinnunen, K. A. Lee, ASVspoof 2019: future horizons in spoofed and fake audio detection, in: Proc. Interspeech, 2019, pp. 1008–1012
2019
-
[6]
Yamagishi, X
J. Yamagishi, X. Wang, M. Todisco, M. Sahidullah, J. Patino, A. Nautsch, X. Liu, K. A. Lee, T. Kinnunen, N. Evans, H. Delgado, ASVspoof 2021: Accelerating progress in spoofed and deepfake speech detection, in: Proc. ASVspoof Challenge Workshop, 2021, pp. 47–54
2021
-
[7]
X. Wang, H. Delgado, H. Tak, J.-w. Jung, H.-j. Shim, M. Todisco, I. Kukanov, X. Liu, M. Sahidullah, T. Kinnunen, N. Evans, K. A. Lee, J. Yamagishi, ASVspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale, in: ASVspoof Workshop 2024, 2024, pp. 1–8
2024
-
[8]
Yamagishi, C
J. Yamagishi, C. Veaux, K. MacDonald, CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92) (2019)
2019
Show all 128 references
-
[9]
X. Wang, J. Yamagishi, M. Todisco, H. Delgado, A. Nautsch, N. Evans, M. Sahidullah, V. Vestman, T. Kinnunen, K. A. Lee, L. Juvela, P. Alku, Y.-H. Peng, H.-T. Hwang, Y. Tsao, H.-M. Wang, S. L. Maguer, M. Becker, F. Henderson, R. Clark, Y. Zhang, Q. Wang, Y. Jia, K. Onuma, K. Mu...
2020
-
[10]
Lorenzo-Trueba, J
J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, T. Kinnunen, Z. Ling, The Voice Conversion Challenge 2018: Promoting development of parallel and nonparallel methods, in: Proc. Odyssey, 2018, pp. 195–202
2018
-
[11]
Zhao, W.-C
Y. Zhao, W.-C. Huang, X. Tian, J. Yamagishi, R. K. Das, T. Kinnunen, Z.-H. Ling, T. Toda, Voice Conversion Challenge 2020 — Intra-lingual semi-parallel and cross-lingual voice conversion —, in: Proc. Joint Workshop for the Blizzard Challenge and Voice Conversion Challenge 2020...
2020
-
[12]
Geirhos, et al., Shortcut learning in deep neural networks, Nature Machine Intelligence 2 (11) (2020) 665–673
R. Geirhos, et al., Shortcut learning in deep neural networks, Nature Machine Intelligence 2 (11) (2020) 665–673
2020
-
[13]
Chettri, E
B. Chettri, E. Benetos, B. L. T. Sturm, Dataset artefacts in anti-spoofing systems: A case study on the ASVspoof 2017 benchmark, IEEE/ACM Transactions on Audio, Speech, and Language Processing 28 (2020) 3018–3028
2020
-
[14]
M¨ uller, F
N. M¨ uller, F. Dieckmann, P. Czempin, R. Canals, K. B¨ ottinger, J. Williams, Speech is silver, silence is golden: What do ASVspoof-trained models really learn?, in: Proc. ASVspoof Challenge workshop, 2021, pp. 55–60
2021
-
[15]
H.-j. Shim, R. G. Hautam¨ aki, M. Sahidullah, T. Kinnunen, How to construct perfect and worse-than-coin-flip spoofing countermeasures: A word of warning on shortcut learning, in: Proc. Interspeech, 2023, pp. 785–789
2023
-
[16]
Zhang, Z
Y. Zhang, Z. Li, J. Lu, H. Hua, W. Wang, P. Zhang, The impact of silence on speech anti-spoofing, IEEE/ACM Transactions on Audio, Speech, and Language Processing 31 (2023) 3374–3389
2023
-
[17]
X. Liu, X. Wang, M. Sahidullah, J. Patino, H. Delgado, T. Kinnunen, M. Todisco, J. Yamagishi, N. Evans, A. Nautsch, K. A. Lee, ASVspoof 2021: Towards spoofed and deepfake speech detection in the wild, IEEE/ACM Transactions on Audio, Speech, and Language Processing 31 (2023) 2507–2522
2023
-
[18]
Pratap, Q
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, R. Collobert, Mls: A large-scale multilingual dataset for speech research, in: Proc. Interspeech, 2020, pp. 2757–2761
2020
-
[19]
Kearns, Librivox: Free public domain audiobooks, Reference Reviews 28 (1) (2014) 7–8
J. Kearns, Librivox: Free public domain audiobooks, Reference Reviews 28 (1) (2014) 7–8
2014
-
[20]
Szegedy, W
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, R. Fergus, Intriguing properties of neural networks, in: Proc. ICLR, 2013
2013
-
[21]
I. J. Goodfellow, J. Shlens, C. Szegedy, Explaining and harnessing adversarial examples, in: Proc. ICLR, 2015
2015
-
[22]
Z. Wu, P. Swietojanski, C. Veaux, S. Renals, S. King, A study of speaker adaptation for DNN-based speech synthesis, in: Proc. Interspeech, 2015, pp. 879–883
2015
-
[23]
S. Arik, J. Chen, K. Peng, W. Ping, Y. Zhou, Neural voice cloning with a few samples, in: Proc.NeurIPS, Vol. 31, 2018
2018
-
[24]
Y. Yan, X. Tan, B. Li, G. Zhang, T. Qin, S. Zhao, Y. Shen, W.-Q. Zhang, T.-Y. Liu, Adaspeech 3: Adaptive text to speech for spontaneous style, in: Proc. Interspeech, 2021, pp. 4668–4672
2021
-
[25]
M. Chen, X. Tan, B. Li, Y. Liu, T. Qin, S. Zhao, T.-Y. Liu, Adaspeech: Adaptive text to speech for custom voice, in: Proc. ICLR, 2021
2021
-
[26]
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. L. Moreno, Y. Wu, et al., Transfer learning from speaker verification to multispeaker text-to-speech synthesis, in: Proc. NeurIPS, 2018, pp. 4480–4490
2018
-
[27]
Cooper, C.-I
E. Cooper, C.-I. Lai, Y. Yasuda, F. Fang, X. Wang, N. Chen, J. Yamagishi, Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings, in: Proc. ICASSP, 2020, pp. 6184–6188
2020
-
[28]
Y. Wu, X. Tan, B. Li, L. He, S. Zhao, R. Song, T. Qin, T.-Y. Liu, Adaspeech 4: Adaptive text to speech in zero-shot scenarios, in: Proc. Interspeech, 2022, pp. 2568–2572
2022
-
[29]
X. Tan, T. Qin, F. Soong, T.-Y. Liu, A survey on neural speech synthesis, arXiv preprint arXiv:2106.15561 (2021). 31
2021 arXiv
-
[30]
Papernot, P
N. Papernot, P. McDaniel, I. Goodfellow, Transferability in machine learning: From phenomena to black-box attacks using adversarial samples, arXiv preprint arXiv:1605.07277 (2016)
2016 arXiv
-
[31]
Panayotov, G
V. Panayotov, G. Chen, D. Povey, S. Khudanpur, Librispeech: An ASR corpus based on public domain audio books, in: Proc. ICASSP, 2015, pp. 5206–5210
2015
-
[32]
Mohamed, H.-y
A. Mohamed, H.-y. Lee, L. Borgholt, J. D. Havtorn, J. Edin, C. Igel, K. Kirchhoff, S.-W. Li, K. Livescu, L. Maaløe, T. N. Sainath, S. Watanabe, Self-supervised speech representation learning: A review, IEEE Journal of Selected Topics in Signal Processing 16 (6) (2022) 1179–1210
2022
-
[33]
Ardila, M
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, G. Weber, Common Voice: A massively-multilingual speech corpus, in: Proc. LREC, 2020, pp. 4218–4222
2020
-
[34]
Delgado, N
H. Delgado, N. Evans, J.-w. Jung, T. Kinnunen, I. Kukanov, K.-A. Lee, X. Liu, H.-j. Shim, M. Sahidullah, H. Tak, M. Todisco, X. Wang, J. Yamagishi, ASVspoof 5 evaluation plan (phase 2),https://www.asvspoof.org/file/ASVspoof5_ __Evaluation_Plan_Phase2.pdf, v0.6, accessed 23-Jul...
2024
-
[35]
Better be computer or I’m dumb
K. Warren, T. Tucker, A. Crowder, D. Olszewski, A. Lu, C. Fedele, M. Pasternak, S. Layton, K. Butler, C. Gates, et al., “Better be computer or I’m dumb”: A large-scale evaluation of humans as audio deepfake detectors, in: Proc. ACM CCS, 2024, pp. 2696–2710
2024
-
[36]
J.-w. Jung, X. Wang, N. Evans, S. Watanabe, H.-j. Shim, H. Tak, S. Arora, J. Yamagishi, J. S. Chung, To what extent can ASV systems naturally defend against spoofing attacks?, in: Proc. Interspeech, 2024, pp. 3240–3244
2024
-
[37]
Taylor, Text-to-speech synthesis, Cambridge University Press, 2009
P. Taylor, Text-to-speech synthesis, Cambridge University Press, 2009
2009
-
[38]
S. Liu, H. Wu, H.-Y. Lee, H. Meng, Adversarial attacks on spoofing countermeasures of automatic speaker verification, in: Proc. ASRU, 2019, pp. 312–319
2019
-
[39]
L. Wang, J. Li, Y. Luo, J. Zheng, L. Wang, H. Li, K. Xu, C. Fang, J. Shi, Z. Wu, Advsv: An over-the-air adversarial attack dataset for speaker verification, in: Proc. ICASSP, IEEE, 2024, pp. 4555–4559
2024
-
[40]
X. Li, J. Zhong, X. Wu, J. Yu, X. Liu, H. Meng, Adversarial attacks on GMM i-vector based speaker verification systems, in: Proc. ICASSP, IEEE, 2020, pp. 6579–6583
2020
-
[41]
Panariello, W
M. Panariello, W. Ge, H. Tak, M. Todisco, N. Evans, Malafide: a novel adversarial convolutive noise attack against deepfake and spoofing detection systems, in: Proc. Interspeech, 2023, pp. 2868–2872
2023
-
[42]
Todisco, M
M. Todisco, M. Panariello, X. Wang, H. Delgado, K.-A. Lee, N. Evans, Malacopula: Adversarial automatic speaker verification attacks using a neural-based generalised hammerstein model, in: Proc. ASVspoof Workshop 2024, 2024, pp. 94–100
2024
-
[43]
J. Kim, S. Kim, J. Kong, S. Yoon, Glow-TTS: A generative flow for text-to-speech via monotonic alignment search, in: Proc. NeurIPS, 2020, pp. 8067–8077
2020
-
[44]
Popov, I
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, M. Kudinov, Grad-TTS: A diffusion probabilistic model for text-to-speech, in: Proc. ICML, 2021, pp. 8599–8608
2021
-
[45]
La´ ncucki, Fastpitch: Parallel text-to-speech with pitch prediction, in: Proc
A. La´ ncucki, Fastpitch: Parallel text-to-speech with pitch prediction, in: Proc. ICASSP, 2021, pp. 6588–6592
2021
-
[46]
J. Kim, J. Kong, J. Son, Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech, in: Proc. ICML, 2021, pp. 5530–5540
2021
-
[47]
F. Lux, J. Koch, S. Meyer, T. Bott, N. Schauffler, P. Denisov, A. Schweitzer, N. T. Vu, The IMS Toucan system for the Blizzard Challenge 2023, in: Proc. Blizzard Challenge Workshop, ISCA, 2023, pp. 40–45
2023
-
[48]
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, et al., Natural TTS synthesis by conditioning wavenet on Mel spectrogram predictions, in: Proc. ICASSP, 2018, pp. 4779–4783
2018
-
[49]
M. Baas, H. Kamper, StarGAN-ZSVC: Towards zero-shot voice conversion in low-resource contexts, in: Artificial Intel- ligence Research. SACAIR 2021. Communications in Computer and Information Science, Vol. 1342, 2020, pp. 69–84
2021
-
[50]
Casanova, J
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. G¨ olge, M. A. Ponti, YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone, in: Proc. ICML, 2022, pp. 2709–2720
2022
-
[51]
E. A. AlBadawy, S. Lyu, Voice conversion using speech-to-speech neuro-style transfer, in: Proc. Interspeech, 2020, pp. 4726–4730
2020
-
[52]
C. Gong, X. Wang, E. Cooper, D. Wells, L. Wang, J. Dang, K. Richmond, J. Yamagishi, ZMM-TTS: Zero-shot mul- tilingual and multispeaker speech synthesis conditioned on self-supervised discrete speech representations, IEEE/ACM Transactions on Audio, Speech, and Language Processi...
2024
-
[53]
Schr¨ oder, M
M. Schr¨ oder, M. Charfuelan, S. Pammi, I. Steiner, Open source voice creation toolkit for the MARY TTS platform, in: Proc. Interspeech, 2011, pp. 3253–3256
2011
-
[54]
Popov, I
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, M. Kudinov, J. Wei, Diffusion-based voice conversion with fast maximum likelihood sampling scheme, in: Proc. ICLR, 2022
2022
-
[55]
Casanova, K
E. Casanova, K. Davis, E. G¨ olge, G. G¨ oknar, I. Gulea, L. Hart, A. Aljafari, J. Meyer, R. Morais, S. Olayemi, et al., XTTS: A massively multilingual zero-shot text-to-speech model, Proc. Interspeech (2024) 4978–4982
2024
-
[56]
D. P. Kingma, P. Dhariwal, Glow: Generative flow with invertible 1x1 convolutions, in: Proc. NeurIPS, 2018, pp. 10236–10245
2018
-
[57]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, Attention is all you need, in: Proc. NeurIPS, 2017, pp. 5998–6008
2017
-
[58]
J. Kong, J. Kim, J. Bae, HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis, in: Proc. NeurIPS, 2020, pp. 17022–17033
2020
-
[59]
Desplanques, J
B. Desplanques, J. Thienpondt, K. Demuynck, ECAPA-TDNN: Emphasized channel attention, propagation and aggre- gation in TDNN based speaker verification, in: Proc. Interspeech, 2020, pp. 3830–3834
2020
-
[60]
K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: Proc. CVPR, 2016, pp. 770–778
2016
-
[61]
G. Zhu, F. Jiang, Z. Duan, Y-vector: Multiscale waveform encoder for speaker embedding, in: Proc. Interspeech, 2021, 32 pp. 96–100
2021
-
[62]
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, B. Poole, Score-based generative modeling through stochastic differential equations, in: Proc. ICLR, 2021
2021
-
[63]
Ronneberger, P
O. Ronneberger, P. Fischer, T. Brox, U-net: Convolutional networks for biomedical image segmentation, in: Proc. MICCAI, Springer, 2015, pp. 234–241
2015
-
[64]
Snyder, D
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, S. Khudanpur, X-vectors: Robust DNN embeddings for speaker recognition, in: Proc. ICASSP, 2018, pp. 5329–5333
2018
-
[65]
Hayashi, R
T. Hayashi, R. Yamamoto, K. Inoue, T. Yoshimura, S. Watanabe, T. Toda, K. Takeda, Y. Zhang, X. Tan, ESPnet- TTS: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit, in: Proc. ICASSP, 2020, pp. 7654–7658
2020
-
[66]
F. Lux, J. Koch, A. Schweitzer, N. T. Vu, The IMS Toucan system for the Blizzard Challenge 2021, in: Proc. Blizzard Challenge Workshop, ISCA, 2021, pp. 14–19
2021
-
[67]
Y. Wang, D. Stanton, Y. Zhang, R.-S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, R. A. Saurous, Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis, in: Proc. ICML, 2018, pp. 5180–5189
2018
-
[68]
Gulati, J
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, et al., Conformer: Convolution-augmented transformer for speech recognition, in: Proc. Interspeech, 2020, pp. 5036–5040
2020
-
[69]
Y. Ren, J. Liu, Z. Zhao, Portaspeech: Portable and high-quality generative text-to-speech, in: Proc. NeurIPS, 2021, pp. 13963–13974
2021
-
[70]
F. Lux, J. Koch, N. T. Vu, Exact prosody cloning in zero-shot multispeaker text-to-speech, in: Proc. SLT, 2022, pp. 962–969
2022
-
[71]
Graves, Supervised sequence labelling with recurrent neural networks, Ph.D
A. Graves, Supervised sequence labelling with recurrent neural networks, Ph.D. thesis, TUM (2008)
2008
-
[72]
L. Wan, Q. Wang, A. Papir, I. L. Moreno, Generalized end-to-end loss for speaker verification, in: Proc. ICASSP, 2018, pp. 4879–4883
2018
-
[73]
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, Y. Bengio, Attention-based models for speech recognition, in: Proc. NeurIPS, 2015, pp. 577–585
2015
-
[74]
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, W. Chan, WaveGrad: Estimating gradients for waveform generation, in: Proc. ICLR, 2021
2021
-
[75]
Miyato, M
T. Miyato, M. Koyama, cGANs with projection discriminator, in: Proc. ICLR, 2018
2018
-
[76]
K. Cho, B. van Merri¨ enboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, Y. Bengio, Learning phrase repre- sentations using RNN encoder–decoder for statistical machine translation, in: Proc. EMNLP, 2014, pp. 1724–1734
2014
-
[77]
Prenger, R
R. Prenger, R. Valle, B. Catanzaro, WaveGlow: A flow-based generative network for speech synthesis, in: Proc. ICASSP, 2019, pp. 3617–3621
2019
-
[78]
Kaneko, H
T. Kaneko, H. Kameoka, K. Tanaka, N. Hojo, StarGAN-VC2: Rethinking conditional methods for StarGAN-based voice conversion, in: Proc. Interspeech, 2019, pp. 679–683
2019
-
[79]
J.-w. Jung, Y. Kim, H.-S. H. Heo, B.-J. Lee, Y. Kwon, J. S. Chung, Pushing the limits of raw waveform speaker recognition, in: Proc. Interspeech, 2022, pp. 2228–2232
2022
-
[80]
D. P. Kingma, M. Welling, Auto-encoding variational bayes, in: Proc. ICLR, 2014
2014
-
[81]
J.-Y. Zhu, T. Park, P. Isola, A. A. Efros, Unpaired image-to-image translation using cycle-consistent adversarial networks, in: Proc. CVPR, 2017, pp. 2223–2232
2017
-
[82]
A. v. d. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, K. Kavukcuoglu, Wavenet: A generative model for raw audio, arXiv preprint arXiv:1609.03499 (2016)
2016 arXiv
-
[83]
L. Sun, H. Wang, S. Kang, K. Li, H. Meng, Personalized, cross-lingual TTS using phonetic posteriorgrams, in: Proc. Interspeech, 2016, pp. 322–326
2016
-
[84]
H. Wang, S. Zheng, Y. Chen, L. Cheng, Q. Chen, Cam++: A fast and efficient network for speaker verification using context-aware masking, in: Proc. Interspeech, 2023, pp. 5301–5305
2023
-
[85]
Baevski, Y
A. Baevski, Y. Zhou, A. Mohamed, M. Auli, Wav2vec 2.0: A framework for self-supervised learning of speech represen- tations, in: Proc. NuerIPS, Vol. 33, 2020, pp. 12449–12460
2020
-
[86]
C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, E. Dupoux, VoxPopuli: A large- scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation, in: Proc. ACL, 2021, pp. 993–1003
2021
-
[87]
Jadoul, B
Y. Jadoul, B. Thompson, B. de Boer, Introducing Parselmouth: A Python interface to Praat, Journal of Phonetics 71 (2018) 1–15
2018
-
[88]
L. T. Nguyen, T. Pham, D. Q. Nguyen, XPhoneBERT: A pre-trained multilingual model for phoneme representations for text-to-speech, in: Proc. Interspeech, 2023, pp. 5506–5510
2023
-
[89]
Van Den Oord, O
A. Van Den Oord, O. Vinyals, et al., Neural discrete representation learning, in: Proc. NeurIPS, 2017, pp. 6309–6318
2017
-
[90]
H. Guo, F. Xie, X. Wu, F. K. Soong, H. Meng, MSMC-TTS: Multi-stage multi-codebook VQ-V AE based neural TTS, IEEE/ACM Transactions on Audio, Speech, and Language Processing 31 (2023) 1811–1824
2023
-
[91]
J. S. Chung, A. Nagrani, A. Zisserman, Voxceleb2: Deep speaker recognition, in: Proc. Interspeech, 2018, pp. 1086–1090
2018
-
[92]
Steiner, S
I. Steiner, S. Le Maguer, Creating new language and voice components for the updated MaryTTS text-to-speech synthesis platform, in: Proc. LREC, 2018, pp. 3171–3175
2018
-
[93]
Sagisaka, Speech synthesis by rule using an optimal selection of non-uniform synthesis units, in: Proc
Y. Sagisaka, Speech synthesis by rule using an optimal selection of non-uniform synthesis units, in: Proc. ICASSP, 1988, pp. 679–682
1988
-
[94]
A. J. Hunt, A. W. Black, Unit selection in a concatenative speech synthesis system using a large speech database, in: Proc. ICASSP, Vol. 1, IEEE, 1996, pp. 373–376
1996
-
[95]
S.-g. Lee, W. Ping, B. Ginsburg, B. Catanzaro, S. Yoon, BigVGAN: A universal neural vocoder with large-scale training, 33 in: Proc. ICLR, 2022
2022
-
[96]
Ziyin, T
L. Ziyin, T. Hartwig, M. Ueda, Neural networks fail to learn periodic functions and how to fix it, in: Proc. NeurIPS, 2020, pp. 1583–1594
2020
-
[97]
Kintzley, A
K. Kintzley, A. Jansen, H. Hermansky, Event selection from phone posteriorgrams using matched filters, in: Proc. Interspeech, 2011, pp. 1905–1908
2011
-
[98]
D´ efossez, G
A. D´ efossez, G. Synnaeve, Y. Adi, Real time speech enhancement in the waveform domain, in: Proc. Interspeech, 2020, pp. 3291–3295
2020
-
[99]
Eren, The Coqui TTS Team, Coqui TTS (Jan
G. Eren, The Coqui TTS Team, Coqui TTS (Jan. 2021). doi:10.5281/zenodo.6334862
2021 doi
-
[100]
H. S. Heo, B.-J. Lee, J. Huh, J. S. Chung, Clova baseline system for the VoxCeleb speaker recognition challenge 2020, arXiv preprint arXiv:2009.14153 (2020)
2020 arXiv
-
[101]
Radford, J
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al., Language models are unsupervised multitask learners, OpenAI blog 1 (8) (2019) 9
2019
-
[102]
D´ efossez, J
A. D´ efossez, J. Copet, G. Synnaeve, Y. Adi, High fidelity neural audio compression, Transactions on Machine Learning Research (2023)
2023
-
[103]
Novak, P
A. Novak, P. Lotton, L. Simon, Synchronized swept-sine: Theory, application, and implementation, J. Audio Eng. Soc 63 (10) (2015) 786–798
2015
-
[104]
Lapuschkin, S
S. Lapuschkin, S. W¨ aldchen, A. Binder, G. Montavon, W. Samek, K.-R. M¨ uller, Unmasking Clever Hans predictors and assessing what machines really learn, Nature communications 10 (1) (2019) 1096
2019
-
[105]
Radford, J
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, I. Sutskever, Robust speech recognition via large-scale weak supervision, in: Proc. ICML, 2023, pp. 28492–28518
2023
-
[106]
Kinnunen, H
T. Kinnunen, H. Li, An overview of text-independent speaker recognition: From features to supervectors, Speech com- munication 52 (1) (2010) 12–40
2010
-
[107]
Van der Maaten, G
L. Van der Maaten, G. Hinton, Visualizing data using t-SNE, JMLR 9 (11) (2008)
2008
-
[108]
Hastie, R
T. Hastie, R. Tibshirani, J. H. Friedman, J. H. Friedman, The elements of statistical learning: data mining, inference, and prediction, Vol. 2, Springer, 2009
2009
-
[109]
M. L. Waskom, Seaborn: statistical data visualization, Journal of Open Source Software 6 (60) (2021) 3021
2021
-
[110]
Ravanelli, T
M. Ravanelli, T. Parcollet, A. Moumen, S. de Langen, C. Subakan, P. Plantinga, Y. Wang, P. Mousavi, L. Della Libera, A. Ploujnikov, et al., Open-source conversational AI with SpeechBrain 1.0, JMLR 25 (333) (2024) 1–11
2024
-
[111]
Nagrani, J
A. Nagrani, J. S. Chung, A. Zisserman, VoxCeleb: A Large-Scale Speaker Identification Dataset, in: Proc. Interspeech, 2017, pp. 2616–2620
2017
-
[112]
Snyder, G
D. Snyder, G. Chen, D. Povey, MUSAN: A Music, Speech, and Noise Corpus, arXiv:1510.08484v1 (2015)
2015 arXiv
-
[113]
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, S. Khudanpur, A study on data augmentation of reverberant speech for robust speech recognition, in: Proc. ICASSP, 2017, pp. 5220–5224
2017
-
[114]
Jung, H.-S
J.-w. Jung, H.-S. Heo, H. Tak, H.-j. Shim, J. S. Chung, B.-J. Lee, H.-J. Yu, N. Evans, AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks, in: Proc. ICASSP, 2022, pp. 6367–6371
2022
-
[115]
H. Tak, J. Patino, M. Todisco, A. Nautsch, N. Evans, A. Larcher, End-to-end anti-spoofing with RawNet2, in: Proc. ICASSP, 2021, pp. 6369–6373
2021
-
[116]
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, Q. V. Le, SpecAugment: A simple data augmentation method for automatic speech recognition, in: Proc. Interspeech, 2019, pp. 2613–2617
2019
-
[117]
Jung, S.-b
J.-w. Jung, S.-b. Kim, H.-j. Shim, J.-h. Kim, H.-J. Yu, Improved RawNet with feature map scaling for text-independent speaker verification using raw waveforms, in: Proc. Interspeech, 2020, pp. 1496–1500
2020
-
[118]
Ravanelli, Y
M. Ravanelli, Y. Bengio, Speaker recognition from raw waveform with sincnet, in: Proc. SLT, 2018, pp. 1021–1028
2018
-
[119]
J.-w. Jung, H. Tak, H.-j. Shim, H.-S. Heo, B.-J. Lee, S.-W. Chung, H.-J. Yu, N. Evans, T. Kinnunen, SASV 2022: The first spoofing-aware speaker verification challenge, in: Proc. Interspeech, 2022, pp. 2893–2897
2022
-
[120]
S. H. Mun, H.-j. Shim, H. Tak, X. Wang, X. Liu, M. Sahidullah, M. Jeong, M. H. Han, M. Todisco, K. A. Lee, et al., Towards single integrated spoofing-aware speaker verification embeddings, in: Proc. Interspeech, 2023, pp. 3989–3993
2023
-
[121]
H.-j. Shim, H. Tak, X. Liu, H.-S. Heo, J.-w. Jung, J. S. Chung, S.-W. Chung, H.-J. Yu, B.-J. Lee, M. Todisco, et al., Baseline systems for the first spoofing-aware speaker verification challenge: Score and embedding fusion, in: Proc. Odyssey, 2022, pp. 330–337
2022
-
[122]
X. Wang, T. Kinnunen, L. Kong Aik, P.-G. Noe, J. Yamagishi, Revisiting and improving scoring fusion for spoofing-aware speaker verification using compositional data analysis, in: Proc. Interspeech, 2024, pp. 1110–1114
2024
-
[123]
Zhang, Z
Y. Zhang, Z. Lv, H. Wu, S. Zhang, P. Hu, Z. Wu, H.-y. Lee, H. Meng, MF A-conformer: Multi-scale feature aggregation conformer for automatic speaker verification, in: Proc. Interspeech, 2022, pp. 306–310
2022
-
[124]
X. Wang, J. Yamagishi, Spoofed training data for speech spoofing countermeasure can be efficiently created using neural vocoders, in: Proc. ICASSP, 2023, pp. 1–5
2023
-
[125]
Cooper, W.-C
E. Cooper, W.-C. Huang, T. Toda, J. Yamagishi, Generalization ability of MOS prediction networks, in: Proc. ICASSP, 2022, pp. 8442–8446
2022
-
[126]
Shim, J.-w
H.-j. Shim, J.-w. Jung, T. Kinnunen, et al., a-DCF: an architecture agnostic metric with application to spoofing-robust speaker verification, in: Proc. Speaker Odyssey, 2024, pp. 158–164
2024
-
[127]
D. A. Van Leeuwen, N. Br¨ ummer, An introduction to application-independent evaluation of speaker recognition systems, in: Speaker Classification I, Springer, 2007, pp. 330–353
2007
-
[128]
A. F. Martin, G. R. Doddington, T. Kamm, M. Ordowski, M. A. Przybocki, The DET curve in assessment of detection task performance., in: Eurospeech, Vol. 4, 1997, pp. 1895–1898. 34
1997
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.