Pith. sign in

REVIEW 2 major objections 5 minor 3 cited by

ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech

T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read ASVspoof 5 is a crowdsourced, 2,000-speaker benchmark for speech deepfake and spoofing detection.

desk verdict The ASVspoof 5 database paper is the definitive description of the next generation benchmark for speech deepfake detection; it deserves a serious referee and will be widely cited. read the letter →

arxiv 2502.08857 v4 pith:UDUTZOYK submitted 2025-02-13 eess.AS

classification eess.AS
keywords ASVspoof5speechdeepfakedetectionspoofingcountermeasurescrowdsourcedcorpusadversarialattacksautomaticspeakerverificationshortcutartefactscodecrobustness
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper presents ASVspoof 5, a new public benchmark for detecting speech deepfakes, spoofing attacks, and adversarial attacks on speaker verification. Its central claim is that a database built from volunteer-recorded audiobook speech collected in homes and offices, from roughly 2,000 speakers rather than about 100, gives a more realistic picture of detector performance than earlier studio-recorded benchmarks. The database contains genuine speech plus speech generated by 32 different algorithms, including adversarial attacks for the first time, and applies encoding and compression conditions that mimic real transmission channels. The authors argue the resource is validated by baseline experiments showing that attacks remain a serious threat and that realistic acoustic and codec variation makes detection harder. If the claim holds, the community gains a testbed whose results transfer better to in-the-wild conditions.

What carries the argument

The load-bearing object is the ASVspoof 5 database itself: a collection of genuine and generated speech split into speaker-disjoint training, development, and evaluation partitions, plus a separate auxiliary set from about 30,000 additional speakers for training speaker encoders. Three mechanisms carry the argument. First, the MLS English source provides acoustic diversity because each volunteer recorded in their own home or office, which is what makes the benchmark more representative than the earlier studio-quality corpus of about 100 speakers. Second, surrogate ASV and detector models are trained and held out so that attack contributors can tune their algorithms against open-source stand-ins, simulating a realistic black-box adversary. Third, a post-processing pipeline rescales peak amplitude, randomly trims leading and trailing non-speech segments, caps utterance duration, and removes metadata, and the evaluation set adds twelve encoding and compression conditions, all to prevent detectors from exploiting accidental cues rather than genuine signs of synthesis.

What would settle it

Train a spoofing detector on the ASVspoof 5 training set and test it on an independent corpus of conversational or telephone speech containing both genuine and synthetic utterances; if the error-rate ranking of attacks on that corpus diverges sharply from the ranking reported on ASVspoof 5, or if the post-processed shortcut-artefact distributions still separate genuine from synthetic audio for a newly contributed attack, the claim that the benchmark reflects in-the-wild conditions would be weakened. A more direct check is to compute the equal error rate of a classifier that uses only the five identified shortcut cues on the released evaluation set, since any cue that still yields an equal error rate well below 50% would show that the cleaning pipeline has not fully removed the artefact.

Watch

Extended reading notes

Core claim

The central discovery the paper seeks to establish is that the ASVspoof 5 database is a valid, more realistic replacement for earlier ASVspoof benchmarks. Genuine (bona fide) utterances are drawn from the MLS English corpus of volunteer audiobook recordings, giving thousands of speakers in varied, non-studio acoustic environments, while spoofed and deepfake utterances are contributed by independent TTS and voice-conversion experts using 32 attack algorithms, seven of them adversarial filtering attacks applied on top of TTS/VC outputs. The protocol splits the MLS source into seven speaker-disjoint partitions so that attack models, surrogate detectors, and the final training, development, and evaluation sets never share speakers. Validation with two spoofing-detector baselines, one automatic speaker verification baseline, and two spoofing-aware verification baselines shows pooled detector equal error rates of 29.49% and 17.83% on development data and 36.04% and 29.12% on evaluation data, with higher error rates under encoding and compression, and shows that legacy unit-selection TTS attacks remain among the most dangerous. The paper also reports a post-processing pipeline that visibly reduces shortcut artefacts such as peak amplitude and utterance-duration cues, so that performance differences are less likely to come from dataset quirks.

Load-bearing premise

The load-bearing premise is that home-recorded volunteer audiobook speech is representative enough of real-world speech that detector performance on ASVspoof 5 will generalize to operational settings such as telephone calls and online media; if that representativeness fails, the main motivation for replacing the studio-quality corpus weakens.

Editorial extensions

If this is right

  • Detectors trained and evaluated on ASVspoof 5 should yield performance estimates that transfer better to non-studio, codec-degraded speech than estimates from studio-quality datasets.
  • Legacy unit-selection TTS attacks such as A12 and A19 remain among the most effective at defeating both speaker verification and spoofing detectors, so countermeasure research cannot concentrate exclusively on modern neural TTS and voice conversion.
  • Adversarial filtering attacks tuned against surrogate ASV and detector models transfer to raise error rates of unseen baselines, particularly Malacopula on verification and Malafide on detectors, and their combination is especially potent.
  • Encoding and compression, including neural codecs and telephone-bandwidth pipelines, consistently degrade all baseline systems and should be part of realistic evaluation protocols.
  • Because the database and baselines are publicly released, future attack algorithms can be benchmarked under the same protocols, making detector progress measurable across time.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the volunteer audiobook acoustics turn out to match operational conditions such as telephone or videoconference speech, the codec conditions in the evaluation set make ASVspoof 5 a natural testbed for encoding-robust detectors, a capability earlier studio benchmarks lacked.
  • The 30,000-speaker auxiliary partition could be used to measure how the scale of speaker-encoder training data changes attack success, a question the paper's protocol enables but does not itself answer.
  • Because adversarial attacks were tuned against open-source surrogate models, real-world adversaries who target a specific deployed system may achieve even higher error rates; the reported threat levels are plausibly a lower bound rather than an upper bound.
  • The shortcut-reduction pipeline could be adopted as a standard cleaning step for future synthetic-speech corpora, since it targets cues that are unrelated to whether speech is genuine.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper reports the design, collection, and validation of the ASVspoof 5 database, the fifth edition of a well-known benchmark for speech spoofing and deepfake detection. The database is built from the MLS/LibriVox source corpus, increasing the number of speakers to roughly 2,000 and introducing crowdsourced acoustic variability, in contrast to earlier editions derived from studio-quality VCTK data. It contains 32 attack algorithms, including TTS, VC, and adversarial attacks, organized into seven speaker-disjoint partitions that separate data for attack training, surrogate model development/evaluation, and the main train/dev/eval sets. The paper also describes a post-processing pipeline that reduces shortcut artefacts (peak amplitude, non-speech durations, total duration, energy) and a set of codec/compression conditions applied to the evaluation set. Validation experiments with one ASV baseline, two CM baselines, and two SASV baselines show that the evaluation set is challenging, especially for unit-selection TTS and adversarial attacks, and that encoding/compression degrades detection performance.

Significance. If the resource is delivered as described, it will be a significant community asset: it is substantially larger in speaker coverage, includes the first adversarial attacks in the ASVspoof series, adopts disjoint protocols designed for attack and surrogate-model tuning, and is made publicly available with open baseline implementations. The shortcut-artifact analysis in Figure 5 provides a concrete, quantitative check that the post-processing reduces known cues, and the A17 'seen' speaker check in Section 3.4 is a transparent treatment of an unavoidable data-overlap risk. The main weakness is that the paper's motivating claim of improved in-the-wild reliability is not directly validated by any cross-dataset experiment; the current experiments only measure performance on data from the same MLS source distribution.

major comments (2)
  1. [Section 1 and Section 7] The paper motivates the new database by arguing that MLS/LibriVox recordings are more representative of in-the-wild conditions than VCTK studio data and 'should help provide more reliable estimates' of attack impact and CM performance (Section 1, third paragraph). However, all validation experiments in Section 7 are performed on the ASVspoof 5 evaluation set, which is drawn from the same MLS source distribution, so the design hypothesis is never tested directly. No cross-dataset experiment (e.g., training a CM on ASVspoof 5 and testing on ASVspoof 2019 LA, or on an independent noisy/conversational corpus) is reported. I recommend either adding such an experiment or explicitly restating the in-the-wild generalization as an untested design goal rather than a validated property.
  2. [Section 2.1 and Section 2.5] The paper states in Section 2.1 that MLS recordings were made in diverse acoustic conditions, and Section 2.5 describes speaker/utterance selection criteria, but no quantitative description of the acoustic variability across the seven partitions is provided (e.g., SNR, background noise type, device distribution, number of recording sessions per speaker). Without such measurements, the central 'more representative than VCTK' claim rests on an appeal to the source corpus's nature rather than on evidence about the actual partitions used in the database. Please add at least a summary of basic acoustic statistics for the train, development, and evaluation partitions.
minor comments (5)
  1. [Section 7.5 and Section 7.1] The MOS-based analysis in Section 7.5 uses a pre-trained MOS predictor without human validation, and Section 7.1 notes that no human listeners were recruited. The comparative statements about perceptual quality (e.g., that A12 and A19 are easily distinguishable by human listeners) therefore rest on an unvalidated proxy; please add a caveat or, if feasible, a small human listening test for a subset of attacks.
  2. [Abstract and Section 2.6] The abstract mentions an auxiliary set from 'an additional 30k speakers', but Section 2.6 describes two subsets of approximately 13k speakers each. Please clarify the total unique speaker count to resolve the apparent discrepancy.
  3. [Section 8] The conclusion calls the validation 'comprehensive', but only two CM baselines and one ASV baseline are used. I suggest softening this wording or adding a third, feature-based CM baseline to make the validation more broadly representative.
  4. [Section 3.4, Section 7.1, Appendix D] There are several small typos: 'MarryTTS' should be 'MaryTTS' in Section 3.4; the baseline B04 is referred to as 'B4' in Section 7.1; and the title of Appendix D reads 'encoding/compressiong'. Please correct these.
  5. [Figure 8 and Section 6] The surrogate optimization results in Figure 8 cover only three of the five data contributors. A sentence explaining why the other two are absent (e.g., no optimization rounds were needed) would make the figure self-contained.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: ASVspoof 5 is a resource/dataset paper whose validation baselines are empirical measurements, not derivations from the database design.

full rationale

The paper's central contribution is the construction and public release of a benchmark database, not a mathematical derivation whose conclusion is equivalent to an input. The motivating claim that MLS/LibriVox speech is more acoustically diverse than VCTK is an empirical assumption about the source corpus, not a result derived from the paper's own outputs. The baseline ASV/CM/SASV experiments in Section 7 are standard supervised evaluations: baselines are trained on the ASVspoof 5 training partition and scored on disjoint development and evaluation partitions, so the reported EERs and a-DCF values are external to the protocol design rather than fitted predictions. The surrogate-model attack tuning described in Sections 2.4 and 6 is openly disclosed as optimization of attack algorithms against open-source stand-in detectors, and the paper does not present the resulting attack EERs as evidence that the attacks generalize to unseen detectors; the architectural overlap between surrogate AASIST/RawNet2 and baseline B01/B02 is a realistic black-box attack scenario, not a hidden reuse of test labels. The shortcut-artifact post-processing in Section 4.2 removes the specific cues it targets, and Figure 5 confirms the distributions overlap after processing; this is a stated design objective, and the paper does not claim the EER reductions as independent predictions. The acknowledged lack of human MOS listening tests is a limitation for perceptual quality assessment, not a circularity. Self-citations, including the Malafide and Malacopula attack implementations, describe externally published tools borrowed for attack generation; they are descriptive and not load-bearing justifications of the database's validity. No circular step can be exhibited, so the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

No free parameters or invented entities are introduced; the database is assembled from existing speech data and known attack algorithms. The main assumptions are domain-level claims about representativeness.

assumptions (2)
  • domain assumption MLS/LibriVox recordings represent diverse, non-studio acoustic conditions similar to real-world deployments
    Invoked in Section 2.1 to justify replacing VCTK. No quantitative evidence is provided that LibriVox matches specific real-world conditions.
  • domain assumption The chosen baselines (ECAPA-TDNN ASV, RawNet2 and AASIST CMs, SASV systems) are representative enough to validate database difficulty
    Used in Section 7 to draw conclusions about attack threat. If baselines are unrepresentative or weak, the validation may mislead.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech." pith.science (2026). https://pith.science/paper/UDUTZOYK

@misc{pith2026250208857,
  author       = {Pith},
  title        = {Pith review of: ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UDUTZOYK}},
  note         = {Machine review of arXiv:2502.08857}
}
read the original abstract

ASVspoof 5 is the fifth edition in a series of challenges which promote the study of speech spoofing and deepfake attacks as well as the design of detection solutions. We introduce the ASVspoof 5 database which is generated in a crowdsourced fashion from data collected in diverse acoustic conditions (cf. studio-quality data for earlier ASVspoof databases) and from ~2,000 speakers (cf. ~100 earlier). The database contains attacks generated with 32 different algorithms, also crowdsourced, and optimised to varying degrees using new surrogate detection models. Among them are attacks generated with a mix of legacy and contemporary text-to-speech synthesis and voice conversion models, in addition to adversarial attacks which are incorporated for the first time. ASVspoof 5 protocols comprise seven speaker-disjoint partitions. They include two distinct partitions for the training of different sets of attack models, two more for the development and evaluation of surrogate detection models, and then three additional partitions which comprise the ASVspoof 5 training, development and evaluation sets. An auxiliary set of data collected from an additional 30k speakers can also be used to train speaker encoders for the implementation of attack algorithms. Also described herein is an experimental validation of the new ASVspoof 5 database using a set of automatic speaker verification and spoof/deepfake baseline detectors. With the exception of protocols and tools for the generation of spoofed/deepfake speech, the resources described in this paper, already used by participants of the ASVspoof 5 challenge in 2024, are now all freely available to the community.

Figures

Figures reproduced from arXiv: 2502.08857 by the authors.

Figure 1
Figure 1. Overview of the ASVspoof 5 database. Bona fide utterances contained in the speaker-disjoint training, development, [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. An illustration of the MLS English source database partitioning scheme and its use in generation of the ASVspoof 5 [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. An illustration of the Track 1 spoof/deepfake detection task which involves the design of a CM and the Track 2 [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Illustration of speech generation using zero-shot (top-middle) and few-shot (bottom-middle) TTS systems. Few [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Distributions of potential shortcut artefacts (different columns) [PITH_FULL_IMAGE:figures/full_fig_p017_5.png]
Figure 6
Figure 6. Figure 6: Illustration of bona fide and spoofed samples from each attack in a t-SNE embedding space (sub-figure on the left). [PITH_FULL_IMAGE:figures/full_fig_p018_6.png]
Figure 7
Figure 7. Figure 7: Dendrogram (hierarchical cluster map) of attacks in the ASVspoof 5 database. For visualisation, heatmap plots the [PITH_FULL_IMAGE:figures/full_fig_p020_7.png]
Figure 8
Figure 8. Figure 8: Progress in TTS and VC system optimisation, measured using the EERs of the surrogate models. [PITH_FULL_IMAGE:figures/full_fig_p021_8.png]
Figure 9
Figure 9. Figure 9: Distribution of ASV scores on the evaluation set when the trials are processed with encoding/compression (top [PITH_FULL_IMAGE:figures/full_fig_p024_9.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models

    cs.SD 2025-09 conditional novelty 6.0 of 10

    Speech DF Arena standardizes audio deepfake detection benchmarking across 14 datasets and 15 systems, showing that most open-source detectors have high error rates on out-of-domain attacks.

  2. Open-Set Source Tracing of Audio Deepfake Systems

    eess.AS 2025-07 conditional novelty 5.0 of 10

    Softmax energy, a modified out-of-distribution score, improves open-set source tracing of audio deepfake systems, achieving a 31% relative FPR95 reduction and best FPR95 of 8.3% with augmentation.

  3. Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis

    eess.AS 2025-06 conditional novelty 5.0 of 10

    Rectified flows with a tunable sampling temperature offer the best naturalness-diversity trade-off among stochastic prosody predictors for text-to-speech.

Reference graph

Works this paper leans on

128 extracted references · 69 canonical work pages · cited by 3 Pith papers

  1. [1]

    Evans, T

    N. Evans, T. Kinnunen, J. Yamagishi, Spoofing and countermeasures for automatic speaker verification, in: Proc. Interspeech, 2013, pp. 925–929

  2. [2]

    Z. Wu, T. Kinnunen, N. Evans, J. Yamagishi, C. Hanil¸ ci, M. Sahidullah, A. Sizov, ASVspoof 2015: the first automatic speaker verification spoofing and countermeasures challenge, in: Proc. Interspeech, 2015, pp. 2037–2041

  3. [3]

    Z. Wu, J. Yamagishi, T. Kinnunen, C. Hanil¸ ci, M. Sahidullah, A. Sizov, N. Evans, M. Todisco, H. Delgado, ASVspoof: The automatic speaker verification spoofing and countermeasures challenge, IEEE Journal of Selected Topics in Signal Processing 11 (4) (2017) 588–604

  4. [4]

    Kinnunen, M

    T. Kinnunen, M. Sahidullah, H. Delgado, M. Todisco, N. Evans, J. Yamagishi, K. Lee, The ASVspoof 2017 challenge: Assessing the limits of replay spoofing attack detection, in: Proc. Interspeech, 2017, pp. 2–6

  5. [5]

    Todisco, X

    M. Todisco, X. Wang, V. Vestman, M. Sahidullah, H. Delgado, A. Nautsch, J. Yamagishi, N. Evans, T. Kinnunen, K. A. Lee, ASVspoof 2019: future horizons in spoofed and fake audio detection, in: Proc. Interspeech, 2019, pp. 1008–1012

  6. [6]

    Yamagishi, X

    J. Yamagishi, X. Wang, M. Todisco, M. Sahidullah, J. Patino, A. Nautsch, X. Liu, K. A. Lee, T. Kinnunen, N. Evans, H. Delgado, ASVspoof 2021: Accelerating progress in spoofed and deepfake speech detection, in: Proc. ASVspoof Challenge Workshop, 2021, pp. 47–54

  7. [7]

    X. Wang, H. Delgado, H. Tak, J.-w. Jung, H.-j. Shim, M. Todisco, I. Kukanov, X. Liu, M. Sahidullah, T. Kinnunen, N. Evans, K. A. Lee, J. Yamagishi, ASVspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale, in: ASVspoof Workshop 2024, 2024, pp. 1–8

  8. [8]

    Yamagishi, C

    J. Yamagishi, C. Veaux, K. MacDonald, CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92) (2019)

Show all 128 references
  1. [9]

    X. Wang, J. Yamagishi, M. Todisco, H. Delgado, A. Nautsch, N. Evans, M. Sahidullah, V. Vestman, T. Kinnunen, K. A. Lee, L. Juvela, P. Alku, Y.-H. Peng, H.-T. Hwang, Y. Tsao, H.-M. Wang, S. L. Maguer, M. Becker, F. Henderson, R. Clark, Y. Zhang, Q. Wang, Y. Jia, K. Onuma, K. Mu...

  2. [10]

    Lorenzo-Trueba, J

    J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, T. Kinnunen, Z. Ling, The Voice Conversion Challenge 2018: Promoting development of parallel and nonparallel methods, in: Proc. Odyssey, 2018, pp. 195–202

  3. [11]

    Zhao, W.-C

    Y. Zhao, W.-C. Huang, X. Tian, J. Yamagishi, R. K. Das, T. Kinnunen, Z.-H. Ling, T. Toda, Voice Conversion Challenge 2020 — Intra-lingual semi-parallel and cross-lingual voice conversion —, in: Proc. Joint Workshop for the Blizzard Challenge and Voice Conversion Challenge 2020...

  4. [12]

    Geirhos, et al., Shortcut learning in deep neural networks, Nature Machine Intelligence 2 (11) (2020) 665–673

    R. Geirhos, et al., Shortcut learning in deep neural networks, Nature Machine Intelligence 2 (11) (2020) 665–673

  5. [13]

    Chettri, E

    B. Chettri, E. Benetos, B. L. T. Sturm, Dataset artefacts in anti-spoofing systems: A case study on the ASVspoof 2017 benchmark, IEEE/ACM Transactions on Audio, Speech, and Language Processing 28 (2020) 3018–3028

  6. [14]

    M¨ uller, F

    N. M¨ uller, F. Dieckmann, P. Czempin, R. Canals, K. B¨ ottinger, J. Williams, Speech is silver, silence is golden: What do ASVspoof-trained models really learn?, in: Proc. ASVspoof Challenge workshop, 2021, pp. 55–60

  7. [15]

    H.-j. Shim, R. G. Hautam¨ aki, M. Sahidullah, T. Kinnunen, How to construct perfect and worse-than-coin-flip spoofing countermeasures: A word of warning on shortcut learning, in: Proc. Interspeech, 2023, pp. 785–789

  8. [16]

    Zhang, Z

    Y. Zhang, Z. Li, J. Lu, H. Hua, W. Wang, P. Zhang, The impact of silence on speech anti-spoofing, IEEE/ACM Transactions on Audio, Speech, and Language Processing 31 (2023) 3374–3389

  9. [17]

    X. Liu, X. Wang, M. Sahidullah, J. Patino, H. Delgado, T. Kinnunen, M. Todisco, J. Yamagishi, N. Evans, A. Nautsch, K. A. Lee, ASVspoof 2021: Towards spoofed and deepfake speech detection in the wild, IEEE/ACM Transactions on Audio, Speech, and Language Processing 31 (2023) 2507–2522

  10. [18]

    Pratap, Q

    V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, R. Collobert, Mls: A large-scale multilingual dataset for speech research, in: Proc. Interspeech, 2020, pp. 2757–2761

  11. [19]

    Kearns, Librivox: Free public domain audiobooks, Reference Reviews 28 (1) (2014) 7–8

    J. Kearns, Librivox: Free public domain audiobooks, Reference Reviews 28 (1) (2014) 7–8

  12. [20]

    Szegedy, W

    C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, R. Fergus, Intriguing properties of neural networks, in: Proc. ICLR, 2013

  13. [21]

    I. J. Goodfellow, J. Shlens, C. Szegedy, Explaining and harnessing adversarial examples, in: Proc. ICLR, 2015

  14. [22]

    Z. Wu, P. Swietojanski, C. Veaux, S. Renals, S. King, A study of speaker adaptation for DNN-based speech synthesis, in: Proc. Interspeech, 2015, pp. 879–883

  15. [23]

    S. Arik, J. Chen, K. Peng, W. Ping, Y. Zhou, Neural voice cloning with a few samples, in: Proc.NeurIPS, Vol. 31, 2018

  16. [24]

    Y. Yan, X. Tan, B. Li, G. Zhang, T. Qin, S. Zhao, Y. Shen, W.-Q. Zhang, T.-Y. Liu, Adaspeech 3: Adaptive text to speech for spontaneous style, in: Proc. Interspeech, 2021, pp. 4668–4672

  17. [25]

    M. Chen, X. Tan, B. Li, Y. Liu, T. Qin, S. Zhao, T.-Y. Liu, Adaspeech: Adaptive text to speech for custom voice, in: Proc. ICLR, 2021

  18. [26]

    Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. L. Moreno, Y. Wu, et al., Transfer learning from speaker verification to multispeaker text-to-speech synthesis, in: Proc. NeurIPS, 2018, pp. 4480–4490

  19. [27]

    Cooper, C.-I

    E. Cooper, C.-I. Lai, Y. Yasuda, F. Fang, X. Wang, N. Chen, J. Yamagishi, Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings, in: Proc. ICASSP, 2020, pp. 6184–6188

  20. [28]

    Y. Wu, X. Tan, B. Li, L. He, S. Zhao, R. Song, T. Qin, T.-Y. Liu, Adaspeech 4: Adaptive text to speech in zero-shot scenarios, in: Proc. Interspeech, 2022, pp. 2568–2572

  21. [29]

    X. Tan, T. Qin, F. Soong, T.-Y. Liu, A survey on neural speech synthesis, arXiv preprint arXiv:2106.15561 (2021). 31

  22. [30]

    Papernot, P

    N. Papernot, P. McDaniel, I. Goodfellow, Transferability in machine learning: From phenomena to black-box attacks using adversarial samples, arXiv preprint arXiv:1605.07277 (2016)

  23. [31]

    Panayotov, G

    V. Panayotov, G. Chen, D. Povey, S. Khudanpur, Librispeech: An ASR corpus based on public domain audio books, in: Proc. ICASSP, 2015, pp. 5206–5210

  24. [32]

    Mohamed, H.-y

    A. Mohamed, H.-y. Lee, L. Borgholt, J. D. Havtorn, J. Edin, C. Igel, K. Kirchhoff, S.-W. Li, K. Livescu, L. Maaløe, T. N. Sainath, S. Watanabe, Self-supervised speech representation learning: A review, IEEE Journal of Selected Topics in Signal Processing 16 (6) (2022) 1179–1210

  25. [33]

    Ardila, M

    R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, G. Weber, Common Voice: A massively-multilingual speech corpus, in: Proc. LREC, 2020, pp. 4218–4222

  26. [34]

    Delgado, N

    H. Delgado, N. Evans, J.-w. Jung, T. Kinnunen, I. Kukanov, K.-A. Lee, X. Liu, H.-j. Shim, M. Sahidullah, H. Tak, M. Todisco, X. Wang, J. Yamagishi, ASVspoof 5 evaluation plan (phase 2),https://www.asvspoof.org/file/ASVspoof5_ __Evaluation_Plan_Phase2.pdf, v0.6, accessed 23-Jul...

  27. [35]

    Better be computer or I’m dumb

    K. Warren, T. Tucker, A. Crowder, D. Olszewski, A. Lu, C. Fedele, M. Pasternak, S. Layton, K. Butler, C. Gates, et al., “Better be computer or I’m dumb”: A large-scale evaluation of humans as audio deepfake detectors, in: Proc. ACM CCS, 2024, pp. 2696–2710

  28. [36]

    J.-w. Jung, X. Wang, N. Evans, S. Watanabe, H.-j. Shim, H. Tak, S. Arora, J. Yamagishi, J. S. Chung, To what extent can ASV systems naturally defend against spoofing attacks?, in: Proc. Interspeech, 2024, pp. 3240–3244

  29. [37]

    Taylor, Text-to-speech synthesis, Cambridge University Press, 2009

    P. Taylor, Text-to-speech synthesis, Cambridge University Press, 2009

  30. [38]

    S. Liu, H. Wu, H.-Y. Lee, H. Meng, Adversarial attacks on spoofing countermeasures of automatic speaker verification, in: Proc. ASRU, 2019, pp. 312–319

  31. [39]

    L. Wang, J. Li, Y. Luo, J. Zheng, L. Wang, H. Li, K. Xu, C. Fang, J. Shi, Z. Wu, Advsv: An over-the-air adversarial attack dataset for speaker verification, in: Proc. ICASSP, IEEE, 2024, pp. 4555–4559

  32. [40]

    X. Li, J. Zhong, X. Wu, J. Yu, X. Liu, H. Meng, Adversarial attacks on GMM i-vector based speaker verification systems, in: Proc. ICASSP, IEEE, 2020, pp. 6579–6583

  33. [41]

    Panariello, W

    M. Panariello, W. Ge, H. Tak, M. Todisco, N. Evans, Malafide: a novel adversarial convolutive noise attack against deepfake and spoofing detection systems, in: Proc. Interspeech, 2023, pp. 2868–2872

  34. [42]

    Todisco, M

    M. Todisco, M. Panariello, X. Wang, H. Delgado, K.-A. Lee, N. Evans, Malacopula: Adversarial automatic speaker verification attacks using a neural-based generalised hammerstein model, in: Proc. ASVspoof Workshop 2024, 2024, pp. 94–100

  35. [43]

    J. Kim, S. Kim, J. Kong, S. Yoon, Glow-TTS: A generative flow for text-to-speech via monotonic alignment search, in: Proc. NeurIPS, 2020, pp. 8067–8077

  36. [44]

    Popov, I

    V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, M. Kudinov, Grad-TTS: A diffusion probabilistic model for text-to-speech, in: Proc. ICML, 2021, pp. 8599–8608

  37. [45]

    La´ ncucki, Fastpitch: Parallel text-to-speech with pitch prediction, in: Proc

    A. La´ ncucki, Fastpitch: Parallel text-to-speech with pitch prediction, in: Proc. ICASSP, 2021, pp. 6588–6592

  38. [46]

    J. Kim, J. Kong, J. Son, Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech, in: Proc. ICML, 2021, pp. 5530–5540

  39. [47]

    F. Lux, J. Koch, S. Meyer, T. Bott, N. Schauffler, P. Denisov, A. Schweitzer, N. T. Vu, The IMS Toucan system for the Blizzard Challenge 2023, in: Proc. Blizzard Challenge Workshop, ISCA, 2023, pp. 40–45

  40. [48]

    J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, et al., Natural TTS synthesis by conditioning wavenet on Mel spectrogram predictions, in: Proc. ICASSP, 2018, pp. 4779–4783

  41. [49]

    M. Baas, H. Kamper, StarGAN-ZSVC: Towards zero-shot voice conversion in low-resource contexts, in: Artificial Intel- ligence Research. SACAIR 2021. Communications in Computer and Information Science, Vol. 1342, 2020, pp. 69–84

  42. [50]

    Casanova, J

    E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. G¨ olge, M. A. Ponti, YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone, in: Proc. ICML, 2022, pp. 2709–2720

  43. [51]

    E. A. AlBadawy, S. Lyu, Voice conversion using speech-to-speech neuro-style transfer, in: Proc. Interspeech, 2020, pp. 4726–4730

  44. [52]

    C. Gong, X. Wang, E. Cooper, D. Wells, L. Wang, J. Dang, K. Richmond, J. Yamagishi, ZMM-TTS: Zero-shot mul- tilingual and multispeaker speech synthesis conditioned on self-supervised discrete speech representations, IEEE/ACM Transactions on Audio, Speech, and Language Processi...

  45. [53]

    Schr¨ oder, M

    M. Schr¨ oder, M. Charfuelan, S. Pammi, I. Steiner, Open source voice creation toolkit for the MARY TTS platform, in: Proc. Interspeech, 2011, pp. 3253–3256

  46. [54]

    Popov, I

    V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, M. Kudinov, J. Wei, Diffusion-based voice conversion with fast maximum likelihood sampling scheme, in: Proc. ICLR, 2022

  47. [55]

    Casanova, K

    E. Casanova, K. Davis, E. G¨ olge, G. G¨ oknar, I. Gulea, L. Hart, A. Aljafari, J. Meyer, R. Morais, S. Olayemi, et al., XTTS: A massively multilingual zero-shot text-to-speech model, Proc. Interspeech (2024) 4978–4982

  48. [56]

    D. P. Kingma, P. Dhariwal, Glow: Generative flow with invertible 1x1 convolutions, in: Proc. NeurIPS, 2018, pp. 10236–10245

  49. [57]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, Attention is all you need, in: Proc. NeurIPS, 2017, pp. 5998–6008

  50. [58]

    J. Kong, J. Kim, J. Bae, HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis, in: Proc. NeurIPS, 2020, pp. 17022–17033

  51. [59]

    Desplanques, J

    B. Desplanques, J. Thienpondt, K. Demuynck, ECAPA-TDNN: Emphasized channel attention, propagation and aggre- gation in TDNN based speaker verification, in: Proc. Interspeech, 2020, pp. 3830–3834

  52. [60]

    K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: Proc. CVPR, 2016, pp. 770–778

  53. [61]

    G. Zhu, F. Jiang, Z. Duan, Y-vector: Multiscale waveform encoder for speaker embedding, in: Proc. Interspeech, 2021, 32 pp. 96–100

  54. [62]

    Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, B. Poole, Score-based generative modeling through stochastic differential equations, in: Proc. ICLR, 2021

  55. [63]

    Ronneberger, P

    O. Ronneberger, P. Fischer, T. Brox, U-net: Convolutional networks for biomedical image segmentation, in: Proc. MICCAI, Springer, 2015, pp. 234–241

  56. [64]

    Snyder, D

    D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, S. Khudanpur, X-vectors: Robust DNN embeddings for speaker recognition, in: Proc. ICASSP, 2018, pp. 5329–5333

  57. [65]

    Hayashi, R

    T. Hayashi, R. Yamamoto, K. Inoue, T. Yoshimura, S. Watanabe, T. Toda, K. Takeda, Y. Zhang, X. Tan, ESPnet- TTS: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit, in: Proc. ICASSP, 2020, pp. 7654–7658

  58. [66]

    F. Lux, J. Koch, A. Schweitzer, N. T. Vu, The IMS Toucan system for the Blizzard Challenge 2021, in: Proc. Blizzard Challenge Workshop, ISCA, 2021, pp. 14–19

  59. [67]

    Y. Wang, D. Stanton, Y. Zhang, R.-S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, R. A. Saurous, Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis, in: Proc. ICML, 2018, pp. 5180–5189

  60. [68]

    Gulati, J

    A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, et al., Conformer: Convolution-augmented transformer for speech recognition, in: Proc. Interspeech, 2020, pp. 5036–5040

  61. [69]

    Y. Ren, J. Liu, Z. Zhao, Portaspeech: Portable and high-quality generative text-to-speech, in: Proc. NeurIPS, 2021, pp. 13963–13974

  62. [70]

    F. Lux, J. Koch, N. T. Vu, Exact prosody cloning in zero-shot multispeaker text-to-speech, in: Proc. SLT, 2022, pp. 962–969

  63. [71]

    Graves, Supervised sequence labelling with recurrent neural networks, Ph.D

    A. Graves, Supervised sequence labelling with recurrent neural networks, Ph.D. thesis, TUM (2008)

  64. [72]

    L. Wan, Q. Wang, A. Papir, I. L. Moreno, Generalized end-to-end loss for speaker verification, in: Proc. ICASSP, 2018, pp. 4879–4883

  65. [73]

    J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, Y. Bengio, Attention-based models for speech recognition, in: Proc. NeurIPS, 2015, pp. 577–585

  66. [74]

    N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, W. Chan, WaveGrad: Estimating gradients for waveform generation, in: Proc. ICLR, 2021

  67. [75]

    Miyato, M

    T. Miyato, M. Koyama, cGANs with projection discriminator, in: Proc. ICLR, 2018

  68. [76]

    K. Cho, B. van Merri¨ enboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, Y. Bengio, Learning phrase repre- sentations using RNN encoder–decoder for statistical machine translation, in: Proc. EMNLP, 2014, pp. 1724–1734

  69. [77]

    Prenger, R

    R. Prenger, R. Valle, B. Catanzaro, WaveGlow: A flow-based generative network for speech synthesis, in: Proc. ICASSP, 2019, pp. 3617–3621

  70. [78]

    Kaneko, H

    T. Kaneko, H. Kameoka, K. Tanaka, N. Hojo, StarGAN-VC2: Rethinking conditional methods for StarGAN-based voice conversion, in: Proc. Interspeech, 2019, pp. 679–683

  71. [79]

    J.-w. Jung, Y. Kim, H.-S. H. Heo, B.-J. Lee, Y. Kwon, J. S. Chung, Pushing the limits of raw waveform speaker recognition, in: Proc. Interspeech, 2022, pp. 2228–2232

  72. [80]

    D. P. Kingma, M. Welling, Auto-encoding variational bayes, in: Proc. ICLR, 2014

  73. [81]

    J.-Y. Zhu, T. Park, P. Isola, A. A. Efros, Unpaired image-to-image translation using cycle-consistent adversarial networks, in: Proc. CVPR, 2017, pp. 2223–2232

  74. [82]

    A. v. d. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, K. Kavukcuoglu, Wavenet: A generative model for raw audio, arXiv preprint arXiv:1609.03499 (2016)

  75. [83]

    L. Sun, H. Wang, S. Kang, K. Li, H. Meng, Personalized, cross-lingual TTS using phonetic posteriorgrams, in: Proc. Interspeech, 2016, pp. 322–326

  76. [84]

    H. Wang, S. Zheng, Y. Chen, L. Cheng, Q. Chen, Cam++: A fast and efficient network for speaker verification using context-aware masking, in: Proc. Interspeech, 2023, pp. 5301–5305

  77. [85]

    Baevski, Y

    A. Baevski, Y. Zhou, A. Mohamed, M. Auli, Wav2vec 2.0: A framework for self-supervised learning of speech represen- tations, in: Proc. NuerIPS, Vol. 33, 2020, pp. 12449–12460

  78. [86]

    C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, E. Dupoux, VoxPopuli: A large- scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation, in: Proc. ACL, 2021, pp. 993–1003

  79. [87]

    Jadoul, B

    Y. Jadoul, B. Thompson, B. de Boer, Introducing Parselmouth: A Python interface to Praat, Journal of Phonetics 71 (2018) 1–15

  80. [88]

    L. T. Nguyen, T. Pham, D. Q. Nguyen, XPhoneBERT: A pre-trained multilingual model for phoneme representations for text-to-speech, in: Proc. Interspeech, 2023, pp. 5506–5510

  81. [89]

    Van Den Oord, O

    A. Van Den Oord, O. Vinyals, et al., Neural discrete representation learning, in: Proc. NeurIPS, 2017, pp. 6309–6318

  82. [90]

    H. Guo, F. Xie, X. Wu, F. K. Soong, H. Meng, MSMC-TTS: Multi-stage multi-codebook VQ-V AE based neural TTS, IEEE/ACM Transactions on Audio, Speech, and Language Processing 31 (2023) 1811–1824

  83. [91]

    J. S. Chung, A. Nagrani, A. Zisserman, Voxceleb2: Deep speaker recognition, in: Proc. Interspeech, 2018, pp. 1086–1090

  84. [92]

    Steiner, S

    I. Steiner, S. Le Maguer, Creating new language and voice components for the updated MaryTTS text-to-speech synthesis platform, in: Proc. LREC, 2018, pp. 3171–3175

  85. [93]

    Sagisaka, Speech synthesis by rule using an optimal selection of non-uniform synthesis units, in: Proc

    Y. Sagisaka, Speech synthesis by rule using an optimal selection of non-uniform synthesis units, in: Proc. ICASSP, 1988, pp. 679–682

  86. [94]

    A. J. Hunt, A. W. Black, Unit selection in a concatenative speech synthesis system using a large speech database, in: Proc. ICASSP, Vol. 1, IEEE, 1996, pp. 373–376

  87. [95]

    S.-g. Lee, W. Ping, B. Ginsburg, B. Catanzaro, S. Yoon, BigVGAN: A universal neural vocoder with large-scale training, 33 in: Proc. ICLR, 2022

  88. [96]

    Ziyin, T

    L. Ziyin, T. Hartwig, M. Ueda, Neural networks fail to learn periodic functions and how to fix it, in: Proc. NeurIPS, 2020, pp. 1583–1594

  89. [97]

    Kintzley, A

    K. Kintzley, A. Jansen, H. Hermansky, Event selection from phone posteriorgrams using matched filters, in: Proc. Interspeech, 2011, pp. 1905–1908

  90. [98]

    D´ efossez, G

    A. D´ efossez, G. Synnaeve, Y. Adi, Real time speech enhancement in the waveform domain, in: Proc. Interspeech, 2020, pp. 3291–3295

  91. [99]

    Eren, The Coqui TTS Team, Coqui TTS (Jan

    G. Eren, The Coqui TTS Team, Coqui TTS (Jan. 2021). doi:10.5281/zenodo.6334862

  92. [100]

    H. S. Heo, B.-J. Lee, J. Huh, J. S. Chung, Clova baseline system for the VoxCeleb speaker recognition challenge 2020, arXiv preprint arXiv:2009.14153 (2020)

  93. [101]

    Radford, J

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al., Language models are unsupervised multitask learners, OpenAI blog 1 (8) (2019) 9

  94. [102]

    D´ efossez, J

    A. D´ efossez, J. Copet, G. Synnaeve, Y. Adi, High fidelity neural audio compression, Transactions on Machine Learning Research (2023)

  95. [103]

    Novak, P

    A. Novak, P. Lotton, L. Simon, Synchronized swept-sine: Theory, application, and implementation, J. Audio Eng. Soc 63 (10) (2015) 786–798

  96. [104]

    Lapuschkin, S

    S. Lapuschkin, S. W¨ aldchen, A. Binder, G. Montavon, W. Samek, K.-R. M¨ uller, Unmasking Clever Hans predictors and assessing what machines really learn, Nature communications 10 (1) (2019) 1096

  97. [105]

    Radford, J

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, I. Sutskever, Robust speech recognition via large-scale weak supervision, in: Proc. ICML, 2023, pp. 28492–28518

  98. [106]

    Kinnunen, H

    T. Kinnunen, H. Li, An overview of text-independent speaker recognition: From features to supervectors, Speech com- munication 52 (1) (2010) 12–40

  99. [107]

    Van der Maaten, G

    L. Van der Maaten, G. Hinton, Visualizing data using t-SNE, JMLR 9 (11) (2008)

  100. [108]

    Hastie, R

    T. Hastie, R. Tibshirani, J. H. Friedman, J. H. Friedman, The elements of statistical learning: data mining, inference, and prediction, Vol. 2, Springer, 2009

  101. [109]

    M. L. Waskom, Seaborn: statistical data visualization, Journal of Open Source Software 6 (60) (2021) 3021

  102. [110]

    Ravanelli, T

    M. Ravanelli, T. Parcollet, A. Moumen, S. de Langen, C. Subakan, P. Plantinga, Y. Wang, P. Mousavi, L. Della Libera, A. Ploujnikov, et al., Open-source conversational AI with SpeechBrain 1.0, JMLR 25 (333) (2024) 1–11

  103. [111]

    Nagrani, J

    A. Nagrani, J. S. Chung, A. Zisserman, VoxCeleb: A Large-Scale Speaker Identification Dataset, in: Proc. Interspeech, 2017, pp. 2616–2620

  104. [112]

    Snyder, G

    D. Snyder, G. Chen, D. Povey, MUSAN: A Music, Speech, and Noise Corpus, arXiv:1510.08484v1 (2015)

  105. [113]

    T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, S. Khudanpur, A study on data augmentation of reverberant speech for robust speech recognition, in: Proc. ICASSP, 2017, pp. 5220–5224

  106. [114]

    Jung, H.-S

    J.-w. Jung, H.-S. Heo, H. Tak, H.-j. Shim, J. S. Chung, B.-J. Lee, H.-J. Yu, N. Evans, AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks, in: Proc. ICASSP, 2022, pp. 6367–6371

  107. [115]

    H. Tak, J. Patino, M. Todisco, A. Nautsch, N. Evans, A. Larcher, End-to-end anti-spoofing with RawNet2, in: Proc. ICASSP, 2021, pp. 6369–6373

  108. [116]

    D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, Q. V. Le, SpecAugment: A simple data augmentation method for automatic speech recognition, in: Proc. Interspeech, 2019, pp. 2613–2617

  109. [117]

    Jung, S.-b

    J.-w. Jung, S.-b. Kim, H.-j. Shim, J.-h. Kim, H.-J. Yu, Improved RawNet with feature map scaling for text-independent speaker verification using raw waveforms, in: Proc. Interspeech, 2020, pp. 1496–1500

  110. [118]

    Ravanelli, Y

    M. Ravanelli, Y. Bengio, Speaker recognition from raw waveform with sincnet, in: Proc. SLT, 2018, pp. 1021–1028

  111. [119]

    J.-w. Jung, H. Tak, H.-j. Shim, H.-S. Heo, B.-J. Lee, S.-W. Chung, H.-J. Yu, N. Evans, T. Kinnunen, SASV 2022: The first spoofing-aware speaker verification challenge, in: Proc. Interspeech, 2022, pp. 2893–2897

  112. [120]

    S. H. Mun, H.-j. Shim, H. Tak, X. Wang, X. Liu, M. Sahidullah, M. Jeong, M. H. Han, M. Todisco, K. A. Lee, et al., Towards single integrated spoofing-aware speaker verification embeddings, in: Proc. Interspeech, 2023, pp. 3989–3993

  113. [121]

    H.-j. Shim, H. Tak, X. Liu, H.-S. Heo, J.-w. Jung, J. S. Chung, S.-W. Chung, H.-J. Yu, B.-J. Lee, M. Todisco, et al., Baseline systems for the first spoofing-aware speaker verification challenge: Score and embedding fusion, in: Proc. Odyssey, 2022, pp. 330–337

  114. [122]

    X. Wang, T. Kinnunen, L. Kong Aik, P.-G. Noe, J. Yamagishi, Revisiting and improving scoring fusion for spoofing-aware speaker verification using compositional data analysis, in: Proc. Interspeech, 2024, pp. 1110–1114

  115. [123]

    Zhang, Z

    Y. Zhang, Z. Lv, H. Wu, S. Zhang, P. Hu, Z. Wu, H.-y. Lee, H. Meng, MF A-conformer: Multi-scale feature aggregation conformer for automatic speaker verification, in: Proc. Interspeech, 2022, pp. 306–310

  116. [124]

    X. Wang, J. Yamagishi, Spoofed training data for speech spoofing countermeasure can be efficiently created using neural vocoders, in: Proc. ICASSP, 2023, pp. 1–5

  117. [125]

    Cooper, W.-C

    E. Cooper, W.-C. Huang, T. Toda, J. Yamagishi, Generalization ability of MOS prediction networks, in: Proc. ICASSP, 2022, pp. 8442–8446

  118. [126]

    Shim, J.-w

    H.-j. Shim, J.-w. Jung, T. Kinnunen, et al., a-DCF: an architecture agnostic metric with application to spoofing-robust speaker verification, in: Proc. Speaker Odyssey, 2024, pp. 158–164

  119. [127]

    D. A. Van Leeuwen, N. Br¨ ummer, An introduction to application-independent evaluation of speaker recognition systems, in: Speaker Classification I, Springer, 2007, pp. 330–353

  120. [128]

    A. F. Martin, G. R. Doddington, T. Kamm, M. Ordowski, M. A. Przybocki, The DET curve in assessment of detection task performance., in: Eurospeech, Vol. 4, 1997, pp. 1895–1898. 34

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.