Pith. sign in

REVIEW 4 major objections 4 minor 80 references

Improving Generalization for AI-Synthesized Voice Detection

T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A disentanglement framework that trains on shared vocoder artifacts cuts unseen-vocoder detection error by up to 7.59 percentage points.

desk verdict A promising empirical recipe whose central mechanism is undermined by a sign error in the mutual information loss, yet the engineering is worth referee time. read the letter →

arxiv 2412.19279 v2 pith:ZFG6575K submitted 2024-12-26 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords AI-synthesizedvoicedetectionaudiodeepfakedomaingeneralizationdisentangledrepresentationlearningvocoderartifactsmutualinformationestimationsharpness-awareminimizationcross-vocoderevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

AI-synthesized voice detectors work well on the voice generators they were trained on but fail on new ones, and this paper aims to fix that. It proposes a training framework that splits each voice into three parts: the spoken content, artifacts unique to a specific vocoder, and artifacts shared across vocoders. The shared artifacts are used for detection, and training is pushed toward a flat loss landscape so the model does not settle into sharp, overfit minima. On the LibriSeVoc benchmark the method lowers equal error rate from 18.67% to 13.55% on seen vocoders and from 27.86% to 20.27% on unseen vocoders compared with the prior best method, which are the up-to-5.12 and 7.59 percentage point improvements cited in the abstract.

What carries the argument

The load-bearing mechanism is the disentanglement-and-flattening pipeline. An encoder backbone is split into content and artifact branches: the artifact branch yields domain-specific features, which identify the vocoder, and domain-agnostic features, which flag synthetic voice regardless of generator. Cross-reconstruction through an AdaIN decoder forces the two artifact types to be separable, a contrastive loss organizes the feature space, and a mutual information term built on the Donsker–Varadhan lower bound aligns the domain-agnostic features with the content distribution. Sharpness-Aware Minimization perturbs weights toward higher loss before each gradient step, flattening the landscape; at inference only the domain-agnostic classification head is used.

What would settle it

Measure, on a held-out test set, the estimated mutual information between the learned domain-agnostic features and the vocoder identity, and the estimated mutual information between those features and the spoken content. The paper's mechanism predicts that after training the artifact features separate real from synthetic while remaining independent of which vocoder made them; if the features stay strongly tied to vocoder identity, or if editing only the mutual information loss leaves the cross-domain equal error rate essentially unchanged on a new-vocoder split, the central claim is falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that domain-agnostic artifact features, the traces left by speech synthesis that are common to many vocoders, can be extracted by disentanglement and then used directly for classification, giving a detector that generalizes across generators. It establishes that a content encoder and an artifact encoder with two classification heads (one for vocoder identity, one for real-versus-synthetic), a cross-reconstruction decoder, contrastive learning, and a mutual information term together produce these features, and that Sharpness-Aware Minimization then flattens the loss landscape and keeps the model out of sharp minima. The reported experiments on LibriSeVoc, ASVspoof2019, WaveFake, and FakeAVCeleb support the claim, with the largest gains on unseen vocoders.

Load-bearing premise

The central bet is that making the shared fake-voice features statistically more dependent on the spoken content turns them into universal fake-voice features; if that step actually does the opposite, the framework's claimed mechanism for cross-domain generalization is not supported.

Editorial extensions

If this is right

  • Detectors trained with this pipeline keep working when new vocoder families appear, because the detector keys on shared artifacts rather than on the six generators it saw.
  • The component ablation shows each piece, reconstruction, classification heads, contrastive loss, mutual information, and sharpness-aware minimization, contributes; removing the mutual information term alone costs about 6.85 equal error rate points on unseen FakeAVCeleb audio.
  • Training-data diversity matters: using more vocoders in the training set monotonically lowers average equal error rate on both seen and unseen test vocoders.
  • The approach also improves cross-dataset detection: trained on LibriSeVoc and tested on WaveFake, mean unseen-vocoder equal error rate drops from 34.06% to 23.38%.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The mutual information objective as written maximizes dependence between content and artifact features, which is the opposite of disentangling them; the reported cross-domain gains could therefore come from sharpness-aware minimization, contrastive learning, or multi-task classification rather than from the claimed alignment. An isolated test would replace the mutual information term with a decorr
  • If flatness is the real driver, then applying sharpness-aware minimization alone to simpler baselines should recover a meaningful fraction of the 7.59-point gain without any disentanglement; that is a direct test the paper does not report.
  • The same recipe, shared artifact extraction plus flat-minimum optimization, is naturally transferable to other deepfake media where content and manipulation artifacts also mix; one could test it on cross-dataset face-swap or audio-visual deepfake benchmarks.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes a disentanglement framework for AI-synthesized voice detection aimed at improving generalization to unseen vocoders. The method trains a RawNet2-based encoder that separates content features, domain-specific artifact features, and domain-agnostic artifact features, guided by a multi-task classification loss, a contrastive loss, a reconstruction loss, a mutual information loss, and sharpness-aware minimization (SAM). The authors evaluate on LibriSeVoc, ASVspoof2019, WaveFake, and FakeAVCeleb, reporting improved equal error rates (EER) in both intra-domain and cross-domain settings compared to several baselines, and they release code. The central claim is that the domain-agnostic artifact features, made 'universally applicable' via the mutual information loss, and the flattened loss landscape from SAM jointly improve cross-domain generalization.

Significance. If the reported results hold, the framework would be a meaningful step toward cross-domain audio deepfake detection, which is a recognized weakness of current detectors. The paper's strengths include evaluation on multiple external benchmarks, a component-wise ablation study, and public code. However, the stated mechanism for the mutual information loss is technically questionable, and the empirical claims are based on single-run point estimates without uncertainty quantification. These issues are central because the proposed mechanism for domain-agnostic features is the paper's main novelty, and the numerical gains are the main evidence for it. With corrections to the MI formulation and stronger statistical validation, the work could be a solid contribution; in its current form, the supporting evidence is not fully convincing.

major comments (4)
  1. [Mutual Information Loss and Eq. (1)] The sign and interpretation of the mutual information loss are internally inconsistent. The paper states that maximizing MI(c; a^g) 'aligns domain-agnostic features with the content feature distribution,' but Eq. (1) subtracts λ4 L_MI from the total loss, so gradient descent on L maximizes the Donsker-Varadhan lower bound L_MI. Maximizing mutual information increases statistical dependence between c and a^g, which is the opposite of disentangling content from artifact features. The DV lower bound does not match marginal distributions; it amplifies per-sample dependence, as seen in Algorithm 2 where E_joint pulls c_i and a^g_i together for the same input. This undermines the claimed mechanism by which a^g becomes 'universally applicable' and vocoder-agnostic. The authors should either minimize MI(c; a^g) to enforce independence, or provide a different, correct justification for why maximizing MI yields domain-agnostic features, along with direct evidence about what a^g encodes. The ablation gain of VD over VC in Table 4 may then be attributable to an unintended regularizer rather than the described alignment, so the central claim is not yet supported.
  2. [Algorithm 2 (MI pseudocode)] The PyTorch-style pseudocode in Algorithm 2 is dimensionally unclear and not reproducible as written. The comment says c is 'B x n0 x dim' and a is 'B x n1 x dim', but u = torch.mm(a, c.t()) operates on 2D tensors; the subsequent reshape to 'B x B x n0 x n1' and the use of mean(2) are not consistent with typical 3D feature tensors. It is also unclear what the 'joint' and 'margin' scores represent after masking, and whether the average is over the batch, the feature dimensions, or both. Because the mutual information loss is a load-bearing component of the framework, the implementation must be specified unambiguously so that the reported results can be reproduced and the behavior of the loss term can be independently checked.
  3. [Experimental Results, Tables 1-4] All reported EER numbers appear to be single-run point estimates with no error bars, confidence intervals, or significance tests for the improvements over baselines. The abstract highlights improvements of 5.12% and 7.59% in EER, but without run-to-run variance it is impossible to assess whether these differences are statistically meaningful, particularly for a model with several interacting loss terms and hyperparameters. The authors should report mean and standard deviation over multiple random seeds, or at least provide significance tests for the main comparisons against Sun et al. and RawNet2. This is load-bearing because the central claim of outperforming state-of-the-art methods rests entirely on these point estimates.
  4. [Ablation Study, Table 4 and Figure 4] The paper claims that the mutual information module 'greatly improves performance in cross-domain evaluation,' but the evidence is mixed. In Table 4, adding MI (VD vs. VC) improves unseen ASP EER from 26.83 to 23.23 and unseen FakeAVCeleb from 27.64 to 20.79, yet worsens seen WF EER from 24.72 to 25.62. Given the MI sign issue described above, the authors should show what the learned a^g actually encodes, for example by measuring vocoder classification accuracy from a^g or by quantifying how much content information remains in a^g. The qualitative UMAP in Figure 4 is suggestive but not quantitative. Without direct evidence that a^g is both vocoder-invariant and content-independent, the claimed disentanglement effect is not established.
minor comments (4)
  1. [Throughout] There are many typographical and spacing errors, such as 'conetent' in Algorithm 2, 'V oice' in several headings, and inconsistent use of 'Vo ice' and 'voice.' These should be corrected in a revised version.
  2. [Eq. (1) and Section Mutual Information Loss] The notation for the mutual information loss is confusing: Eq. (1) subtracts L_MI, but the text refers to it as a 'loss' and claims it 'aligns' distributions. It would be clearer to call it a regularization term with an explicit sign convention, and to define whether the reported hyperparameter λ4 controls the magnitude of maximization or minimization.
  3. [Implementation Details] The hyperparameters for the method are given (λ1=0.1, λ2=0.3, λ3=0.05, λ4=0.03, b=3, γ=0.07), but the corresponding hyperparameter tuning procedure is not described. It is also unclear how the baselines were tuned for their own hyperparameters; a brief explanation would help fairness of comparison.
  4. [Appendix, Algorithm 1] Algorithm 1's update step computes ϵ* from ∇θL and then updates θ using the gradient at θ+ϵ*, but the line 'Update θ: θl+1 ← θl − β∇θL|θl+ϵ*' overloads ∇θL; this should be written more explicitly to avoid ambiguity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central generalization claims are evaluated against external baselines and held-out vocoder datasets; the only self-citation supplies a contrastive-learning technique, not the target result.

full rationale

The claimed derivation chain is self-contained and externally checked. The disentanglement objective (Eq. (1)) combines a classification loss, a contrastive loss, a reconstruction loss, and a mutual-information term; all are defined on the training data and labels, and their contributions are isolated in the ablation study (Table 4). The mutual-information term uses the Donsker-Varadhan lower bound from Belghazi et al. (2018), an external method, and its effect is measured, not assumed. The optimization-level component is SAM (Foret et al. 2020), an external technique adapted to this task. The paper's headline improvements (5.12% intra-domain and 7.59% cross-domain EER reductions) are comparisons against external baselines (LCNN, RawNet2, WavLM, XLS-R, Sun et al.) on external benchmarks (LibriSeVoc, ASVspoof2019, WaveFake, FakeAVCeleb), including held-out vocoders. The only self-citation is Lin et al. 2024, cited as 'Inspired by' for the contrastive loss; that prior work contributes a technique, not the paper's predicted outcome, and no uniqueness theorem or fitted parameter is imported from it. A reviewer-level concern that maximizing MI(c; ag) may not produce the stated 'alignment' with content features is a mechanistic correctness issue, not a circularity: the loss is not defined in terms of the test metric or the target result. Accordingly, no load-bearing step reduces to its own input.

Assumptions & free parameters 6 free parameters · 5 assumptions · 3 invented entities

The method relies on tuned hyperparameters (six listed), on the Donsker-Varadhan MI estimator, on the domain-label assumption, and on a poorly specified alignment-to-content assumption. It introduces two novel latent feature spaces and a vague 'content feature distribution' reference without external falsifiable handles. These are the main things the reader pays for beyond standard ML machinery.

free parameters (6)
  • lambda_1 = 0.1
    Balances the domain-specific and domain-agnostic classification losses in L_cls; chosen by hand, no sensitivity analysis reported.
  • lambda_2 = 0.3
    Balances the contrastive loss; chosen by hand, no sensitivity analysis reported.
  • lambda_3 = 0.05
    Balances the reconstruction loss; chosen by hand, no sensitivity analysis reported.
  • lambda_4 = 0.03
    Balances the mutual information loss; sensitivity analysis in Figure 4(a) shows optimum at 0.03.
  • contrastive margin b = 3
    Hinge margin in L_con; chosen by hand, no sensitivity analysis reported.
  • SAM perturbation gamma = 0.07
    Controls perturbation magnitude in Eq. (2); sensitivity analysis in Figure 4(b) shows optimum at 0.07.
assumptions (5)
  • standard math Donsker-Varadhan representation provides a valid lower bound on mutual information that can be estimated and maximized with a neural network.
    Relied on for L_MI in Eq. (4)-(5); follows Belghazi et al. (2018), which is an accepted background result.
  • domain assumption Vocoder identity is a sufficient domain label; all generalization-relevant variation is captured by the vocoder type in the training set.
    Used in the domain classification head and in splitting datasets into seen/unseen vocoders (Table 7). If speaker, recording condition, or other factors act as domains, the disentanglement may miss them.
  • ad hoc to paper Maximizing MI(c; ag) aligns domain-agnostic features with the content feature distribution and makes them universally applicable.
    This is the stated mechanism in 'Mutual Information Loss'; the equations maximize mutual information, which measures dependence, not distribution alignment. The assumption is unsupported and internally inconsistent with the disentanglement goal.
  • standard math SAM's first-order approximation and dual norm solution for the adversarial perturbation are valid for flattening the loss landscape.
    Eq. (2) follows Foret et al. (2020); accepted background result.
  • domain assumption Common domain-agnostic artifacts exist across different vocoders and are learnable from a finite set of seen vocoders.
    Central hypothesis behind the disentanglement framework; if false, cross-domain generalization would not improve.
invented entities (3)
  • domain-agnostic artifact feature space (a^g)
    purpose: Represent shared forgery artifacts across vocoders, used by the detection head at inference.
    No out-of-sample falsifiable handle; evidence is internal UMAP visualizations and ablation studies.
  • domain-specific artifact feature space (a^s)
    purpose: Capture vocoder-specific artifacts for the auxiliary domain classification task.
    Internal latent variable with no external handle; supported only by in-paper ablations.
  • content feature distribution as a benchmark
    purpose: Reference distribution to align domain-agnostic features via the MI loss.
    The paper never defines this distribution explicitly, and the MI objective as written does not perform distribution alignment, so the entity is vague and lacks external evidence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Improving Generalization for AI-Synthesized Voice Detection." pith.science (2026). https://pith.science/paper/ZFG6575K

@misc{pith2026241219279,
  author       = {Pith},
  title        = {Pith review of: Improving Generalization for AI-Synthesized Voice Detection},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZFG6575K}},
  note         = {Machine review of arXiv:2412.19279}
}
read the original abstract

AI-synthesized voice technology has the potential to create realistic human voices for beneficial applications, but it can also be misused for malicious purposes. While existing AI-synthesized voice detection models excel in intra-domain evaluation, they face challenges in generalizing across different domains, potentially becoming obsolete as new voice generators emerge. Current solutions use diverse data and advanced machine learning techniques (e.g., domain-invariant representation, self-supervised learning), but are limited by predefined vocoders and sensitivity to factors like background noise and speaker identity. In this work, we introduce an innovative disentanglement framework aimed at extracting domain-agnostic artifact features related to vocoders. Utilizing these features, we enhance model learning in a flat loss landscape, enabling escape from suboptimal solutions and improving generalization. Extensive experiments on benchmarks show our approach outperforms state-of-the-art methods, achieving up to 5.12% improvement in the equal error rate metric in intra-domain and 7.59% in cross-domain evaluations.

Figures

Figures reproduced from arXiv: 2412.19279 by the authors.

Figure 1
Figure 1. Comparison of audio deepfake detection methods. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Experimental results for Motivation. (Left) The differences of mel-spectrogram between human voices and AI￾synthesized ones (e.g., based on WaveGrad (Chen et al. 2020a) and WaveRNN (Kalchbrenner et al. 2018) vocoders). The red circles highlight AI vocoder artifacts. More details about these differences can be found in Appendix. (Middle) The UMAP (McInnes, Healy, and Melville 2018) visualization of features from rela… view at source ↗
Figure 3
Figure 3. The overall architecture of our proposed approach. (a) In the encoder module, two RawNet2 (Tak et al. 2021)- [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: (Left-Top) The virtualizations of model with MI (Vc) and model without MI (Vd). (Left-Bottom) The loss landscape visualization of our method with and without flattening the loss landscape. (Middle) (a) The effect of λ4 for balancing mutual information term. (b) The eff…
Figure 5
Figure 5. Figure 5: The artifacts introduced by the AI vocoders to a [PITH_FULL_IMAGE:figures/full_fig_p014_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

80 extracted references · 67 canonical work pages

  1. [1]

    Aloufi, R.; Haddadi, H.; and Boyle, D. 2020. Privacy-preserving voice analysis via disentangled representations. In Proceedings of the 2020 ACM SIGSAC Conference on Cloud Computing Security Workshop, 1--14

  2. [2]

    ASVspoof. 2023. Automatic Speaker Verification and Spoofing Countermeasures Challenge. In https://www.asvspoof.org/

  3. [3]

    Babu, A.; Wang, C.; Tjandra, A.; Lakhotia, K.; Xu, Q.; Goyal, N.; Singh, K.; von Platen, P.; Saraf, Y.; Pino, J.; et al. 2021. XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale. arXiv e-prints, arXiv--2111

  4. [4]

    Barrington, S.; Barua, R.; Koorma, G.; and Farid, H. 2023. Single and Multi-Speaker Cloned Voice Detection: From Perceptual to Learned Features. IEEE International Workshop on Information Forensics and Security

  5. [5]

    I.; Baratin, A.; Rajeshwar, S.; Ozair, S.; Bengio, Y.; Courville, A.; and Hjelm, D

    Belghazi, M. I.; Baratin, A.; Rajeshwar, S.; Ozair, S.; Bengio, Y.; Courville, A.; and Hjelm, D. 2018. Mutual information neural estimation. In International conference on machine learning, 531--540. PMLR

  6. [6]

    Champion, P.; Jouvet, D.; and Larcher, A. 2022. Are disentangled representations all you need to build speaker anonymization systems?

  7. [7]

    Chen, M.; Zhou, Y.; Huang, H.; and Hain, T. 2022 a . Efficient Non-Autoregressive GAN Voice Conversion using VQWav2vec Features and Dynamic Convolution. arXiv:2203.17172

  8. [8]

    J.; Norouzi, M.; and Chan, W

    Chen, N.; Zhang, Y.; Zen, H.; Weiss, R. J.; Norouzi, M.; and Chan, W. 2020 a . WaveGrad: Estimating Gradients for Waveform Generation. In International Conference on Learning Representations

Show all 80 references
  1. [9]

    Chen, S.; Wang, C.; Chen, Z.; Wu, Y.; Liu, S.; Chen, Z.; Li, J.; Kanda, N.; Yoshioka, T.; Xiao, X.; et al. 2022 b . Wavlm: Large-scale self-supervised pre-training for full stack speech processing. IEEE Journal of Selected Topics in Signal Processing, 16(6): 1505--1518

  2. [10]

    Chen, T.; Kumar, A.; Nagarsheth, P.; Sivaraman, G.; and Khoury, E. 2020 b . Generalization of Audio Deepfake Detection . In Proc. The Speaker and Language Recognition Workshop (Odyssey 2020), 132--137

  3. [11]

    Chettri, B.; Stoller, D.; Morfi, V.; Ram \'i rez, M. A. M.; Benetos, E.; and Sturm, B. L. 2019. Ensemble Models for Spoofing Detection in Automatic Speaker Verification. In Interspeech

  4. [12]

    Ding, S.; Zhang, Y.; and Duan, Z. 2023. SAMO: Speaker Attractor Multi-Center One-Class Learning For Voice Anti-Spoofing. In ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1--5

  5. [13]

    Donahue, J.; Dieleman, S.; Binkowski, M.; Elsen, E.; and Simonyan, K. 2021. End-to-End Adversarial Text-to-Speech. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net

  6. [14]

    D.; and Varadhan, S

    Donsker, M. D.; and Varadhan, S. S. 1983. Asymptotic evaluation of certain Markov process expectations for large time. IV. Communications on pure and applied mathematics, 36(2): 183--212

  7. [15]

    Forbes. 2019. A Voice Deepfake Was Used To Scam A CEO Out Of \ 243,000. In https://www.forbes.com/sites/jessedamiani/2019/09/03/a-voice-deepfake-was-used-to-scam-a-ceo-out-of-243000/?sh=432e93d12241

  8. [16]

    Foret, P.; Kleiner, A.; Mobahi, H.; and Neyshabur, B. 2020. Sharpness-aware Minimization for Efficiently Improving Generalization. In International Conference on Learning Representations

  9. [17]

    Frank, J.; and Sch \"o nherr, L. 2021. WaveFake: A Data Set to Facilitate Audio Deepfake Detection. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)

  10. [18]

    He, K.; Zhang, X.; Ren, S.; and Sun, J. 2015. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision, 1026--1034

  11. [19]

    He, R.; Wu, X.; Sun, Z.; and Tan, T. 2017. Learning invariant deep representation for nir-vis face recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 31

  12. [20]

    He, R.; Wu, X.; Sun, Z.; and Tan, T. 2018. Wasserstein CNN: Learning invariant features for NIR-VIS face recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(7): 1761--1773

  13. [21]

    D.; Fedorov, A.; Lavoie-Marchildon, S.; Grewal, K.; Bachman, P.; Trischler, A.; and Bengio, Y

    Hjelm, R. D.; Fedorov, A.; Lavoie-Marchildon, S.; Grewal, K.; Bachman, P.; Trischler, A.; and Bengio, Y. 2018. Learning deep representations by mutual information estimation and maximization. arXiv preprint arXiv:1808.06670

  14. [22]

    H.; Lakhotia, K.; Salakhutdinov, R.; and Mohamed, A

    Hsu, W.-N.; Bolte, B.; Tsai, Y.-H. H.; Lakhotia, K.; Salakhutdinov, R.; and Mohamed, A. 2021. HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units. IEEE/ACM Trans. Audio, Speech and Lang. Proc., 29: 3451–3460

  15. [23]

    Huang, X.; and Belongie, S. 2017. Arbitrary style transfer in real-time with adaptive instance normalization. In Proceedings of the IEEE international conference on computer vision, 1501--1510

  16. [24]

    Jia, Y.; Zhang, Y.; Weiss, R.; Wang, Q.; Shen, J.; Ren, F.; Nguyen, P.; Pang, R.; Lopez Moreno, I.; Wu, Y.; et al. 2018. Transfer learning from speaker verification to multispeaker text-to-speech synthesis. Advances in neural information processing systems, 31

  17. [25]

    Joyce, J. M. 2011. Kullback-leibler divergence. In International Encyclopedia of Statistical Science, 720--722

  18. [26]

    Kalchbrenner, N.; Elsen, E.; Simonyan, K.; Noury, S.; Casagrande, N.; Lockhart, E.; Stimberg, F.; Oord, A.; Dieleman, S.; and Kavukcuoglu, K. 2018. Efficient neural audio synthesis. In International Conference on Machine Learning, 2410--2419. PMLR

  19. [27]

    Khalid, H.; Tariq, S.; Kim, M.; and Woo, S. S. 2021. FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)

  20. [28]

    Kim, J.; Kong, J.; and Son, J. 2021. Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech. In Meila, M.; and Zhang, T., eds., Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learn...

  21. [29]

    Kim, S.; Lee, S.-G.; Song, J.; Kim, J.; and Yoon, S. 2019. F lo W ave N et : A Generative Flow for Raw Audio. In Chaudhuri, K.; and Salakhutdinov, R., eds., Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Resea...

  22. [30]

    P.; and Ba, J

    Kingma, D. P.; and Ba, J. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980

  23. [31]

    B.; and Atwal, G

    Kinney, J. B.; and Atwal, G. S. 2014. Equitability, mutual information, and the maximal information coefficient. Proceedings of the National Academy of Sciences, 111(9): 3354--3359

  24. [32]

    Kong, J.; Kim, J.; and Bae, J. 2020. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. Advances in Neural Information Processing Systems, 33: 17022--17033

  25. [33]

    Kong, Z.; Ping, W.; Huang, J.; Zhao, K.; and Catanzaro, B. 2020. DiffWave: A Versatile Diffusion Model for Audio Synthesis. In International Conference on Learning Representations

  26. [34]

    Z.; Sotelo, J.; De Brebisson, A.; Bengio, Y.; and Courville, A

    Kumar, K.; Kumar, R.; De Boissiere, T.; Gestin, L.; Teoh, W. Z.; Sotelo, J.; De Brebisson, A.; Bengio, Y.; and Courville, A. C. 2019. Melgan: Generative adversarial networks for conditional waveform synthesis. Advances in neural information processing systems, 32

  27. [35]

    Lavrentyeva, G.; Tseren, A.; Volkova, M.; Gorlanov, A.; Kozlov, A.; and Novoselov, S. 2019. STC antispoofing systems for the AsVspoof2019 challenge. In Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, 1033--1037

  28. [36]

    Lee, S.-H.; Kim, J.-H.; Lee, K.-E.; and Lee, S.-W. 2022. FRE-GAN 2: Fast and Efficient Frequency-Consistent Audio Synthesis. In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6192--6196

  29. [37]

    A.; Zare, A

    Li, Y. A.; Zare, A. A.; and Mesgarani, N. 2021. StarGANv2-VC: A Diverse, Unsupervised, Non-parallel Framework for Natural-Sounding Voice Conversion. In Interspeech

  30. [38]

    Lian, J.; Zhang, C.; and Yu, D. 2022. Robust Disentangled Variational Speech Representation Learning for Zero-Shot Voice Conversion. In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6572--6576

  31. [39]

    Lin, L.; He, X.; Ju, Y.; Wang, X.; Ding, F.; and Hu, S. 2024. Preserving fairness generalization in deepfake detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 16815--16825

  32. [40]

    Liu, H.; Chen, Z.; Yuan, Y.; Mei, X.; Liu, X.; Mandic, D.; Wang, W.; and Plumbley, M. D. 2023 a . A udio LDM : Text-to-Audio Generation with Latent Diffusion Models. In Krause, A.; Brunskill, E.; Cho, K.; Engelhardt, B.; Sabato, S.; and Scarlett, J., eds., Proceedings of the 4...

  33. [41]

    A.; Zhang, H.; and Dang, J

    Liu, X.; Liu, M.; Wang, L.; Lee, K. A.; Zhang, H.; and Dang, J. 2023 b . Leveraging Positional-Related Local-Global Dependency for Synthetic Speech Detection. In ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1--5

  34. [42]

    Long, Z.; Zheng, Y.; Yu, M.; and Xin, J. 2022. Enhancing Zero-Shot Many to Many Voice Conversion via Self-Attention VAE with Structurally Regularized Layers. In 2022 5th International Conference on Artificial Intelligence for Industries (AI4I), 59--63

  35. [43]

    Lorenzo-Trueba, J.; Drugman, T.; Latorre, J.; Merritt, T.; Putrycz, B.; Barra-Chicote, R.; Moinet, A.; and Aggarwal, V. 2019. Towards Achieving Robust Universal Neural Vocoding . In Proc. Interspeech 2019, 181--185

  36. [44]

    Luong, M.; and Tran, V. A. 2021. Many-to-many voice conversion based feature disentanglement using variational autoencoder. In Interspeech 2021, 22nd Annual Conference of the International Speech Communication Association

  37. [45]

    McInnes , L.; Healy , J.; and Melville , J. 2018. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction . ArXiv e-prints

  38. [46]

    Miao, C.; Shuang, L.; Liu, Z.; Minchuan, C.; Ma, J.; Wang, S.; and Xiao, J. 2021. EfficientTTS: An Efficient and High-Quality Text-to-Speech Architecture. In Meila, M.; and Zhang, T., eds., Proceedings of the 38th International Conference on Machine Learning, volume 139 of Pro...

  39. [47]

    Müller, N.; Czempin, P.; Diekmann, F.; Froghyar, A.; and Böttinger, K. 2022. Does Audio Deepfake Detection Generalize? In Proc. Interspeech 2022, 2783--2787

  40. [48]

    Oord, A. v. d.; Dieleman, S.; Zen, H.; Simonyan, K.; Vinyals, O.; Graves, A.; Kalchbrenner, N.; Senior, A.; and Kavukcuoglu, K. 2016. Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499

  41. [49]

    Oord, A. v. d.; Li, Y.; and Vinyals, O. 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748

  42. [50]

    Paul, D.; Pantazis, Y.; and Stylianou, Y. 2020. Speaker Conditional WaveRNN: Towards Universal Neural Vocoder for Unseen Speaker and Recording Conditions . In Proc. Interspeech 2020, 235--239

  43. [51]

    Peng, K.; Ping, W.; Song, Z.; and Zhao, K. 2020. Non-Autoregressive Neural Text-to-Speech. In III, H. D.; and Singh, A., eds., Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, 7586--7598. PMLR

  44. [52]

    Ping, W.; Peng, K.; and Chen, J. 2019. ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net

  45. [53]

    Prenger, R.; Valle, R.; and Catanzaro, B. 2019. Waveglow: A Flow-based Generative Network for Speech Synthesis. In ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 3617--3621

  46. [54]

    Ren, Y.; Hu, C.; Tan, X.; Qin, T.; Zhao, S.; Zhao, Z.; and Liu, T. 2021. FastSpeech 2: Fast and High-Quality End-to-End Text to Speech. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net

  47. [55]

    Salvi, D.; Bestagini, P.; and Tubaro, S. 2023. Reliability Estimation for Synthetic Speech Detection. In ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1--5

  48. [56]

    Schneider, S.; Baevski, A.; Collobert, R.; and Auli, M. 2019. wav2vec: Unsupervised Pre-training for Speech Recognition. In Interspeech

  49. [57]

    F.; Kastner, K.; Courville, A

    Sotelo, J.; Mehri, S.; Kumar, K.; Santos, J. F.; Kastner, K.; Courville, A. C.; and Bengio, Y. 2017. Char2Wav: End-to-End Speech Synthesis. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Workshop Track Proceedings . O...

  50. [58]

    Sun, C.; Jia, S.; Hou, S.; and Lyu, S. 2023. AI-Synthesized Voice Detection Using Neural Vocoder Artifacts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshop, 904--912

  51. [59]

    Tak, H.; Patino, J.; Nautsch, A.; Evans, N. W. D.; and Todisco, M. 2020. An explainability study of the constant Q cepstral coefficient spoofing countermeasure for automatic speaker verification. In The Speaker and Language Recognition Workshop

  52. [60]

    Tak, H.; Patino, J.; Todisco, M.; Nautsch, A.; Evans, N.; and Larcher, A. 2021. End-to-end anti-spoofing with RawNet2 . In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6369--6373. IEEE

  53. [61]

    K.; and Liu, T

    Tan, X.; Qin, T.; Soong, F. K.; and Liu, T. 2021. A Survey on Neural Speech Synthesis. CoRR, abs/2106.15561

  54. [62]

    A.; Sahidullah, M.; Evans, N.; Kinnunen, T.; and Yamagishi, J

    Todisco, M.; Delgado, H.; Lee, K. A.; Sahidullah, M.; Evans, N.; Kinnunen, T.; and Yamagishi, J. 2018. Integrated Presentation Attack Detection and Automatic Speaker Verification: Common Features and Gaussian Back-end Fusion . In Proc. Interspeech 2018, 77--81

  55. [63]

    Veaux, C.; Yamagishi, J.; MacDonald, K.; et al. 2016. Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit

  56. [64]

    Y.; Zhang, S.; and Chen, X

    Wang, C.; Yi, J.; Tao, J.; Zhang, C. Y.; Zhang, S.; and Chen, X. 2023. Detection of Cross-Dataset Fake Audio Based on Prosodic and Pronunciation Features . In Proc. INTERSPEECH 2023, 3844--3848

  57. [65]

    Wang, X.; Chen, H.; Tang, S.; Wu, Z.; and Zhu, W. 2022. Disentangled representation learning. arXiv preprint arXiv:2211.11695

  58. [66]

    Wang, X.; and Yamagishi, J. 2023. Can large-scale vocoded spoofed data improve speech spoofing countermeasure with a self-supervised front end? arXiv preprint arXiv:2309.06014

  59. [67]

    J.; Skerry-Ryan, R.; Battenberg, E.; Mariooryad, S.; and Kingma, D

    Weiss, R. J.; Skerry-Ryan, R.; Battenberg, E.; Mariooryad, S.; and Kingma, D. P. 2021. Wave-Tacotron: Spectrogram-Free End-to-End Text-to-Speech Synthesis. In ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 5679--5683

  60. [68]

    S.; Lee, B.-J.; jin Yu, H.; and Evans, N

    weon Jung, J.; Heo, H.-S.; Tak, H.; jin Shim, H.; Chung, J. S.; Lee, B.-J.; jin Yu, H.; and Evans, N. W. D. 2021. AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and S...

  61. [69]

    Xie, Y.; Cheng, H.; Wang, Y.; and Ye, L. 2023. Domain Generalization Via Aggregation and Separation for Audio Deepfake Detection. IEEE Transactions on Information Forensics and Security

  62. [70]

    Yadav, A. K. S.; Bhagtani, K.; Xiang, Z.; Bestagini, P.; Tubaro, S.; and Delp, E. J. 2023. Dsvae: Interpretable disentangled representation for synthetic speech detection. arXiv preprint arXiv:2304.03323

  63. [71]

    A.; Kinnunen, T.; Evans, N.; et al

    Yamagishi, J.; Wang, X.; Todisco, M.; Sahidullah, M.; Patino, J.; Nautsch, A.; Liu, X.; Lee, K. A.; Kinnunen, T.; Evans, N.; et al. 2021. ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection. arXiv preprint arXiv:2109.00537

  64. [72]

    Yamamoto, R.; Song, E.; and Kim, J.-M. 2020. Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 61...

  65. [73]

    Yan, Z.; Zhang, Y.; Fan, Y.; and Wu, B. 2023. Ucf: Uncovering common features for generalizable deepfake detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 22412--22423

  66. [74]

    Yu, C.; Lu, H.; Hu, N.; Yu, M.; Weng, C.; Xu, K.; Liu, P.; Tuo, D.; Kang, S.; Lei, G.; Su, D.; and Yu, D. 2020. DurIAN: Duration Informed Attention Network for Speech Synthesis . In Proc. Interspeech 2020, 2027--2031

  67. [75]

    Zhai, B.; Gao, T.; Xue, F.; Rothchild, D.; Wu, B.; Gonzalez, J.; and Keutzer, K. 2020. SqueezeWave: Extremely Lightweight Vocoders for On-device Speech Synthesis. ArXiv, abs/2001.05685

  68. [76]

    Zhang, X.; Yi, J.; Tao, J.; Wang, C.; and Zhang, C. Y. 2023. Do You Remember? Overcoming Catastrophic Forgetting for Fake Audio Detection. In International conference on machine learning. PMLR

  69. [77]

    Zhang, Y.; Wang, W.; and Zhang, P. 2021. The Effect of Silence and Dual-Band Fusion in Anti-Spoofing System . In Proc. Interspeech 2021, 4279--4283

  70. [78]

    Zhang, Y.-J.; Pan, S.; He, L.; and Ling, Z.-H. 2019. Learning latent representations for style control and transfer in end-to-end speech synthesis. In IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE

  71. [79]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...

  72. [80]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.