Pith. sign in

REVIEW 12 cited by

Audio Deepfake Detection: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.14970 v1 pith:FRFOETAL submitted 2023-08-29 cs.SD eess.AS

classification cs.SDeess.AS
keywords detectiondeepfakeaudiosurveydatasetsdevelopmentsevaluationfeatures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio deepfake detection is an emerging active topic. A growing number of literatures have aimed to study deepfake detection algorithms and achieved effective performance, the problem of which is far from being solved. Although there are some review literatures, there has been no comprehensive survey that provides researchers with a systematic overview of these developments with a unified evaluation. Accordingly, in this survey paper, we first highlight the key differences across various types of deepfake audio, then outline and analyse competitions, datasets, features, classifications, and evaluation of state-of-the-art approaches. For each aspect, the basic techniques, advanced developments and major challenges are discussed. In addition, we perform a unified comparison of representative features and classifiers on ASVspoof 2021, ADD 2023 and In-the-Wild datasets for audio deepfake detection, respectively. The survey shows that future research should address the lack of large scale datasets in the wild, poor generalization of existing detection methods to unknown fake attacks, as well as interpretability of detection results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Meta-learning training concentrates loss-relevant LoRA updates in query/key projections and spreads them in output projections, relative to standard empirical-risk training.

  2. Probing Speaker Identity Sensitivity in Audio Deepfake Detectors

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A new Identity Sensitivity Score flags misclassified audio deepfake detections with AUC up to 0.954, but its claim to isolate speaker-identity behavior from plain confidence is not yet controlled.

  3. What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Underrepresented gender in training suffers higher deepfake-detection error; WavLM gaps stay large under balance, and all post-hoc calibrations leave the EER gap fixed at 1.317 pp.

  4. Over-the-Air Adversarial Attack Detection: from Datasets to Defenses

    eess.AS 2025-09 conditional novelty 6.0 of 10

    AdvSV 2.0 provides 629k adversarial and bona fide audio samples for automatic speaker verification, with a neural replay simulator attack and a one-class contrastive domain-aligned detector reaching 11.2% EER.

  5. Emoanti: audio anti-deepfake with refined emotion-guided representations

    cs.SD 2025-09 conditional novelty 5.0 of 10

    EmoAnti fine-tunes Wav2Vec2 on emotion recognition and refines the resulting features with a convolutional residual extractor, achieving low EER on ASVspoof LA but worse performance on DF than its own no-finetuning baseline.

  6. Generalizable Audio Spoofing Detection using Non-Semantic Representations

    cs.SD 2025-08 conditional novelty 5.0 of 10

    Frozen non-semantic TRILLson embeddings with a lightweight backend beat prior spoofing detectors on out-of-domain datasets while staying competitive in-domain.

  7. RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection

    cs.SD 2025-05 conditional novelty 5.0 of 10

    A reconstruction-perception-reinforcement-attention framework for audio deepfake detection reports state-of-the-art EERs on ASVspoof 2019/2021 and competitive results on sound and singing benchmarks.

  8. Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes

    cs.SD 2025-05 conditional novelty 5.0 of 10

    ADD-GP, a Gaussian Process classifier with XLS-R speech embeddings, adapts to unseen TTS models with as few as 5 samples and achieves state-of-the-art low error rates on the new LibriFake benchmark.

  9. CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    CAD combines cross-modal lip-speech alignment with per-modality artifact distillation and reports 99.96% AUC on IDForge-v2, with strong cross-dataset results.

  10. Segment Transformer: AI-Generated Music Detection via Music Structural Analysis

    cs.SD 2025-09 conditional novelty 4.0 of 10

    A two-stage transformer framework classifies AI-generated music from short clips and beat-segmented full tracks, reporting 99.9% accuracy on SONICS without releasing code or ablations.

  11. Unmasking Synthetic Realities in Generative AI: A Comprehensive Review of Adversarially Robust Deepfake Detection Systems

    cs.CR 2025-07 conditional novelty 3.0 of 10

    A systematic review of deepfake detection finds a pervasive lack of adversarial robustness evaluation across all modalities and calls for resilient, modality-agnostic detectors.

  12. Speaker Privacy and Security in the Big Data Era: Protection and Defense against Deepfake

    eess.AS 2025-09 accept novelty 1.0 of 10

    A concise survey of voice anonymization, deepfake detection, and speech watermarking as defenses against deepfake speech, with current challenges.

Pith tools