Pith. sign in

REVIEW 6 cited by

A Survey on Speech Deepfake Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.13914 v2 pith:4KAUWBF7 submitted 2024-04-22 cs.SD cs.CRcs.MMeess.AS

classification cs.SDcs.CRcs.MMeess.AS
keywords deepfakedetectionspeechsurveyavailabilitychallengescontentevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The availability of smart devices leads to an exponential increase in multimedia content. However, advancements in deep learning have also enabled the creation of highly sophisticated Deepfake content, including speech Deepfakes, which pose a serious threat by generating realistic voices and spreading misinformation. To combat this, numerous challenges have been organized to advance speech Deepfake detection techniques. In this survey, we systematically analyze more than 200 papers published up to March 2024. We provide a comprehensive review of each component in the detection pipeline, including model architectures, optimization techniques, generalizability, evaluation metrics, performance comparisons, available datasets, and open source availability. For each aspect, we assess recent progress and discuss ongoing challenges. In addition, we explore emerging topics such as partial Deepfake detection, cross-dataset evaluation, and defences against adversarial attacks, while suggesting promising research directions. This survey not only identifies the current state of the art to establish strong baselines for future experiments but also offers clear guidance for researchers aiming to enhance speech Deepfake detection systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Performance and Complexity Trade-off Optimization of Speech Models During Training

    cs.SD 2026-01 conditional novelty 5.0 of 10

    By turning each layer's width into a continuous, noise-smoothed parameter, the authors train speech models whose sizes shrink during training, reducing FLOPs and size by roughly 80–90% in their case studies.

  2. Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's Alternative

    eess.AS 2025-08 unverdicted novelty 5.0 of 10

    Fake-Mamba reports EERs of 0.97%, 1.74%, and 5.85% on three speech deepfake benchmarks, but the provided full text is an unrelated paper, so the claims cannot be verified.

  3. CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    CAD combines cross-modal lip-speech alignment with per-modality artifact distillation and reports 99.96% AUC on IDForge-v2, with strong cross-dataset results.

  4. Hybrid Audio Detection Using Fine-Tuned Audio Spectrogram Transformers: A Dataset-Driven Evaluation of Mixed AI-Human Speech

    cs.SD 2025-05 reject novelty 5.0 of 10

    Fine-tuned Audio Spectrogram Transformers achieve 97% accuracy on a new, unreleased hybrid human-AI speech dataset, but the evaluation is in-domain and internally inconsistent.

  5. Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection

    eess.AS 2025-02 conditional novelty 5.0 of 10

    Across six self-supervised speech models and ten deepfake datasets, the first 4-12 transformer layers match full-model fake audio detection performance, reducing parameters by at least half.

  6. When Fine-Tuning is Not Enough: Lessons from HSAD on Hybrid and Adversarial Audio Spoof Detection

    cs.SD 2025-09 reject novelty 4.0 of 10

    A new hybrid spoofed-audio benchmark is claimed to show that fine-tuning on it reaches 97%+ accuracy, but the reported numbers are internally inconsistent.

Pith tools