Pith. sign in

REVIEW 16 cited by

ADD 2023: the Second Audio Deepfake Detection Challenge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.13774 v1 pith:AOPINLM3 submitted 2023-05-23 cs.SD eess.AS

classification cs.SDeess.AS
keywords audiodeepfakefakedetectionchallengeevaluationgameincludes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio deepfake detection is an emerging topic in the artificial intelligence community. The second Audio Deepfake Detection Challenge (ADD 2023) aims to spur researchers around the world to build new innovative technologies that can further accelerate and foster research on detecting and analyzing deepfake speech utterances. Different from previous challenges (e.g. ADD 2022), ADD 2023 focuses on surpassing the constraints of binary real/fake classification, and actually localizing the manipulated intervals in a partially fake speech as well as pinpointing the source responsible for generating any fake audio. Furthermore, ADD 2023 includes more rounds of evaluation for the fake audio game sub-challenge. The ADD 2023 challenge includes three subchallenges: audio fake game (FG), manipulation region location (RL) and deepfake algorithm recognition (AR). This paper describes the datasets, evaluation metrics, and protocols. Some findings are also reported in audio deepfake detection tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks

    eess.AS 2026-08 conditional novelty 7.0 of 10

    A 260-hour emotional deepfake benchmark spanning 21 attack systems shows state-of-the-art speech deepfake detectors degrade badly on emotionally expressive and LALM-based spoofing.

  2. Probing Speaker Identity Sensitivity in Audio Deepfake Detectors

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A new Identity Sensitivity Score flags misclassified audio deepfake detections with AUC up to 0.954, but its claim to isolate speaker-identity behavior from plain confidence is not yet controlled.

  3. Time-Frequency Consistency Learning for Robust Speech Deepfake Detection

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Under a simulated real-time front-end pipeline, TFCL — temporal cross-attention alignment plus frequency CKA consistency — cuts EER roughly in half relative to matched baselines on ASVspoof2019 LA.

  4. Large Audio Language Models for Spoofing-Aware Speaker Verification

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Adapted LALMs can reach competitive spoofing-aware speaker verification (89.3% accuracy, 0.19 min a-DCF on an ASVspoof5 subset), though zero-shot performance is near chance.

  5. Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models

    cs.SD 2025-09 conditional novelty 6.0 of 10

    Speech DF Arena standardizes audio deepfake detection benchmarking across 14 datasets and 15 systems, showing that most open-source detectors have high error rates on out-of-domain attacks.

  6. Can Emotion Fool Anti-spoofing?

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Emotional synthetic speech from zero-shot TTS fools the pre-trained RawNet2 anti-spoofing model, and a gated ensemble of emotion-specialized detectors reduces the error and the emotion gap.

  7. Hidden-Domain Routing for All-Type Audio Deepfake Detection

    cs.SD 2026-08 accept novelty 5.0 of 10

    A router-then-specialist audio deepfake detector, which classifies audio type first and then applies type-specific models and thresholds, achieved 96.10% Macro-F1 and first place on AT-ADD Track2.

  8. Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection

    eess.AS 2026-07 conditional novelty 5.0 of 10

    Adding dataset identity as an auxiliary task or adversarial label improves aggregate audio-deepfake detection EER on the 2025 Speech DeepFake Arena benchmark.

  9. Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation

    cs.SD 2025-07 conditional novelty 5.0 of 10

    Segment-level EER overstates deployment readiness for partial fake speech localizers, which drop from 7.6% to above 40% EER on out-of-domain test sets.

  10. Hybrid Audio Detection Using Fine-Tuned Audio Spectrogram Transformers: A Dataset-Driven Evaluation of Mixed AI-Human Speech

    cs.SD 2025-05 reject novelty 5.0 of 10

    Fine-tuned Audio Spectrogram Transformers achieve 97% accuracy on a new, unreleased hybrid human-AI speech dataset, but the evaluation is in-domain and internally inconsistent.

  11. Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection

    eess.AS 2025-02 conditional novelty 5.0 of 10

    Across six self-supervised speech models and ten deepfake datasets, the first 4-12 transformer layers match full-model fake audio detection performance, reducing parameters by at least half.

  12. Teffic-Audio: Tell Fact from Fiction

    cs.SD 2026-07 conditional novelty 4.0 of 10

    A simple Conformer deepfake detector trained with multi-source balanced sampling and diverse augmentation reaches 1.454% pooled EER on Speech-DF-Arena, first among public systems.

  13. Segment Transformer: AI-Generated Music Detection via Music Structural Analysis

    cs.SD 2025-09 conditional novelty 4.0 of 10

    A two-stage transformer framework classifies AI-generated music from short clips and beat-segmented full tracks, reporting 99.9% accuracy on SONICS without releasing code or ablations.

  14. When Fine-Tuning is Not Enough: Lessons from HSAD on Hybrid and Adversarial Audio Spoof Detection

    cs.SD 2025-09 reject novelty 4.0 of 10

    A new hybrid spoofed-audio benchmark is claimed to show that fine-tuning on it reaches 97%+ accuracy, but the reported numbers are internally inconsistent.

  15. Enhancing In-Domain and Out-Domain EmoFake Detection via Cooperative Multilingual Speech Foundation Models

    eess.AS 2025-07 conditional novelty 4.0 of 10

    Multilingual speech foundation models, fused with a Tucker-Hadamard module, are reported to reach about 1% equal error rate on emotion fake audio detection, a large drop from prior benchmarks.

  16. Speaker Privacy and Security in the Big Data Era: Protection and Defense against Deepfake

    eess.AS 2025-09 accept novelty 1.0 of 10

    A concise survey of voice anonymization, deepfake detection, and speech watermarking as defenses against deepfake speech, with current challenges.

Pith tools