REVIEW 16 cited by
ADD 2023: the Second Audio Deepfake Detection Challenge
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Audio deepfake detection is an emerging topic in the artificial intelligence community. The second Audio Deepfake Detection Challenge (ADD 2023) aims to spur researchers around the world to build new innovative technologies that can further accelerate and foster research on detecting and analyzing deepfake speech utterances. Different from previous challenges (e.g. ADD 2022), ADD 2023 focuses on surpassing the constraints of binary real/fake classification, and actually localizing the manipulated intervals in a partially fake speech as well as pinpointing the source responsible for generating any fake audio. Furthermore, ADD 2023 includes more rounds of evaluation for the fake audio game sub-challenge. The ADD 2023 challenge includes three subchallenges: audio fake game (FG), manipulation region location (RL) and deepfake algorithm recognition (AR). This paper describes the datasets, evaluation metrics, and protocols. Some findings are also reported in audio deepfake detection tasks.
Forward citations
Cited by 16 Pith papers
-
AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks
A 260-hour emotional deepfake benchmark spanning 21 attack systems shows state-of-the-art speech deepfake detectors degrade badly on emotionally expressive and LALM-based spoofing.
-
Probing Speaker Identity Sensitivity in Audio Deepfake Detectors
A new Identity Sensitivity Score flags misclassified audio deepfake detections with AUC up to 0.954, but its claim to isolate speaker-identity behavior from plain confidence is not yet controlled.
-
Time-Frequency Consistency Learning for Robust Speech Deepfake Detection
Under a simulated real-time front-end pipeline, TFCL — temporal cross-attention alignment plus frequency CKA consistency — cuts EER roughly in half relative to matched baselines on ASVspoof2019 LA.
-
Large Audio Language Models for Spoofing-Aware Speaker Verification
Adapted LALMs can reach competitive spoofing-aware speaker verification (89.3% accuracy, 0.19 min a-DCF on an ASVspoof5 subset), though zero-shot performance is near chance.
-
Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models
Speech DF Arena standardizes audio deepfake detection benchmarking across 14 datasets and 15 systems, showing that most open-source detectors have high error rates on out-of-domain attacks.
-
Can Emotion Fool Anti-spoofing?
Emotional synthetic speech from zero-shot TTS fools the pre-trained RawNet2 anti-spoofing model, and a gated ensemble of emotion-specialized detectors reduces the error and the emotion gap.
-
Hidden-Domain Routing for All-Type Audio Deepfake Detection
A router-then-specialist audio deepfake detector, which classifies audio type first and then applies type-specific models and thresholds, achieved 96.10% Macro-F1 and first place on AT-ADD Track2.
-
Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection
Adding dataset identity as an auxiliary task or adversarial label improves aggregate audio-deepfake detection EER on the 2025 Speech DeepFake Arena benchmark.
-
Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation
Segment-level EER overstates deployment readiness for partial fake speech localizers, which drop from 7.6% to above 40% EER on out-of-domain test sets.
-
Hybrid Audio Detection Using Fine-Tuned Audio Spectrogram Transformers: A Dataset-Driven Evaluation of Mixed AI-Human Speech
Fine-tuned Audio Spectrogram Transformers achieve 97% accuracy on a new, unreleased hybrid human-AI speech dataset, but the evaluation is in-domain and internally inconsistent.
-
Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection
Across six self-supervised speech models and ten deepfake datasets, the first 4-12 transformer layers match full-model fake audio detection performance, reducing parameters by at least half.
-
Teffic-Audio: Tell Fact from Fiction
A simple Conformer deepfake detector trained with multi-source balanced sampling and diverse augmentation reaches 1.454% pooled EER on Speech-DF-Arena, first among public systems.
-
Segment Transformer: AI-Generated Music Detection via Music Structural Analysis
A two-stage transformer framework classifies AI-generated music from short clips and beat-segmented full tracks, reporting 99.9% accuracy on SONICS without releasing code or ablations.
-
When Fine-Tuning is Not Enough: Lessons from HSAD on Hybrid and Adversarial Audio Spoof Detection
A new hybrid spoofed-audio benchmark is claimed to show that fine-tuning on it reaches 97%+ accuracy, but the reported numbers are internally inconsistent.
-
Enhancing In-Domain and Out-Domain EmoFake Detection via Cooperative Multilingual Speech Foundation Models
Multilingual speech foundation models, fused with a Tucker-Hadamard module, are reported to reach about 1% equal error rate on emotion fake audio detection, a large drop from prior benchmarks.
-
Speaker Privacy and Security in the Big Data Era: Protection and Defense against Deepfake
A concise survey of voice anonymization, deepfake detection, and speech watermarking as defenses against deepfake speech, with current challenges.
Discussion (0). Continue with ORCID to comment.