REVIEW 16 cited by
ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
ASVspoof 5 is the fifth edition in a series of challenges that promote the study of speech spoofing and deepfake attacks, and the design of detection solutions. Compared to previous challenges, the ASVspoof 5 database is built from crowdsourced data collected from a vastly greater number of speakers in diverse acoustic conditions. Attacks, also crowdsourced, are generated and tested using surrogate detection models, while adversarial attacks are incorporated for the first time. New metrics support the evaluation of spoofing-robust automatic speaker verification (SASV) as well as stand-alone detection solutions, i.e., countermeasures without ASV. We describe the two challenge tracks, the new database, the evaluation metrics, baselines, and the evaluation platform, and present a summary of the results. Attacks significantly compromise the baseline systems, while submissions bring substantial improvements.
Forward citations
Cited by 16 Pith papers
-
Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry
A four-backbone SSL ensemble achieves near-perfect deepfake detection in the ImageCLEF 2026 track, while the same team's generated audio ranks first in the generation track, revealing an OR-versus-AND asymmetry betwee...
-
How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection
Meta-learning training concentrates loss-relevant LoRA updates in query/key projections and spreads them in output projections, relative to standard empirical-risk training.
-
Large Audio Language Models for Spoofing-Aware Speaker Verification
Adapted LALMs can reach competitive spoofing-aware speaker verification (89.3% accuracy, 0.19 min a-DCF on an ASVspoof5 subset), though zero-shot performance is near chance.
-
What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection
Underrepresented gender in training suffers higher deepfake-detection error; WavLM gaps stay large under balance, and all post-hoc calibrations leave the EER gap fixed at 1.317 pp.
-
An Intervention-Based Framework for Shortcut Diagnosis in Spoofing Countermeasures
Controlled non-speech interventions cause the largest detection-cost spikes, confirming non-speech structure as the dominant confound-driven shortcut in ASVspoof-trained XLS-R + RawGAT-ST models.
-
Open-Set Source Tracing as Compositional Factors via Structured Prototypes
Factorized orthonormal prototypes with architecture/data/residual subspaces improve few-shot open-set source tracing of synthetic speech over ArcFace on MLAAD.
-
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.
-
Can Emotion Fool Anti-spoofing?
Emotional synthetic speech from zero-shot TTS fools the pre-trained RawNet2 anti-spoofing model, and a gated ensemble of emotion-specialized detectors reduces the error and the emotion gap.
-
Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection
Multi-ratio diffusion reconstruction residuals, added as a scalar-gated correction to a frozen WavLM anchor, lower ITW EER to 15.3% vs 18.3% for a separately optimized reference.
-
Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection
Adding dataset identity as an auxiliary task or adversarial label improves aggregate audio-deepfake detection EER on the 2025 Speech DeepFake Arena benchmark.
-
Wav2DF-TSL: Two-stage Learning with Efficient Pre-training and Hierarchical Experts Fusion for Robust Audio Deepfake Detection
A two-stage training scheme (spoofed-speech adapter pre-training plus hierarchical mixture-of-experts fusion) improves audio deepfake detection EER across four benchmarks, including a 27.5% relative gain on In-the-wild.
-
Multi-Granularity Adaptive Time-Frequency Attention Framework for Audio Deepfake Detection under Real-World Communication Degradations
A multi-granularity adaptive attention model for audio deepfake detection remains accurate across six speech codecs and five packet-loss levels, reportedly outperforming baselines.
-
Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation
Segment-level EER overstates deployment readiness for partial fake speech localizers, which drop from 7.6% to above 40% EER on out-of-domain test sets.
-
KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features
A challenge entry combining Wav2Vec-AASIST audio scores with lightweight handcrafted-feature video scores via calibration and maxout reports 92.78% AUC on AV-Deepfake1M++ testA.
-
Two Views, One Truth: Spectral and Self-Supervised Features Fusion for Robust Speech Deepfake Detection
Fusing CQCC spectral features with Wav2Vec2.0 embeddings via cross-attention lowers average equal error rate from 10.87% to 6.80% across four speech deepfake benchmarks.
-
Generalizable Detection of Audio Deepfakes
A Wav2Vec2 XLS-R based detector with augmented training and teacher-model knowledge transfer reports an ASVspoof 5 test EER of 4.48%, below the 5.56% best single-system reference.
Discussion (0). Sign in to comment.