Pith. sign in

REVIEW 14 cited by

Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.12233 v2 pith:FZKXDZJ7 submitted 2022-02-24 eess.AS cs.SD

classification eess.AScs.SD
keywords dataattacksaugmentationdeepfakespoofingwav2vecaccessalmost
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The performance of spoofing countermeasure systems depends fundamentally upon the use of sufficiently representative training data. With this usually being limited, current solutions typically lack generalisation to attacks encountered in the wild. Strategies to improve reliability in the face of uncontrolled, unpredictable attacks are hence needed. We report in this paper our efforts to use self-supervised learning in the form of a wav2vec 2.0 front-end with fine tuning. Despite initial base representations being learned using only bona fide data and no spoofed data, we obtain the lowest equal error rates reported in the literature for both the ASVspoof 2021 Logical Access and Deepfake databases. When combined with data augmentation,these results correspond to an improvement of almost 90% relative to our baseline system.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Less is More: Modality-Decoupling for General AIGC Audio-Video Detection

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A decoupled audio-video AIGC detector that fuses independent audio and visual predictions at decision level ranks first in the DDL 2.0 general AIGC detection challenge with a final score of 0.8460.

  2. Large Audio Language Models for Spoofing-Aware Speaker Verification

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Adapted LALMs can reach competitive spoofing-aware speaker verification (89.3% accuracy, 0.19 min a-DCF on an ASVspoof5 subset), though zero-shot performance is near chance.

  3. SARA: Stress Test Reasoning in Audio Deepfake Detection

    cs.CL 2026-01 conditional novelty 6.0 of 10

    Reasoning traces in audio deepfake detectors fail in two distinct ways — incoherent panic under acoustic attacks, confident false reasoning under linguistic attacks — and those failure signatures may be detectable.

  4. SONAR: Spectral-Contrastive Audio Residuals for Generalizable Deepfake Detection

    cs.SD 2025-11 conditional novelty 6.0 of 10

    SONAR improves audio deepfake detection by explicitly aligning low- and high-frequency representations for real speech and repelling them for fakes, setting new benchmark EERs on ASVspoof 2021 and in-the-wild data.

  5. Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models

    cs.SD 2025-09 conditional novelty 6.0 of 10

    Speech DF Arena standardizes audio deepfake detection benchmarking across 14 datasets and 15 systems, showing that most open-source detectors have high error rates on out-of-domain attacks.

  6. Evaluating Fake Music Detection Performance Under Audio Augmentations

    cs.SD 2025-07 conditional novelty 6.0 of 10

    SONICS, a recent fake-music detector, suffers large accuracy drops under light audio augmentations and fails to generalize to unseen generative models.

  7. What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Small semantic-preserving changes to transcripts, passed through text-to-speech, significantly reduce the accuracy of both open-source and commercial audio anti-spoofing detectors.

  8. Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection

    eess.AS 2026-07 conditional novelty 5.0 of 10

    Adding dataset identity as an auxiliary task or adversarial label improves aggregate audio-deepfake detection EER on the 2025 Speech DeepFake Arena benchmark.

  9. Multi-Granularity Adaptive Time-Frequency Attention Framework for Audio Deepfake Detection under Real-World Communication Degradations

    eess.AS 2025-08 unverdicted novelty 5.0 of 10

    A multi-granularity adaptive attention model for audio deepfake detection remains accurate across six speech codecs and five packet-loss levels, reportedly outperforming baselines.

  10. Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes

    cs.SD 2025-05 conditional novelty 5.0 of 10

    ADD-GP, a Gaussian Process classifier with XLS-R speech embeddings, adapts to unseen TTS models with as few as 5 samples and achieves state-of-the-art low error rates on the new LibriFake benchmark.

  11. Teffic-Audio: Tell Fact from Fiction

    cs.SD 2026-07 conditional novelty 4.0 of 10

    A simple Conformer deepfake detector trained with multi-source balanced sampling and diverse augmentation reaches 1.454% pooled EER on Speech-DF-Arena, first among public systems.

  12. Segment Transformer: AI-Generated Music Detection via Music Structural Analysis

    cs.SD 2025-09 conditional novelty 4.0 of 10

    A two-stage transformer framework classifies AI-generated music from short clips and beat-segmented full tracks, reporting 99.9% accuracy on SONICS without releasing code or ablations.

  13. Two Views, One Truth: Spectral and Self-Supervised Features Fusion for Robust Speech Deepfake Detection

    cs.SD 2025-07 conditional novelty 4.0 of 10

    Fusing CQCC spectral features with Wav2Vec2.0 embeddings via cross-attention lowers average equal error rate from 10.87% to 6.80% across four speech deepfake benchmarks.

  14. Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models

    cs.SD 2025-06 conditional novelty 4.0 of 10

    Swapping the MLP projector for a GR-KAN layer in XLSR-Conformer reduces equal error rates on ASVspoof 2021 LA and DF, reaching 0.70% EER on the variable-length LA set.

Pith tools