Pith. sign in

REVIEW 5 cited by

The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.04880 v3 pith:HMYMOHJI submitted 2024-05-08 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords audiodeepfakealm-baseddetectiondatasetcodecfakecountermeasuredomain
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

With the proliferation of Audio Language Model (ALM) based deepfake audio, there is an urgent need for generalized detection methods. ALM-based deepfake audio currently exhibits widespread, high deception, and type versatility, posing a significant challenge to current audio deepfake detection (ADD) models trained solely on vocoded data. To effectively detect ALM-based deepfake audio, we focus on the mechanism of the ALM-based audio generation method, the conversion from neural codec to waveform. We initially constructed the Codecfake dataset, an open-source, large-scale collection comprising over 1 million audio samples in both English and Chinese, focus on ALM-based audio detection. As countermeasure, to achieve universal detection of deepfake audio and tackle domain ascent bias issue of original sharpness aware minimization (SAM), we propose the CSAM strategy to learn a domain balanced and generalized minima. In our experiments, we first demonstrate that ADD model training with the Codecfake dataset can effectively detects ALM-based audio. Furthermore, our proposed generalization countermeasure yields the lowest average equal error rate (EER) of 0.616% across all test conditions compared to baseline models. The dataset and associated code are available online.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AUDDT: A Unified Benchmark Toolkit for Audio and Speech Deepfake Detectors

    eess.AS 2025-09 conditional novelty 6.0 of 10

    AUDDT packages 28 audio deepfake datasets into a unified benchmarking pipeline and shows that a popular ASVspoof-trained detector's accuracy ranges from 96% to 2% depending on the dataset.

  2. Neural Codec Source Tracing: Toward Comprehensive Attribution in Open-Set Condition

    cs.SD 2025-01 conditional novelty 6.0 of 10

    A new open-set neural codec source tracing benchmark and dataset shows strong in-distribution classification and OOD detection, but poor generalization to unseen real audio.

  3. Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A speech-deepfake dataset for ten public figures built with transcription-based segmentation reports high synthetic naturalness (NISQA 3.69) and a human misclassification rate of 61.9%.

  4. Manipulated Regions Localization For Partially Deepfake Audio: A Survey

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A literature review of manipulated region localization for partially deepfake audio, organized into four method families with a comparison of reported performances.

  5. Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection

    eess.AS 2025-02 conditional novelty 5.0 of 10

    Across six self-supervised speech models and ten deepfake datasets, the first 4-12 transformer layers match full-model fake audio detection performance, reducing parameters by at least half.

Pith tools