Pith. sign in

REVIEW 3 cited by

Audio Deepfake Attribution: An Initial Dataset and Investigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.10489 v4 pith:CF2RZUHJ submitted 2022-08-21 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords audiodeepfakeattributionclassescrmldatasetknownbinary
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The rapid progress of deep speech synthesis models has posed significant threats to society such as malicious manipulation of content. This has led to an increase in studies aimed at detecting so-called deepfake audio. However, existing works focus on the binary detection of real audio and fake audio. In real-world scenarios such as model copyright protection and digital evidence forensics, binary classification alone is insufficient. It is essential to identify the source of deepfake audio. Therefore, audio deepfake attribution has emerged as a new challenge. To this end, we designed the first deepfake audio dataset for the attribution of audio generation tools, called Audio Deepfake Attribution (ADA), and conducted a comprehensive investigation on system fingerprints. To address the challenges of attribution of continuously emerging unknown audio generation tools in the real world, we propose the Class-Representation Multi-Center Learning (CRML) method for open-set audio deepfake attribution (OSADA). CRML enhances the global directional variation of representations, ensuring the learning of discriminative representations with strong intra-class similarity and inter-class discrepancy among known classes. Finally, the strong class discrimination capability learned from known classes is extended to both known and unknown classes. Experimental results demonstrate that the CRML method effectively addresses open-set risks in real-world scenarios. The dataset is publicly available at: https://zenodo.org/records/13318702, and https://zenodo.org/records/13340666.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Source Tracing of Synthetic Speech Systems Through Paralinguistic Pre-Trained Representations

    eess.AS 2025-06 conditional novelty 5.0 of 10

    Paralinguistic speech representations, especially TRILLsson, are the most effective single features for tracing synthetic speech to its source generator, and the TRIO fusion with x-vector reports new accuracy highs.

  2. Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution

    eess.AS 2024-12 conditional novelty 5.0 of 10

    x-vector embeddings and a Rényi divergence fusion loss achieve the best audio deepfake source attribution on ASVspoof 2019 and CFAD, though the benchmark protocol is non-standard.

  3. Reject Threshold Adaptation for Open-Set Model Attribution of Deepfake Audio

    cs.SD 2024-12 conditional novelty 3.0 of 10

    ReTA improves open-set deepfake audio source attribution by learning reconstruction error distributions and computing per-class reject thresholds automatically.

Pith tools