Pith. sign in

REVIEW 11 cited by

DF40: Toward Next-Generation Deepfake Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.13495 v2 pith:YG5OWEYL submitted 2024-06-19 cs.CV

classification cs.CV
keywords deepfakedetectionforgerydatasetevaluationsdetectorsdf40techniques
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a new comprehensive benchmark to revolutionize the current deepfake detection field to the next generation. Predominantly, existing works identify top-notch detection algorithms and models by adhering to the common practice: training detectors on one specific dataset (e.g., FF++) and testing them on other prevalent deepfake datasets. This protocol is often regarded as a "golden compass" for navigating SoTA detectors. But can these stand-out "winners" be truly applied to tackle the myriad of realistic and diverse deepfakes lurking in the real world? If not, what underlying factors contribute to this gap? In this work, we found the dataset (both train and test) can be the "primary culprit" due to: (1) forgery diversity: Deepfake techniques are commonly referred to as both face forgery and entire image synthesis. Most existing datasets only contain partial types of them, with limited forgery methods implemented; (2) forgery realism: The dominated training dataset, FF++, contains out-of-date forgery techniques from the past four years. "Honing skills" on these forgeries makes it difficult to guarantee effective detection generalization toward nowadays' SoTA deepfakes; (3) evaluation protocol: Most detection works perform evaluations on one type, which hinders the development of universal deepfake detectors. To address this dilemma, we construct a highly diverse deepfake detection dataset called DF40, which comprises 40 distinct deepfake techniques. We then conduct comprehensive evaluations using 4 standard evaluation protocols and 8 representative detection methods, resulting in over 2,000 evaluations. Through these evaluations, we provide an extensive analysis from various perspectives, leading to 7 new insightful findings. We also open up 4 valuable yet previously underexplored research questions to inspire future works. Our project page is https://github.com/YZY-stack/DF40.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A continuously refreshed, incentive-driven deepfake detector beats static detectors on in-the-wild benchmarks and improves on post-export AI-generated media.

  2. When Generative Replay Meets Evolving Deepfakes: Dual Confusion-Aware Regularization for Incremental Face Forgery Detection

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Adaptively weighting direct supervision versus a relative-separation loss makes generative replay work for incremental deepfake detection.

  3. AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences

    cs.CV 2025-08 conditional novelty 6.0 of 10

    AEGIS is a large-scale benchmark for detecting AI-generated videos, with a hard test set of Sora and KLing clips that current vision-language models detect at near-chance accuracy.

  4. HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A frozen-CLIP plugin with bidirectional visual-text fusion reaches 90.07% average AUC on seven unseen deepfake benchmarks, up 6.68 points over prior work.

  5. Practical Manipulation Model for Robust Deepfake Detection

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A data-augmentation method for deepfake detection that adds diverse pseudo-fakes and strong degradations during training, increasing robustness and low-quality benchmark AUC at a slight cost on clean high-quality data.

  6. AuthGuard: Generalizable Deepfake Detection via Language Guidance

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AuthGuard trains a deepfake vision encoder with MLLM-generated text descriptions plus uncertainty-weighted contrastive learning, improving cross-dataset deepfake detection and adding interpretable LLM reasoning.

  7. InfoDense: Density-Aware Regional Decisive Replay for Memory-Efficient Incremental Face Forgery Detection

    cs.CV 2026-07 conditional novelty 5.0 of 10

    InfoDense replays only density-ranked, forgery-decisive face fragments rather than full images, cutting memory use and improving incremental deepfake detection.

  8. Generalizable Audio Spoofing Detection using Non-Semantic Representations

    cs.SD 2025-08 conditional novelty 5.0 of 10

    Frozen non-semantic TRILLson embeddings with a lightweight backend beat prior spoofing detectors on out-of-domain datasets while staying competitive in-domain.

  9. From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A pipeline combining a deepfake classifier, Grad-CAM heatmaps, image captioning, and an LLM generates layered explanations of deepfake verdicts for non-expert users.

  10. CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    CAD combines cross-modal lip-speech alignment with per-modality artifact distillation and reports 99.96% AUC on IDForge-v2, with strong cross-dataset results.

  11. De-Fake: Style based Anomaly Deepfake Detection

    cs.CV 2025-07 reject novelty 3.0 of 10

    A style-feature face-swap detector that requires a reference photo, with flawed threshold arithmetic and invalid external tests.

Pith tools