REVIEW 11 cited by
DF40: Toward Next-Generation Deepfake Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose a new comprehensive benchmark to revolutionize the current deepfake detection field to the next generation. Predominantly, existing works identify top-notch detection algorithms and models by adhering to the common practice: training detectors on one specific dataset (e.g., FF++) and testing them on other prevalent deepfake datasets. This protocol is often regarded as a "golden compass" for navigating SoTA detectors. But can these stand-out "winners" be truly applied to tackle the myriad of realistic and diverse deepfakes lurking in the real world? If not, what underlying factors contribute to this gap? In this work, we found the dataset (both train and test) can be the "primary culprit" due to: (1) forgery diversity: Deepfake techniques are commonly referred to as both face forgery and entire image synthesis. Most existing datasets only contain partial types of them, with limited forgery methods implemented; (2) forgery realism: The dominated training dataset, FF++, contains out-of-date forgery techniques from the past four years. "Honing skills" on these forgeries makes it difficult to guarantee effective detection generalization toward nowadays' SoTA deepfakes; (3) evaluation protocol: Most detection works perform evaluations on one type, which hinders the development of universal deepfake detectors. To address this dilemma, we construct a highly diverse deepfake detection dataset called DF40, which comprises 40 distinct deepfake techniques. We then conduct comprehensive evaluations using 4 standard evaluation protocols and 8 representative detection methods, resulting in over 2,000 evaluations. Through these evaluations, we provide an extensive analysis from various perspectives, leading to 7 new insightful findings. We also open up 4 valuable yet previously underexplored research questions to inspire future works. Our project page is https://github.com/YZY-stack/DF40.
Forward citations
Cited by 11 Pith papers
-
Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System
A continuously refreshed, incentive-driven deepfake detector beats static detectors on in-the-wild benchmarks and improves on post-export AI-generated media.
-
When Generative Replay Meets Evolving Deepfakes: Dual Confusion-Aware Regularization for Incremental Face Forgery Detection
Adaptively weighting direct supervision versus a relative-separation loss makes generative replay work for incremental deepfake detection.
-
AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences
AEGIS is a large-scale benchmark for detecting AI-generated videos, with a hard test set of Sora and KLing clips that current vision-language models detect at near-chance accuracy.
-
HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection
A frozen-CLIP plugin with bidirectional visual-text fusion reaches 90.07% average AUC on seven unseen deepfake benchmarks, up 6.68 points over prior work.
-
Practical Manipulation Model for Robust Deepfake Detection
A data-augmentation method for deepfake detection that adds diverse pseudo-fakes and strong degradations during training, increasing robustness and low-quality benchmark AUC at a slight cost on clean high-quality data.
-
AuthGuard: Generalizable Deepfake Detection via Language Guidance
AuthGuard trains a deepfake vision encoder with MLLM-generated text descriptions plus uncertainty-weighted contrastive learning, improving cross-dataset deepfake detection and adding interpretable LLM reasoning.
-
InfoDense: Density-Aware Regional Decisive Replay for Memory-Efficient Incremental Face Forgery Detection
InfoDense replays only density-ranked, forgery-decisive face fragments rather than full images, cutting memory use and improving incremental deepfake detection.
-
Generalizable Audio Spoofing Detection using Non-Semantic Representations
Frozen non-semantic TRILLson embeddings with a lightweight backend beat prior spoofing detectors on out-of-domain datasets while staying competitive in-domain.
-
From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users
A pipeline combining a deepfake classifier, Grad-CAM heatmaps, image captioning, and an LLM generates layered explanations of deepfake verdicts for non-expert users.
-
CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation
CAD combines cross-modal lip-speech alignment with per-modality artifact distillation and reports 99.96% AUC on IDForge-v2, with strong cross-dataset results.
-
De-Fake: Style based Anomaly Deepfake Detection
A style-feature face-swap detector that requires a reference photo, with flawed threshold arithmetic and invalid external tests.
Discussion (0). Continue with ORCID to comment.