Pith. sign in

REVIEW 10 cited by

DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.01426 v2 pith:Z7O77JUS submitted 2023-07-04 cs.CV

classification cs.CV
keywords detectiondeepfakebenchmarkcomprehensivedatadeepfakebenchevaluationevaluations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A critical yet frequently overlooked challenge in the field of deepfake detection is the lack of a standardized, unified, comprehensive benchmark. This issue leads to unfair performance comparisons and potentially misleading results. Specifically, there is a lack of uniformity in data processing pipelines, resulting in inconsistent data inputs for detection models. Additionally, there are noticeable differences in experimental settings, and evaluation strategies and metrics lack standardization. To fill this gap, we present the first comprehensive benchmark for deepfake detection, called DeepfakeBench, which offers three key contributions: 1) a unified data management system to ensure consistent input across all detectors, 2) an integrated framework for state-of-the-art methods implementation, and 3) standardized evaluation metrics and protocols to promote transparency and reproducibility. Featuring an extensible, modular-based codebase, DeepfakeBench contains 15 state-of-the-art detection methods, 9 deepfake datasets, a series of deepfake detection evaluation protocols and analysis tools, as well as comprehensive evaluations. Moreover, we provide new insights based on extensive analysis of these evaluations from various perspectives (e.g., data augmentations, backbones). We hope that our efforts could facilitate future research and foster innovation in this increasingly critical domain. All codes, evaluations, and analyses of our benchmark are publicly available at https://github.com/SCLBD/DeepfakeBench.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 19 citations worldwide. Full citation record

  1. SphereVideo: Prototype-anchored Hyperspherical Boundary for Continual AI-generated Video Detection

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A hyperspherical prototype boundary with temporal-coherence losses improves continual AI-generated video detection by about 3 to 4 percentage points over prior methods.

  2. Abductive Corroboration of Probabilistic AI Models for Forensic Synthetic Media Detection

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Multi-detector corroboration reduces FP/TP from ~0.22 to 0.02 (two models) or 0 (three models) while first measuring OpenAI SynthID production rollout and detector complementarity.

  3. When Generative Replay Meets Evolving Deepfakes: Dual Confusion-Aware Regularization for Incremental Face Forgery Detection

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Adaptively weighting direct supervision versus a relative-separation loss makes generative replay work for incremental deepfake detection.

  4. Practical Manipulation Model for Robust Deepfake Detection

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A data-augmentation method for deepfake detection that adds diverse pseudo-fakes and strong degradations during training, increasing robustness and low-quality benchmark AUC at a slight cost on clean high-quality data.

  5. AuthGuard: Generalizable Deepfake Detection via Language Guidance

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AuthGuard trains a deepfake vision encoder with MLLM-generated text descriptions plus uncertainty-weighted contrastive learning, improving cross-dataset deepfake detection and adding interpretable LLM reasoning.

  6. InfoDense: Density-Aware Regional Decisive Replay for Memory-Efficient Incremental Face Forgery Detection

    cs.CV 2026-07 conditional novelty 5.0 of 10

    InfoDense replays only density-ranked, forgery-decisive face fragments rather than full images, cutting memory use and improving incremental deepfake detection.

  7. A Spatial-Frequency Aware Multi-Scale Fusion Network for Real-Time Deepfake Detection

    cs.CV 2025-08 conditional novelty 5.0 of 10

    SFMFNet uses wavelet-frequency gating, token-selective cross-attention, and blur pooling to reach 0.8682 average cross-dataset AUC with only 1.27 GFLOPs and 6.64M parameters.

  8. Think Twice before Adaptation: Improving Adaptability of DeepFake Detection via Online Test-Time Adaptation

    cs.CV 2025-05 reject novelty 5.0 of 10

    A test-time adaptation method using uncertainty-aware negative learning and gradient masking improves deepfake detector performance under unknown postprocessing and distribution shifts.

  9. LLM Harms: A Taxonomy and Discussion

    cs.CY 2025-12 unverdicted novelty 3.0 of 10

    This paper proposes a taxonomy of LLM harms in five categories and suggests mitigation strategies plus a dynamic auditing system for responsible development.

  10. Ensemble-Based Deepfake Detection using State-of-the-Art Models with Robust Cross-Dataset Generalisation

    cs.CV 2025-07 conditional novelty 3.0 of 10

    Averaging the outputs of six pretrained deepfake detectors achieves stable near-best accuracy on two out-of-domain face forgery datasets.

Pith tools