Pith. sign in

REVIEW 13 cited by

FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.05080 v4 pith:77SMH234 submitted 2021-08-11 cs.CV cs.MMcs.SDeess.AS

classification cs.CVcs.MMcs.SDeess.AS
keywords deepfakedatasetaudiodevelopmultimodalpersonvideovideos
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While the significant advancements have made in the generation of deepfakes using deep learning technologies, its misuse is a well-known issue now. Deepfakes can cause severe security and privacy issues as they can be used to impersonate a person's identity in a video by replacing his/her face with another person's face. Recently, a new problem of generating synthesized human voice of a person is emerging, where AI-based deep learning models can synthesize any person's voice requiring just a few seconds of audio. With the emerging threat of impersonation attacks using deepfake audios and videos, a new generation of deepfake detectors is needed to focus on both video and audio collectively. To develop a competent deepfake detector, a large amount of high-quality data is typically required to capture real-world (or practical) scenarios. Existing deepfake datasets either contain deepfake videos or audios, which are racially biased as well. As a result, it is critical to develop a high-quality video and audio deepfake dataset that can be used to detect both audio and video deepfakes simultaneously. To fill this gap, we propose a novel Audio-Video Deepfake dataset, FakeAVCeleb, which contains not only deepfake videos but also respective synthesized lip-synced fake audios. We generate this dataset using the most popular deepfake generation methods. We selected real YouTube videos of celebrities with four ethnic backgrounds to develop a more realistic multimodal dataset that addresses racial bias, and further help develop multimodal deepfake detectors. We performed several experiments using state-of-the-art detection methods to evaluate our deepfake dataset and demonstrate the challenges and usefulness of our multimodal Audio-Video deepfake dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 23 citations worldwide. Full citation record

  1. Tell me Habibi, is it Real or Fake?

    cs.CV 2025-05 conditional novelty 7.0 of 10

    ArEnAV, the first large-scale Arabic-English code-switched audio-visual deepfake dataset, makes current state-of-the-art detectors fail much more than on monolingual data.

  2. Less is More: Modality-Decoupling for General AIGC Audio-Video Detection

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A decoupled audio-video AIGC detector that fuses independent audio and visual predictions at decision level ranks first in the DDL 2.0 general AIGC detection challenge with a final score of 0.8460.

  3. How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Meta-learning training concentrates loss-relevant LoRA updates in query/key projections and spreads them in output projections, relative to standard empirical-risk training.

  4. Detecting AI-Generated Video: A Vision-Language Dual-View Survey

    cs.CV 2026-07 conditional novelty 6.0 of 10

    AIGC-V detection should be treated as factual fidelity verification and organized by a four-layer vision-language dual-view taxonomy spanning cues, motion, cross-modal consistency, and world-level reasoning.

  5. SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation

    cs.SD 2026-07 conditional novelty 6.0 of 10

    SynSFX provides a multi-generator sound-effect deepfake corpus showing speech detectors fail, joint training mitigates forgetting, but generalization to unseen generators remains poor due to artifact overfitting.

  6. Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    The paper introduces semantic mismatch between authentic audio and video as a new DeepFake detection challenge via the RARV-SMM class and demonstrates that a semantic reinforcement strategy with ImageBind embeddings i...

  7. HOLA: Enhancing Audio-visual Deepfake Detection via Hierarchical Contextual Aggregations and Efficient Pre-training

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A two-stage audio-visual deepfake detector, HOLA, uses 1.81M pre-training samples and hierarchical cross-modal fusion modules to achieve first place and near-perfect AUC on AV-Deepfake1M++ video-level detection.

  8. Bona fide Cross Testing Reveals Weak Spot in Audio Deepfake Detection Systems

    cs.SD 2025-09 reject novelty 5.0 of 10

    A new evaluation protocol exhaustively pairs 164 speech synthesizers with nine bona fide speech types and reports max-pooled EERs, revealing larger failures than pooled averages show.

  9. Lightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A 0.48M-parameter single-stream network with iterative audio-visual fusion outperforms larger two-stream baselines on DF-TIMIT, FakeAVCeleb, and DFDC deepfake detection benchmarks.

  10. CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    CAD combines cross-modal lip-speech alignment with per-modality artifact distillation and reports 99.96% AUC on IDForge-v2, with strong cross-dataset results.

  11. Deepfake News Detection: A Multimodal Framework Integrating LipNet, DeepSpeech and ResNET for Enhanced Audio-Visual Analysis

    cs.CR 2026-07 reject novelty 4.0 of 10

    Using LipNet, DeepSpeech2, BlazeFace, and ResNet18 features with Random Forest, the authors report 94% accuracy on FakeAVCeleb audio features, but evaluation flaws make that result unsupported.

  12. SocialDF: Benchmark Dataset and Detection Model for Mitigating Harmful Deepfake Content on Social Media Platforms

    cs.LG 2025-06 reject novelty 4.0 of 10

    A benchmark of 2,126 Instagram videos labeled real or deepfake by uploader disclosure, evaluated with an LLM fact-checking pipeline that reaches 90.4% accuracy but conflates authenticity with factualness.

  13. Ensemble Deep Learning Approaches for AI-Altered Video Detection

    cs.CV 2026-07 conditional novelty 3.0 of 10

    Merging one audio and three video deepfake detectors with voting-based fusion gives ~70–73% accuracy on FakeAVCeleb, with the audio branch performing at chance in the wild.

Pith tools