REVIEW 13 cited by
FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While the significant advancements have made in the generation of deepfakes using deep learning technologies, its misuse is a well-known issue now. Deepfakes can cause severe security and privacy issues as they can be used to impersonate a person's identity in a video by replacing his/her face with another person's face. Recently, a new problem of generating synthesized human voice of a person is emerging, where AI-based deep learning models can synthesize any person's voice requiring just a few seconds of audio. With the emerging threat of impersonation attacks using deepfake audios and videos, a new generation of deepfake detectors is needed to focus on both video and audio collectively. To develop a competent deepfake detector, a large amount of high-quality data is typically required to capture real-world (or practical) scenarios. Existing deepfake datasets either contain deepfake videos or audios, which are racially biased as well. As a result, it is critical to develop a high-quality video and audio deepfake dataset that can be used to detect both audio and video deepfakes simultaneously. To fill this gap, we propose a novel Audio-Video Deepfake dataset, FakeAVCeleb, which contains not only deepfake videos but also respective synthesized lip-synced fake audios. We generate this dataset using the most popular deepfake generation methods. We selected real YouTube videos of celebrities with four ethnic backgrounds to develop a more realistic multimodal dataset that addresses racial bias, and further help develop multimodal deepfake detectors. We performed several experiments using state-of-the-art detection methods to evaluate our deepfake dataset and demonstrate the challenges and usefulness of our multimodal Audio-Video deepfake dataset.
Forward citations
Cited by 13 Pith papers
-
Tell me Habibi, is it Real or Fake?
ArEnAV, the first large-scale Arabic-English code-switched audio-visual deepfake dataset, makes current state-of-the-art detectors fail much more than on monolingual data.
-
Less is More: Modality-Decoupling for General AIGC Audio-Video Detection
A decoupled audio-video AIGC detector that fuses independent audio and visual predictions at decision level ranks first in the DDL 2.0 general AIGC detection challenge with a final score of 0.8460.
-
How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection
Meta-learning training concentrates loss-relevant LoRA updates in query/key projections and spreads them in output projections, relative to standard empirical-risk training.
-
Detecting AI-Generated Video: A Vision-Language Dual-View Survey
AIGC-V detection should be treated as factual fidelity verification and organized by a four-layer vision-language dual-view taxonomy spanning cues, motion, cross-modal consistency, and world-level reasoning.
-
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation
SynSFX provides a multi-generator sound-effect deepfake corpus showing speech detectors fail, joint training mitigates forgetting, but generalization to unseen generators remains poor due to artifact overfitting.
-
Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge
The paper introduces semantic mismatch between authentic audio and video as a new DeepFake detection challenge via the RARV-SMM class and demonstrates that a semantic reinforcement strategy with ImageBind embeddings i...
-
HOLA: Enhancing Audio-visual Deepfake Detection via Hierarchical Contextual Aggregations and Efficient Pre-training
A two-stage audio-visual deepfake detector, HOLA, uses 1.81M pre-training samples and hierarchical cross-modal fusion modules to achieve first place and near-perfect AUC on AV-Deepfake1M++ video-level detection.
-
Bona fide Cross Testing Reveals Weak Spot in Audio Deepfake Detection Systems
A new evaluation protocol exhaustively pairs 164 speech synthesizers with nine bona fide speech types and reports max-pooled EERs, revealing larger failures than pooled averages show.
-
Lightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework
A 0.48M-parameter single-stream network with iterative audio-visual fusion outperforms larger two-stream baselines on DF-TIMIT, FakeAVCeleb, and DFDC deepfake detection benchmarks.
-
CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation
CAD combines cross-modal lip-speech alignment with per-modality artifact distillation and reports 99.96% AUC on IDForge-v2, with strong cross-dataset results.
-
Deepfake News Detection: A Multimodal Framework Integrating LipNet, DeepSpeech and ResNET for Enhanced Audio-Visual Analysis
Using LipNet, DeepSpeech2, BlazeFace, and ResNet18 features with Random Forest, the authors report 94% accuracy on FakeAVCeleb audio features, but evaluation flaws make that result unsupported.
-
SocialDF: Benchmark Dataset and Detection Model for Mitigating Harmful Deepfake Content on Social Media Platforms
A benchmark of 2,126 Instagram videos labeled real or deepfake by uploader disclosure, evaluated with an LLM fact-checking pipeline that reaches 90.4% accuracy but conflates authenticity with factualness.
-
Ensemble Deep Learning Approaches for AI-Altered Video Detection
Merging one audio and three video deepfake detectors with voting-based fusion gives ~70–73% accuracy on FakeAVCeleb, with the audio branch performing at chance in the wild.
Discussion (0). Continue with ORCID to comment.