REVIEW 2 major objections 5 minor 278 references
Detecting AI-Generated Video: A Vision-Language Dual-View Survey
T0 review · 2 major / 5 minor · reviewed 2026-07-14 · grok-4.5
Pith's one-line read AI-generated video detection should judge whether a video’s events and entities match real-world facts, not only whether frames look fake.
desk verdict Solid organizational survey: factual-fidelity framing plus a usable four-layer dual-view map of AIGC-V detection, with the usual taxonomy-assignment caveats. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Factual Fidelity Verification, carried by a Vision-Language Dual-View four-layer taxonomy. Methods are assigned by dominant evidence pathway: Layer 1 frame-level distributional cues, Layer 2 inter-frame spatiotemporal relations, Layer 3 within-video multimodal verification, and Layer 4 external-knowledge or world-level claim checks. The taxonomy turns a fragmented literature into one evidence hierarchy and defines a detector as returning both a suspicion score and grounded, traceable evidence items.
What would settle it
If independent annotators cannot reproducibly assign multi-cue detectors to layers by the stated evidence-pathway rules, and if strong low-level or temporal-only detectors match or beat world-level systems on high-fidelity generated clips that are internally consistent yet factually false, the taxonomy fails as an operational map of the field.
Extended reading notes
Core claim
The paper’s central claim is that AIGC-V detection is best understood as Factual Fidelity Verification—judging whether fact-level propositions implied by a video and its metadata remain consistent with the real world—and that existing methods form a hierarchical Vision-Language Dual-View landscape of four evidence layers: intrinsic cue analysis, spatiotemporal consistency, cross-modal consistency, and language-guided world-level reasoning. This framing marks a shift from artifact matching in classic deepfake detection to evidence-based semantic verification enabled by vision-language models and agentic pipelines.
Load-bearing premise
The survey treats its four-layer dual-view taxonomy as a natural operational partition of methods by dominant evidence pathway, rather than a convenient filing system that forces multi-cue systems into single-layer labels.
Editorial extensions
If this is right
- Detectors should return structured, grounded evidence—segments, claims, tool outputs—not only a real/fake score.
- Evaluation must stress claim correctness, evidence grounding, localization, and transfer under generator and compression shift, not closed-set AUC alone.
- Benchmarks should track generation paradigms (local manipulation, audio-visual editing, full synthesis) and move toward claim-level and continuously refreshed testbeds.
- Practical systems need cross-layer fusion with abstention when evidence conflicts, plus provenance checks when available.
- Language-view methods become a major recent axis of work while low-level visual forensics remain complementary rather than obsolete.
Reading between the lines
- If generators keep suppressing artifacts while remaining factually loose, systems without external claim-checking will miss the most damaging forgeries: fabricated events that look and sound coherent.
- Claim-level, timestamped annotations may become the scarce resource that bottlenecks progress more than model size, because the taxonomy only becomes testable when “what is fake” is labeled as propositions, not binary clip tags.
- Newsrooms, platforms, and courts that still treat detector scores as black-box priors risk misleading downstream claim verification unless workflows are rebuilt around inspectable evidence objects.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey reframes AI-generated video (AIGC-V) detection as Factual Fidelity Verification (Defs. 3.1–3.2): judging whether the fact-level propositions implied by a video and its metadata are consistent with the real world, and requiring detectors to return a suspicion score together with grounded, traceable evidence. It organizes the field via a Vision-Language Dual-View four-layer taxonomy—intrinsic cues, spatiotemporal consistency, cross-modal consistency, and language-guided world-level reasoning—with explicit evidence-oriented assignment rules. Based on a review of 221 works, the paper synthesizes three generation paradigms (local manipulation, audio-visual editing, generative video synthesis), surveys methods layer by layer with representative tables, provides year-wise layer counts (Table 5) and a protocol-aware AUC snapshot (Table 6), aligns metrics and benchmarks with the dual-view framing, and outlines challenges for diagnostic evaluation, claim-level supervision, unified explainable detection, and evidence-first trustworthiness.
Significance. If the reframing and taxonomy hold as a usable map, the paper supplies a timely organizing principle for a field that has outgrown face-swap forensics and is fragmenting across diffusion generators, VLMs, and agentic pipelines. Strengths include formal task definitions, operational layer boundaries, a large literature base, quantitative year-wise shift evidence (Table 5), paradigm-aligned benchmark tables, and a concrete evidence-first agenda that connects CV forensics with multimodal reasoning. The contribution is organizational rather than empirical; its value is as a shared roadmap and evaluation checklist for the next generation of explainable AIGC-V detectors.
major comments (2)
- §3.2 and the opening of §4 state that layer assignment follows the dominant evidence pathway and that multi-cue modules are mapped by primary evidence type, but the manuscript never reports how many of the 221 works required multi-layer mapping, how borderline Layer 3 vs. Layer 4 cases were resolved, or inter-annotator agreement. Table 5’s year-wise counts and the claim of a decisive 2025 language-view shift therefore rest on an unvalidated assignment procedure. A short appendix with assignment criteria examples, multi-layer counts, and a sensitivity check would make the quantitative summary load-bearing rather than illustrative.
- Table 6 is presented as a ‘layer-wise performance snapshot,’ yet Layer 3 is represented only by a single in-domain result (MDS on DFDC) while Layers 1, 2, and 4 use cross-dataset protocols. The text correctly warns that CD and ID are not comparable, but the table still invites cross-layer reading that the data cannot support. Either expand Layer 3 with CD numbers where available, drop the table to an appendix of protocol-tagged citations, or replace it with a qualitative protocol-coverage matrix so the performance claim does not overreach the evidence.
minor comments (5)
- §4.3: the Layer 3–4 boundary is stated clearly in prose, but a one-sentence decision rule (e.g., ‘external knowledge required for the decisive claim’) repeated at the start of §4.4 would reduce residual ambiguity for multi-agent systems that also do within-video fusion.
- Figure 1 and Figure 3 are dense; ensuring that the four-layer labels and the visual/language split remain legible in print (or providing a simplified schematic) would help readers who first encounter the taxonomy via the figures.
- Table 5’s cumulative denominator includes three pre-2020 papers that are not broken out by year; a footnote listing those three would make the cumulative percentages fully auditable.
- Scattered notation inconsistencies (e.g., ‘Tex t-Visual’ spacing in Figure 3; occasional ‘AIGC-V .’ with a space before the period) should be cleaned in copyediting.
- §7’s comparison table is useful; adding one sentence on how the present survey’s generation-paradigm split (LMV/AVE/GVS) differs from prior face-centric taxonomies would sharpen the novelty claim without lengthening the section.
Circularity Check
No load-bearing circular derivation; mild definitional self-reference only in the author-proposed taxonomy used for counting.
-
self definitional
[§3.2 Dual-view Four-Layer Taxonomy and opening of §4 (layer-assignment rules); Table 5]
"we propose a Vision-Language Dual-View approach and organize existing methods into a four-layer methodological landscape... Our taxonomy is evidence-oriented rather than task-exclusive: layer assignment follows the dominant evidence pathway used for decision-making... When a method combines cues from multiple layers, we map each module to its primary evidence type... Table 5 summarizes the surveyed detection papers... The yearly language-view share rises from 7.7% in 2020... to 62.5% in 2025."
The four-layer partition and the assignment criteria are defined by the authors; the subsequent year-wise counts and ‘shift toward language view’ are then obtained simply by applying those same criteria to the corpus. This is definitional self-reference inherent to any new taxonomy, not a derivation that reduces a claimed prediction or uniqueness result to its inputs by construction. No fitted parameter is renamed a prediction, and the counts remain descriptive rather than forced.
full rationale
This is an organizational survey, not a paper that derives quantitative predictions, uniqueness theorems, or first-principles results from fitted parameters or self-cited lemmas. Definitions 3.1–3.2 introduce Factual Fidelity Verification as the reframing objective; §3.2 and the opening of §4 then state explicit, operational layer-assignment rules (dominant evidence pathway; multi-cue modules mapped by primary evidence type; Layer 1 frame-level distributional vs. Layer 2 sequence relations vs. Layer 3 within-video multimodal vs. Layer 4 external knowledge). Table 5 and the year-wise counts simply tally the 221 surveyed works under those rules. That is ordinary survey taxonomy construction, not a prediction that reduces to its inputs by construction, nor a uniqueness claim imported from overlapping authors, nor an ansatz smuggled via self-citation. Self-citations to the authors’ own prior agent/VLM work appear only as ordinary related-work pointers and are not load-bearing for any central claim. No equations equate a fitted quantity to a claimed prediction. The single mild self-reference is that the taxonomy is author-defined and then used to organize and count; this does not force any result and does not raise the score above 1. The paper is self-contained as a literature synthesis.
Assumptions & free parameters
free parameters (2)
- 221-work corpus inclusion cutoff and selection rules
- Dominant-evidence layer assignment rule
assumptions (4)
- ad hoc to paper AIGC-V detection is best defined as factual fidelity verification of the propositions implied by a video and its metadata against the real world.
- ad hoc to paper Methods can be partitioned into four hierarchical evidence layers: intrinsic cues, spatiotemporal consistency, cross-modal consistency, and language-guided world-level reasoning.
- domain assumption Audio can legitimately enter the language view either as non-linguistic cross-modal alignment or as language-grounded speech evidence.
- domain assumption Generation paradigms can be grouped into local manipulation, audio-visual editing, and generative video synthesis.
invented entities (2)
-
Factual Fidelity Verification
-
Vision-Language Dual-View four-layer taxonomy
Cite this review
Pith. "Pith review of Detecting AI-Generated Video: A Vision-Language Dual-View Survey." pith.science (2026). https://pith.science/paper/UT6VKUYL
@misc{pith2026260710787,
author = {Pith},
title = {Pith review of: Detecting AI-Generated Video: A Vision-Language Dual-View Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/UT6VKUYL}},
note = {Machine review of arXiv:2607.10787}
}
read the original abstract
The evolving realism of AI-generated Videos (AIGC-V) is rapidly rendering traditional artifact-centric detection insufficient, necessitating a paradigm shift from low-level inspection to high-level semantic verification. This paper presents a comprehensive survey of AIGC-V detection, reframing the task as Factual Fidelity Verification, which asks whether the events, entities, and physical processes depicted in a video are consistent with real-world facts. To systematize this rapidly evolving field, we propose a Vision-Language Dual-View taxonomy that organizes existing methods into a hierarchical, four-layer landscape, spanning intrinsic cue analysis, spatiotemporal consistency modeling, cross-modal consistency reasoning, and language-guided world-level reasoning. This dual-view framing highlights a fundamental transition from artifact matching in traditional deepfake detection to evidence-based semantic verification enabled by vision-language models and agentic reasoning pipelines. Based on a systematic review of 221 works, we synthesize AIGC-V generation paradigms, survey the landscape of detection methods, and review evaluation metrics and benchmarks in line with proposed views. Finally, we discuss current challenges and identify promising directions toward robust, explainable, and trustworthy detection.
Reference graph
Works this paper leans on
-
[1]
Advances in Neural Information Processing Systems , volume=
Freqblender: Enhancing deepfake detection by blending frequency knowledge , author=. Advances in Neural Information Processing Systems , volume=
-
[2]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =
Chen, Zhikai and Xie, Lingxi and Pang, Shanmin and He, Yong and Zhang, Bo , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =. 2021 , pages =
2021
-
[3]
2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages =
Zhiyuan Yan and Tao Liu and Mei-Lin Zhang and Wenbin Li and Jianmin Zhang and Jun Lu , title =. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages =. 2024 , doi =
2024
-
[4]
Advances in Neural Information Processing Systems , volume=
Can we leave deepfake data behind in training deepfake detector? , author=. Advances in Neural Information Processing Systems , volume=
-
[5]
Proceedings of the 37th International Conference on Machine Learning , pages =
Leveraging Frequency Analysis for Deep Fake Image Recognition , author =. Proceedings of the 37th International Conference on Machine Learning , pages =. 2020 , editor =
2020
-
[6]
European conference on computer vision , pages=
Explaining deepfake detection by analysing image matching , author=. European conference on computer vision , pages=. 2022 , organization=
2022
-
[7]
2021 IEEE/CVF International Conference on Computer Vision (ICCV) , pages =
Davide Cozzolino and Justus Thies and Gernot Riegler and Oliver Wang and Matthias Niessner and Luisa Verdoliva , title =. 2021 IEEE/CVF International Conference on Computer Vision (ICCV) , pages =. 2021 , doi =
2021
-
[8]
Proceedings of the AAAI conference on artificial intelligence , volume=
Delving into the local: Dynamic inconsistency learning for deepfake video detection , author=. Proceedings of the AAAI conference on artificial intelligence , volume=
Show all 278 references
-
[9]
Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
Exploiting style latent flows for generalizing deepfake video detection , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
-
[10]
European Conference on Computer Vision , pages=
Fake It till You Make It: Curricular Dynamic Forgery Augmentations Towards General Deepfake Detection , author=. European Conference on Computer Vision , pages=
-
[11]
European conference on computer vision , pages=
Hierarchical contrastive inconsistency learning for deepfake video detection , author=. European conference on computer vision , pages=. 2022 , organization=
2022
-
[12]
2019 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) , pages =
Steven Lawrence Fernandes and Raghavender Reddy O and Satoshi Oishi and Niusha Vosoughi and Subrahmanyam Mupparaju and Utkarsh Mittal , title =. 2019 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) , pages =. 2019 , doi =
2019
-
[13]
2020 IEEE international joint conference on biometrics (IJCB) , pages=
How do the hearts of deep fakes beat? Deep fake source detection via interpreting residuals with biological signals , author=. 2020 IEEE international joint conference on biometrics (IJCB) , pages=. 2020 , organization=
2020
-
[14]
2018 IEEE International workshop on information forensics and security (WIFS) , pages=
In ictu oculi: Exposing ai created fake videos by detecting eye blinking , author=. 2018 IEEE International workshop on information forensics and security (WIFS) , pages=. 2018 , organization=
2018
-
[15]
The Visual Computer , volume=
Local attention and long-distance interaction of rPPG for deepfake detection , author=. The Visual Computer , volume=. 2024 , publisher=
2024
-
[16]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Improving the efficiency and robustness of deepfakes detection through precise geometric features , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
-
[17]
2021 International Conference on Information and Communication Technology Convergence (ICTC) , pages=
A study on effective use of bpm information in deepfake detection , author=. 2021 International Conference on Information and Communication Technology Convergence (ICTC) , pages=. 2021 , organization=
2021
-
[18]
IEEE Access , volume=
Real, forged or deep fake? Enabling the ground truth on the internet , author=. IEEE Access , volume=. 2021 , publisher=
2021
-
[19]
arXiv preprint arXiv:2207.08380 , year=
Visual Representations of Physiological Signals for Fake Video Detection , author=. arXiv preprint arXiv:2207.08380 , year=
-
[20]
arXiv preprint arXiv:2110.15561 , year=
Exposing deepfake with pixel-wise ar and ppg correlation from faint signals , author=. arXiv preprint arXiv:2110.15561 , year=
-
[21]
IEEE transactions on pattern analysis and machine intelligence , year=
Fakecatcher: Detection of synthetic portrait videos using biological signals , author=. IEEE transactions on pattern analysis and machine intelligence , year=
-
[22]
Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
Seeable: Soft discrepancies and bounded contrastive learning for exposing deepfakes , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
-
[23]
arXiv preprint arXiv:2406.09601 , year=
Turns Out I'm Not Real: Towards Robust Detection of AI-Generated Videos , author=. arXiv preprint arXiv:2406.09601 , year=
-
[24]
arXiv preprint arXiv:2010.00400 , year=
Deepfakeson-phys: Deepfakes detection based on heart rate estimation , author=. arXiv preprint arXiv:2010.00400 , year=
2010 arXiv
-
[25]
Proceedings of the 28th ACM international conference on multimedia , pages=
Deeprhythm: Exposing deepfakes with attentional visual heartbeat rhythms , author=. Proceedings of the 28th ACM international conference on multimedia , pages=
-
[26]
Advances in Neural Information Processing Systems , volume=
Ost: Improving generalization of deepfake detection via one-shot test-time training , author=. Advances in Neural Information Processing Systems , volume=
-
[27]
Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
Quality-agnostic deepfake detection with intra-model collaborative learning , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
-
[28]
European Conference on Computer Vision , pages=
Real appearance modeling for more general deepfake detection , author=. European Conference on Computer Vision , pages=. 2024 , organization=
2024
-
[29]
The 20th Irish machine vision and image processing conference (IMVIP) , pages=
Detection of deepfake video manipulation , author=. The 20th Irish machine vision and image processing conference (IMVIP) , pages=
-
[30]
Proceedings of the 29th ACM international conference on information & knowledge management , pages=
Towards generalizable deepfake detection with locality-aware autoencoder , author=. Proceedings of the 29th ACM international conference on information & knowledge management , pages=
-
[31]
Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Noise based deepfake detection via multi-head relative-interaction , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=
-
[32]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Deepfake disrupter: The detector of deepfake is my friend , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
-
[33]
ICASSP 2019-2019 IEEE international conference on acoustics, speech and signal processing (ICASSP) , pages=
Exposing deep fakes using inconsistent head poses , author=. ICASSP 2019-2019 IEEE international conference on acoustics, speech and signal processing (ICASSP) , pages=. 2019 , organization=
2019
-
[34]
APSIPA Transactions on Signal and Information Processing , volume =
How Good is ChatGPT at Audiovisual Deepfake Detection: A Comparative Study of ChatGPT, AI Models and Human Perception , author =. APSIPA Transactions on Signal and Information Processing , volume =. 2025 , doi =
2025
-
[35]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) , year =
Can ChatGPT Detect DeepFakes? A Study of Using Multimodal Large Language Models for Media Forensics , author =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) , year =
-
[36]
35th British Machine Vision Conference 2024,
Ching-Yi Lai and Chiou-ting Hsu and Chih-Chung Hsu and Chia-Wen Lin , title =. 35th British Machine Vision Conference 2024,. 2024 , url =
2024
-
[37]
Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence , pages=
From skepticism to acceptance: simulating the attitude dynamics toward fake news , author=. Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence , pages=
-
[38]
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=
The stepwise deception: Simulating the evolution from true news to fake news with llm agents , author=. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=
2025
-
[39]
Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
The truth becomes clearer through debate! multi-agent systems with large language models unmask fake news , author=. Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
-
[40]
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Weaving context across images: Improving vision-language models through focus-centric visual chains , author=. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[41]
arXiv preprint arXiv:2510.24285 , year=
Viper: Empowering the self-evolution of visual perception abilities in vision-language model , author=. arXiv preprint arXiv:2510.24285 , year=
-
[42]
arXiv preprint arXiv:2601.06803 , year=
Forest Before Trees: Latent Superposition for Efficient Visual Reasoning , author=. arXiv preprint arXiv:2601.06803 , year=
-
[43]
Proceedings of the AAAI Conference on Artificial Intelligence , year =
Standing on the Shoulders of Giants: Reprogramming Visual-Language Model for General Deepfake Detection , author =. Proceedings of the AAAI Conference on Artificial Intelligence , year =
-
[44]
Proceedings of the 42nd International Conference on Machine Learning , series =
Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection , author =. Proceedings of the 42nd International Conference on Machine Learning , series =. 2025 , month =
2025
-
[45]
arXiv preprint arXiv:2506.04501 , year =
AuthGuard: Generalizable Deepfake Detection via Language Guidance , author =. arXiv preprint arXiv:2506.04501 , year =
-
[46]
arXiv preprint arXiv:2502.14994 , year =
LAVID: An Agentic LVLM Framework for Diffusion-Generated Video Detection , author =. arXiv preprint arXiv:2502.14994 , year =
-
[48]
arXiv preprint arXiv:2510.02282 , year =
VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL , author =. arXiv preprint arXiv:2510.02282 , year =
-
[49]
arXiv preprint arXiv:2508.21048 , year =
Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning , author =. arXiv preprint arXiv:2508.21048 , year =
-
[51]
arXiv preprint arXiv:2509.22646 , year =
Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs , author =. arXiv preprint arXiv:2509.22646 , year =
-
[52]
arXiv preprint arXiv:2108.05080 , year =
FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset , author =. arXiv preprint arXiv:2108.05080 , year =
-
[53]
arXiv preprint arXiv:2006.07397 , year =
The DeepFake Detection Challenge (DFDC) Dataset , author =. arXiv preprint arXiv:2006.07397 , year =
2006 arXiv
-
[54]
Proceedings of the IEEE/CVF international conference on computer vision , pages=
Faceforensics++: Learning to detect manipulated facial images , author=. Proceedings of the IEEE/CVF international conference on computer vision , pages=
-
[55]
Image Analysis and Processing -- ICIAP 2023 , series =
On Using rPPG Signals for DeepFake Detection: A Cautionary Note , author =. Image Analysis and Processing -- ICIAP 2023 , series =. 2023 , publisher =
2023
-
[56]
The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=
AI-Generated Video Detection via Perceptual Straightening , author=. The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=
-
[57]
arXiv preprint arXiv:2510.07550 , year=
TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility , author=. arXiv preprint arXiv:2510.07550 , year=
-
[58]
arXiv preprint arXiv:2504.02918 , year=
Morpheus: Benchmarking Physical Reasoning of Video Generative Models with Real Physical Experiments , author=. arXiv preprint arXiv:2504.02918 , year=
-
[59]
arXiv preprint arXiv:2505.00337 , year=
T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation , author=. arXiv preprint arXiv:2505.00337 , year=
-
[60]
arXiv preprint arXiv:2507.13428 , year=
PhyWorldBench: A Physical Realism Benchmark for Text-to-Video Generation , author=. arXiv preprint arXiv:2507.13428 , year=
-
[61]
arXiv preprint arXiv:2512.01843 , year=
PhyDetEx: A Benchmark Dataset and Method for Detecting and Explaining Physical Plausibility in Text-to-Video Models , author=. arXiv preprint arXiv:2512.01843 , year=
-
[62]
arXiv preprint arXiv:2411.02385 , year=
How Far is Video Generation from World Model: A Physical Law Perspective , author=. arXiv preprint arXiv:2411.02385 , year=
-
[63]
The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding , author=. The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=
-
[64]
arXiv preprint arXiv:2511.18102 , year=
SPOTLIGHT: Identifying and Localizing Video Generation Errors Using VLMs , author=. arXiv preprint arXiv:2511.18102 , year=
-
[65]
Proceedings of the 29th ACM international conference on multimedia , pages=
Spatiotemporal inconsistency learning for deepfake video detection , author=. Proceedings of the 29th ACM international conference on multimedia , pages=
-
[66]
, author=
Region-Aware Temporal Inconsistency Learning for DeepFake Video Detection. , author=. IJCAI , volume=
-
[67]
International Journal of Computer Vision , volume=
Learning spatiotemporal inconsistency via thumbnail layout for face deepfake detection , author=. International Journal of Computer Vision , volume=. 2024 , publisher=
2024
-
[68]
, author=
Dynamic Inconsistency-aware DeepFake Video Detection. , author=. IJCAI , pages=
-
[69]
European Conference on Computer Vision , pages=
Learning natural consistency representation for face forgery video detection , author=. European Conference on Computer Vision , pages=. 2024 , organization=
2024
-
[70]
2020 IEEE international conference on multimedia and expo (ICME) , pages=
Fsspotter: Spotting face-swapped video by spatial and temporal clues , author=. 2020 IEEE international conference on multimedia and expo (ICME) , pages=. 2020 , organization=
2020
-
[71]
, author=
Detecting Deepfake Videos with Temporal Dropout 3DCNN. , author=. IJCAI , pages=
-
[72]
Proceedings of the IEEE/CVF international conference on computer vision , pages=
Tall: Thumbnail layout for deepfake video detection , author=. Proceedings of the IEEE/CVF international conference on computer vision , pages=
-
[73]
European conference on computer vision , pages=
Two-branch recurrent network for isolating deepfakes in videos , author=. European conference on computer vision , pages=. 2020 , organization=
2020
-
[74]
CoRR , year=
Decof: Generated video detection via frame consistency , author=. CoRR , year=
-
[75]
arXiv preprint arXiv:2501.13435 , year=
GC-ConsFlow: Leveraging Optical Flow Residuals and Global Context for Robust Deepfake Detection , author=. arXiv preprint arXiv:2501.13435 , year=
-
[76]
IEEE Transactions on Information Forensics and Security , volume=
Dynamic difference learning with spatio--temporal correlation for deepfake video detection , author=. IEEE Transactions on Information Forensics and Security , volume=. 2023 , publisher=
2023
-
[77]
IEEE Transactions on Circuits and Systems for Video Technology , volume=
Exploiting complementary dynamic incoherence for deepfake video detection , author=. IEEE Transactions on Circuits and Systems for Video Technology , volume=. 2023 , publisher=
2023
-
[78]
IEEE Transactions on Multimedia , year=
DIP: diffusion learning of inconsistency pattern for general deepfake detection , author=. IEEE Transactions on Multimedia , year=
-
[79]
arXiv preprint arXiv:2507.02398 , year=
Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video Detection , author=. arXiv preprint arXiv:2507.02398 , year=
-
[80]
Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
D3: Training-Free AI-Generated Video Detection Using Second-Order Features , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
-
[81]
arXiv preprint arXiv:2507.03334 , year=
De-Fake: Style based Anomaly Deepfake Detection , author=. arXiv preprint arXiv:2507.03334 , year=
-
[82]
arXiv preprint arXiv:2503.07607 , year=
VoD: Learning Volume of Differences for Video-Based Deepfake Detection , author=. arXiv preprint arXiv:2503.07607 , year=
-
[83]
Proceedings of the Winter Conference on Applications of Computer Vision , pages=
DiffFake: Exposing Deepfakes using Differential Anomaly Detection , author=. Proceedings of the Winter Conference on Applications of Computer Vision , pages=
-
[84]
The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=
Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection , author=. The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=
-
[85]
Proceedings of the Computer Vision and Pattern Recognition Conference , pages=
Detecting Localized Deepfake Manipulations Using Action Unit-Guided Video Representations , author=. Proceedings of the Computer Vision and Pattern Recognition Conference , pages=
-
[86]
Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
Detecting deep-fake videos from aural and oral dynamics , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
-
[87]
arXiv preprint arXiv:2509.25503 , year=
DeepFake Detection in Dyadic Video Calls using Point of Gaze Tracking , author=. arXiv preprint arXiv:2509.25503 , year=
-
[88]
arXiv preprint arXiv:2012.03930 , year=
Identity-driven deepfake detection , author=. arXiv preprint arXiv:2012.03930 , year=
2012 arXiv
-
[89]
2020 IEEE international workshop on information forensics and security (WIFS) , pages=
Detecting deep-fake videos from appearance and behavior , author=. 2020 IEEE international workshop on information forensics and security (WIFS) , pages=. 2020 , organization=
2020
-
[90]
Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops , pages=
Towards untrusted social video verification to combat deepfakes via face geometry consistency , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops , pages=
-
[91]
arXiv preprint arXiv:2501.01184 , year=
Vulnerability-Aware Spatio-Temporal Learning for Generalizable and Interpretable Deepfake Video Detection , author=. arXiv preprint arXiv:2501.01184 , year=
-
[92]
Proceedings of the Computer Vision and Pattern Recognition Conference , pages=
Generalizing deepfake video detection with plug-and-play: Video-level blending and spatiotemporal adapter tuning , author=. Proceedings of the Computer Vision and Pattern Recognition Conference , pages=
-
[93]
Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages=
Interpretable and trustworthy deepfake detection via dynamic prototypes , author=. Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages=
-
[94]
arXiv preprint arXiv:2507.13224 , year=
Leveraging Pre-Trained Visual Models for AI-Generated Video Detection , author=. arXiv preprint arXiv:2507.13224 , year=
-
[95]
International conference on image analysis and processing , pages=
Combining efficientnet and vision transformers for video deepfake detection , author=. International conference on image analysis and processing , pages=. 2022 , organization=
2022
-
[96]
arXiv preprint arXiv:2508.05526 , year=
When Deepfake Detection Meets Graph Neural Network: a Unified and Lightweight Learning Framework , author=. arXiv preprint arXiv:2508.05526 , year=
-
[97]
arXiv preprint arXiv:2506.16802 , year=
Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented Augmentation , author=. arXiv preprint arXiv:2506.16802 , year=
-
[98]
Proceedings of the web conference 2021 , pages=
One detector to rule them all: Towards a general deepfake attack detection framework , author=. Proceedings of the web conference 2021 , pages=
2021
-
[99]
arXiv preprint arXiv:2411.05335 , year=
A quality-centric framework for generic deepfake detection , author=. arXiv preprint arXiv:2411.05335 , year=
-
[100]
Pattern Recognition , volume=
Video anomaly detection via self-supervised and spatio-temporal proxy tasks learning , author=. Pattern Recognition , volume=. 2025 , publisher=
2025
-
[101]
proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
Self-supervised video forensics by audio-visual anomaly detection , author=. proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
-
[102]
Proceedings of the 28th ACM international conference on multimedia , pages=
Sharp multiple instance learning for deepfake video detection , author=. Proceedings of the 28th ACM international conference on multimedia , pages=
-
[103]
Proceedings of the KDD Undergraduate and Master’s Consortium (KDD-UMC’25) , volume=
Deepfake Detection Using Spatiotemporal Methods and Vision-Language Models , author=. Proceedings of the KDD Undergraduate and Master’s Consortium (KDD-UMC’25) , volume=
-
[104]
Artificial Intelligence Review , volume =
Achhardeep Kaur and Azadeh Noori Hoshyar and Vidya Saikrishna and Selena Firmin and Feng Xia , title =. Artificial Intelligence Review , volume =. 2024 , doi =
2024
-
[105]
ACM Computing Surveys , year =
Tianyi Wang and Xin Liao and Kam Pui Chow and Xiaodong Lin and Yinglong Wang , title =. ACM Computing Surveys , year =. doi:10.1145/3699710 , url =
-
[106]
arXiv preprint , eprint =
Gan Pei and Jiangning Zhang and Menghan Hu and Zhenyu Zhang and Chengjie Wang and Yunsheng Wu and Guangtao Zhai and Jian Yang and Chunhua Shen and Dacheng Tao , title =. arXiv preprint , eprint =. 2024 , url =
2024
-
[107]
arXiv preprint , eprint =
Ammarah Hashmi and Sahibzada Adil Shahzad and Chia-Wen Lin and Yu Tsao and Hsin-Min Wang , title =. arXiv preprint , eprint =. 2024 , url =
2024
-
[108]
arXiv preprint , eprint =
Ping Liu and Qiqi Tao and Joey Tianyi Zhou , title =. arXiv preprint , eprint =. 2025 , url =
2025
-
[109]
arXiv preprint , eprint =
Hong-Hanh Nguyen-Le and Van-Tuan Tran and Dinh-Thuc Nguyen and Nhien-An Le-Khac , title =. arXiv preprint , eprint =. 2025 , url =
2025
-
[110]
Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook , journal =
Florinel-Alin Croitoru and Andrei-Iulian H. Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook , journal =. 2024 , url =. 2411.19537 , archivePrefix=
2024 arXiv
-
[111]
Not made for each other- Audio-Visual Dissonance-based Deepfake Detection and Localization
Komal Chugh and Parul Gupta and Abhinav Dhall and Ramanathan Subramanian. Not made for each other- Audio-Visual Dissonance-based Deepfake Detection and Localization. MM 2020 - Proceedings of the 28th ACM International Conference on Multimedia. 2020. doi:10.1145/3394171.3413700
2020 doi
-
[112]
IEEE Transactions on Information Forensics and Security , year=
Preventing DeepFake Attacks on Speaker Authentication by Dynamic Lip Movement Analysis , author=. IEEE Transactions on Information Forensics and Security , year=
-
[113]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops , month =
Agarwal, Shruti and Farid, Hany , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops , month =. 2021 , pages =
2021
-
[114]
Lip Sync Matters: A Novel Multimodal Forgery Detector , year=
Shahzad, Sahibzada Adil and Hashmi, Ammarah and Khan, Sarwar and Peng, Yan-Tsung and Tsao, Yu and Wang, Hsin-Min , booktitle=. Lip Sync Matters: A Novel Multimodal Forgery Detector , year=
-
[115]
2025 , eprint=
SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection , author=. 2025 , eprint=
2025
-
[116]
2025 , eprint=
PIA: Deepfake Detection Using Phoneme-Temporal and Identity-Dynamic Analysis , author=. 2025 , eprint=
2025
-
[117]
2025 , eprint=
Detecting Lip-Syncing Deepfakes: Vision Temporal Transformer for Analyzing Mouth Inconsistencies , author=. 2025 , eprint=
2025
-
[118]
2023 , eprint=
Self-Supervised Video Forensics by Audio-Visual Anomaly Detection , author=. 2023 , eprint=
2023
-
[119]
DeepFake Videos Detection Using Self-Supervised Decoupling Network , year=
Zhang, Jian and Ni, Jiangqun and Xie, Hao , booktitle=. DeepFake Videos Detection Using Self-Supervised Decoupling Network , year=
-
[120]
2022 , eprint=
Voice-Face Homogeneity Tells Deepfake , author=. 2022 , eprint=
2022
-
[121]
2023 , eprint=
Audio-Visual Person-of-Interest DeepFake Detection , author=. 2023 , eprint=
2023
-
[122]
2023 , eprint=
Integrating Audio-Visual Features for Multimodal Deepfake Detection , author=. 2023 , eprint=
2023
-
[123]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops , month =
Bohacek, Matyas and Farid, Hany , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops , month =. 2024 , pages =
2024
-
[124]
2024 , eprint=
Zero-Shot Fake Video Detection by Audio-Visual Consistency , author=. 2024 , eprint=
2024
-
[125]
2025 , eprint=
DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization , author=. 2025 , eprint=
2025
-
[126]
Joint Audio-Visual Deepfake Detection , year=
Zhou, Yipin and Lim, Ser-Nam , booktitle=. Joint Audio-Visual Deepfake Detection , year=
-
[127]
2022 , eprint=
An Audio-Visual Attention Based Multimodal Network for Fake Talking Face Videos Detection , author=. 2022 , eprint=
2022
-
[128]
2024 , eprint=
Unsupervised Multimodal Deepfake Detection Using Intra- and Cross-Modal Inconsistencies , author=. 2024 , eprint=
2024
-
[129]
2024 , eprint=
AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection , author=. 2024 , eprint=
2024
-
[130]
2024 , eprint=
AVT2-DWF: Improving Deepfake Detection with Audio-Visual Fusion and Dynamic Weighting Strategies , author=. 2024 , eprint=
2024
-
[131]
2024 , eprint=
Contextual Cross-Modal Attention for Audio-Visual Deepfake Detection and Localization , author=. 2024 , eprint=
2024
-
[132]
2024 , eprint=
Cross-Modality and Within-Modality Regularization for Audio-Visual DeepFake Detection , author=. 2024 , eprint=
2024
-
[133]
2024 , eprint=
Statistics-aware Audio-visual Deepfake Detector , author=. 2024 , eprint=
2024
-
[134]
A Multimodal Deviation Perceiving Framework for Weakly-Supervised Temporal Forgery Localization , url=
Xu, Wenbo and Wu, Junyan and Lu, Wei and Luo, Xiangyang and Wang, Qian , year=. A Multimodal Deviation Perceiving Framework for Weakly-Supervised Temporal Forgery Localization , url=. doi:10.1145/3746027.3755534 , booktitle=
-
[135]
2025 , eprint=
Audio-Visual Deepfake Detection With Local Temporal Inconsistencies , author=. 2025 , eprint=
2025
-
[136]
2025 , eprint=
FauForensics: Boosting Audio-Visual Deepfake Detection with Facial Action Units , author=. 2025 , eprint=
2025
-
[137]
2025 , eprint=
KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features , author=. 2025 , eprint=
2025
-
[138]
2023 , issue_date =
Ilyas, Hafsa and Javed, Ali and Malik, Khalid Mahmood , title =. 2023 , issue_date =. doi:10.1016/j.asoc.2023.110124 , journal =
2023 doi
-
[139]
2020 , eprint=
Emotions Don't Lie: An Audio-Visual Deepfake Detection Method Using Affective Cues , author=. 2020 , eprint=
2020
-
[140]
2022 , eprint=
M2TR: Multi-modal Multi-scale Transformers for Deepfake Detection , author=. 2022 , eprint=
2022
-
[141]
2025 , eprint=
Weakly Supervised Multimodal Temporal Forgery Localization via Multitask Learning , author=. 2025 , eprint=
2025
-
[142]
2025 , eprint=
CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation , author=. 2025 , eprint=
2025
-
[143]
2025 , eprint=
Consistency-aware Fake Videos Detection on Short Video Platforms , author=. 2025 , eprint=
2025
-
[144]
2024 , eprint=
FKA-Owl: Advancing Multimodal Fake News Detection through Knowledge-Augmented LVLMs , author=. 2024 , eprint=
2024
-
[145]
2025 , eprint=
HOLA: Enhancing Audio-visual Deepfake Detection via Hierarchical Contextual Aggregations and Efficient Pre-training , author=. 2025 , eprint=
2025
-
[146]
arXiv preprint arXiv:2504.21495 , year =
Consistency-aware Fake Videos Detection on Short Video Platforms , author =. arXiv preprint arXiv:2504.21495 , year =
-
[147]
arXiv preprint arXiv:2506.05890 , year =
Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation , author =. arXiv preprint arXiv:2506.05890 , year =
-
[148]
Proceedings of the AAAI Conference on Artificial Intelligence , year =
Multi-modal Deepfake Detection via Multi-task Audio-Visual Prompt Learning , author =. Proceedings of the AAAI Conference on Artificial Intelligence , year =
-
[149]
arXiv preprint arXiv:2306.05241 , year =
Two Heads Are Better Than One: Improving Fake News Video Detection by Correlating with Neighbors , author =. arXiv preprint arXiv:2306.05241 , year =
-
[150]
arXiv preprint arXiv:2507.20286 , year =
T ^3 SVFND: Towards an Evolving Fake News Detector for Emergencies with Test-time Training on Short Video Platforms , author =. arXiv preprint arXiv:2507.20286 , year =
-
[151]
arXiv preprint arXiv:2510.05839 , year =
Towards Robust and Reliable Multimodal Fake News Detection with Incomplete Modality , author =. arXiv preprint arXiv:2510.05839 , year =
-
[152]
arXiv preprint arXiv:2402.11943 , year =
LEMMA: Multi-turn Reasoning with External Knowledge for Multimodal Misinformation Detection , author =. arXiv preprint arXiv:2402.11943 , year =
-
[153]
arXiv preprint arXiv:2505.14714 , year =
KGAlign: Joint Semantic-Structural Knowledge Encoding for Multimodal Fake News Detection , author =. arXiv preprint arXiv:2505.14714 , year =
-
[154]
arXiv preprint arXiv:2508.10444 , year =
DiFaR: Enhancing Multimodal Misinformation Detection with Diverse, Factual, and Relevant Rationales , author =. arXiv preprint arXiv:2508.10444 , year =
-
[155]
2025 , eprint=
Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation , author=. 2025 , eprint=
2025
-
[156]
Miao, Hui and Guo, Yuanfang and Liu, Zeming and Wang, Yunhong , title =. Proceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence and Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence and Fifteenth Symposium on Educational Advanc...
2025 doi
-
[157]
2023 , eprint=
Two Heads Are Better Than One: Improving Fake News Video Detection by Correlating with Neighbors , author=. 2023 , eprint=
2023
-
[158]
2025 , eprint=
Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench , author=. 2025 , eprint=
2025
-
[159]
2025 , eprint=
Perception, Understanding and Reasoning, A Multimodal Benchmark for Video Fake News Detection , author=. 2025 , eprint=
2025
-
[160]
Sora: Creating Video from Text , year =
-
[161]
Official Launch of Seedance 2.0 , year =
-
[162]
Introducing Gen-3 Alpha: A New Frontier for Video Generation , year =
-
[163]
Dream Machine , year =
-
[164]
Kling O1: Unified Multimodal Video Model , year =
-
[165]
2025 , eprint=
Video Reality Test: Can AI-Generated ASMR Videos fool VLMs and Humans? , author=. 2025 , eprint=
2025
-
[166]
arXiv preprint arXiv:2405.19707 , year=
DeMamba: AI-Generated Video Detection on Million-Scale GenVideo Benchmark , author=. arXiv preprint arXiv:2405.19707 , year=
-
[167]
ICLR , year=
LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models , author=. ICLR , year=
-
[168]
arXiv preprint arXiv:2506.14827 , year=
DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning , author=. arXiv preprint arXiv:2506.14827 , year=
-
[169]
2025 , eprint=
Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning , author=. 2025 , eprint=
2025
-
[170]
arXiv preprint arXiv:2501.11340 , year=
Genvidbench: A challenging benchmark for detecting ai-generated video , author=. arXiv preprint arXiv:2501.11340 , year=
-
[171]
arXiv preprint arXiv:2510.16442 , year=
EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning , author=. arXiv preprint arXiv:2510.16442 , year=
-
[172]
arXiv preprint arXiv:2503.14421 , year=
Exddv: A new dataset for explainable deepfake detection in video , author=. arXiv preprint arXiv:2503.14421 , year=
-
[173]
Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
Celeb-df: A large-scale challenging dataset for deepfake forensics , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
-
[174]
arXiv preprint arXiv:2508.14581 , year=
Memory-Anchored Multimodal Reasoning for Explainable Video Forensics , author=. arXiv preprint arXiv:2508.14581 , year=
-
[175]
2020 , isbn =
Zi, Bojia and Chang, Minghao and Chen, Jingjing and Ma, Xingjun and Jiang, Yu-Gang , title =. 2020 , isbn =. doi:10.1145/3394171.3413769 , booktitle =
2020 doi
-
[176]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =
Jiang, Liming and Li, Ren and Wu, Wayne and Qian, Chen and Loy, Chen Change , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =
-
[177]
arXiv preprint , eprint =
Li Lin and Neeraj Gupta and Yue Zhang and Hainan Ren and Chun-Hao Liu and Feng Ding and Xin Wang and Xin Li and Luisa Verdoliva and Shu Hu , title =. arXiv preprint , eprint =. 2024 , doi =
2024
-
[178]
arXiv preprint , eprint =
Yueying Zou and Peipei Li and Zekun Li and Huaibo Huang and Xing Cui and Xuannan Liu and Chenghanyu Zhang and Ran He , title =. arXiv preprint , eprint =. 2025 , doi =
2025
-
[179]
ACM Computing Surveys , volume =
Jingyi Deng and Chenhao Lin and Zhengyu Zhao and Shuai Liu and Zhe Peng and Qian Wang and Chao Shen , title =. ACM Computing Surveys , volume =. 2025 , doi =
2025
-
[180]
ACM Computing Surveys , volume =
Qiang Xu and Wenpeng Mu and Jianing Li and Tanfeng Sun and Xinghao Jiang , title =. ACM Computing Surveys , volume =. 2025 , doi =
2025
-
[181]
2024 , eprint=
X2-DFD: A framework for eXplainable and eXtendable Deepfake Detection , author=. 2024 , eprint=
2024
-
[182]
DeepFake-Adapter: Dual-Level Adapter for DeepFake Detection , volume=
Shao, Rui and Wu, Tianxing and Nie, Liqiang and Liu, Ziwei , year=. DeepFake-Adapter: Dual-Level Adapter for DeepFake Detection , volume=. International Journal of Computer Vision , publisher=. doi:10.1007/s11263-024-02274-6 , number=
-
[183]
2025 , eprint=
GenWorld: Towards Detecting AI-generated Real-world Simulation Videos , author=. 2025 , eprint=
2025
-
[184]
2025 , eprint =
DeepAgent: A Dual Stream Multi Agent Fusion for Robust Multimodal Deepfake Detection , author =. 2025 , eprint =
2025
-
[185]
2024 , note =
C2PA Technical Specification , author =. 2024 , note =
2024
-
[186]
IEEE Journal of Selected Topics in Signal Processing , volume =
Luisa Verdoliva , title =. IEEE Journal of Selected Topics in Signal Processing , volume =. 2020 , doi =
2020
-
[187]
European conference on computer vision , pages=
Common sense reasoning for deepfake detection , author=. European conference on computer vision , pages=. 2024 , organization=
2024
-
[188]
DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection , url =
Yan, Zhiyuan and Zhang, Yong and Yuan, Xinhang and Lyu, Siwei and Wu, Baoyuan , booktitle =. DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection , url =
-
[189]
arXiv preprint arXiv:2503.02857 , year=
Deepfake-eval-2024: A multi-modal in-the-wild benchmark of deepfakes circulated in 2024 , author=. arXiv preprint arXiv:2503.02857 , year=
2024 arXiv
-
[190]
Proceedings of the IEEE/CVF international conference on computer vision , pages=
Kodf: A large-scale korean deepfake detection dataset , author=. Proceedings of the IEEE/CVF international conference on computer vision , pages=
-
[191]
arXiv preprint arXiv:2402.02085 , year=
Detecting AI-Generated Video via Frame Consistency , author=. arXiv preprint arXiv:2402.02085 , year=
-
[192]
2022 International Conference on Digital Image Computing: Techniques and Applications (DICTA) , pages=
Do you really mean that? content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization , author=. 2022 International Conference on Digital Image Computing: Techniques and Applications (DICTA) , pages=. 2022 , organization=
2022
-
[193]
arXiv preprint arXiv:2505.22581 , year=
Tell me Habibi, is it Real or Fake? , author=. arXiv preprint arXiv:2505.22581 , year=
-
[194]
Proceedings of the 32nd ACM International Conference on Multimedia , pages=
AV-Deepfake1M: A large-scale LLM-driven audio-visual deepfake dataset , author=. Proceedings of the 32nd ACM International Conference on Multimedia , pages=
-
[195]
On Learning Multi-Modal Forgery Representation for Diffusion Generated Video Detection , url =
Song, Xiufeng and Guo, Xiao and Zhang, Jiache and Li, Qirui and Bai, Lei and Liu, Xiaoming and Zhai, Guangtao and Liu, Xiaohong , booktitle =. On Learning Multi-Modal Forgery Representation for Diffusion Generated Video Detection , url =. doi:10.52202/079017-3878 , editor =
-
[196]
arXiv preprint arXiv:2505.12620 , year=
BusterX: MLLM-Powered AI-Generated Video Forgery Detection and Explanation , author=. arXiv preprint arXiv:2505.12620 , year=
-
[197]
arXiv preprint arXiv:2507.14632 , year=
Busterx++: Towards unified cross-modal ai-generated content detection and explanation with mllm , author=. arXiv preprint arXiv:2507.14632 , year=
-
[198]
arXiv preprint arXiv:2405.15343 , year=
Distinguish any fake videos: Unleashing the power of large-scale data and motion features , author=. arXiv preprint arXiv:2405.15343 , year=
-
[199]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =
He, Yinan and Gan, Bei and Chen, Siyu and Zhou, Yichun and Yin, Guojun and Song, Luchuan and Sheng, Lu and Shao, Jing and Liu, Ziwei , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =. 2021 , pages =
2021
-
[200]
arXiv preprint arXiv:2505.11109 , year=
MAVOS-DD: Multilingual Audio-Video Open-Set Deepfake Detection Benchmark , author=. arXiv preprint arXiv:2505.11109 , year=
-
[201]
Proceedings of the Computer Vision and Pattern Recognition Conference , pages=
Ai-face: A million-scale demographically annotated ai-generated face dataset and fairness benchmark , author=. Proceedings of the Computer Vision and Pattern Recognition Conference , pages=
-
[202]
Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , month =
Li, Chuqiao and Huang, Zhiwu and Paudel, Danda Pani and Wang, Yabin and Shahbazi, Mohamad and Hong, Xiaopeng and Van Gool, Luc , title =. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , month =. 2023 , pages =
2023
-
[203]
arXiv preprint arXiv:2506.00979 , year=
Ivy-fake: A unified explainable framework and benchmark for image and video aigc detection , author=. arXiv preprint arXiv:2506.00979 , year=
-
[204]
Proceedings of the 4th ACM International Workshop on Multimedia AI against Disinformation , pages=
SocialDF: Benchmark Dataset and Detection Model for Mitigating Harmful Deepfake Content on Social Media Platforms , author=. Proceedings of the 4th ACM International Workshop on Multimedia AI against Disinformation , pages=
-
[205]
Proceedings of the 33rd ACM International Conference on Multimedia , pages=
Av-deepfake1m++: A large-scale audio-visual deepfake benchmark with real-world perturbations , author=. Proceedings of the 33rd ACM International Conference on Multimedia , pages=
-
[206]
International Conference on Intelligent Computing , pages=
FMNV: A Dataset of Media-Published News Videos for Fake News Detection , author=. International Conference on Intelligent Computing , pages=. 2025 , organization=
2025
-
[207]
VCapAV: A Video-Caption Based Audio-Visual Deepfake Detection Dataset , author=. Proc. Interspeech 2025 , pages=
2025
-
[208]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =
Smeu, Stefan and Boldisor, Dragos-Alexandru and Oneata, Dan and Oneata, Elisabeta , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =. 2025 , pages =
2025
-
[209]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =
Narayan, Kartik and Agarwal, Harsh and Thakral, Kartik and Mittal, Surbhi and Vatsa, Mayank and Singh, Richa , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =. 2023 , pages =
2023
-
[210]
Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Fakesv: A multimodal benchmark with rich social context for fake news detection on short video platforms , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=
-
[211]
arXiv preprint arXiv:2408.02954 , year=
WWW: Where, Which and Whatever Enhancing Interpretability in Multimodal Deepfake Detection , author=. arXiv preprint arXiv:2408.02954 , year=
-
[212]
arXiv preprint arXiv:2505.16512 , year=
Beyond Face Swapping: A Diffusion-Based Digital Human Benchmark for Multimodal Deepfake Detection , author=. arXiv preprint arXiv:2505.16512 , year=
-
[213]
2025 , isbn =
Li, Jieyu and Zhang, Xin and Zhou, Joey Tianyi , title =. 2025 , isbn =. doi:10.1145/3746027.3758295 , booktitle =
2025 doi
-
[214]
Proceedings of the Computer Vision and Pattern Recognition Conference , pages=
Towards a universal synthetic video detector: From face or background manipulations to fully ai-generated content , author=. Proceedings of the Computer Vision and Pattern Recognition Conference , pages=
-
[215]
arXiv preprint arXiv:2511.10212 , year=
Next-Frame Feature Prediction for Multimodal Deepfake Detection and Temporal Localization , author=. arXiv preprint arXiv:2511.10212 , year=
-
[216]
2022 , eprint=
High-Resolution Image Synthesis with Latent Diffusion Models , author=. 2022 , eprint=
2022
-
[217]
arXiv preprint arXiv:2508.21052 , year=
FakeParts: a New Family of AI-Generated DeepFakes , author=. arXiv preprint arXiv:2508.21052 , year=
-
[218]
Proceedings of the 34th ACM International Conference on Information and Knowledge Management , pages=
FakeChain: Exposing Shallow Cues in Multi-Step Deepfake Detection , author=. Proceedings of the 34th ACM International Conference on Information and Knowledge Management , pages=
-
[219]
Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
DynamicFace: High-quality and consistent face swapping for image and video using composable 3D facial priors , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
-
[220]
Advances in Neural Information Processing Systems , volume=
FuseAnyPart: Diffusion-Driven Facial Parts Swapping via Multiple Reference Images , author=. Advances in Neural Information Processing Systems , volume=
-
[221]
SIGGRAPH Asia 2022 Conference Papers , year =
VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing in the Wild , author =. SIGGRAPH Asia 2022 Conference Papers , year =. 2211.14758 , archivePrefix =
2022 arXiv
-
[222]
2025 , eprint =
SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion , author =. 2025 , eprint =
2025
-
[223]
2024 , eprint =
Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis , author =. 2024 , eprint =
2024
-
[224]
2025 , eprint =
Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation , author =. 2025 , eprint =
2025
-
[225]
2023 , eprint =
GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis , author =. 2023 , eprint =
2023
-
[226]
2022 , eprint =
Video Diffusion Models , author =. 2022 , eprint =
2022
-
[227]
2022 , eprint =
Imagen Video: High Definition Video Generation with Diffusion Models , author =. 2022 , eprint =
2022
-
[228]
2022 , eprint =
Make-A-Video: Text-to-Video Generation without Text-Video Data , author =. 2022 , eprint =
2022
-
[229]
2023 , eprint =
MagicVideo: Efficient Video Generation with Latent Diffusion Models , author =. 2023 , eprint =
2023
-
[230]
2023 , eprint =
ModelScope Text-to-Video Technical Report , author =. 2023 , eprint =
2023
-
[231]
2022 , eprint =
CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers , author =. 2022 , eprint =
2022
-
[232]
2022 , eprint =
Phenaki: Variable Length Video Generation From Open Domain Textual Description , author =. 2022 , eprint =
2022
-
[233]
2024 , howpublished =
Video generation models as world simulators , author =. 2024 , howpublished =
2024
-
[234]
2024 , howpublished =
Sora System Card , author =. 2024 , howpublished =
2024
-
[235]
2025 , howpublished =
Sora 2 System Card , author =. 2025 , howpublished =
2025
-
[236]
The Thirteenth International Conference on Learning Representations , year =
VideoPhy: Evaluating Physical Commonsense for Video Generation , author =. The Thirteenth International Conference on Learning Representations , year =
-
[237]
Forty-Second International Conference on Machine Learning , year =
WorldSimBench: Towards Video Generation Models as World Simulators , author =. Forty-Second International Conference on Machine Learning , year =
-
[238]
Proceedings of the 42nd International Conference on Machine Learning , year =
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation , author =. Proceedings of the 42nd International Conference on Machine Learning , year =
-
[239]
The Fourteenth International Conference on Learning Representations , year =
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation , author =. The Fourteenth International Conference on Learning Representations , year =
-
[240]
2025 , eprint =
DeepShield: Fortifying Deepfake Video Detection with Local and Global Forgery Analysis , author =. 2025 , eprint =
2025
-
[241]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Lips Don't Lie: A Generalisable and Robust Approach To Face Forgery Detection , author =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[242]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Beyond deepfake images: Detecting ai-generated videos , author =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[243]
Proceedings of the 42nd International Conference on Machine Learning , series=
Impossible Videos , author=. Proceedings of the 42nd International Conference on Machine Learning , series=. 2025 , month=
2025
-
[244]
arXiv preprint arXiv:2501.09038 , year=
Do generative video models understand physical principles? , author=. arXiv preprint arXiv:2501.09038 , year=
-
[245]
International Journal of Computer Vision , volume=
Show-1: Marrying pixel and latent diffusion models for text-to-video generation , author=. International Journal of Computer Vision , volume=. 2025 , publisher=
2025
-
[246]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Grid diffusion models for text-to-video generation , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=. 2024 , url =
2024
-
[247]
Information Fusion , volume =
Ruben Tolosana and Ruben Vera-Rodriguez and Julian Fierrez and Aythami Morales and Javier Ortega-Garcia , title =. Information Fusion , volume =. 2020 , doi =
2020
-
[248]
ACM Computing Surveys , volume =
Yisroel Mirsky and Wenke Lee , title =. ACM Computing Surveys , volume =. 2021 , doi =
2021
-
[249]
Nguyen , title =
Thanh Thi Nguyen and Quoc Viet Hung Nguyen and Dung Tien Nguyen and Duc Thanh Nguyen and Thien Huynh-The and Saeid Nahavandi and Thanh Tam Nguyen and Quoc-Viet Pham and Cuong M. Nguyen , title =. Computer Vision and Image Understanding , volume =. 2022 , doi =
2022
-
[250]
CoRR , volume =
Long Ma and Zihao Xue and Yan Wang and Zhiyuan Yan and Jin Xu and Xiaorui Jiang and Haiyang Yu and Yong Liao and Zhen Bi , title =. CoRR , volume =. 2026 , doi =
2026
-
[251]
CoRR , volume =
Xinan He and Kaiqing Lin and Yue Zhou and Jiaming Zhong and Wei Ye and Wenhui Yi and Bing Fan and Feng Ding and Haodong Li and Bo Cao and Bin Li , title =. CoRR , volume =. 2026 , doi =
2026
-
[252]
CoRR , volume =
Gautam Siddharth Kashyap and Harsh Joshi and Niharika Jain and Ebad Shabbir and Jiechao Gao and Nipun Joshi and Usman Naseem , title =. CoRR , volume =. 2026 , doi =
2026
-
[253]
CoRR , volume =
Ning Jiang and Dingheng Zeng and Yanhong Liu and Haiyang Yi and Shijie Yu and Minghe Weng and Haifeng Shen and Ying Li , title =. CoRR , volume =. 2026 , doi =
2026
-
[254]
CoRR , volume =
Lord Sen and Shyamapada Mukherjee , title =. CoRR , volume =. 2026 , doi =
2026
-
[255]
CoRR , volume =
Adrian Serrano and Erwan Umlil and Ronan Thomas , title =. CoRR , volume =. 2026 , doi =
2026
-
[256]
Doloriel and Habib Ullah and Kristian Hovde Liland and Fadi Al Machot and Ngai-Man Cheung , title =
Chandler Timm C. Doloriel and Habib Ullah and Kristian Hovde Liland and Fadi Al Machot and Ngai-Man Cheung , title =. CoRR , volume =. 2025 , doi =
2025
-
[257]
Al Imran , title =
Sifatullah Sheikh Urmi and Nuzath Tabassum Arthi and Md. Al Imran , title =. CoRR , volume =. 2026 , doi =
2026
-
[258]
CoRR , volume =
Roberto Leotta and Salvatore Alfio Sambataro and Claudio Vittorio Ragaglia and Mirko Casu and Yuri Petralia and Francesco Guarnera and Luca Guarnera and Sebastiano Battiato , title =. CoRR , volume =. 2026 , doi =
2026
-
[259]
CoRR , volume =
Qingcao Li and Miao He and Liang Yi and Qing Wen and Yitao Zhang and Hongshuo Jin and Peng Cheng and Zhongjie Ba and Li Lu and Kui Ren , title =. CoRR , volume =. 2026 , doi =
2026
-
[260]
CoRR , volume =
Hao Tan and Jun Lan and Senyuan Shi and Zichang Tan and Zijian Yu and Huijia Zhu and Weiqiang Wang and Jun Wan and Zhen Lei , title =. CoRR , volume =. 2026 , doi =
2026
-
[261]
Tarek Hasan and Sanjay Saha and Shaojing Fan and Swakkhar Shatabda and Terence Sim , title =
Md. Tarek Hasan and Sanjay Saha and Shaojing Fan and Swakkhar Shatabda and Terence Sim , title =. CoRR , volume =. 2026 , doi =
2026
-
[262]
CoRR , volume =
Chen Chen and Dion Hoe-Lian Goh , title =. CoRR , volume =. 2026 , doi =
2026
-
[263]
A. S. M. Sharifuzzaman Sagar and Mohammed Bennamoun and Farid Boussaid and Naeha Sharif and Lian Xu and Shaaban Sahmoud and Ali Kishk , title =. CoRR , volume =. 2026 , doi =
2026
-
[264]
Collins and Shangzhe Wu and Ayush Tewari and Miri Zilka , title =
Danqing Shi and Lan Jiang and Katherine M. Collins and Shangzhe Wu and Ayush Tewari and Miri Zilka , title =. CoRR , volume =. 2026 , doi =
2026
-
[265]
2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =
Youngseo Kim and Kwan Yun and Seokhyeon Hong and Sihun Cha and Colette Suhjung Koo and Junyong Noh , title =. 2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =. doi:10.48550/ARXIV.2603.08483 , url =
2026 doi
-
[266]
CoRR , volume =
Jianfeng Liao and Yichen Wei and Raymond Chan Ching Bon and Shulan Wang and Kam-Pui Chow and Kwok-Yan Lam , title =. CoRR , volume =. 2026 , doi =
2026
-
[267]
2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , year =
Ashutosh Anshul and Eng Chng and Deepu Rajan , title =. 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , year =
2026
-
[268]
2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , year =
Songjun Cao and Yuqi Li and Yunpeng Luo and Jianjun Yin and Long Ma , title =. 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , year =. doi:10.48550/ARXIV.2602.23393 , url =
2026 doi
-
[269]
2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =
Zheyuan Gu and Qingsong Zhao and Yusong Wang and Zhaohong Huang and Xinqi Li and Cheng Yuan and Jiaowei Shao and Chi Zhang and Xuelong Li , title =. 2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =. doi:10.48550/ARXIV.2602.21779 , url =
2026 doi
-
[270]
Multimedia Tools and Applications , volume =
Sujata Bahadure and Vanita Mane , title =. Multimedia Tools and Applications , volume =. 2026 , doi =
2026
-
[271]
Scientific Reports , year =
Fahima Hajjej and Muhammad Hamid and Ala Saleh Alluhaidan , title =. Scientific Reports , year =. doi:10.1038/s41598-026-40166-6 , url =
-
[272]
Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pages =
William Peebles and Saining Xie , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pages =. 2023 , url =
2023
-
[273]
arXiv preprint arXiv:2503.21755 , year=
VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness , author=. arXiv preprint arXiv:2503.21755 , year=
-
[274]
arXiv preprint arXiv:2412.16211 , year=
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation , author=. arXiv preprint arXiv:2412.16211 , year=
-
[275]
arXiv preprint arXiv:2507.18107 , year=
T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation , author=. arXiv preprint arXiv:2507.18107 , year=
-
[276]
arXiv preprint arXiv:2510.08398 , year=
VideoVerse: How Far is Your T2V Generator from a World Model? , author=. arXiv preprint arXiv:2510.08398 , year=
-
[277]
arXiv preprint arXiv:2512.21507 , year=
SVBench: Evaluation of Video Generation Models on Social Reasoning , author=. arXiv preprint arXiv:2512.21507 , year=
-
[278]
arXiv preprint arXiv:2603.19607 , year=
Physion-Eval: Evaluating Physical Realism in Generated Video via Human Reasoning , author=. arXiv preprint arXiv:2603.19607 , year=
-
[279]
arXiv preprint arXiv:2509.17550 , year=
Is It Certainly a Deepfake? Reliability Analysis in Detection & Generation Ecosystem , author=. arXiv preprint arXiv:2509.17550 , year=
-
[280]
2024 , eprint=
CoAct: A Global-Local Hierarchy for Autonomous Agent Collaboration , author=. 2024 , eprint=
2024
Reviewed July 14, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.