REVIEW 4 major objections 3 minor 44 references
Nearly Solved? Robust Deepfake Detection Requires More than Visual Forensics
T0 review · 4 major / 3 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Recent deepfake detectors that top benchmarks are vulnerable to query-only adversarial attacks, while a zero-shot GPT-4o prompt detects Celeb-DF deepfakes more accurately.
desk verdict A novel typographic attack on VLM deepfake detection is buried under unsupported comparisons against weak reimplementations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing objects are two detector families and the attacks that expose their limits. LaDeDa is a ResNet50 variant whose 1x1 convolutions restrict the receptive field to 9x9 patches, producing per-patch deepfake scores that are average-pooled; this local scoring is what makes it brittle. CLIPping the Deception is a frozen CLIP image-text model adapted by CoOp prompt tuning, where 16 learnable context tokens are optimized while the encoders stay frozen; this is what gives it semantic, less locally brittle features. Against both, the paper uses a genetic black-box attack with population, elites, crossover and mutation, bounded by epsilon, requiring only query access and averaging 1,010 queries per image. The novel typographic attack overlays faint source-path-looking text to manipulate GPT-4o's semantic reading of the image without altering its low-level artifacts.
What would settle it
Run the same bounded genetic black-box attack against the official released weights of LaDeDa and of the CLIP-based detector, with identical epsilon and query budget: if the official LaDeDa shows near-perfect benign AUC on Celeb-DF and its attack success rate is not around 70%, the paper's central robustness comparison fails. Separately, test the typographic overlay with the text removed or placed elsewhere on the same Celeb-DF frames to confirm that the verdict flips are caused by the text rather than by cropping or compression.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that recent gains in deepfake detection do not transfer to adversarial settings. The local-patch detector LaDeDa, which scores 9x9 patches independently and pools them, reaches 69.41% accuracy and AUC 79.63% on FaceForensics++ but only AUC 48.8% on Celeb-DF, and a black-box genetic attack flips 70% of its true-positive detections. The same attack against CLIPping the Deception, a prompt-tuned CLIP detector, flips only about 30%, evidence that semantic embeddings are harder to perturb. A zero-shot GPT-4o prompt achieves AUC 73.18% on Celeb-DF, above the dedicated detectors, but adding white 7%-opacity text that looks like a file path succeeds on 6.64% of its verdict flips. The paper therefore positions hybrid low-level plus high-level detection as the necessary path.
Load-bearing premise
The argument depends on the assumption that the authors' re-trained LaDeDa and CLIP detectors faithfully reproduce the published state-of-the-art models; if the training pipeline did not reproduce them, the 70% versus 30% attack gap may say nothing about the real detectors.
Editorial extensions
If this is right
- Benchmark accuracy on standard deepfake datasets does not imply robustness: a detector that posts near-perfect results can be reduced to about 30% accuracy by a query-only adversarial attack.
- Detectors built on high-level semantic embeddings, such as CLIP-derived features, are substantially more resistant to bounded pixel perturbations than local patch-based detectors.
- Large visuo-lingual models can perform useful zero-shot deepfake detection with no fine-tuning, and here GPT-4o outperforms the dedicated detectors evaluated on Celeb-DF.
- Semantic detectors are vulnerable to a new attack surface: text overlaid on the image can flip verdicts even when the model's stated reason does not mention the text.
- A hybrid detector combining visual artifacts and high-level semantics should be more robust because the two paradigms fail in complementary ways.
Reading between the lines
- An implicit corollary is that the 70% versus 30% attack gap should be re-measured on the official released checkpoints, since the paper's re-trained LaDeDa falls below chance on Celeb-DF and far below the published benign performance.
- A testable extension is to run the typographic attack against a human baseline and against other vision-language models, to see whether the 6.64% success rate reflects a general semantic vulnerability or a GPT-4o-specific quirk.
- The complementarity argument predicts that adversarial examples crafted against a low-level detector will transfer poorly to a semantic detector and vice versa; that transfer experiment is the natural next measurement.
- If the paper's framing is right, deepfake detection should be treated as an adversarial game rather than a fixed benchmark, and evaluation protocols should include perturbation budgets and semantic-manipulation probes from the start.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper investigates the adversarial robustness of deepfake detectors and argues that high-level semantic representations are more robust than low-level visual forensics. The authors retrain LaDeDa and CLIPping the Deception on FaceForensics++, attack them with a genetic black-box algorithm, evaluate zero-shot GPT-4o on a subset of Celeb-DF, and introduce a typographic text-overlay attack on GPT-4o. They conclude that recently developed state-of-the-art detectors are vulnerable to classical black-box attacks, that GPT-4o outperforms current state-of-the-art deepfake detectors, and that hybridising low-level and high-level detectors is a promising direction.
Significance. If the claims were fully supported, the paper would make a useful contribution by connecting adversarial robustness of deepfake detectors to the distinction between semantic and forensic features, by demonstrating a realistic black-box attack, and by introducing a proof-of-concept typographic semantic attack. The authors are transparent about their hypotheses and report concrete numerical results, which makes their claims falsifiable. However, the central comparisons rest on reimplementations that are far below published benchmark performance, so the headline conclusions about state-of-the-art detectors and about GPT-4o superiority are not supported as stated. The typographic attack and the hybridisation discussion are interesting but preliminary.
major comments (4)
- [Section 6.2, Table 1] The retrained LaDeDa model achieves an AUC of 0.488 on Celeb-DF, which is at chance level, and 0.7963 on FaceForensics++, while the authors themselves note that the original method reports near-perfect benchmark performance. The genetic black-box attack with 70% ASR is run on this retrained model, not on the released LaDeDa model. Because the reproduction is not a faithful stand-in for the published state-of-the-art detector, the attack results do not transfer to the actual method, and the Abstract's claim that "recently developed state-of-the-art detectors are susceptible" is unsupported. The training in Section 6.1 uses FaceForensics++ rather than the original WildRF training set, and the terminology argument in Section 6.2 about deepfakes versus AI-generated content does not repair this mismatch.
- [Section 7.2, Table 1] The CLIPping the Deception reproduction scores only 63.82% AUC on FaceForensics++ and 58.84% AUC on Celeb-DF, values far below the published results for that method. The comparison between the 30% ASR for this model and the 70% ASR for the LaDeDa reproduction therefore confounds the detector paradigm (semantic versus forensic) with training quality and reproduction fidelity. Without a faithful and comparably trained baseline, the paper cannot attribute the robustness gap to reliance on high-level semantic embeddings.
- [Section 8.2, Table 1] The claim that GPT-4o performs zero-shot deepfake detection "better than current state-of-the-art methods" is not supported by the evidence. The comparison is made only against the authors' weak reimplementations, not against the published results of LaDeDa or other state-of-the-art detectors. Moreover, on the Celeb-DF subset with 570 fake and 300 real frames, the majority-class baseline accuracy is 570/870 = 0.6552, while GPT-4o achieves 0.6425, which is below the trivial always-fake baseline. The stated superiority claim therefore collapses even relative to the paper's own evaluation set.
- [Section 9] The typographic attack is presented as a proof of concept, and the 6.64% attack success rate is appropriately modest. However, the paper also states that the attack is "without being obvious to a human observer," yet no human evaluation, no ablation over text content or opacity, and no confidence intervals are provided. As a result, the imperceptibility claim is unsupported, and the later discussion in Section 10.1 that overlayed text is "easily detectable by a low-level forensics model" is speculative rather than demonstrated.
minor comments (3)
- [Section 5, Section 6.2] The genetic algorithm description omits the mutation probability "p" and mutation weight "w" used in the experiments, and the phrase "performing on average 1,010 queries per input image batch" is ambiguous: it should state whether queries are counted per image or per batch and how many images were attacked.
- [Section 8.1] The zero-shot GPT-4o evaluation averages five outputs per image, but the paper does not report variance across the five samples, the number of times the model refused to respond, or the number of retries; these details matter because the model is stochastic and the evaluation subset is small.
- [Table 1, Section 10.2] In Table 1, the "NQ" entry for the typographic attack is listed as 0, which is unclear because the zero-shot evaluation is not an iterative attack; the column should be defined more carefully. In Section 10.2, "their potentially susceptibility" should read "their potential susceptibility."
Circularity Check
No circularity: the empirical claims are measured comparisons on external benchmarks, not derivations that reduce to their own inputs.
full rationale
The paper contains no derived equations and no parameter that is fitted and then renamed as a prediction; its central claims are empirical measurements on external benchmarks (FaceForensics++ and Celeb-DF). The black-box attack ASRs are measured on the authors' retrained LaDeDa and CoOp models, and the paper explicitly acknowledges that these retrained models underperform the published versions, even reporting chance-level Celeb-DF AUC for LaDeDa and declining to attack Celeb-DF because of poor benign performance. That is a candid reproducibility/validity limitation, not a circular construction: the attack results are not defined by the training objective, and no equation or definition forces the tested model to be identical to the published detectors. The 'higher semantics' robustness explanation is an interpretation of the observed ASR gap, not an input to the experiments. The GPT-4o zero-shot comparison is likewise an external evaluation, and the typographic attack success rate is a measured outcome after a small-scale pilot, not a quantity implied by the attack design. The cited detectors are from other research groups, with no self-citation chain carrying the argument. Because the evaluation is self-contained against external benchmarks and the limitations are stated rather than hidden, there is no self-definitional, fitted-input, or imported-uniqueness circularity.
Assumptions & free parameters
free parameters (8)
- GA population size =
n=10
- GA generations =
m=100
- GA elites =
k=5
- Maximum perturbation bound =
epsilon = 10/255
- Mutation probability and weight =
p, w (values not reported)
- Typographic text opacity and size =
7% opacity, 29pt
- Number of GPT-4o outputs averaged =
5
- CoOp context tokens =
16
assumptions (5)
- domain assumption The retrained LaDeDa and CoOp models faithfully represent the published state-of-the-art detectors.
- domain assumption Patch-based 1x1 convolutions make the model rely on 'non-robust' local features, increasing adversarial susceptibility.
- domain assumption Celeb-DF is a challenging and representative real-world deepfake benchmark.
- domain assumption Faint overlay text at 7% opacity is imperceptible to human observers.
- domain assumption GPT-4o's knowledge of celebrities in Celeb-DF is a potential confounder.
Cite this review
Pith. "Pith review of Nearly Solved? Robust Deepfake Detection Requires More than Visual Forensics." pith.science (2026). https://pith.science/paper/BDMHIGBL
@misc{pith2026241205676,
author = {Pith},
title = {Pith review of: Nearly Solved? Robust Deepfake Detection Requires More than Visual Forensics},
year = {2026},
howpublished = {\url{https://pith.science/paper/BDMHIGBL}},
note = {Machine review of arXiv:2412.05676}
}
read the original abstract
Deepfakes are on the rise, with increased sophistication and prevalence allowing for high-profile social engineering attacks. Detecting them in the wild is therefore important as ever, giving rise to new approaches breaking benchmark records in this task. In line with previous work, we show that recently developed state-of-the-art detectors are susceptible to classical adversarial attacks, even in a highly-realistic black-box setting, putting their usability in question. We argue that crucial 'robust features' of deepfakes are in their higher semantics, and follow that with evidence that a detector based on a semantic embedding model is less susceptible to black-box perturbation attacks. We show that large visuo-lingual models like GPT-4o can perform zero-shot deepfake detection better than current state-of-the-art methods, and introduce a novel attack based on high-level semantic manipulation. Finally, we argue that hybridising low- and high-level detectors can improve adversarial robustness, based on their complementary strengths and weaknesses.
Figures
Reference graph
Works this paper leans on
-
[1]
Retrieved from https://github.com/deepfakes/ faceswap
DeepFakes Github. Retrieved from https://github.com/deepfakes/ faceswap
-
[2]
Retrieved from https://github.com/ MarekKowalski/FaceSwap/
FaceSwap Github. Retrieved from https://github.com/ MarekKowalski/FaceSwap/
-
[3]
Moustafa Alzantot, Yash Sharma, Supriyo Chakraborty, Huan Zhang, Cho-Jui Hsieh, and Mani Srivastava. 2019. GenAttack: Practical Black-box Attacks with Gradient-Free Optimization. Retrieved from https://arxiv.org/abs/1805.11090
arXiv 2019
-
[4]
Illya Bakurov, Marco Buzzelli, Raimondo Schettini, Mauro Castelli, and Leonardo Vanneschi. 2022. Structural similarity index (SSIM) revisited: A data-driven approach. Expert Systems with Applications 189, (2022), 116087. https://doi.org/ https://doi.org/10.1016/j.eswa. 2021.116087
arXiv 2022
-
[5]
Dylan Butts. 2024. Deepfake scams have robbed companies of millions. Experts warn it could get worse. Retrieved 10 August 2024 from https://www.cnbc.com/2024/05/28/deepfake-scams-have- looted-millions-experts-warn-it-could-get-worse.html
work page 2024
-
[6]
Nicholas Carlini and Hany Farid. 2020. Evading Deepfake-Image De- tectors with White- and Black-Box Attacks. Retrieved from https:// arxiv.org/abs/2004.00622
work page Pith review arXiv 2020
-
[7]
Bar Cavia, Eliahu Horwitz, Tal Reiss, and Yedid Hoshen. 2024. Real- Time Deepfake Detection in the Real-World. Retrieved from https:// arxiv.org/abs/2406.09398
arXiv 2024
-
[8]
Jinyin Chen, Mengmeng Su, Shijing Shen, Hui Xiong, and Haibin Zheng. 2019. POBA-GA: Perturbation optimized black-box adversar- ial attacks via genetic algorithm. Computers & Security 85, (August 2019), 89–106. https://doi.org/10.1016/j.cose.2019.04.014
Show all 44 references
-
[9]
Hao Cheng, Erjia Xiao, Jindong Gu, Le Yang, Jinhao Duan, Jize Zhang, Jiahang Cao, Kaidi Xu, and Renjing Xu. 2024. Unveiling Typographic Deceptions: Insights of the Typographic Vulnerability in Large Vision-Language Model. Retrieved from https://arxiv.org/ abs/2402.19150
2024 arXiv
-
[10]
François Chollet. 2017. Xception: Deep Learning with Depthwise Separable Convolutions. Retrieved from https://arxiv.org/abs/1610. 02357
2017
-
[11]
Gabriel Goh, Nick Cammarata †, Chelsea Voss †, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah. 2021. Multimodal Neurons in Artificial Neural Networks. Distill (2021). https://doi.org/10.23915/distill.00030
2021 doi
-
[12]
Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio
-
[13]
Chenhui Gou, Abdulwahab Felemban, Faizan Farooq Khan, Deyao Zhu, Jianfei Cai, Hamid Rezatofighi, and Mohamed Elhoseiny. 2024. How Well Can Vision Language Models See Image Details?. Re- trieved from https://arxiv.org/abs/2408.03940
2024 arXiv
-
[14]
Yang Hou, Qing Guo, Yihao Huang, Xiaofei Xie, Lei Ma, and Jianjun Zhao. 2023. Evading DeepFake Detectors via Adversarial Statistical Consistency. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2023. 12271–12280. https://doi.org/ 10. 1109/CVPR52...
2023
-
[15]
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, and others. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276 (2024)
2024 arXiv
-
[16]
Shehzeen Hussain, Paarth Neekhara, Malhar Jere, Farinaz Koushan- far, and Julian McAuley. 2021. Adversarial deepfakes: Evaluating vulnerability of deepfake detectors to adversarial examples. In Pro- ceedings of the IEEE/CVF winter conference on applications of computer vision,...
2021
-
[17]
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry. 2019. Adversarial examples are not bugs, they are features. Advances in neural information pro- cessing systems 32, (2019)
2019
-
[18]
Shan Jia, Reilin Lyu, Kangran Zhao, Yize Chen, Zhiyuan Yan, Yan Ju, Chuanbo Hu, Xin Li, Baoyuan Wu, and Siwei Lyu. 2024. Can chatgpt detect deepfakes? a study of using multimodal large language models for media forensics. In Proceedings of the IEEE/CVF Conference on Computer V...
2024
-
[19]
Huo Jingnan. 2024. It’s quick and easy to clone famous politicians’ voices, despite safeguards. Retrieved from https://www.npr.org/ 2024/05/30/nx-s1-4986088/deepfake-audio-elections-politics-ai
2024
-
[20]
Sohail Ahmed Khan and Duc-Tien Dang-Nguyen. 2024. CLIPping the Deception: Adapting Vision-Language Models for Universal Deep- fake Detection. In Proceedings of the 2024 International Conference on Multimedia Retrieval, 2024. 1006–1015
2024
-
[21]
Pavel Korshunov and Sebastien Marcel. 2018. DeepFakes: a New Threat to Face Recognition? Assessment and Detection. Retrieved from https://arxiv.org/abs/1812.08685
2018 arXiv
-
[22]
Romeo Lanzino, Federico Fontana, Anxhelo Diko, Marco Raoul Marini, and Luigi Cinque. 2024. Faster Than Lies: Real-time Deepfake Detection using Binary Neural Networks. In Proceedings of the IEEE/ CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June ...
2024
-
[23]
Tony Lee, Haoqin Tu, Chi Heem Wong, Wenhao Zheng, Yiyang Zhou, Yifan Mai, Josselin Somerville Roberts, Michihiro Yasunaga, Huaxiu Yao, Cihang Xie, and Percy Liang. 2024. VHELM: A Holistic Evalu- ation of Vision Language Models. Retrieved from https://arxiv.org/ abs/2410.07112
2024 arXiv
-
[24]
Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu. 2020. Celeb-DF: A Large-scale Challenging Dataset for DeepFake Foren- sics. Retrieved from https://arxiv.org/abs/1909.12962
2020 arXiv
-
[25]
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2019. Towards Deep Learning Models Resistant to Adversarial Attacks. Retrieved from https://arxiv.org/ abs/1706.06083
2019 arXiv
-
[26]
Utkarsh Ojha, Yuheng Li, and Yong Jae Lee. 2023. Towards universal fake image detectors that generalize across generative models. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. 24480–24489
2023
-
[27]
OpenAI. 2024. Introducing Structured Outputs in the API. Retrieved from https://openai.com/index/introducing-structured- outputs-in-the-api/
2024
-
[28]
Gan Pei, Jiangning Zhang, Menghan Hu, Zhenyu Zhang, Chengjie Wang, Yunsheng Wu, Guangtao Zhai, Jian Yang, Chunhua Shen, and Dacheng Tao. 2024. Deepfake Generation and Detection: A Bench- mark and Survey. Retrieved from https://arxiv.org/abs/2403.17881
2024
-
[29]
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language Models as Knowledge Bases?. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th Inter- national Jo...
2019 doi
-
[30]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, and others. 2021. Learning transferable visual models from natural language supervision. In International conference on machine l...
2021
-
[31]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Super- vision. Retrieved fr...
2021 arXiv
-
[32]
Nick Robins-Early. 2024. CEO of world’s biggest ad firm tar- geted by deepfake scam. Retrieved 10 August 2024 from https://www.theguardian.com/technology/article/2024/may/10/ceo- wpp-deepfake-scam
2024
-
[33]
Vincenzo De Rosa, Fabrizio Guillaro, Giovanni Poggi, Davide Coz- zolino, and Luisa Verdoliva. 2024. Exploring the Adversarial Robust- ness of CLIP for AI-generated Image Detection. Retrieved from https://arxiv.org/abs/2407.19553
2024 arXiv
-
[34]
Andreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. 2018. FaceForensics: A Large- scale Video Dataset for Forgery Detection in Human Faces. Retrieved from https://arxiv.org/abs/1803.09179
2018 arXiv
-
[35]
Andreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. 2019. FaceForensics++: Learning to Detect Manipulated Facial Images. In International Conference on Computer Vision (ICCV), 2019
2019
-
[36]
Sakib Shahriar, Brady Lund, Nishith Reddy Mannuru, Muhammad Arbab Arshad, Kadhim Hayawi, Ravi Varma Kumar Bevara, Aashrith Mannuru, and Laiba Batool. 2024. Putting GPT-4o to the Sword: A Comprehensive Evaluation of Language, Vision, Speech, and Multi- modal Proficiency. Retrie...
2024 arXiv
-
[37]
Haoyu Song, Li Dong, Wei-Nan Zhang, Ting Liu, and Furu Wei. 2022. CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment. Retrieved from https://arxiv.org/abs/2203.07190
2022 arXiv
-
[38]
Justus Thies, Michael Zollhöfer, and Matthias Nießner. 2019. Deferred Neural Rendering: Image Synthesis using Neural Textures. Retrieved from https://arxiv.org/abs/1904.12356
2019 arXiv
-
[39]
Justus Thies, Michael Zollhöfer, Marc Stamminger, Christian Theobalt, and Matthias Nießner. 2020. Face2Face: Real-time Face Capture and Reenactment of RGB Videos. Retrieved from https:// arxiv.org/abs/2007.14808
2020 arXiv
-
[40]
Stuart A. Thomson. 2024. How ‘Deepfake Elon Musk’ Became the Internet’s Biggest Scammer. Retrieved 17 August 2024 from https://www.nytimes.com/interactive/2024/08/14/technology/ elon-musk-ai-deepfake-scam.html
2024
-
[41]
Mika Westerlund. 2019. The emergence of deepfake technology: A review. Technology innovation management review 9, 11 (2019)
2019
-
[42]
Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Chongxuan Li, Ngai-Man Man Cheung, and Min Lin. 2024. On evaluating adversar- ial robustness of large vision-language models. Advances in Neural Information Processing Systems 36, (2024)
2024
-
[43]
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. 2022. Learning to prompt for vision-language models. International Journal of Computer Vision 130, 9 (2022), 2337–2348
2022
-
[2014]
Retrieved from https://arxiv
Generative Adversarial Networks. Retrieved from https://arxiv. org/abs/1406.2661
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.