Pith. sign in

REVIEW 3 major objections 5 minor 93 references

Foundation Models are Implicit Deepfake Detectors

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Fake images and videos systematically yield lower-magnitude, sparser representations in pretrained self-supervised foundation models, and a score based on those statistics detects deepfakes competitively without any learned classifier.

desk verdict A strong training-free deepfake baseline from feature norms, with a causal interpretation the evidence doesn't yet support. read the letter →

arxiv 2608.09427 v1 pith:JVAG2XFK submitted 2026-08-10 cs.CV

classification cs.CV
keywords deepfakedetectionself-supervisedlearningfeaturenormsparsityanomalyout-of-distributionfoundationmodelsgeneratedimage
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that pretrained self-supervised foundation models are already implicit deepfake detectors: fake images and videos produce feature vectors with systematically lower magnitude and higher sparsity than real ones, across models, datasets, and both image and video domains. On this basis the authors propose NormFake, which scores a sample by either the $\ell^1$ norm of its representation or the $\ell^1/\ell^2$ sparsity ratio, and show that these two statistics rival far more complex trained detectors, reaching 95.1% mean ROC-AUC on the GenImage image benchmark and 82.4% on six audio-visual video benchmarks. They attribute the effect mainly to semantic shift: generated content lies outside the real-world distribution the encoder was pretrained on, so the model encodes it with weaker, more concentrated activations, while low-level generative fingerprints play a comparatively small role. If the claim is right, representation learning itself yields a scalable, classifier-free deepfake-detection signal that strengthens as foundation models grow.

What carries the argument

The load-bearing object is the representation-statistic score, defined directly on the frozen features of a pretrained self-supervised encoder: magnitude($h$) = $\|h\|_1$ and sparsity($h$) = $\|h\|_1/\|h\|_2$, where $h$ is the feature vector of an image (the CLS token for DINOv3) or of a video frame (the visual AV-HuBERT embedding). The sparsity ratio is the Hoyer sparsity measure: invariant to global rescaling and a numerically stable lower bound on true $\ell^0$ sparsity; both statistics are lower for fake than for real samples, so the fakeness score is their negative. These statistics carry the entire argument because no classifier, learning step, or calibration is applied, and the implicit anomaly-detection capability of the foundation model is read off directly. The causal analysis additionally rests on two controlled perturbations, out-of-domain real inputs and autoencoder-reconstructed real inputs run through the same encoders, used to attribute the norm gap to semantic shift rather than low-level fingerprints.

What would settle it

A decisive experiment would hold semantics fixed while varying only the generative pipeline: take real images from a domain the encoder knows well, create a 'fake' set with a generator fine-tuned on that same domain (so content is semantically in-distribution), and check whether the $\ell^1$-norm gap appears; a gap here would mean low-level fingerprints, not semantic shift, carry the signal, overturning the paper's causal explanation.

Watch

Extended reading notes

Core claim

The central discovery is a property of frozen self-supervised representations rather than a new detector architecture: across diverse backbones (including DINOv3-7B for images and AV-HuBERT for video), fake samples consistently yield features with lower $\ell^1$ norm than real samples, and the ratio $\|h\|_1/\|h\|_2$ shows that fake representations are also sparser, meaning their activation energy concentrates in fewer dimensions. The paper operationalizes this as anomaly detection: because foundation models are pretrained on real media, real inputs are in-distribution and produce large, diffuse activations, while fake inputs are out-of-distribution and produce small, concentrated activations, so simply negating the magnitude or sparsity score separates the classes without any learned classifier, and the separation widens with model scale. The authors trace the cause to semantic shift using two controls: out-of-domain real data (EuroSAT satellite imagery and MAVOS-DD Arabic videos) that also show reduced norms but not as low as fakes, and SD1.5 autoencoder reconstructions of real images that inject generative fingerprints while preserving semantics and stay close to the real distribution. They conclude that semantic deviation, not pixel-level fingerprints, is the dominant driver of the norm gap.

Load-bearing premise

The load-bearing premise is that fake media get smaller feature norms because they fall outside the encoder's pretraining distribution in a semantic sense, not because of low-level generative fingerprints or coincidental dataset differences; if the norm gap is actually driven by low-level artifacts, NormFake's generalization to unseen generators and domains is not assured.

Editorial extensions

If this is right

  • A frozen foundation model becomes a zero-shot deepfake detector: NormFake (sparsity) reaches 95.1% mean ROC-AUC across eight generators in GenImage without ever seeing a fake sample during training.
  • The signal transfers across domains and generators: on six audio-visual video benchmarks NormFake (sparsity) averages 82.4% ROC-AUC using only visual AV-HuBERT features, outperforming several fake-aware and real-only methods and landing about one point behind the best multimodal baseline.
  • The discriminative strength scales with backbone size: average GenImage performance rises steadily from the 21M-parameter ViT-S to the 6.7B-parameter ViT-7B, with gains saturating near ViT-H+, while two prior real-only baselines degrade when moved to the larger DINOv3-7B backbone.
  • Because the effect is dominated by semantic shift, NormFake works best on generators whose outputs deviate semantically (continuous latent diffusion models such as SD1.4/1.5, Wukong, Midjourney) and is comparatively weaker on generators whose fakes are semantically close to real images (BigGAN, VQDM), where pixel-level artifact detectors excel.
  • Per-layer analysis shows the deepest representations separate real from fake best, and for AV-HuBERT the final LayerNorm substantially restores separability that the raw last transformer block loses, indicating the signal is carried by semantic features rather than early low-level cues.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: if the norm gap is really tracking pretraining-distribution shift, then NormFake's reliability is hostage to pretraining data; a future self-supervised encoder trained on corpora already containing large amounts of generated media should show a compressed norm gap, exactly the contamination failure the paper lists as a limitation.
  • Editorial inference: the same magnitude/sparsity statistics may extend to audio-only deepfake detection, but the paper's own audio experiments on FakeAVCeleb (sparsity near or below chance) suggest the phenomenon is far weaker outside the visual domain, so any such extension would need fresh evidence rather than an assumption of transfer.
  • Editorial inference: NormFake could serve as a prior or regularizer for trainable detectors, for example by penalizing a learned classifier whenever its features' $\ell^1/\ell^2$ ratio leaves the range typical of real samples, which the paper names as future work but does not test.
  • Editorial inference: the strong correlation between NormFake and audio-visual synchronization methods (SpeechForensics, FACTOR) hints that a single underlying failure, semantic drift of generated media, produces both visual feature shrinkage and audio-visual desynchronization, so detectors across modalities may be measuring facets of the same defect rather than independent cues.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports a consistent empirical phenomenon: frozen self-supervised foundation models produce lower-magnitude feature representations for fake media than for real media, across image and video domains and several backbones. The authors operationalize this as two parameter-free scores, magnitude(h)=||h||_1 and sparsity(h)=||h||_1/||h||_2, and show that these scores alone achieve competitive deepfake detection ROC-AUC on GenImage and on six audio-visual video benchmarks, without training a classifier. They also report analyses aimed at attributing the norm gap to semantic shift rather than low-level generative fingerprints, and a scaling study showing that larger DINOv3 backbones yield stronger detection. The central detection claim is straightforward, falsifiable, and largely supported by the reported tables; the causal attribution and the universality of the phenomenon are less well supported.

Significance. If the core phenomenon holds, this is a valuable and surprising result: a zero-shot, parameter-free statistic of frozen SSL features can compete with trained deepfake detectors, and the connection to the familiarity hypothesis in OOD detection is conceptually useful. The method has no fitted parameters and no calibration stage, which rules out a circular-fitting concern for the detector itself; the paper also evaluates on a broad set of benchmarks and generators and includes informative comparisons with prior real-only and fake-aware methods. The main significance risk is that the paper's explanatory claim—that the effect is primarily semantic—rests on a confounded control experiment, and that the reported test-set layer selection and absence of error bars make the precise quantitative claims less secure. With additional controls and more careful statistical reporting, the contribution would be a strong baseline for the field.

major comments (3)
  1. [Section 3, 'Why does magnitude differ?' and Figure 2] The experiment intended to isolate semantic shift from low-level generative fingerprints is confounded. The SD1.5 autoencoder reconstruction injects only the VAE encoder-decoder path, omitting the diffusion sampling noise, prompt conditioning, and generator-specific frequency characteristics that are present in actual SD1.5 fakes; a small norm shift for VAE reconstructions therefore does not bound the low-level contribution in real fakes. Similarly, the out-of-distribution real controls (EuroSAT and MAVOS-DD Arabic) differ from the encoder's pretraining distribution in sensor, resolution, compression, and language, not only in semantics. Since the qualitative examples in Section 5 show that blurry or cluttered real frames receive low norms, the data are consistent with a low-level mechanism as well as a semantic one. The conclusion that reduced feature magnitude is 'primarily associated with semantic shifts' is therefore underdetermined. I would ask for additional controls, for example full-pipeline generation with fixed semantics, low-level-only perturbations such as compression or blur, and out-of-distribution real data matched in acquisition statistics, before the attribution claim is accepted.
  2. [Section 4, Tables 1 and 2, and supplementary Figure 8] The quantitative comparison is weakened by test-set layer selection and the absence of error bars. The caption of Table 2 states that for NormFake 'we report the best-performing layer for each variant,' and the text explains that the sparsity variant uses the penultimate block; this layer choice is made after inspecting test performance, so the reported 95.1% mean may be optimistic. In addition, all ROC-AUC values in Tables 1-3 and Figures 6-7 are single point estimates with no standard errors or confidence intervals, so statements such as 'trailing by just one percentage point' (Table 1) or 'outperforming' prior methods (Tables 1 and 2) cannot be evaluated for statistical significance. I recommend reporting confidence intervals and a validation-based or pre-specified layer-selection rule, or explicitly labeling the reported numbers as oracle-layer results.
  3. [Section 5, Table 3] The claim that NormFake provides 'consistently strong performance across all evaluated backbones' is not supported by the numbers in Table 3. On GenImage, NormFake (magnitude) achieves 57.3% for BEiT-L, 61.9% for OpenCLIP-G/14, and 73.5% for SigLIP2-Giant, while PE-Core-G14 and DINOv3-7B reach 87.7% and 86.9%. The abstract's statement that the phenomenon holds 'across multiple pretrained models' should therefore be qualified: the effect is strong for some backbones but weak or near-chance for others, and the conditions under which the norm gap appears remain to be characterized. This matters because the paper's framing as an 'implicit' property of foundation models depends on the breadth of the empirical generalization.
minor comments (5)
  1. [Section 3, paragraph after Figure 2] The sentence 'rather the the low-level fingerprints' contains a typo and should read 'rather than the low-level fingerprints.'
  2. [Section 4, 'Evaluation: Video deepfake detection'] The sentence 'This proves that NormFake, despite its simplicity, acts as a strong baseline' is too strong given the single point estimates in Table 1; 'demonstrates' or 'suggests' would be more appropriate.
  3. [Figure 6 caption and axis labels] The caption contains '6,7B' and the horizontal axis labels include spaces such as 'ViT -S'; these formatting issues should be corrected.
  4. [Supplementary material, Section 9 and main text] There are several typos, including 'Evaluted' in the Table 3 caption, 'seperability' in Section 5, and 'supplmentary' in Section 4; a careful proofreading pass is needed.
  5. [Section 6, Limitations] The limitations paragraph on pretraining-data contamination is welcome, but it does not mention the potential sensitivity of NormFake to low-level image quality factors such as compression or blur; adding this to the limitations would align the paper with its own qualitative observations.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: NormFake is a parameter-free statistic of frozen features; the causal-attribution experiment is confounded but not definitionally circular.

full rationale

The detection pipeline is not circular: NormFake's scores (Eqs. 1-2) are direct L1 and L1/L2 statistics of frozen features, with no fitted parameters, calibration stage, or learned classifier, and the ROC-AUC metric is threshold-independent. The central phenomenon (lower-magnitude features for fakes) is an empirical observation supported by histograms and per-layer sweeps across independent benchmarks, not a consequence of the score definition. The main interpretive claim—that the norm gap is driven by semantic shift rather than low-level fingerprints (Sec. 3, Fig. 2)—is underdetermined: EuroSAT and MAVOS-DD Arabic differ from the encoders' pretraining data in low-level characteristics as well as semantics, and the SD1.5 autoencoder reconstruction injects only part of the generative pipeline. However, this is a validity/confound concern, not a circularity: 'semantic shift' is not defined via the norm, and the conclusion is an empirical inference rather than an equation. Self-citations (Smeu et al. 2025; Boldisor et al. 2026) are used only for baselines and preprocessing protocols, not to justify the magnitude/sparsity claim. The practice of reporting the best-performing layer per variant is a test-set selection caveat, but the reported AUCs are measured values, not quantities forced by construction. No definitional reduction or fitted-parameter-renamed-as-prediction was found.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

No fitted weights or calibrated thresholds are introduced; the score is computed from frozen features. The manual choices are the layer index, the score statistic, and the evaluation subsets. The causal interpretation rests on assumptions about familiarity, real-only pretraining, and the isolation of semantic shift from low-level fingerprints, none of which is proven quantitatively.

free parameters (3)
  • Layer index for NormFake variants = magnitude: final output embeddings; sparsity: penultimate transformer block for DINOv3-7B, final output for AV-HuBERT
    The paper reports the best-performing layer for each variant in Table 2 and Figure 8, with the choice made after inspecting test-set ROC-AUC. This is a hand-selected hyperparameter that can inflate reported performance.
  • Score statistic = L1 norm and L1/L2 ratio
    The two statistics are the method itself and were chosen after observing norm separability in Figure 1. The choice is data-driven and not derived from a theory that predicts which statistic should work.
  • Evaluation subset and filtering choices = DFE filtered to 577 videos; MAVOS-DD English-only; DFDC last two partitions (48 and 49)
    Dataset pruning rules follow prior work but are manually selected and affect aggregate scores. They are reasonable, yet they are not justified by a formal criterion.
assumptions (4)
  • domain assumption Foundation model encoders are pretrained on real media only, which makes fake media out-of-distribution.
    Used in Sections 1 and 3 to explain why fakes have lower norms. The paper's own Limitations section notes that future pretraining data contamination would weaken this assumption.
  • domain assumption Feature norm reflects familiarity: higher norms for in-distribution inputs and lower norms for out-of-distribution inputs.
    Borrowed from the familiarity hypothesis of Dietterich and Guyer, cited in Section 3. The paper does not prove this monotonic relationship for self-supervised features.
  • ad hoc to paper The SD1.5 autoencoder reconstruction injects generative fingerprints while preserving semantics, so the norm gap between reconstructed reals and fakes isolates the semantic component.
    This is the crux of the semantic-shift conclusion in Section 3. The protocol does not control for low-level changes introduced by reconstruction, so the isolation claim is not demonstrated.
  • ad hoc to paper EuroSAT and MAVOS-DD Arabic are out-of-distribution only semantically, not in low-level statistics.
    Used in Section 3 to attribute lower norms to semantic shift. Satellite imagery and Arabic speech differ from pretraining data in sensor, resolution, compression, and language, so low-level differences are also present.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Foundation Models are Implicit Deepfake Detectors." pith.science (2026). https://pith.science/paper/JVAG2XFK

@misc{pith2026260809427,
  author       = {Pith},
  title        = {Pith review of: Foundation Models are Implicit Deepfake Detectors},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JVAG2XFK}},
  note         = {Machine review of arXiv:2608.09427}
}
read the original abstract

Pretrained self-supervised representations have emerged as a core component of current deepfake detection methods, yet it remains unclear which of their properties make real and fake media distinguishable. In this work, we uncover a surprisingly consistent phenomenon: across multiple pretrained models, datasets, and both image and video domains, fake samples systematically produce lower-magnitude representations than their real counterparts. Motivated by this finding, we formulate deepfake detection as an anomaly detection problem and show that simple statistics of feature magnitude achieve competitive performance with far more sophisticated deepfake detection methods. We further investigate the origin of this effect and demonstrate that reduced feature magnitude is primarily associated with semantic shifts introduced by fake content, while low-level generative fingerprints play a comparatively smaller role. Finally, we show that this discriminative signal strengthens as the size of the underlying foundation model grows, suggesting that advances in representation learning naturally translate into stronger zero-shot deepfake detectors.

Figures

Figures reproduced from arXiv: 2608.09427 by the authors.

Figure 1
Figure 1. Histograms of ℓ1 feature norm. They are shown for real (blue) and fake (orange) samples extracted from (a) visual-only AV-HuBERT on FakeAVCeleb and (b) DINOv3- 7B on GenImage SD1.5. In both settings, fake samples exhibit features with lower ℓ1 norm. performance by fine-tuning or adapting the pretrained repre￾sentations (Khan and Dang-Nguyen 2024; Park and Owens 2025), or by learning representations on data and prete… view at source ↗
Figure 3
Figure 3. Per-dimension mean absolute activations for (a) AV [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 5
Figure 5. Qualitative examples of DINOv3-7B features with [PITH_FULL_IMAGE:figures/full_fig_p005_5.png] view at source ↗
Figures from the paper (6 more)
Figure 6
Figure 6. Figure 6: ROC-AUC of NormFake (magnitude) using DI [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: Per transformer block and output embeddings ROC [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: Per transformer block and output embeddings ROC [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 9
Figure 9. Figure 9: Pearson correlation between the sample-level pre [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]
Figure 10
Figure 10. Figure 10: Qualitative examples from FakeAVCeleb, ordered by the NormFake scores. The x-axis corresponds to the magnitude [PITH_FULL_IMAGE:figures/full_fig_p015_10.png]
Figure 11
Figure 11. Figure 11: Qualitative examples from GenImage, ordered by the NormFake scores. The x-axis corresponds to the magnitude [PITH_FULL_IMAGE:figures/full_fig_p016_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

93 extracted references · 62 canonical work pages

  1. [1]

    S.; and Zisserman, A

    Afouras, T.; Chung, J. S.; and Zisserman, A. 2018. LRS3-TED: A Large-Scale Dataset for Visual Speech Recognition. arXiv preprint arXiv:1809.00496

  2. [2]

    Baevski, A.; Zhou, Y.; Mohamed, A.; and Auli, M. 2020. Wav2vec w.0: A Framework for Self-Supervised Learning of Speech Representations. In NeurIPS

  3. [3]

    Bao, H.; Dong, L.; Piao, S.; and Wei, F. 2022. BEiT: BERT Pre-Training of Image Transformers. In ICLR

  4. [4]

    Y.; Kassel, L.; and Gilboa, G

    Ben Hayun, O.; Betser, R.; Levi, M. Y.; Kassel, L.; and Gilboa, G. 2026. Training-free detection of generated videos via spatial-temporal likelihoods. In CVPR

  5. [5]

    Boldisor, D.-A.; Smeu, S.; Oneata, D.; and Oneata, E. 2026. Investigating Self-Supervised Representations for Audio-Visual Deepfake Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

  6. [6]

    H.; Madotto, A.; Wei, C.; Ma, T.; Zhi, J.; Rajasegaran, J.; Rasheed, H.; Wang, J.; Monteiro, M.; Xu, H.; Dong, S.; Ravi, N.; Li, D.; Doll \'a r, P.; and Feichtenhofer, C

    Bolya, D.; Huang, P.-Y.; Sun, P.; Cho, J. H.; Madotto, A.; Wei, C.; Ma, T.; Zhi, J.; Rajasegaran, J.; Rasheed, H.; Wang, J.; Monteiro, M.; Xu, H.; Dong, S.; Ravi, N.; Li, D.; Doll \'a r, P.; and Feichtenhofer, C. 2025. Perception Encoder: The best visual embeddings are not at the output of the network. arXiv:2504.13181

  7. [7]

    Brock, A.; Donahue, J.; and Simonyan, K. 2019. Large Scale GAN Training for High Fidelity Natural Image Synthesis. In ICLR

  8. [8]

    Brokman, J.; Giloni, A.; Hofman, O.; Vainshtein, R.; Kojima, H.; and Gilboa, G. 2025. Manifold Induced Biases for Zero-Shot and Few-Shot Detection of Generated Images. In ICLR

Show all 93 references
  1. [9]

    Carvalho, T.; Farid, H.; and Kee, E. R. 2015. Exposing photo manipulation from user-guided 3d lighting analysis. In Media Watermarking, Security, and Forensics, volume 9409, 940902. SPIE

  2. [10]

    Chai, L.; Bau, D.; Lim, S.-N.; and Isola, P. 2020. What makes fake images detectable? U nderstanding properties that generalize. In ECCV

  3. [11]

    A.; Lee, H.; Murtfeldt, R.; Qiu, L.; Karmakar, A.; Tanumihardja, E.; Farhat, K.; Caffee, B.; Lee, C.; Choi, J.; Paik, S.; Kim, A.; and Etzioni, O

    Chandra, N. A.; Lee, H.; Murtfeldt, R.; Qiu, L.; Karmakar, A.; Tanumihardja, E.; Farhat, K.; Caffee, B.; Lee, C.; Choi, J.; Paik, S.; Kim, A.; and Etzioni, O. 2026. Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024. In CVPRW

  4. [12]

    Chen, Y.; Liang, S.; Zhou, Z.; Huang, Z.; Ma, Y.; Tang, J.; Lin, Q.; Zhou, Y.; and Lu, Q. 2025. HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters. arXiv preprint arXiv:2505.20156

  5. [13]

    Chen, Z.; Cao, J.; Chen, Z.; Li, Y.; and Ma, C. 2024. EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions. In AAAI

  6. [14]

    Cherti, M.; Beaumont, R.; Wightman, R.; Wortsman, M.; Ilharco, G.; Gordon, C.; Schuhmann, C.; Schmidt, L.; and Jitsev, J. 2023. Reproducible Scaling Laws for Contrastive Language-Image Learning. In CVPR

  7. [15]

    J.; and Lee, M

    Choi, S.; Lee, H.; Lee, J.; Kim, R.; Choi, S. J.; and Lee, M. 2026. A Debiased Reconstruction-based Framework for Training-Free Detection of AI-Generated Images. In CVPR

  8. [16]

    Choi, S.; Lee, H.; and Lee, M. 2025. Training-free Detection of AI-generated Images via Cropping Robustness. In Advances in Neural Information Processing Systems, volume 38. Curran Associates, Inc. NeurIPS 2025

  9. [17]

    S.; Nagrani, A.; and Zisserman, A

    Chung, J. S.; Nagrani, A.; and Zisserman, A. 2018. VoxCeleb2: Deep Speaker Recognition. In INTERSPEECH

  10. [18]

    A.; Demir, I.; and Yin, L

    Ciftci, U. A.; Demir, I.; and Yin, L. 2020. Fake C atcher: Detection of synthetic portrait videos using biological signals. IEEE Trans. Pattern Anal. Mach. Intell

  11. [19]

    T.; Khan, F

    Croitoru, F.-A.; Hondru, V.; Popescu, M.; Ionescu, R. T.; Khan, F. S.; and Shah, M. 2025. MAVOS-DD : Multilingual Audio-Video Open-Set Deepfake Detection Benchmark. arXiv preprint arXiv:2505.11109

  12. [20]

    R.; G \"u nther, M.; and Boult, T

    Dhamija, A. R.; G \"u nther, M.; and Boult, T. 2018. Reducing network agnostophobia. In NeurIPS

  13. [21]

    Dhariwal, P.; and Nichol, A. 2021. Diffusion Models Beat GANs on Image Synthesis. In NeurIPS

  14. [22]

    G.; and Guyer, A

    Dietterich, T. G.; and Guyer, A. 2022. The familiarity hypothesis: Explaining the behavior of deep open set methods. Pattern Recognition, 132: 108931

  15. [23]

    Dolhansky, B.; Bitton, J.; Pflaum, B.; Lu, J.; Howes, R.; Wang, M.; and Ferrer, C. C. 2020. The DeepFake Detection Challenge (DFDC) Dataset. arXiv:2006.07397

  16. [24]

    Dufour, N.; and Gully, A. 2019. Contributing Data to Deepfake Detection Research. Google AI Blog

  17. [25]

    Feng, C.; Chen, Z.; and Owens, A. 2023. Self-supervised video forensics by audio-visual anomaly detection. In CVPR

  18. [26]

    Frank, J.; Eisenhofer, T.; Sch \"o nherr, L.; Fischer, A.; Kolossa, D.; and Holz, T. 2020. Leveraging frequency analysis for deep fake image recognition. In ICML

  19. [27]

    Gu, S.; Chen, D.; Bao, J.; Wen, F.; Zhang, B.; Chen, D.; Yuan, L.; and Guo, B. 2022. Vector Quantized Diffusion Model for Text-to-Image Synthesis. In CVPR

  20. [28]

    Guo, J.; Zhang, D.; Liu, X.; Zhong, Z.; Zhang, Y.; Wan, P.; and Zhang, D. 2024. LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control. arXiv preprint arXiv:2407.03168

  21. [29]

    Haliassos, A.; Ma, P.; Mira, R.; Petridis, S.; and Pantic, M. 2025. Jointly Learning Visual and Auditory Speech Representations from Raw Data. In ICLR

  22. [30]

    Haliassos, A.; Mira, R.; Petridis, S.; and Pantic, M. 2022. Leveraging real talking faces via self-supervision for robust forgery detection. In CVPR

  23. [31]

    Haliassos, A.; Zinonos, A.; Mira, R.; Petridis, S.; and Pantic, M. 2024. BRAVEn: Improving Self-Supervised Pre-training for Visual and Auditory Speech Recognition. In ICASSP

  24. [32]

    He, Z.; Chen, P.-Y.; and Ho, T.-Y. 2024. RIGID: A Training-free and Model-Agnostic Framework for Robust AI-Generated Image Detection. CoRR, abs/2405.20112

  25. [33]

    Helber, P.; Bischke, B.; Dengel, A.; and Borth, D. 2019. EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 12(7): 2217--2226

  26. [34]

    Hoyer, P. O. 2004. Non-negative Matrix Factorization with Sparseness Constraints. Journal of Machine Learning Research, 5: 1457--1469

  27. [35]

    Huang, D.; and De la Torre, F. 2012. Facial Action Transfer with Personalized Bilinear Regression. In ECCV

  28. [36]

    Huawei Noah's Ark Lab . 2022. Wukong. https://xihe.mindspore.cn/modelzoo/wukong

  29. [37]

    Ilharco, G.; Wortsman, M.; Wightman, R.; Gordon, C.; Carlini, N.; Taori, R.; Dave, A.; Shankar, V.; Namkoong, H.; Miller, J.; Hajishirzi, H.; Farhadi, A.; and Schmidt, L. 2021. OpenCLIP

  30. [38]

    Ji, X.; Hu, X.; Xu, Z.; Zhu, J.; Lin, C.; He, Q.; Zhang, J.; Luo, D.; Chen, Y.; Lin, Q.; Lu, Q.; and Wang, C. 2025. SONIC: Shifting Focus to Global Audio Perception in Portrait Animation. In CVPR

  31. [39]

    J.; Wang, Q.; Shen, J.; Ren, F.; Chen, Z.; Nguyen, P.; Pang, R.; Moreno, I

    Jia, Y.; Zhang, Y.; Weiss, R. J.; Wang, Q.; Shen, J.; Ren, F.; Chen, Z.; Nguyen, P.; Pang, R.; Moreno, I. L.; and Wu, Y. 2018. Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis. In NeurIPS

  32. [40]

    Jiang, L.; Wu, W.; Li, R.-C.; Qian, C.; and Loy, C. C. 2020. DeeperForensics-1.0: A Large-Scale Dataset for Real-World Face Forgery Detection. CVPR

  33. [41]

    Karras, T.; Laine, S.; and Aila, T. 2019. A Style-Based Generator Architecture for Generative Adversarial Networks. In CVPR

  34. [42]

    Khalid, H.; Tariq, S.; Kim, M.; and Woo, S. S. 2021. Fake AVC eleb: A Novel Audio-Video Multimodal Deepfake Dataset. In NeurIPS Datasets and Benchmarks Track

  35. [43]

    A.; and Dang-Nguyen, D.-T

    Khan, S. A.; and Dang-Nguyen, D.-T. 2024. CLIP ping the deception: Adapting vision-language models for universal deepfake detection. In ICMR

  36. [44]

    Kim, T.; Choi, J.; Jeong, Y.; Noh, H.; Yoo, J.; Baek, S.; and Choi, J. 2025. Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video Detection. In ICCV

  37. [45]

    S.; and Noh, J

    Kim, Y.; Yun, K.; Hong, S.; Cha, S.; Koo, C. S.; and Noh, J. 2026. X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection. In CVPR

  38. [46]

    Korshunova, I.; Shi, W.; Dambre, J.; and Theis, L. 2017. Fast Face-Swap Using Convolutional Neural Networks. In ICCV

  39. [47]

    Koutlis, C.; and Papadopoulos, S. 2024. Leveraging representations from intermediate encoder-blocks for synthetic image detection. In ECCV

  40. [48]

    Koutlis, C.; and Papadopoulos, S. 2026. AuViRe: Audio-visual Speech Representation Reconstruction for Deepfake Temporal Localization. In WACV

  41. [49]

    Layton, S.; De Andrade, T.; Olszewski, D.; Warren, K.; Gates, C.; Butler, K.; and Traynor, P. 2025. Every breath you don't take: Deepfake speech detection using breath. Digital Threats: Research and Practice, 6(3): 1--18

  42. [50]

    Li, Y.; Chang, M.-C.; and Lyu, S. 2018. In ictu oculi: Exposing AI created fake videos by detecting eye blinking. In WIFS

  43. [51]

    Li, Y.; Sun, P.; Qi, H.; and Lyu, S. 2020. Celeb-DF: A Large-scale Challenging Dataset for DeepFake Forensics . In CVPR

  44. [52]

    Liang, Y.; Yu, M.; Li, G.; Jiang, J.; Li, B.; Yu, F.; Zhang, N.; Meng, X.; and Huang, W. 2024. Speech F orensics: Audio-Visual Speech Representation Learning for Face Forgery Detection. In NeurIPS

  45. [53]

    Liu, H.; Tan, Z.; Tan, C.; Wei, Y.; Wang, J.; and Zhao, Y. 2024 a . Forgery-aware Adaptive Transformer for Generalizable Synthetic Image Detection. In CVPR

  46. [54]

    Liu, W.; She, T.; Liu, J.; Li, B.; Yao, D.; Liang, Z.; and Wang, R. 2024 b . Lips Are Lying: Spotting the Temporal Inconsistency between Audio and Visual in Lip-Syncing DeepFakes. In NeurIPS

  47. [55]

    Lopes, M. 2013. Estimating unknown sparsity in compressed sensing. In ICML

  48. [56]

    Marra, F.; Gragnaniello, D.; Verdoliva, L.; and Poggi, G. 2019. Do GAN s leave artificial fingerprints? In Multimedia Information Processing and Retrieval

  49. [57]

    Midjourney, Inc. 2022. Midjourney. https://www.midjourney.com/

  50. [58]

    Nichol, A.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; McGrew, B.; Sutskever, I.; and Chen, M. 2021. GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models. In ICML

  51. [59]

    Nirkin, Y.; Keller, Y.; and Hassner, T. 2019. FSGAN: Subject Agnostic Face Swapping and Reenactment. In ICCV

  52. [60]

    Ojha, U.; Li, Y.; and Lee, Y. J. 2023. Towards Universal Fake Image Detectors that Generalize Across Generative Models. In CVPR

  53. [61]

    Oorloff, T.; Koppisetti, S.; Bonettini, N.; Solanki, D.; Colman, B.; Yacoob, Y.; Shahriyari, A.; and Bharaj, G. 2024. AVFF : Audio-Visual Feature Fusion for Video Deepfake Detection. In CVPR

  54. [63]

    Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; et al. 2023. DINO v2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193

  55. [64]

    Park, J.; Chai, J. C. L.; Yoon, J.; and Teoh, A. B. J. 2023. Understanding the Feature Norm for Out-of-Distribution Detection. In ICCV

  56. [65]

    Park, J.; and Owens, A. 2025. Community forensics: Using thousands of generators to train fake image detectors. In CVPR

  57. [66]

    S.; RP, L.; Jiang, J.; et al

    Perov, I.; Gao, D.; Chervoniy, N.; Liu, K.; Marangonda, S.; Um \'e , C.; Dpfks, M.; Facenheim, C. S.; RP, L.; Jiang, J.; et al. 2020. DeepFaceLab: Integrated, Flexible and Extensible Face-Swapping Framework. arXiv preprint arXiv:2005.05535

  58. [67]

    R.; Mukhopadhyay, R.; Namboodiri, V

    Prajwal, K. R.; Mukhopadhyay, R.; Namboodiri, V. P.; and Jawahar, C. 2020. A Lip Sync Expert Is All You Need for Speech to Lip Generation In the Wild. In ACM International Conference on Multimedia

  59. [68]

    W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al

    Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021. Learning transferable visual models from natural language supervision. In ICML

  60. [69]

    Reiss, T.; Cavia, B.; and Hoshen, Y. 2023. Detecting Deepfakes Without Seeing Any. CoRR, abs/2311.01458

  61. [70]

    Ricker, J.; Damm, S.; Holz, T.; and Fischer, A. 2024. Towards the detection of diffusion model deepfakes. In VISAPP

  62. [71]

    Ricker, J.; Lukovnikov, D.; and Fischer, A. 2024. AEROBLADE : Training-Free Detection of Latent Diffusion Images Using Autoencoder Reconstruction Error. In CVPR

  63. [72]

    Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 a . High-Resolution Image Synthesis with Latent Diffusion Models. In CVPR

  64. [73]

    Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 b . High-Resolution Image Synthesis with Latent Diffusion Models. In CVPR

  65. [74]

    R \"o ssler, A.; Cozzolino, D.; Verdoliva, L.; Riess, C.; Thies, J.; and Nie ner, M. 2019. FaceForensics++: Learning to Detect Manipulated Facial Images. ICCV

  66. [75]

    A.; and Bhattad, A

    Sarkar, A.; Mai, H.; Mahapatra, A.; Lazebnik, S.; Forsyth, D. A.; and Bhattad, A. 2024. Shadows don't lie and lines can't bend! G enerative models don't know projective geometry... for now. In CVPR

  67. [76]

    Shi, B.; Hsu, W.; Lakhotia, K.; and Mohamed, A. 2022. Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction. In ICLR

  68. [77]

    V.; Seitzer, M.; Baldassarre, F.; Oquab, M.; Jose, C.; Khalidov, V.; Szafraniec, M.; Yi, S.; Ramamonjisoa, M.; et al

    Sim \'e oni, O.; Vo, H. V.; Seitzer, M.; Baldassarre, F.; Oquab, M.; Jose, C.; Khalidov, V.; Szafraniec, M.; Yi, S.; Ramamonjisoa, M.; et al. 2025. DINO v3. arXiv preprint arXiv:2508.10104

  69. [78]

    Smeu, S.; Boldisor, D.-A.; Oneata, D.; and Oneata, E. 2025. Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning localization. In CVPR

  70. [79]

    Sun, Y.; Guo, C.; and Li, Y. 2021. ReAct : Out-of-distribution Detection With Rectified Activations. In NeurIPS

  71. [80]

    Tan, C.; Zhao, Y.; Wei, S.; Gu, G.; Liu, P.; and Wei, Y. 2024. Frequency-aware deepfake detection: Improving generalizability through frequency space domain learning. In AAAI

  72. [81]

    F.; and Chen, P.-Y

    Tsai, C.-T.; Ko, C.-Y.; Chung, I.-H.; Wang, Y.-C. F.; and Chen, P.-Y. 2024. Understanding and Improving Training-Free AI-Generated Image Detections with Vision Foundation Models. CoRR, abs/2411.19117

  73. [82]

    F.; Alabdulmohsin, I.; Parthasarathy, N.; Evans, T.; Beyer, L.; Xia, Y.; Mustafa, B.; Hénaff, O.; Harmsen, J.; Steiner, A.; and Zhai, X

    Tschannen, M.; Gritsenko, A.; Wang, X.; Naeem, M. F.; Alabdulmohsin, I.; Parthasarathy, N.; Evans, T.; Beyer, L.; Xia, Y.; Mustafa, B.; Hénaff, O.; Harmsen, J.; Steiner, A.; and Zhai, X. 2025. SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding...

  74. [83]

    T.; and Li, H

    Wang, J.; Qian, X.; Zhang, M.; Tan, R. T.; and Li, H. 2023. Seeing What You Said: Talking Face Generation Guided by a Lip Reading Expert. In CVPR

  75. [84]

    Wang, Y.; Chen, X.; Zhu, J.; Chu, W.; Tai, Y.; Wang, C.; Li, J.; Wu, Y.; Huang, F.; and Ji, R. 2021. HifiFace: 3D Shape and Semantic Prior Guided High Fidelity Face Swapping. In IJCAI

  76. [85]

    Wei, H.; Yang, Z.; and Wang, Z. 2024. AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation. arXiv preprint arXiv:2403.17694

  77. [86]

    Yan, S.; Li, O.; Cai, J.; Hao, Y.; Jiang, X.; Hu, Y.; and Xie, W. 2025. A Sanity Check for AI-generated Image Detection. In ICLR

  78. [87]

    Yang, S.; Li, H.; Wu, J.; Jing, M.; Li, L.; Ji, R.; Liang, J.; Fan, H.; and Wang, J. 2025. MegaActor-Sigma: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer. In AAAI

  79. [88]

    Yang, X.; Li, Y.; and Lyu, S. 2019. Exposing deep fakes using inconsistent head poses. In ICASSP

  80. [89]

    Yu, Y.; Shin, S.; Lee, S.; Jun, C.; and Lee, K. 2023. Block Selection Method for Using Feature Norm in Out-of-Distribution Detection. In CVPR

  81. [90]

    Zakharov, E.; Shysheya, A.; Burkov, E.; and Lempitsky, V. 2019. Few-Shot Adversarial Learning of Realistic Neural Talking Head Models. In ICCV

  82. [91]

    Zhang, W.; Cun, X.; Wang, X.; Zhang, Y.; Shen, X.; Guo, Y.; Shan, Y.; and Wang, F. 2023. SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation. In CVPR

  83. [92]

    Zheng, L.; Zhang, Y.; Guo, H.; Pan, J.; Tan, Z.; Lu, J.; Tang, C.; An, B.; and Yan, S. 2026. MEMO : Memory-Guided Diffusion for Expressive Talking Video Generation. TMLR

  84. [93]

    Zhou, Y.; Han, X.; Shechtman, E.; Echevarria, J.; Kalogerakis, E.; and Li, D. 2020. MakeItTalk: Speaker-Aware Talking-Head Animation. ACM Transactions on Graphics

  85. [94]

    Zhu, M.; Chen, H.; Yan, Q.; Huang, X.; Lin, G.; Li, W.; Tu, Z.; Hu, H.; Hu, J.; and Wang, Y. 2023. GenImage: A Million-Scale Benchmark for Detecting AI-Generated Image. In NeurIPS

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.