REVIEW 4 major objections 4 minor 60 references
FauForensics: Boosting Audio-Visual Deepfake Detection with Facial Action Units
T0 review · 4 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper claims that facial action units—quantitative descriptors of facial muscle activity—make audio-visual deepfake detection more accurate and more transferable across datasets than prior clip-level methods.
desk verdict Solid engineering paper with a plausible FAU-guided audio-visual architecture, but the headline cross-dataset SOTA claim is undercut by an under-specified evaluation protocol and a numeric inconsistency. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the facial action unit (FAU), a quantitative descriptor of facial muscle activity such as AU25 for lip parts; the paper treats FAUs as physiologically invariant and therefore harder for generators to fake than raw pixels. The mechanism runs through four stages: FAU-enhanced feature learning fuses features from a frozen FAU encoder with features from a trainable video encoder; implicit feature alignment maps audio and visual latents $\mathbf{Z}_a, \mathbf{Z}_v$ into key/value pairs and attends to them with shared learnable queries $\mathbf{Q} \in \mathbb{R}^{T\times L}$, yielding $\mathbf{Z}_{aq}$ and $\mathbf{Z}_{vq}$; a temporal attentional pooler builds dense matrices $\mathbf{M}_{av}=f_{\mathrm{norm}}(\sigma_{av}*f_{\mathrm{mp}}(\mathbf{Z}_{aq},\mathbf{Z}_{vq}))$, $\mathbf{M}_a$, and $\mathbf{M}_v$; and separate MLPs turn the flattened matrices into audio, visual, and audio-visual forgery scores. The device converts the physiological premise into an architecture: frame-by-frame lip-audio coordination is scored explicitly, which is what the authors claim transfers across datasets and forgery styles.
What would settle it
A decisive experiment is to take the trained model and run the LAV-DF test set with the audio track delayed by about 80–200 milliseconds relative to the video frames. The paper's mechanism predicts a large drop in AUC because frame-wise lip-audio alignment is disrupted; if AUC stays near the reported 99.97, then temporal alignment is not the source of the performance. The same test applied to the cross-dataset setting would show whether the FAU cue alone, without exact alignment, still generalizes.
Extended reading notes
Core claim
The central claim is that forged audio-visual content breaks the temporal correlation intensity of facial action units: across the more than 20,000 samples examined, real videos exhibit significantly higher FAU consistency than fake ones, making FAUs a forgery-resistant representation tied to facial physiology. The architecture combines a frozen FAU encoder with a trainable video encoder, aligns the two modalities behind shared learnable queries in a transformer, and computes dense frame-by-frame attention matrices for audio-audio, video-video, and audio-video consistency. The authors report state-of-the-art results under both binary and four-class settings on FakeAVCeleb and LAV-DF, with the strongest gains under cross-dataset evaluation. Their ablation attributes the gain to the FAU stream: removing the FAU encoder drops cross-dataset AUC from 83.77 to 78.01 when training on LAV-DF and testing on FakeAVCeleb.
Load-bearing premise
The load-bearing premise is that the audio track is produced by the face visible on screen; when an off-screen speaker, dubbing, or background music provides the audio, the lip-audio correspondence this detector is built to measure is absent, so its inputs stop carrying the signal it relies on.
Editorial extensions
If this is right
- The model outputs separate audio-only, visual-only, and audio-visual scores, so a deployed detector can report which stream was forged rather than a single fake/real label.
- Within-database binary performance reaches an average AUC of 99.94 on the two benchmark datasets, matching the strongest prior detector while adding four-class capability.
- Cross-dataset binary AUC improves by an average of 4.83 points over the previous state of the art, and the average four-class AUC improvement is larger, indicating the cue transfers rather than memorizing dataset artifacts.
- Ablation evidence shows the FAU stream contributes roughly 5.8 points of cross-dataset AUC, so the physiological representation is doing the work, not just the multimodal fusion.
- The method keeps high AUC under video/JPEG compression, blur, noise, contrast, and saturation perturbations, supporting its use after realistic post-processing.
Reading between the lines
- The frame-wise attention matrices could be reused as temporal forgery localization heatmaps, a use the paper's design suggests but does not evaluate with localization metrics.
- A stress test the authors did not run is dubbed or re-voiced content: when audio is professionally re-synced to a new speaker, the FAU-alignment cue may weaken, so the claimed superiority over other detectors should be re-measured on such data.
- Since FAUs are defined by facial anatomy, the same inductive bias could be ported to 3D avatars or synthetic characters whose muscle dynamics are approximated rather than physically generated, extending the detector beyond photorealistic video.
- The paper's future direction of phoneme-to-FAU consistency is a natural refinement: aligning frame-level phoneme labels with AU25 motion could localize the exact moment of lip-audio mismatch, which the current dense attention matrix already makes possible.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes FauForensics, an audio-visual deepfake detection framework that combines a frozen facial action unit (FAU) encoder pretrained on DISFA with a trainable video encoder, a query-shared multimodal transformer for implicit feature alignment, and a temporal attentional pooler that computes frame-wise intra- and inter-modality attention matrices. The final predictions are produced by modality-specific MLPs, with the multimodal score used at inference. The authors report within-database and cross-database experiments on FakeAVCeleb and LAV-DF, perturbation robustness studies, ablations of the core modules and encoders, t-SNE visualizations, and a phoneme-FAU case study. The paper claims state-of-the-art performance and a 4.83% average cross-dataset AUC improvement over existing methods.
Significance. If the empirical claims hold, the paper would make a useful contribution by showing that biologically grounded facial action units can serve as a transferable representation for audio-visual deepfake detection and that frame-wise temporal correlation modeling can outperform clip-level approaches. The paper has real strengths: the FAU encoder is pretrained on an external dataset and frozen, so the detector is not circularly fitting the motivating temporal-correlation observation; the ablation study isolates the contributions of the proposed modules; the perturbation robustness experiments cover several realistic distortions; and the limitations paragraph honestly identifies the reliance on aligned audio-visual streams. However, the central cross-dataset generalization claim is currently under-supported because the comparison protocol is not fully specified, there is an unresolved numeric inconsistency in the reported gain, and some load-bearing components of the architecture are left undefined.
major comments (4)
- [III-B and III-D, Eqs. (1), (4)-(6)] The fusion operation f_phi in Eq. (1) and the normalization function f_norm in Eqs. (4)-(6) are never defined, and it is not stated whether the learnable scaling factors sigma_av, sigma_a, and sigma_v are scalars or per-channel vectors. Since the paper attributes the cross-dataset improvement specifically to FAU-enhanced frame-wise audio-visual similarity and temporal attentional pooling, these unspecified operations are load-bearing for the central claim and make the method impossible to reimplement or fully check as written.
- [IV-C / Table III and IV-B] The cross-dataset comparison does not control for the identity overlap between FakeAVCeleb and LAV-DF, which are both constructed from VoxCeleb2-derived talking-head videos, and the manuscript does not state whether FakeAV test identities or raw clips overlap with LAV-DF training data or vice versa. Additionally, the text describes the FGMDF-matched test list and the 4-second clip-length filter only for the within-database confusion matrix in Fig. 4, not for Table III. If shared identities are present, the reported 83.77/95.44 AUCs could reflect memorization rather than FAU-driven generalization. The authors should report the identity-overlap statistics and specify the exact evaluation lists and clip filters used for every method in Table III.
- [IV-C vs. Table III] The text in Section IV-C states that the method outperforms FGMDF by an average binary AUC improvement of 5.74% (90.52% vs. 84.78%), but Table III reports FauForensics's binary average as 89.61 and FGMDF's as 84.78, which is a difference of 4.83 percentage points, and the value 90.52 does not appear in Table III. This unresolved numeric mismatch directly affects the headline 4.83% average improvement claim and must be corrected before the results can be assessed.
- [IV-E, Table IV] The ablation study is reported only for training on LAV-DF and testing on both LAV-DF and FakeAV. The complementary cross-dataset direction, training on FakeAV and testing on LAV-DF, is where Table III shows the largest gain (95.44 AUC), but it is not ablated. Without the complementary direction, it is impossible to attribute the strongest cross-dataset result to the FAU-enhanced feature learning, implicit feature alignment, and temporal attentional pooling modules.
minor comments (4)
- [III-A, Eq. (1)] The notation in Eq. (1) is difficult to parse because the audio branch applies f_q_at to the audio encoder output while the visual branch applies f_q_vt to the fused FAU-video features; a short table or a more explicit composition would improve readability.
- [III-B] The claim that the FAU-enhanced visual features are 'dimensionally aligned with the audio latent features Za in R^{T x L}' is not supported because the dimension L is never defined and the relationship between L and the channel dimensions of the audio, video, and FAU encoders is not specified.
- [IV-B, Fig. 4] The caption of Fig. 4 states that all four detectors are evaluated with FGMDF's test list, but the text does not clarify whether the same test list and the same exclusion of shorter videos are used for the numbers in Tables I and II; this should be stated explicitly.
- [IV-F, Fig. 8] The phoneme symbol '℧' in Fig. 8 is not standard IPA and should be replaced with the intended phoneme symbol, and the sentence describing the delayed mouth closure should be reworded for clarity.
Circularity Check
No significant circularity: the FAU encoder is externally pretrained on DISFA, the detector is trained on labeled target data, and no reported metric reduces to a fitted parameter or to a self-citation.
full rationale
The paper's load-bearing design choices are not circular. The FAU encoder is a frozen feature extractor pretrained on DISFA, an external dataset with facial action unit annotations, and the audio/video encoders and fusion modules are trained with cross-entropy losses on the audio-visual deepfake datasets. The reported cross-dataset AUC values are the result of training on labeled data and evaluating on held-out datasets; no parameter is fitted to the reported AUC and then renamed as a prediction. The temporal-correlation observation in Fig. 1(b) is presented as motivating evidence, not as a training target or as a fitted output, so it does not make the evaluation self-confirming. The equations in Sections III-D and III-E define the model's forward pass and loss; none of them reduces a reported experimental result to an input by construction. The self-citations in the reference list (e.g., prior work by co-authors on generalizable deepfake detection) are contextual related-work citations and are not used to justify the central claim or to import a uniqueness theorem. Concerns about protocol fairness, possible identity overlap between FakeAVCeleb and LAV-DF, and the unresolved numeric mismatch in the cross-dataset section are correctness and benchmarking issues, not circularity. Under the stated criteria, this is a normal non-circular empirical paper, so the appropriate score is 0.
Assumptions & free parameters
free parameters (1)
- Loss weights αa, αv, αav =
0.1, 0.1, 0.8
assumptions (4)
- domain assumption A frozen FAU encoder pre-trained on DISFA provides transferable, domain-invariant descriptors of facial muscle motion that survive across deepfake datasets.
- domain assumption Real videos exhibit stronger temporal correlation of FAUs than fake videos, used to justify the design.
- domain assumption The audio track is synchronized with the on-screen speaker's lips.
- standard math Standard attention and softmax operations work as expected.
Cite this review
Pith. "Pith review of FauForensics: Boosting Audio-Visual Deepfake Detection with Facial Action Units." pith.science (2026). https://pith.science/paper/YCHX2G3I
@misc{pith2026250508294,
author = {Pith},
title = {Pith review of: FauForensics: Boosting Audio-Visual Deepfake Detection with Facial Action Units},
year = {2026},
howpublished = {\url{https://pith.science/paper/YCHX2G3I}},
note = {Machine review of arXiv:2505.08294}
}
read the original abstract
The rapid evolution of generative AI has increased the threat of realistic audio-visual deepfakes, demanding robust detection methods. Existing solutions primarily address unimodal (audio or visual) forgeries but struggle with multimodal manipulations due to inadequate handling of heterogeneous modality features and poor generalization across datasets. To this end, we propose a novel framework called FauForensics by introducing biologically invariant facial action units (FAUs), which is a quantitative descriptor of facial muscle activity linked to emotion physiology. It serves as forgery-resistant representations that reduce domain dependency while capturing subtle dynamics often disrupted in synthetic content. Besides, instead of comparing entire video clips as in prior works, our method computes fine-grained frame-wise audiovisual similarities via a dedicated fusion module augmented with learnable cross-modal queries. It dynamically aligns temporal-spatial lip-audio relationships while mitigating multi-modal feature heterogeneity issues. Experiments on FakeAVCeleb and LAV-DF show state-of-the-art (SOTA) performance and superior cross-dataset generalizability with up to an average of 4.83\% than existing methods.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
J. U. Blasberg, M. Gallistl, M. Degering, F. Baierlein, and V . Engert, “You look stressed: A pilot study on facial action unit activity in the con- text of psychosocial stress,” Comprehensive Psychoneuroendocrinology, vol. 15, p. 100187, 2023
work page 2023
-
[2]
Not made for each other-audio-visual dissonance-based deepfake detection and lo- calization,
K. Chugh, P. Gupta, A. Dhall, and R. Subramanian, “Not made for each other-audio-visual dissonance-based deepfake detection and lo- calization,” in Proceedings of the ACM International Conference on Multimedia, 2020, pp. 439–447
work page 2020
-
[3]
Multimodal forgery detection using ensemble learning,
A. Hashmi, S. A. Shahzad, W. Ahmad, C. W. Lin, Y . Tsao, and H.-M. Wang, “Multimodal forgery detection using ensemble learning,” in Asia- Pacific Signal and Information Processing Association Annual Summit and Conference. IEEE, 2022, pp. 1524–1532
work page 2022
-
[4]
A. Hashmi, S. A. Shahzad, C.-W. Lin, Y . Tsao, and H.-M. Wang, “Avtenet: A human-cognition-inspired audio-visual transformer-based ensemble network for video deepfake detection,” IEEE Transactions on Cognitive and Developmental Systems , 2025
work page 2025
-
[5]
Pvass- mdd: predictive visual-audio alignment self-supervision for multimodal deepfake detection,
Y . Yu, X. Liu, R. Ni, S. Yang, Y . Zhao, and A. C. Kot, “Pvass- mdd: predictive visual-audio alignment self-supervision for multimodal deepfake detection,” IEEE Transactions on Circuits and Systems for Video Technology, 2023
work page 2023
-
[6]
Mcl: multimodal contrastive learning for deepfake detection,
X. Liu, Y . Yu, X. Li, and Y . Zhao, “Mcl: multimodal contrastive learning for deepfake detection,” IEEE Transactions on Circuits and Systems for Video Technology, 2023
work page 2023
-
[7]
Cross- modality and within-modality regularization for audio-visual deepfake detection,
H. Zou, M. Shen, Y . Hu, C. Chen, E. S. Chng, and D. Rajan, “Cross- modality and within-modality regularization for audio-visual deepfake detection,” in IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 2024, pp. 4900–4904
work page 2024
-
[8]
Avff: Audio-visual feature fusion for video deepfake detection,
T. Oorloff, S. Koppisetti, N. Bonettini, D. Solanki, B. Colman, Y . Ya- coob, A. Shahriyari, and G. Bharaj, “Avff: Audio-visual feature fusion for video deepfake detection,” in Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition , 2024, pp. 27 102– 27 112
work page 2024
Show all 60 references
-
[9]
Joint audio-visual attention with con- trastive learning for more general deepfake detection,
Y . Zhang, W. Lin, and J. Xu, “Joint audio-visual attention with con- trastive learning for more general deepfake detection,” ACM Trans- actions on Multimedia Computing, Communications and Applications , vol. 20, no. 5, pp. 1–23, 2024
2024
-
[10]
Fine- grained multimodal deepfake classification via heterogeneous graphs,
Q. Yin, W. Lu, X. Cao, X. Luo, Y . Zhou, and J. Huang, “Fine- grained multimodal deepfake classification via heterogeneous graphs,” International Journal of Computer Vision , pp. 1–15, 2024
2024
-
[11]
Emotions don’t lie: An audio-visual deepfake detection method using affective cues,
T. Mittal, U. Bhattacharya, R. Chandra, A. Bera, and D. Manocha, “Emotions don’t lie: An audio-visual deepfake detection method using affective cues,” in Proceedings of the ACM International Conference on Multimedia, 2020, pp. 2823–2832
2020
-
[12]
Face x-ray for more general face forgery detection,
L. Li, J. Bao, T. Zhang, H. Yang, D. Chen, F. Wen, and B. Guo, “Face x-ray for more general face forgery detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 5001–5010
2020
-
[13]
Detect and locate: Exposing face manipulation by semantic-and noise-level telltales,
C. Kong, B. Chen, H. Li, S. Wang, A. Rocha, and S. Kwong, “Detect and locate: Exposing face manipulation by semantic-and noise-level telltales,” IEEE Transactions on Information Forensics and Security , vol. 17, pp. 1741–1756, 2022
2022
-
[14]
Lisiam: Localization invariance siamese network for deepfake detection,
J. Wang, Y . Sun, and J. Tang, “Lisiam: Localization invariance siamese network for deepfake detection,” IEEE Transactions on Information Forensics and Security, vol. 17, pp. 2425–2436, 2022
2022
-
[15]
Masked relation learning for deepfake detection,
Z. Yang, J. Liang, Y . Xu, X.-Y . Zhang, and R. He, “Masked relation learning for deepfake detection,” IEEE Transactions on Information Forensics and Security, vol. 18, pp. 1696–1708, 2023
2023
-
[16]
Ucf: Uncovering common features for generalizable deepfake detection,
Z. Yan, Y . Zhang, Y . Fan, and B. Wu, “Ucf: Uncovering common features for generalizable deepfake detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 22 412–22 423
2023
-
[17]
Robust ai- synthesized speech detection using feature decomposition learning and synthesizer feature augmentation,
K. Zhang, Z. Hua, Y . Zhang, Y . Guo, and T. Xiang, “Robust ai- synthesized speech detection using feature decomposition learning and synthesizer feature augmentation,” IEEE Transactions on Information Forensics and Security, 2024
2024
-
[18]
Preventing deepfake attacks on speaker authentication by dynamic lip movement analysis,
C.-Z. Yang, J. Ma, S. Wang, and A. W.-C. Liew, “Preventing deepfake attacks on speaker authentication by dynamic lip movement analysis,” IEEE Transactions on Information Forensics and Security , vol. 16, pp. 1841–1854, 2020
2020
-
[19]
Diff2lip: Audio conditioned diffusion models for lip-synchronization,
S. Mukhopadhyay, S. Suri, R. T. Gadde, and A. Shrivastava, “Diff2lip: Audio conditioned diffusion models for lip-synchronization,” in Pro- ceedings of the IEEE/CVF Winter Conference on Applications of Com- puter Vision, 2024, pp. 5292–5302
2024
-
[20]
Latentsync: Taming audio-conditioned latent diffusion models for lip sync with syncnet supervision,
C. Li, C. Zhang, W. Xu, J. Lin, J. Xie, W. Feng, B. Peng, C. Chen, and W. Xing, “Latentsync: Taming audio-conditioned latent diffusion models for lip sync with syncnet supervision,” arXiv preprint arXiv:2412.09262, 2024
2024 arXiv
-
[21]
Loopy: Taming audio-driven portrait avatar with long-term motion dependency,
J. Jiang, C. Liang, J. Yang, G. Lin, T. Zhong, and Y . Zheng, “Loopy: Taming audio-driven portrait avatar with long-term motion dependency,” in The International Conference on Learning Representations , 2025
2025
-
[22]
Joint audio-visual deepfake detection,
Y . Zhou and S.-N. Lim, “Joint audio-visual deepfake detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14 800–14 809
2021
-
[23]
Lip sync matters: A novel multimodal forgery detector,
S. A. Shahzad, A. Hashmi, S. Khan, Y .-T. Peng, Y . Tsao, and H.-M. Wang, “Lip sync matters: A novel multimodal forgery detector,” in Asia- Pacific Signal and Information Processing Association Annual Summit and Conference. IEEE, 2022, pp. 1885–1892
2022
-
[24]
Fakeout: Leveraging out-of-domain self- supervision for multi-modal video deepfake detection,
G. Knafo and O. Fried, “Fakeout: Leveraging out-of-domain self- supervision for multi-modal video deepfake detection,” arXiv preprint arXiv:2212.00773, 2022
2022 arXiv
-
[25]
Audio-visual person-of-interest deepfake detection,
D. Cozzolino, A. Pianese, M. Nießner, and L. Verdoliva, “Audio-visual person-of-interest deepfake detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 943– 952
2023
-
[26]
Avfakenet: A unified end-to-end dense swin transformer deep learning model for audio–visual deepfakes detection,
H. Ilyas, A. Javed, and K. M. Malik, “Avfakenet: A unified end-to-end dense swin transformer deep learning model for audio–visual deepfakes detection,” Applied Soft Computing , vol. 136, p. 110124, 2023
2023
-
[27]
Multimodaltrace: Deepfake detec- tion using audiovisual representation learning,
M. A. Raza and K. M. Malik, “Multimodaltrace: Deepfake detec- tion using audiovisual representation learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June 2023, pp. 993–1000
2023
-
[28]
Detecting deep- fake videos from phoneme-viseme mismatches,
S. Agarwal, H. Farid, O. Fried, and M. Agrawala, “Detecting deep- fake videos from phoneme-viseme mismatches,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, pp. 660–661
2020
-
[29]
Leveraging real talking faces via self-supervision for robust forgery detection,
A. Haliassos, R. Mira, S. Petridis, and M. Pantic, “Leveraging real talking faces via self-supervision for robust forgery detection,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 14 950–14 962
2022
-
[30]
Audio-visual contrastive pre-train for face forgery detection,
H. Zhao, W. Zhou, D. Chen, W. Zhang, Y . Guo, Z. Cheng, P. Yan, and N. Yu, “Audio-visual contrastive pre-train for face forgery detection,” ACM Transactions on Multimedia Computing, Communications and Applications, 2024
2024
-
[31]
Lips are lying: Spotting the temporal inconsistency between audio and visual in lip- syncing deepfakes,
W. Liu, T. She, J. Liu, B. Li, D. Yao, and R. Wang, “Lips are lying: Spotting the temporal inconsistency between audio and visual in lip- syncing deepfakes,” in Advances in Neural Information Processing Systems, vol. 37, 2024
2024
-
[32]
Where deepfakes gaze at? spatial-temporal gaze inconsistency analysis for video face forgery detection,
C. Peng, Z. Miao, D. Liu, N. Wang, R. Hu, and X. Gao, “Where deepfakes gaze at? spatial-temporal gaze inconsistency analysis for video face forgery detection,” IEEE Transactions on Information Forensics and Security, 2024
2024
-
[33]
Protecting world leader using facial speaking pattern against deepfakes,
B. Chu, W. You, Z. Yang, L. Zhou, and R. Wang, “Protecting world leader using facial speaking pattern against deepfakes,” IEEE Signal Processing Letters, vol. 29, pp. 2078–2082, 2022
2022
-
[34]
Aunet: Learning relations between action units for face forgery detection,
W. Bai, Y . Liu, Z. Zhang, B. Li, and W. Hu, “Aunet: Learning relations between action units for face forgery detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 24 709–24 719
2023
-
[35]
Listen to your face: Inferring facial action units from audio channel,
Z. Meng, S. Han, and Y . Tong, “Listen to your face: Inferring facial action units from audio channel,” IEEE Transactions on Affective Computing, vol. 10, no. 4, pp. 537–551, 2017
2017
-
[36]
Lips don’t lie: A generalisable and robust approach to face forgery detection,
A. Haliassos, K. V ougioukas, S. Petridis, and M. Pantic, “Lips don’t lie: A generalisable and robust approach to face forgery detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 5039–5049
2021
-
[37]
Spatiotemporal inconsistency learning for deepfake video detection,
Z. Gu, Y . Chen, T. Yao, S. Ding, J. Li, F. Huang, and L. Ma, “Spatiotemporal inconsistency learning for deepfake video detection,” in Proceedings of the ACM International Conference on Multimedia , 2021, pp. 3473–3481
2021
-
[38]
Istvt: interpretable spatial-temporal video transformer for deepfake detection,
C. Zhao, C. Wang, G. Hu, H. Chen, C. Liu, and J. Tang, “Istvt: interpretable spatial-temporal video transformer for deepfake detection,” IEEE Transactions on Information Forensics and Security , vol. 18, pp. 1335–1348, 2023
2023
-
[39]
Mintime: multi-identity size-invariant video deepfake detection,
D. A. Coccomini, G. K. Zilos, G. Amato, R. Caldelli, F. Falchi, S. Pa- padopoulos, and C. Gennaro, “Mintime: multi-identity size-invariant video deepfake detection,” IEEE Transactions on Information Forensics and Security, 2024
2024
-
[40]
Deepfake video detection using audio-visual consistency,
Y . Gu, X. Zhao, C. Gong, and X. Yi, “Deepfake video detection using audio-visual consistency,” in International Workshop on Digital Watermarking. Springer, 2021, pp. 168–180. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 12
2021
-
[41]
Self-supervised video forensics by audio-visual anomaly detection,
C. Feng, Z. Chen, and A. Owens, “Self-supervised video forensics by audio-visual anomaly detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 10 491–10 503
2023
-
[42]
V oice- face homogeneity tells deepfake,
H. Cheng, Y . Guo, T. Wang, Q. Li, X. Chang, and L. Nie, “V oice- face homogeneity tells deepfake,” ACM Transactions on Multimedia Computing, Communications and Applications , vol. 20, no. 3, pp. 1– 22, 2023
2023
-
[43]
Lost in translation: Lip-sync deepfake detection from audio-video mismatch,
M. Bohacek and H. Farid, “Lost in translation: Lip-sync deepfake detection from audio-video mismatch,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 4315–4323
2024
-
[44]
Avoid-df: Audio-visual joint learning for detecting deepfake,
W. Yang, X. Zhou, Z. Chen, B. Guo, Z. Ba, Z. Xia, X. Cao, and K. Ren, “Avoid-df: Audio-visual joint learning for detecting deepfake,” IEEE Transactions on Information Forensics and Security , vol. 18, pp. 2015– 2029, 2023
2015
-
[45]
Speechforensics: Audio-visual speech representation learning for face forgery detection,
Y . Liang, M. Yu, G. Li, J. Jiang, B. Li, F. Yu, N. Zhang, X. Meng, and W. Huang, “Speechforensics: Audio-visual speech representation learning for face forgery detection,” in Advances in Neural Information Processing Systems, 2024
2024
-
[46]
Glcf: A global- local multimodal coherence analysis framework for talking face gener- ation detection,
X. Chen, Q. Yin, J. Liu, W. Lu, X. Luo, and J. Zhou, “Glcf: A global- local multimodal coherence analysis framework for talking face gener- ation detection,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 1, 2025, pp. 75–83
2025
-
[47]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
-
[48]
Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,” in International Conference on Machine Learning . PMLR, 2023, pp. 19 730–19 742
2023
-
[49]
Video classification with channel-separated convolutional networks,
D. Tran, H. Wang, L. Torresani, and M. Feiszli, “Video classification with channel-separated convolutional networks,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 5552–5561
2019
-
[50]
Learning multi- dimensional edge feature-based AU relation graph for facial action unit recognition,
C. Luo, S. Song, W. Xie, L. Shen, and H. Gunes, “Learning multi- dimensional edge feature-based AU relation graph for facial action unit recognition,” in International Joint Conference on Artificial Intelligence, 2022, pp. 1239–1246
2022
-
[51]
Robust speech recognition via large-scale weak super- vision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak super- vision,” in International Conference on Machine Learning . PMLR, 2023, pp. 28 492–28 518
2023
-
[52]
Disfa: A spontaneous facial action intensity database,
S. M. Mavadati, M. H. Mahoor, K. Bartlett, P. Trinh, and J. F. Cohn, “Disfa: A spontaneous facial action intensity database,” IEEE Transactions on Affective Computing , vol. 4, no. 2, pp. 151–160, 2013
2013
-
[53]
Do you really mean that? content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization,
Z. Cai, K. Stefanov, A. Dhall, and M. Hayat, “Do you really mean that? content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization,” in International Conference on Digital Image Computing: Techniques and Applications. IEEE, 2022, pp. 1–10
2022
-
[54]
Fakeavceleb: A novel audio-video multimodal deepfake dataset,
H. Khalid, S. Tariq, M. Kim, and S. S. Woo, “Fakeavceleb: A novel audio-video multimodal deepfake dataset,” arXiv preprint arXiv:2108.05080, 2021
2021 arXiv
-
[55]
End-to-end spectro-temporal graph attention networks for speaker ver- ification anti-spoofing and speech deepfake detection,
H. Tak, J.-W. Jung, J. Patino, M. Kamble, M. Todisco, and N. Evans, “End-to-end spectro-temporal graph attention networks for speaker ver- ification anti-spoofing and speech deepfake detection,” in ASVSPOOF , Automatic Speaker Verification and Spoofing Countermeasures Chal- le...
2021
-
[56]
Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation,
H. Tak, M. Todisco, X. Wang, J.-w. Jung, J. Yamagishi, and N. Evans, “Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation,” in The Speaker and Language Recognition Workshop. ISCA, 2022
2022
-
[57]
Hearing and seeing abnormal- ity: Self-supervised audio-visual mutual learning for deepfake detection,
C.-S. Sung, J.-C. Chen, and C.-S. Chen, “Hearing and seeing abnormal- ity: Self-supervised audio-visual mutual learning for deepfake detection,” in IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 2023, pp. 1–5
2023
-
[58]
Adamp: Slowing down the slowdown for momentum optimizers on scale-invariant weights,
B. Heo, S. Chun, S. J. Oh, D. Han, S. Yun, G. Kim, Y . Uh, and J.-W. Ha, “Adamp: Slowing down the slowdown for momentum optimizers on scale-invariant weights,” arXiv preprint arXiv:2006.08217 , 2020
2006 arXiv
-
[59]
Deeperforensics- 1.0: A large-scale dataset for real-world face forgery detection,
L. Jiang, R. Li, W. Wu, C. Qian, and C. C. Loy, “Deeperforensics- 1.0: A large-scale dataset for real-world face forgery detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 2889–2898
2020
-
[60]
Visualizing data using t-sne
L. Van der Maaten and G. Hinton, “Visualizing data using t-sne.” Journal of Machine Learning Research , vol. 9, no. 11, 2008. Jian Wang received the Ph.D. degree at Intelligent Media Analysis Group (IMAG), Nanjing University of Science and Technology, China. He is currently a ...
2008
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.