Pith. sign in

REVIEW 3 cited by

Towards General Visual-Linguistic Face Forgery Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.16545 v2 pith:6GMLBVKK submitted 2023-07-31 cs.CV

classification cs.CV
keywords detectionforgeryfacefine-grainedmethodvlffddatadeepfakes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deepfakes are realistic face manipulations that can pose serious threats to security, privacy, and trust. Existing methods mostly treat this task as binary classification, which uses digital labels or mask signals to train the detection model. We argue that such supervisions lack semantic information and interpretability. To address this issues, in this paper, we propose a novel paradigm named Visual-Linguistic Face Forgery Detection(VLFFD), which uses fine-grained sentence-level prompts as the annotation. Since text annotations are not available in current deepfakes datasets, VLFFD first generates the mixed forgery image with corresponding fine-grained prompts via Prompt Forgery Image Generator (PFIG). Then, the fine-grained mixed data and coarse-grained original data and is jointly trained with the Coarse-and-Fine Co-training framework (C2F), enabling the model to gain more generalization and interpretability. The experiments show the proposed method improves the existing detection models on several challenging benchmarks. Furthermore, we have integrated our method with multimodal large models, achieving noteworthy results that demonstrate the potential of our approach. This integration not only enhances the performance of our VLFFD paradigm but also underscores the versatility and adaptability of our method when combined with advanced multimodal technologies, highlighting its potential in tackling the evolving challenges of deepfake detection.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A frozen-CLIP plugin with bidirectional visual-text fusion reaches 90.07% average AUC on seven unseen deepfake benchmarks, up 6.68 points over prior work.

  2. MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM

    cs.CV 2025-07 reject novelty 5.0 of 10

    MGFFD-VLM combines quality-aware LoRA experts, multi-granularity prompts, and a three-stage training plan to improve explainable deepfake detection on the extended DD-VQA+ dataset.

  3. Visual Language Models as Zero-Shot Deepfake Detectors

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Zero-shot VLMs scored by normalized yes/no token probabilities beat most trained deepfake detectors on a new SimSwap dataset, and a lightly fine-tuned InstructBLIP is near-perfect on DFDC-P.

Pith tools