Pith. sign in

REVIEW 5 cited by

Raising the Bar of AI-generated Image Detection with CLIP

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.00195 v2 pith:EVAWSR4K submitted 2023-11-30 cs.CV

classification cs.CV
keywords datadetectionai-generatedclipcontrarygeneralizationimagesrobustness
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The aim of this work is to explore the potential of pre-trained vision-language models (VLMs) for universal detection of AI-generated images. We develop a lightweight detection strategy based on CLIP features and study its performance in a wide variety of challenging scenarios. We find that, contrary to previous beliefs, it is neither necessary nor convenient to use a large domain-specific dataset for training. On the contrary, by using only a handful of example images from a single generative model, a CLIP-based detector exhibits surprising generalization ability and high robustness across different architectures, including recent commercial tools such as Dalle-3, Midjourney v5, and Firefly. We match the state-of-the-art (SoTA) on in-distribution data and significantly improve upon it in terms of generalization to out-of-distribution data (+6% AUC) and robustness to impaired/laundered data (+13%). Our project is available at https://grip-unina.github.io/ClipBased-SyntheticImageDetection/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    Introduces a multi-domain benchmark for detecting AI-generated text-rich images from GPT-Image-2 and evaluates five detectors showing domain-dependent performance and JPEG sensitivity.

  2. LEGO: LoRA-Enabled Generator-Oriented Framework for Synthetic Image Detection

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    LEGO uses multiple generator-specific LoRA modules modulated by an MLP and fused with attention to detect synthetic images, achieving better performance than prior methods while using under 10% of the training data.

  3. GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification

    cs.CV 2026-07 conditional novelty 6.0 of 10

    GenSyn10 provides 60k CIFAR-10-aligned images from FLUX.2, HunyuanImage-3.0, and Qwen-Image-2512, showing detectors lose 4–18 points of accuracy on an unseen generator.

  4. VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A 36-model cross-paradigm benchmark on a hard 100-image corpus shows commercial APIs lead on MCC, open-source detectors trail on average, and a subset of strong rankers are miscalibrated at their default threshold.

  5. Rethinking Individual Fairness in Deepfake Detection

    cs.LG 2025-07 conditional novelty 5.0 of 10

    The authors show that a naive individual fairness penalty harms deepfake detection, and propose a semantic-agnostic similarity metric plus anchor learning to reconcile the two.

Pith tools