Pith. sign in

REVIEW 1 cited by

Face Recognition in the age of CLIP & Billion image datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.07315 v1 pith:PVRCO36U submitted 2023-01-18 cs.CV

classification cs.CV
keywords clipmodelsfacerecognitionopenaiperformancetasksdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

CLIP (Contrastive Language-Image Pre-training) models developed by OpenAI have achieved outstanding results on various image recognition and retrieval tasks, displaying strong zero-shot performance. This means that they are able to perform effectively on tasks for which they have not been explicitly trained. Inspired by the success of OpenAI CLIP, a new publicly available dataset called LAION-5B was collected which resulted in the development of open ViT-H/14, ViT-G/14 models that outperform the OpenAI L/14 model. The LAION-5B dataset also released an approximate nearest neighbor index, with a web interface for search & subset creation. In this paper, we evaluate the performance of various CLIP models as zero-shot face recognizers. Our findings show that CLIP models perform well on face recognition tasks, but increasing the size of the CLIP model does not necessarily lead to improved accuracy. Additionally, we investigate the robustness of CLIP models against data poisoning attacks by testing their performance on poisoned data. Through this analysis, we aim to understand the potential consequences and misuse of search engines built using CLIP models, which could potentially function as unintentional face recognition engines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Foundation versus Domain-specific Models: Performance Comparison, Fusion, and Explainability in Face Recognition

    cs.CV 2025-07 conditional novelty 5.0 of 10

    On standard face benchmarks, domain-specific face recognition models beat zero-shot foundation models, adding context or fusing scores improves performance at low false-match rates, and GPT-4o can explain and sometime...

Pith tools