Pith. sign in

REVIEW 2 cited by

AutoCLIP: Auto-tuning Zero-Shot Classifiers for Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.16414 v3 pith:Z72L7Q26 submitted 2023-09-28 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords classautocliptemplatesclassifiersdescriptorszero-shotimagemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Classifiers built upon vision-language models such as CLIP have shown remarkable zero-shot performance across a broad range of image classification tasks. Prior work has studied different ways of automatically creating descriptor sets for every class based on prompt templates, ranging from manually engineered templates over templates obtained from a large language model to templates built from random words and characters. Up until now, deriving zero-shot classifiers from the respective encoded class descriptors has remained nearly unchanged, i.e., classify to the class that maximizes cosine similarity between its averaged encoded class descriptors and the image encoding. However, weighing all class descriptors equally can be suboptimal when certain descriptors match visual clues on a given image better than others. In this work, we propose AutoCLIP, a method for auto-tuning zero-shot classifiers. AutoCLIP tunes per-image weights to each prompt template at inference time, based on statistics of class descriptor-image similarities. AutoCLIP is fully unsupervised, has only a minor additional computation overhead, and can be easily implemented in few lines of code. We show that AutoCLIP outperforms baselines across a broad range of vision-language models, datasets, and prompt templates consistently and by up to 3 percent point accuracy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    LoRA-TTT improves CLIP's zero-shot accuracy under distribution shift by test-time training only low-rank adapters in the image encoder, using entropy and masked-class-token consistency losses.

  2. Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models

    cs.CV 2025-05 conditional novelty 5.0 of 10

    AHNPL improves compositional reasoning in VLMs by generating image-side hard negatives from text embedding shifts and adding adaptive contrastive losses.

Pith tools