Pith. sign in

REVIEW 3 cited by

Learning to Prompt for Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.01134 v6 pith:Q36RX6HG submitted 2021-09-02 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords coopmodelspromptcontextlearningvision-languagedownstreamshots
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large pre-trained vision-language models like CLIP have shown great potential in learning representations that are transferable across a wide range of downstream tasks. Different from the traditional representation learning that is based mostly on discretized labels, vision-language pre-training aligns images and texts in a common feature space, which allows zero-shot transfer to a downstream task via prompting, i.e., classification weights are synthesized from natural language describing classes of interest. In this work, we show that a major challenge for deploying such models in practice is prompt engineering, which requires domain expertise and is extremely time-consuming -- one needs to spend a significant amount of time on words tuning since a slight change in wording could have a huge impact on performance. Inspired by recent advances in prompt learning research in natural language processing (NLP), we propose Context Optimization (CoOp), a simple approach specifically for adapting CLIP-like vision-language models for downstream image recognition. Concretely, CoOp models a prompt's context words with learnable vectors while the entire pre-trained parameters are kept fixed. To handle different image recognition tasks, we provide two implementations of CoOp: unified context and class-specific context. Through extensive experiments on 11 datasets, we demonstrate that CoOp requires as few as one or two shots to beat hand-crafted prompts with a decent margin and is able to gain significant improvements over prompt engineering with more shots, e.g., with 16 shots the average gain is around 15% (with the highest reaching over 45%). Despite being a learning-based approach, CoOp achieves superb domain generalization performance compared with the zero-shot model using hand-crafted prompts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 2,607 citations worldwide. Full citation record

  1. $\Delta \mathrm{Energy}$: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization

    cs.CV 2025-10 reject novelty 5.0 of 10

    ΔEnergy, an energy-change OOD score for CLIP, and its EBM fine-tuning loss simultaneously improve OOD detection and covariate-shift generalization.

  2. SynBridge: Bridging Reaction States via Discrete Flow for Bidirectional Reaction Prediction

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A bidirectional discrete flow matching model, SynBridge, predicts reaction products and reactants on graph representations and reports state-of-the-art Top-k accuracy on USPTO-50K, USPTO-MIT, and Pistachio.

  3. StackCLIP: Clustering-Driven Stacked Prompt in Zero-Shot Industrial Anomaly Detection

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Stacking multiple category names in a CLIP text prompt, along with cluster-specific alignment layers, improves zero-shot industrial defect detection and localization.

Pith tools