Pith. sign in

REVIEW 3 cited by

Concept-Guided Prompt Learning for Generalization in Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.07457 v1 pith:BAFCOJBB submitted 2024-01-15 cs.CV

classification cs.CV
keywords concept-guidedfeaturesgeneralizationpromptvisualcliplearningmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contrastive Language-Image Pretraining (CLIP) model has exhibited remarkable efficacy in establishing cross-modal connections between texts and images, yielding impressive performance across a broad spectrum of downstream applications through fine-tuning. However, for generalization tasks, the current fine-tuning methods for CLIP, such as CoOp and CoCoOp, demonstrate relatively low performance on some fine-grained datasets. We recognize the underlying reason is that these previous methods only projected global features into the prompt, neglecting the various visual concepts, such as colors, shapes, and sizes, which are naturally transferable across domains and play a crucial role in generalization tasks. To address this issue, in this work, we propose Concept-Guided Prompt Learning (CPL) for vision-language models. Specifically, we leverage the well-learned knowledge of CLIP to create a visual concept cache to enable concept-guided prompting. In order to refine the text features, we further develop a projector that transforms multi-level visual features into text features. We observe that this concept-guided prompt learning approach is able to achieve enhanced consistency between visual and linguistic modalities. Extensive experimental results demonstrate that our CPL method significantly improves generalization capabilities compared to the current state-of-the-art methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Cache Enhanced Prototype Learning for Test-Time Generalization of Vision-Language Models

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    The submitted full text does not match the abstract, so the manuscript cannot be assessed as a coherent preprint.

  2. Text and Image Are Mutually Beneficial: Enhancing Training-Free Few-Shot Classification with CLIP

    cs.CV 2024-12 conditional novelty 5.0 of 10

    TIMO improves training-free CLIP few-shot classification by mutually guiding text and image features, and a tuned variant TIMO-S reports state-of-the-art accuracy with roughly 100x less time than training-required methods.

  3. SenCLIP: Enhancing zero-shot land-use mapping for Sentinel-2 with ground-level prompting

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Matching Sentinel-2 images to co-located ground-level photos lets a CLIP model do zero-shot land-use mapping from free-form aerial and ground-view text prompts.

Pith tools