REVIEW 7 cited by
PLOT: Prompt Learning with Optimal Transport for Vision-Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the increasing attention to large vision-language models such as CLIP, there has been a significant amount of effort dedicated to building efficient prompts. Unlike conventional methods of only learning one single prompt, we propose to learn multiple comprehensive prompts to describe diverse characteristics of categories such as intrinsic attributes or extrinsic contexts. However, directly matching each prompt to the same visual feature is problematic, as it pushes the prompts to converge to one point. To solve this problem, we propose to apply optimal transport to match the vision and text modalities. Specifically, we first model images and the categories with visual and textual feature sets. Then, we apply a two-stage optimization strategy to learn the prompts. In the inner loop, we optimize the optimal transport distance to align visual features and prompts by the Sinkhorn algorithm, while in the outer loop, we learn the prompts by this distance from the supervised data. Extensive experiments are conducted on the few-shot recognition task and the improvement demonstrates the superiority of our method. The code is available at https://github.com/CHENGY12/PLOT.
Forward citations
Cited by 7 Pith papers
-
Improving Personalized Search with Regularized Low-Rank Parameter Updates
Regularized rank-one LoRA updates to the final value transform of CLIP's text encoder beat textual inversion for personalized retrieval while preserving general knowledge.
-
Few-Shot Adaptation Benchmark for Remote Sensing Vision-Language Models
A first benchmark of few-shot adaptation for remote-sensing VLMs shows zero-shot accuracy does not predict few-shot gains, with no single method dominating.
-
MADPOT: Medical Anomaly Detection with CLIP Adaptation and Partial Optimal Transport
MADPOT adapts CLIP with learnable prompts, partial optimal transport, and contrastive learning, reporting state-of-the-art anomaly detection AUC on the BMAD benchmark in few-shot and zero-shot settings.
-
Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation
Mettle distills frozen transformer layer features into compact meta-tokens via parallel cross-attention and linear projection, cutting training memory dramatically while retaining competitive accuracy on three audio-v...
-
OV-COAST: Cost Aggregation with Optimal Transport for Open-Vocabulary Semantic Segmentation
Applying Sinkhorn optimal transport to the CAT-Seg cost volume yields a small mIoU improvement on the MESS benchmark, but the training mechanism is under-specified.
-
DiSa: Directional Saliency-Aware Prompt Learning for Generalizable Vision-Language Models
DiSa improves CLIP prompt-learning generalization by combining saliency-guided masking with cross-modal KL regularization and directional class-prototype alignment.
-
Learning Clustering-based Prototypes for Compositional Zero-shot Learning
ClusPro improves compositional zero-shot learning by representing each primitive with multiple online-clustered prototypes and adding prototype-anchored contrastive and decorrelation losses.
Discussion (0). Continue with ORCID to comment.