Pith. sign in

REVIEW 9 cited by

Unified Vision and Language Prompt Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.07225 v1 pith:ZA27ZAYE submitted 2022-10-13 cs.CV cs.AI

classification cs.CVcs.AI
keywords prompttuninglearningvisionvisualbenchmarksmethodsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Prompt tuning, a parameter- and data-efficient transfer learning paradigm that tunes only a small number of parameters in a model's input space, has become a trend in the vision community since the emergence of large vision-language models like CLIP. We present a systematic study on two representative prompt tuning methods, namely text prompt tuning and visual prompt tuning. A major finding is that none of the unimodal prompt tuning methods performs consistently well: text prompt tuning fails on data with high intra-class visual variances while visual prompt tuning cannot handle low inter-class variances. To combine the best from both worlds, we propose a simple approach called Unified Prompt Tuning (UPT), which essentially learns a tiny neural network to jointly optimize prompts across different modalities. Extensive experiments on over 11 vision datasets show that UPT achieves a better trade-off than the unimodal counterparts on few-shot learning benchmarks, as well as on domain generalization benchmarks. Code and models will be released to facilitate future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization

    cs.CV 2026-08 conditional novelty 6.0 of 10

    MuRA improves label-free test-time adaptation of CLIP by routing each image's tokens to a weighted mix of rank-2 through rank-32 LoRA experts at the deepest visual layer.

  2. ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A hierarchical prompt pyramid over CLIP with ancestor-descendant attention improves partially relevant video retrieval.

  3. DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DynImg represents a video snippet as a keyframe plus four resized neighboring frames as temporal prompts, with a 4D rotary position embedding, and reports improved video QA accuracy.

  4. HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    HOLa achieves state-of-the-art zero-shot human-object interaction detection on HICO-DET by low-rank decomposing VLM text features and using LLM-generated action descriptions to regularize weight adaptation.

  5. One Last Attention for Your Vision-Language Model

    cs.CV 2025-07 conditional novelty 6.0 of 10

    RAda attaches one lightweight attention layer to the end of a VLM to learn a mask that reweights the final fused image-text representation, improving fine-tuning across FFT, EFT, and TTT settings.

  6. Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MuGCP adapts CLIP by decoding instance-specific prompts from a frozen MLLM's KV cache and fusing them with visual prompts, achieving state-of-the-art few-shot classification on 14 datasets.

  7. C-GAP: Class-Aware and Online Prompting Improves Vision-Language Models on Imbalanced Classes

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Iterative LLM caption refinement guided by minority-class AP@0.5 lifts rare-object detection on frozen open-vocabulary detectors without labels or weight updates.

  8. MMLoP: Multi-Modal Low-Rank Prompting for Efficient Vision-Language Adaptation

    cs.CV 2026-02 conditional novelty 5.0 of 10

    MMLoP compresses deep multi-modal prompts into a rank-1 shared subspace, reaching a 79.70% base-to-novel harmonic mean with 11.5K trainable parameters.

  9. Generalizing vision-language models to novel domains: A comprehensive survey

    cs.CV 2025-06 conditional novelty 3.0 of 10

    A survey of VLM generalization literature organized by transferred module, with benchmark tables and a review of multimodal LLMs.

Pith tools