Pith. sign in

REVIEW 3 cited by

Visual Prompt Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.12119 v2 pith:LSJXXDE7 submitted 2022-03-23 cs.CV

classification cs.CV
keywords tuningfine-tuningfullmodelmodelsparametersbackboneefficient
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The current modus operandi in adapting pre-trained models involves updating all the backbone parameters, ie, full fine-tuning. This paper introduces Visual Prompt Tuning (VPT) as an efficient and effective alternative to full fine-tuning for large-scale Transformer models in vision. Taking inspiration from recent advances in efficiently tuning large language models, VPT introduces only a small amount (less than 1% of model parameters) of trainable parameters in the input space while keeping the model backbone frozen. Via extensive experiments on a wide variety of downstream recognition tasks, we show that VPT achieves significant performance gains compared to other parameter efficient tuning protocols. Most importantly, VPT even outperforms full fine-tuning in many cases across model capacities and training data scales, while reducing per-task storage cost.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Yourself: Class-Incremental Semantic Segmentation with Language-Inspired Bootstrapped Disentanglement

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A language-guided bootstrapped disentanglement framework reduces class and background entanglement in continual semantic segmentation, improving state-of-the-art on VOC and ADE20k.

  2. Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking Design

    cs.CR 2025-08 conditional novelty 6.0 of 10

    Water4MU tunes an invisible watermark on data so that machine unlearning algorithms can remove requested images more effectively, beating prior methods on 'challenging forgets'.

  3. LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation

    cs.CV 2025-02 conditional novelty 5.0 of 10

    LoR-VP adapts frozen vision models by adding a rank-4 low-rank prompt across the full image, outperforming prior visual prompting methods while using far fewer prompt parameters.

Pith tools