Pith. sign in

REVIEW 1 cited by

Facing the Elephant in the Room: Visual Prompt Tuning or Full Finetuning?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.12902 v1 pith:JVBFSJE7 submitted 2024-01-23 cs.CV

classification cs.CV
keywords whendatadistributionsobjectivesoriginalprompttasktasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As the scale of vision models continues to grow, the emergence of Visual Prompt Tuning (VPT) as a parameter-efficient transfer learning technique has gained attention due to its superior performance compared to traditional full-finetuning. However, the conditions favoring VPT (the ``when") and the underlying rationale (the ``why") remain unclear. In this paper, we conduct a comprehensive analysis across 19 distinct datasets and tasks. To understand the ``when" aspect, we identify the scenarios where VPT proves favorable by two dimensions: task objectives and data distributions. We find that VPT is preferrable when there is 1) a substantial disparity between the original and the downstream task objectives (e.g., transitioning from classification to counting), or 2) a similarity in data distributions between the two tasks (e.g., both involve natural images). In exploring the ``why" dimension, our results indicate VPT's success cannot be attributed solely to overfitting and optimization considerations. The unique way VPT preserves original features and adds parameters appears to be a pivotal factor. Our study provides insights into VPT's mechanisms, and offers guidance for its optimal utilization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Multi-modal Long Context Learning for Training-free Adaptation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    EMLoC prunes and compresses demonstration examples in multimodal long contexts, reducing inference cost up to 77% while matching or slightly beating full-context accuracy.

Pith tools