Pith. sign in

REVIEW 3 cited by

AutoVP: An Automated Visual Prompting Framework and Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.08381 v2 pith:7TXEKHHQ submitted 2023-10-12 cs.CV cs.LG

classification cs.CVcs.LG
keywords autovpbenchmarkdesignchoicesdownstreamframeworkimage-classificationincluding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual prompting (VP) is an emerging parameter-efficient fine-tuning approach to adapting pre-trained vision models to solve various downstream image-classification tasks. However, there has hitherto been little systematic study of the design space of VP and no clear benchmark for evaluating its performance. To bridge this gap, we propose AutoVP, an end-to-end expandable framework for automating VP design choices, along with 12 downstream image-classification tasks that can serve as a holistic VP-performance benchmark. Our design space covers 1) the joint optimization of the prompts; 2) the selection of pre-trained models, including image classifiers and text-image encoders; and 3) model output mapping strategies, including nonparametric and trainable label mapping. Our extensive experimental results show that AutoVP outperforms the best-known current VP methods by a substantial margin, having up to 6.7% improvement in accuracy; and attains a maximum performance increase of 27.5% compared to linear-probing (LP) baseline. AutoVP thus makes a two-fold contribution: serving both as an efficient tool for hyperparameter tuning on VP design choices, and as a comprehensive benchmark that can reasonably be expected to accelerate VP's development. The source code is available at https://github.com/IBM/AutoVP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models

    cs.CV 2025-05 conditional novelty 7.0 of 10

    DINO-R1 trains visual-prompt detectors with group-relative query rewards and KL regularization, improving zero-shot and fine-tuned detection over supervised fine-tuning.

  2. Parameter-Efficient Adaptation of mPLUG-Owl2 via Pixel-Level Visual Prompts for NR-IQA

    cs.CV 2025-09 conditional novelty 5.0 of 10

    With a learned 30-pixel border prompt added to input images, a frozen mPLUG-Owl2-7B reaches 0.932 SRCC on KADID-10k using about 156K trainable parameters.

  3. Model Reprogramming Demystified: A Neural Tangent Kernel Perspective

    cs.LG 2025-05 reject novelty 5.0 of 10

    The paper claims the minimum eigenvalue of the source model's NTK matrix controls both source and reprogrammed target model performance.

Pith tools