REVIEW 25 cited by
Exploring Visual Prompts for Adapting Large-Scale Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We investigate the efficacy of visual prompting to adapt large-scale models in vision. Following the recent approach from prompt tuning and adversarial reprogramming, we learn a single image perturbation such that a frozen model prompted with this perturbation performs a new task. Through comprehensive experiments, we demonstrate that visual prompting is particularly effective for CLIP and robust to distribution shift, achieving performance competitive with standard linear probes. We further analyze properties of the downstream dataset, prompt design, and output transformation in regard to adaptation performance. The surprising effectiveness of visual prompting provides a new perspective on adapting pre-trained models in vision. Code is available at http://hjbahng.github.io/visual_prompting .
Forward citations
Cited by 25 Pith papers
-
Visual Textualization for Image Prompted Object Detection
Visual textualization projects support images into the text feature space and prompts an unmodified OVLM, achieving strong few-shot and open-set detection results.
-
Visual prompt engineering for video models
Automatically converting task images to photorealistic variants (visual prompt engineering) improves video-model reasoning performance, often beating text prompt engineering and test-time scaling.
-
AttriPrompt: Dynamic Prompt Composition Learning for CLIP
AttriPrompt dynamically composes text prompts for CLIP by retrieving them from a learned pool using clustered intermediate visual features, improving base-to-novel and cross-domain accuracy.
-
Few-Shot Query Intent Detection via Relation-Aware Prompt Learning
SAID pretrains language models with query-query and query-answer relation-aware soft prompts, then transfers them via intent-specific prompts for few-shot intent detection, reporting up to 27% relative accuracy gains.
-
CLIPSym: Delving into Symmetry Detection with CLIP
CLIPSym achieves state-of-the-art reflection and rotation symmetry detection on DENDI, SDRW, and LDRS by fine-tuning CLIP with a semantic prompt-grouping scheme and an equivariant decoder.
-
Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking Design
Water4MU tunes an invisible watermark on data so that machine unlearning algorithms can remove requested images more effectively, beating prior methods on 'challenging forgets'.
-
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
DynImg represents a video snippet as a keyframe plus four resized neighboring frames as temporal prompts, with a 4D rotary position embedding, and reports improved video QA accuracy.
-
Casper: Inferring Diverse Intents for Assistive Teleoperation with Vision Language Models
A VLM-powered assistive teleoperation system infers diverse user intents from teleoperation snippets and executes them with a skill library, outperforming baselines on real-world mobile manipulation tasks.
-
Exploring Visual Prompting: Robustness Inheritance and Beyond
Visual prompts built on robust source models inherit adversarial robustness but lose standard accuracy; a max-pooling over logit blocks (PBL) improves accuracy while keeping most robustness.
-
Dual-Path Stable Soft Prompt Generation for Domain Generalization
DPSPG trains a negative-prompt branch alongside the normal positive prompt generator for CLIP, improving domain generalization accuracy and reducing prompt variability across random seeds.
-
Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding
A disease-aware prompting method that reweights chest X-ray features using the model's own explainability map improves weakly-supervised visual grounding on three benchmarks.
-
Revisiting the Auxiliary Data in Backdoor Purification
Guided Input Calibration aligns any auxiliary dataset with a victim model's learned features before backdoor purification, consistently improving clean accuracy across dataset types with variable effects on attack suc...
-
MMLoP: Multi-Modal Low-Rank Prompting for Efficient Vision-Language Adaptation
MMLoP compresses deep multi-modal prompts into a rank-1 shared subspace, reaching a 79.70% base-to-novel harmonic mean with 11.5K trainable parameters.
-
Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting
CAW adds a confidence-weighted KL loss and feature-alignment regularization to CLIP adversarial fine-tuning, raising average AutoAttack robust accuracy from 31.6% to 33.5% on 15 datasets.
-
Parameter-Efficient Adaptation of mPLUG-Owl2 via Pixel-Level Visual Prompts for NR-IQA
With a learned 30-pixel border prompt added to input images, a frozen mPLUG-Owl2-7B reaches 0.932 SRCC on KADID-10k using about 156K trainable parameters.
-
Adaptive Point-Prompt Tuning: Fine-Tuning Heterogeneous Foundation Models for 3D Point Cloud Analysis
APPT adapts frozen foundation models of any modality to 3D point clouds using a shared point-prompt generator and permutation-invariant position signals, improving 3D classification and segmentation with minimal training.
-
DepthDark: Robust Monocular Depth Estimation for Low-Light Environments
DepthDark obtains state-of-the-art low-light depth estimates by jointly introducing synthetic nighttime data generation and an efficient fine-tuning strategy for a pretrained depth foundation model.
-
Visual Instance-aware Prompt Tuning
ViaPT generates instance-aware prompts per image, fuses them with dataset-level prompts, and applies PCA compression to outperform VPT-Deep and other PEFT baselines on FGVC, HTA, and VTAB-1k.
-
Model Reprogramming Demystified: A Neural Tangent Kernel Perspective
The paper claims the minimum eigenvalue of the source model's NTK matrix controls both source and reprogrammed target model performance.
-
MetaWriter: Personalized Handwritten Text Recognition Using Meta-Learned Prompt Tuning
MetaWriter uses meta-learned prompt tuning with an image-reconstruction auxiliary task to adapt a handwritten text recognizer to new writers from unlabeled examples, reporting lower error rates on IAM and RIMES with l...
-
Polarization-Resolved Chlorophyll Imaging for Non-Invasive Plant Tissue Assessment Using a Silicon-Rich Nitride Metalens Array
A compact metalens array is claimed to enable label-free polarization-resolved imaging of plant tissue for stress assessment, but only the abstract was available for verification.
-
Progressive Homeostatic and Plastic Prompt Tuning for Audio-Visual Multi-Task Incremental Learning
A three-stage prompt-tuning method for audio-visual multi-task incremental learning is proposed, reporting state-of-the-art results on AVE, AVVP, AVS, and AVQA, with caveats about its evaluation metric and ablations.
-
CoCoA-Mix: Confusion-and-Confidence-Aware Mixture Model for Context Optimization
CoCoA-Mix combines cross-entropy with a confidence penalty and a scalar-weighted mixture of specialized and generalized prompts, beating prior prompt-tuning baselines on base-to-new, cross-dataset, and few-shot increm...
-
Generalizing vision-language models to novel domains: A comprehensive survey
A survey of VLM generalization literature organized by transferred module, with benchmark tables and a review of multimodal LLMs.
-
Twistronics and moir\'e superlattice physics in 2D transition metal dichalcogenides
The provided full text does not match the abstract, so the review's content cannot be assessed.
Discussion (0). Continue with ORCID to comment.