Pith. sign in

REVIEW 4 cited by

Visual Prompting via Image Inpainting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.00647 v1 pith:L27N2X5Z submitted 2022-09-01 cs.CV

classification cs.CV
keywords imagevisualpromptinginpaintingdetectiondownstreamgivenmodel
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

How does one adapt a pre-trained visual model to novel downstream tasks without task-specific finetuning or any model modification? Inspired by prompting in NLP, this paper investigates visual prompting: given input-output image example(s) of a new task at test time and a new input image, the goal is to automatically produce the output image, consistent with the given examples. We show that posing this problem as simple image inpainting - literally just filling in a hole in a concatenated visual prompt image - turns out to be surprisingly effective, provided that the inpainting algorithm has been trained on the right data. We train masked auto-encoders on a new dataset that we curated - 88k unlabeled figures from academic papers sources on Arxiv. We apply visual prompting to these pretrained models and demonstrate results on various downstream image-to-image tasks, including foreground segmentation, single object detection, colorization, edge detection, etc.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pattern Analogies: Learning to Perform Programmatic Image Edits by Analogy

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A diffusion model, trained on synthetic pattern quartets generated by the SplitWeave DSL, can apply a program-level edit demonstrated on one pattern pair to a new real-world pattern.

  2. In-Context Collapse in Vision-Language Models and How to Mitigate it?

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Many-shot in-context learning in vision-language models can collapse accuracy as demonstrations accumulate, and the failure is causally localized to the vision-language integration pathway, where a small adapter repairs it.

  3. From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Frozen CogVideoX1.5, adapted with LoRA on 3 to 30 input-output videos, performs segmentation, pose estimation, and abstract reasoning (ARC-AGI 16.75%) with modest but real generalization.

  4. GraphTheft: Quantifying Privacy Risks in Graph Prompt Learning

    cs.CR 2024-11 conditional novelty 6.0 of 10

    An empirical study showing that graph prompt learning exposes node attributes and links to inference attacks, with prompt tuning adding little extra risk over frozen GNN baselines.

Pith tools