Pith. sign in

REVIEW 10 cited by

A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.12980 v1 pith:WKYYBME3 submitted 2023-07-24 cs.CV

classification cs.CV
keywords modelsengineeringpromptmodelvision-languagelanguagenaturalpre-trained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Prompt engineering is a technique that involves augmenting a large pre-trained model with task-specific hints, known as prompts, to adapt the model to new tasks. Prompts can be created manually as natural language instructions or generated automatically as either natural language instructions or vector representations. Prompt engineering enables the ability to perform predictions based solely on prompts without updating model parameters, and the easier application of large pre-trained models in real-world tasks. In past years, Prompt engineering has been well-studied in natural language processing. Recently, it has also been intensively studied in vision-language modeling. However, there is currently a lack of a systematic overview of prompt engineering on pre-trained vision-language models. This paper aims to provide a comprehensive survey of cutting-edge research in prompt engineering on three types of vision-language models: multimodal-to-text generation models (e.g. Flamingo), image-text matching models (e.g. CLIP), and text-to-image generation models (e.g. Stable Diffusion). For each type of model, a brief model summary, prompting methods, prompting-based applications, and the corresponding responsibility and integrity issues are summarized and discussed. Furthermore, the commonalities and differences between prompting on vision-language models, language models, and vision models are also discussed. The challenges, future directions, and research opportunities are summarized to foster future research on this topic.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 63 citations worldwide. Full citation record

  1. Visual Textualization for Image Prompted Object Detection

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Visual textualization projects support images into the text feature space and prompts an unmodified OVLM, achieving strong few-shot and open-set detection results.

  2. Visual prompt engineering for video models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Automatically converting task images to photorealistic variants (visual prompt engineering) improves video-model reasoning performance, often beating text prompt engineering and test-time scaling.

  3. AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models

    cs.RO 2025-11 conditional novelty 6.0 of 10

    Structured prompting plus a discrete skill library lets a frozen VLM direct aerial manipulation, reaching 87.5% simulated and 80% hardware success in pick-and-place tasks.

  4. Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds

    cs.CV 2025-10 reject novelty 6.0 of 10

    A VLM method aligns hierarchical image and text feature trees across hyperbolic manifolds of different curvatures via an intermediate manifold, but the theoretical justification is flawed.

  5. True Multimodal In-Context Learning Needs Attention to the Visual Context

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 160-parameter attention-scaling method, DARA, improves true multimodal in-context learning on a new dataset, TrueMICL, that forces models to use demo images rather than copy text patterns.

  6. DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DynImg represents a video snippet as a keyframe plus four resized neighboring frames as temporal prompts, with a 4D rotary position embedding, and reports improved video QA accuracy.

  7. CAD2DMD-SET: Synthetic Generation Tool of Digital Measurement Device CAD Model Datasets for fine-tuning Large Vision-Language Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A synthetic data pipeline for digital measurement devices plus a real-image benchmark improves LVLM reading performance from 32.92% to 96.04% ANLS for InternVL.

  8. Adapting Vision-Language Models Without Labels: A Comprehensive Survey

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A survey that organizes unsupervised vision-language model adaptation by unlabeled-data availability into four paradigms: data-free transfer, domain transfer, episodic test-time, and online test-time adaptation.

  9. A Survey of AIOps in the Era of Large Language Models

    cs.SE 2025-06 conditional novelty 3.0 of 10

    A systematic survey that categorizes LLM-based AIOps research into four dimensions: data sources, tasks, methods, and evaluation, claiming to be the first comprehensive such overview.

  10. Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI

    cs.SE 2025-05 conditional novelty 3.0 of 10

    A qualitative taxonomy positions vibe coding and agentic coding as complementary paradigms rather than rivals in AI-assisted software development.

Pith tools