Pith. sign in

REVIEW 5 cited by

PromptCap: Prompt-Guided Task-Aware Image Captioning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.09699 v4 pith:EPV64S4J submitted 2022-11-15 cs.CV cs.CL

classification cs.CVcs.CL
keywords promptcapimagevisualcaptioningcaptionscaptiongenericgpt-3
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Knowledge-based visual question answering (VQA) involves questions that require world knowledge beyond the image to yield the correct answer. Large language models (LMs) like GPT-3 are particularly helpful for this task because of their strong knowledge retrieval and reasoning capabilities. To enable LM to understand images, prior work uses a captioning model to convert images into text. However, when summarizing an image in a single caption sentence, which visual entities to describe are often underspecified. Generic image captions often miss visual details essential for the LM to answer visual questions correctly. To address this challenge, we propose PromptCap (Prompt-guided image Captioning), a captioning model designed to serve as a better connector between images and black-box LMs. Different from generic captions, PromptCap takes a natural-language prompt to control the visual entities to describe in the generated caption. The prompt contains a question that the caption should aid in answering. To avoid extra annotation, PromptCap is trained by examples synthesized with GPT-3 and existing datasets. We demonstrate PromptCap's effectiveness on an existing pipeline in which GPT-3 is prompted with image captions to carry out VQA. PromptCap outperforms generic captions by a large margin and achieves state-of-the-art accuracy on knowledge-based VQA tasks (60.4% on OK-VQA and 59.6% on A-OKVQA). Zero-shot results on WebQA show that PromptCap generalizes well to unseen domains.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing

    cs.CV 2026-08 conditional novelty 6.0 of 10

    An output-blind, rubric-grounded evaluator and benchmark that ranks generated videos closer to human preferences than five existing benchmark-specific evaluators.

  2. Salient Knowledge Pathways: Sparse Cross-Modal Routing for Efficient Knowledge-Intensive Multimodal Question Answering

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A sparse cross-modal routing system matches or exceeds dense retrieval-augmented baselines on five KI-MMQA benchmarks at 3.4–6.8× lower FLOPs.

  3. MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MMIG-Bench is a unified benchmark of 4,850 prompts and 1,750 reference images with a three-level evaluation suite, including the VQA-based Aspect Matching Score that correlates with human ratings.

  4. Prompt-Driven Continual Graph Learning

    cs.LG 2025-02 conditional novelty 6.0 of 10

    PromptCGL learns per-task prompts on a frozen graph neural network and achieves near-joint-training accuracy on four continual graph learning benchmarks with constant memory and near-zero forgetting.

  5. GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A modular zero-shot KB-VQA framework using Grounding DINO, dual captioners, semantic caption filtering, and LLM prompting reports new state-of-the-art numbers on OK-VQA, A-OKVQA, and VQAv2.

Pith tools