Pith. sign in

REVIEW 1 cited by

ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic Arithmetic

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.14447 v2 pith:YHS4S3AI submitted 2021-11-29 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords modelsimageimageszero-shotarithmeticgivenlargelearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent text-to-image matching models apply contrastive learning to large corpora of uncurated pairs of images and sentences. While such models can provide a powerful score for matching and subsequent zero-shot tasks, they are not capable of generating caption given an image. In this work, we repurpose such models to generate a descriptive text given an image at inference time, without any further training or tuning steps. This is done by combining the visual-semantic model with a large language model, benefiting from the knowledge in both web-scale models. The resulting captions are much less restrictive than those obtained by supervised captioning methods. Moreover, as a zero-shot learning method, it is extremely flexible and we demonstrate its ability to perform image arithmetic in which the inputs can be either images or text, and the output is a sentence. This enables novel high-level vision capabilities such as comparing two images or solving visual analogy tests. Our code is available at: https://github.com/YoadTew/zero-shot-image-to-text.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pattern Analogies: Learning to Perform Programmatic Image Edits by Analogy

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A diffusion model, trained on synthetic pattern quartets generated by the SplitWeave DSL, can apply a program-level edit demonstrated on one pattern pair to a new real-world pattern.

Pith tools