Pith. sign in

REVIEW 2 cited by

Unconstrained Open Vocabulary Image Classification: Zero-Shot Transfer from Text to Image via CLIP Inversion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.11211 v4 pith:OTL6UJTM submitted 2024-07-15 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords imageclassificationclipobjectlabelsopentextvocabulary
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce NOVIC, an innovative real-time uNconstrained Open Vocabulary Image Classifier that uses an autoregressive transformer to generatively output classification labels as language. Leveraging the extensive knowledge of CLIP models, NOVIC harnesses the embedding space to enable zero-shot transfer from pure text to images. Traditional CLIP models, despite their ability for open vocabulary classification, require an exhaustive prompt of potential class labels, restricting their application to images of known content or context. To address this, we propose an "object decoder" model that is trained on a large-scale 92M-target dataset of templated object noun sets and LLM-generated captions to always output the object noun in question. This effectively inverts the CLIP text encoder and allows textual object labels from essentially the entire English language to be generated directly from image-derived embedding vectors, without requiring any a priori knowledge of the potential content of an image, and without any label biases. The trained decoders are tested on a mix of manually and web-curated datasets, as well as standard image classification benchmarks, and achieve fine-grained prompt-free prediction scores of up to 87.5%, a strong result considering the model must work for any conceivable image and without any contextual clues.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Just Leaf It: Accelerating Diffusion Classifiers with Hierarchical Class Pruning

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A hierarchical label-tree pruning method reduces diffusion-classifier inference time by up to about 60% while keeping accuracy roughly unchanged or slightly higher.

  2. Expanding Event Modality Applications through a Robust CLIP-Based Encoder

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A CLIP-based encoder for event cameras, trained with contrastive, consistency, and KL losses, improves zero-shot and few-shot object recognition and extends to video anomaly detection and cross-modal retrieval.

Pith tools