Pith. sign in

REVIEW 9 cited by

Visual Classification via Description from Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.07183 v2 pith:HA7S6BDK submitted 2022-10-13 cs.CV cs.LG

classification cs.CVcs.LG
keywords classificationvlmscategorydescriptorsfeatureslanguagemodelsadditional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-language models (VLMs) such as CLIP have shown promising performance on a variety of recognition tasks using the standard zero-shot classification procedure -- computing similarity between the query image and the embedded words for each category. By only using the category name, they neglect to make use of the rich context of additional information that language affords. The procedure gives no intermediate understanding of why a category is chosen, and furthermore provides no mechanism for adjusting the criteria used towards this decision. We present an alternative framework for classification with VLMs, which we call classification by description. We ask VLMs to check for descriptive features rather than broad categories: to find a tiger, look for its stripes; its claws; and more. By basing decisions on these descriptors, we can provide additional cues that encourage using the features we want to be used. In the process, we can get a clear idea of what features the model uses to construct its decision; it gains some level of inherent explainability. We query large language models (e.g., GPT-3) for these descriptors to obtain them in a scalable way. Extensive experiments show our framework has numerous advantages past interpretability. We show improvements in accuracy on ImageNet across distribution shifts; demonstrate the ability to adapt VLMs to recognize concepts unseen during training; and illustrate how descriptors can be edited to effectively mitigate bias compared to the baseline.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Controlling Embedding Spaces with Text-Conditioned Transformations

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A single hypernetwork turns text descriptions of attributes into affine maps of frozen CLIP embeddings, making those attributes control retrieval and clustering without re-encoding the gallery.

  2. Spatially Grounded Concept-Based Image Classification

    cs.CV 2025-10 conditional novelty 6.0 of 10

    SEG-MIL-CBM uses CLIP-guided segmentation with attention-based multiple instance learning to build a concept bottleneck model that produces spatially grounded explanations and improves worst-group accuracy on spurious...

  3. Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MuGCP adapts CLIP by decoding instance-specific prompts from a frozen MLLM's KV cache and fusing them with visual prompts, achieving state-of-the-art few-shot classification on 14 datasets.

  4. GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A zero-shot geospatial classifier that turns satellite images into text descriptions and uses a language model to assign labels, plus an optional LLM-based hierarchical clustering for many-class datasets.

  5. Improving Recommendation Fairness without Sensitive Attributes Using Multi-Persona LLMs

    cs.IR 2025-05 conditional novelty 6.0 of 10

    LLMFOSA improves recommendation fairness by using multi-persona LLMs to infer and then remove sensitive attribute information from user embeddings, without using true sensitive labels during training.

  6. EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A retrieval pipeline that augments CLIP queries with LLM-written entity visual descriptions, selected by a retriever-trained rewriter, improves image-text retrieval over CLIP baselines.

  7. Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Using wrong/correct hard-sample head comparisons, LTC finds spurious CLIP attention heads and corrects them to raise worst-group accuracy on biased benchmarks.

  8. Adapting Vision-Language Models Without Labels: A Comprehensive Survey

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A survey that organizes unsupervised vision-language model adaptation by unlabeled-data availability into four paradigms: data-free transfer, domain transfer, episodic test-time, and online test-time adaptation.

  9. Enhancing Zero-Shot Brain Tumor Subtype Classification via Fine-Grained Patch-Text Alignment

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    FG-PAN improves zero-shot brain tumor subtype classification by aligning refined visual patch features with LLM-generated fine-grained text prototypes.

Pith tools