Pith. sign in

REVIEW 1 cited by

Task Bias in Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.04412 v1 pith:SYOL5GFZ submitted 2022-12-08 cs.CV cs.LG

classification cs.CVcs.LG
keywords taskvisualtowardsrepresentationbiasbiasedrepresentationstasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Incidental supervision from language has become a popular approach for learning generic visual representations that can be prompted to perform many recognition tasks in computer vision. We conduct an in-depth exploration of the CLIP model and show that its visual representation is often strongly biased towards solving some tasks more than others. Moreover, which task the representation will be biased towards is unpredictable, with little consistency across images. To resolve this task bias, we show how to learn a visual prompt that guides the representation towards features relevant to their task of interest. Our results show that these visual prompts can be independent of the input image and still effectively provide a conditioning mechanism to steer visual representations towards the desired task.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Richer text prompts from LLM synonyms and cleaner image regions from activation maps improve zero-shot vision-language classification.

Pith tools