Pith. sign in

REVIEW 1 cited by

Test-time Contrastive Concepts for Open-world Semantic Segmentation with Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.05061 v3 pith:4BVRMK4D submitted 2024-07-06 cs.CV

classification cs.CV
keywords conceptscontrastivetextualgivenimagemodelsproposescenario
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent CLIP-like Vision-Language Models (VLMs), pre-trained on large amounts of image-text pairs to align both modalities with a simple contrastive objective, have paved the way to open-vocabulary semantic segmentation. Given an arbitrary set of textual queries, image pixels are assigned the closest query in feature space. However, this works well when a user exhaustively lists all possible visual concepts in an image that contrast against each other for the assignment. This corresponds to the current evaluation setup in the literature, which relies on having access to a list of in-domain relevant concepts, typically classes of a benchmark dataset. Here, we consider the more challenging (and realistic) scenario of segmenting a single concept, given a textual prompt and nothing else. To achieve good results, besides contrasting with the generic 'background' text, we propose two different approaches to automatically generate, at test time, query-specific textual contrastive concepts. We do so by leveraging the distribution of text in the VLM's training set or crafted LLM prompts. We also propose a metric designed to evaluate this scenario and show the relevance of our approach on commonly used datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DynamicEarth: How Far are We from Open-Vocabulary Change Detection?

    cs.CV 2025-01 reject novelty 4.0 of 10

    The paper shows that composing mask proposal, feature comparison, and open-vocabulary classification models can detect arbitrary-category changes in satellite images without training.

Pith tools