REVIEW 3 cited by
Diffusion Models for Open-Vocabulary Segmentation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Open-vocabulary segmentation is the task of segmenting anything that can be named in an image. Recently, large-scale vision-language modelling has led to significant advances in open-vocabulary segmentation, but at the cost of gargantuan and increasing training and annotation efforts. Hence, we ask if it is possible to use existing foundation models to synthesise on-demand efficient segmentation algorithms for specific class sets, making them applicable in an open-vocabulary setting without the need to collect further data, annotations or perform training. To that end, we present OVDiff, a novel method that leverages generative text-to-image diffusion models for unsupervised open-vocabulary segmentation. OVDiff synthesises support image sets for arbitrary textual categories, creating for each a set of prototypes representative of both the category and its surrounding context (background). It relies solely on pre-trained components and outputs the synthesised segmenter directly, without training. Our approach shows strong performance on a range of benchmarks, obtaining a lead of more than 5% over prior work on PASCAL VOC.
Forward citations
Cited by 3 Pith papers
-
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
Stable Diffusion features, especially when conditioned on the question, improve vision-centric multimodal question answering when fused with CLIP.
-
G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models
Using the discrepancy between an image and its mask-conditioned Stable Diffusion reconstruction, G4Seg refines coarse segmentation masks by aligning pixels in CLIP feature space and mixing foreground probabilities.
-
Segment Anyword: Mask Prompt Inversion for Open-Set Grounded Segmentation
A training-free pipeline uses per-image textual inversion in a frozen diffusion model, then feeds linguistic-guided cross-attention prompts to SAM, achieving state-of-the-art open-set grounded segmentation on several ...
Discussion (0). Continue with ORCID to comment.