Pith. sign in

REVIEW 2 cited by

Ref-Diff: Zero-shot Referring Image Segmentation with Generative Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.16777 v2 pith:P56GJR4C submitted 2023-08-31 cs.CV

classification cs.CV
keywords modelsgenerativereferringref-diffsegmentationtaskdiscriminativezero-shot
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Zero-shot referring image segmentation is a challenging task because it aims to find an instance segmentation mask based on the given referring descriptions, without training on this type of paired data. Current zero-shot methods mainly focus on using pre-trained discriminative models (e.g., CLIP). However, we have observed that generative models (e.g., Stable Diffusion) have potentially understood the relationships between various visual elements and text descriptions, which are rarely investigated in this task. In this work, we introduce a novel Referring Diffusional segmentor (Ref-Diff) for this task, which leverages the fine-grained multi-modal information from generative models. We demonstrate that without a proposal generator, a generative model alone can achieve comparable performance to existing SOTA weakly-supervised models. When we combine both generative and discriminative models, our Ref-Diff outperforms these competing methods by a significant margin. This indicates that generative models are also beneficial for this task and can complement discriminative models for better referring segmentation. Our code is publicly available at https://github.com/kodenii/Ref-Diff.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GS: Generative Segmentation via Label Diffusion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Direct label-space diffusion, conditioned on image and text, achieves state-of-the-art 69.7 Average Recall on Panoptic Narrative Grounding.

  2. Computationally Efficient Diffusion Models in Medical Imaging: A Comprehensive Review

    eess.IV 2025-05 reject novelty 3.0 of 10

    A survey of DDPM, LDM, and WDM diffusion models for medical imaging, organized around training and inference efficiency.

Pith tools