Pith. sign in

REVIEW 2 cited by

ViewCo: Discovering Text-Supervised Segmentation Masks via Multi-View Semantic Consistency

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.10307 v1 pith:CJ6IV632 submitted 2023-01-31 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords segmentationconsistencysemantictextmodelingproposesamesupervision
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, great success has been made in learning visual representations from text supervision, facilitating the emergence of text-supervised semantic segmentation. However, existing works focus on pixel grouping and cross-modal semantic alignment, while ignoring the correspondence among multiple augmented views of the same image. To overcome such limitation, we propose multi-\textbf{View} \textbf{Co}nsistent learning (ViewCo) for text-supervised semantic segmentation. Specifically, we first propose text-to-views consistency modeling to learn correspondence for multiple views of the same input image. Additionally, we propose cross-view segmentation consistency modeling to address the ambiguity issue of text supervision by contrasting the segment features of Siamese visual encoders. The text-to-views consistency benefits the dense assignment of the visual features by encouraging different crops to align with the same text, while the cross-view segmentation consistency modeling provides additional self-supervision, overcoming the limitation of ambiguous text supervision for segmentation masks. Trained with large-scale image-text data, our model can directly segment objects of arbitrary categories in a zero-shot manner. Extensive experiments show that ViewCo outperforms state-of-the-art methods on average by up to 2.9\%, 1.6\%, and 2.4\% mIoU on PASCAL VOC2012, PASCAL Context, and COCO, respectively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Using the discrepancy between an image and its mask-conditioned Stable Diffusion reconstruction, G4Seg refines coarse segmentation masks by aligning pixels in CLIP feature space and mixing foreground probabilities.

  2. SAB3R: Semantic-Augmented Backbone in 3D Reconstruction

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SAB3R unifies 3D reconstruction and open-vocabulary segmentation in a single feed-forward network trained by distilling CLIP and DINOv2 features into MASt3R.

Pith tools