REVIEW 3 major objections 5 minor 71 references
ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read ITACLIP turns frozen CLIP into a top training-free segmenter
desk verdict Systematically combines known CLIP-for-segmentation tricks into a training-free system that posts SOTA mIoU on five benchmarks, but per-dataset validation tuning of the text type and coefficients makes the generality claim weaker than the abstract implies. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the modified attention map of a ViT-based CLIP image encoder: $\mathrm{Attn}(X)=\mathrm{softmax}(XW_QW_Q^\top X^\top/\sqrt{d})+\mathrm{softmax}(XW_KW_K^\top X^\top/\sqrt{d})$, i.e., query-query plus key-key self-self attention instead of query-key attention. The paper pairs this with two fusion rules: the final-layer map is averaged with the mean of maps from layers $l'=\{7,8,10\}$, and the feed-forward block of the last layer is removed so the residual block is $X^{(L)}=X^{(L-1)}+\mathrm{SA}(\mathrm{LN}(X^{(L-1)}))$. Around this core sit two input-enrichment mechanisms: image engineering, which averages first-category augmentation features and combines flip logits after reversing them with weight $\lambda$; and LLM-based auxiliary text, which forms $X^{\mathrm{text}}_{\mathrm{refined}}=\alpha X^{\mathrm{text}}_{\mathrm{aux}}+(1-\alpha)X^{\mathrm{text}}$. The argument is that each piece adds localization or representation diversity that a frozen CLIP lacks, and the ablations attribute the final gain to their combination.
What would settle it
Take a held-out benchmark whose class list was not used in this paper, run ITACLIP as shipped and a version whose layers, text type, and $\lambda,\alpha$ are tuned on that benchmark's validation split; if the untuned defaults lose to NACLIP or fall well short of the tuned version, the general-superiority claim is falsified.
Extended reading notes
Core claim
The central claim is that CLIP's image-level representations carry enough spatial information for accurate segmentation once the encoder's last block is surgically modified and the inputs are enriched. Specifically, the paper argues that replacing the final self-attention with the sum of query-query and key-key self-self attention, deleting the feed-forward network in the last layer, and averaging the resulting attention map with maps from layers 7, 8, and 10 produces more localized and semantically coherent patch features. On the text side, blending the original class-name embedding with an LLM-generated definition or synonym using coefficient $\alpha$ exploits CLIP's open-vocabulary text space. On the image side, features from the original image and two structure-preserving augmentations are averaged, while logits from flip augmentations are computed separately, un-flipped, and combined with weight $\lambda$. With PAMR refinement and a stride of 28, the paper reports mIoU of 27.0 on COCO-Stuff, 37.7 on COCO-Object, 67.9 on Pascal VOC, 37.5 on Pascal Context, and 40.2 on Cityscapes, each above the previous training-free state of the art.
Load-bearing premise
The recipe's specific settings—fused layers {7,8,10}, definitions versus synonyms, weights $\lambda$ and $\alpha$, and stride 28—were chosen on the validation splits of the five benchmarks, so the claim that ITACLIP is generally superior depends on those choices transferring to unseen data.
Editorial extensions
If this is right
- Any frozen CLIP-ViT-B/16 can be upgraded to a stronger segmenter by swapping the attention formula, dropping the final FFN, and fusing middle-layer maps, with no gradient updates or segmentation labels.
- The image-engineering and LLM-text modules are drop-in components that the paper says can be attached to other CLIP-based vision tasks, so gains may transfer beyond segmentation.
- Lower stride consistently improves mIoU across all five datasets, and stride 112 retains state-of-the-art results on most datasets, so users can trade compute for accuracy.
- The method stays competitive on several benchmarks even without PAMR post-processing, so the core gains are not an artifact of refinement.
- Small variations in the fusion coefficients $\lambda$ and $\alpha$ change Pascal Context mIoU by only a few tenths, suggesting the recipe is not acutely sensitive to those two knobs.
Reading between the lines
- The paper tunes the choice of intermediate layers, definition-versus-synonym, $\lambda$, $\alpha$, and stride on the validation splits it reports; an untested consequence is that these exact settings may not be optimal on a new dataset, and a fairer generalization test would tune nothing on the target set.
- Because the middle-layer fusion uses a plain average, an obvious extension the paper does not explore is learning or image-adaptive weights over layers; the ablation shows layer choice matters, so such weights could yield further gains.
- The LLM-generated texts could be replaced by a non-learned lexical source such as WordNet synonyms; if performance held, the mechanism would be shown to be about enriching text variants rather than about the specific LLM.
- The augmentation averaging is effectively test-time augmentation; a natural stress test is whether the same Image Engineering module improves other dense CLIP tasks such as open-vocabulary detection, as the conclusion hints.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes ITACLIP, a training-free extension of CLIP for open-vocabulary semantic segmentation. The method modifies the ViT image encoder by using q-q plus k-k self-self attention in the last layer, removing the final feed-forward block, and averaging the final-layer attention map with attention maps from selected intermediate layers (l' = {7, 8, 10}). It also enriches text inputs with LLM-generated definitions or synonyms and enriches image inputs by averaging features from blur/grayscale augmentations and merging logits from horizontally and vertically flipped views. On the official validation splits of COCO-Stuff, COCO-Object, Pascal VOC, Pascal Context, and Cityscapes, Table 1 reports mIoU gains over previous training-free methods, and the ablations cover each proposed component.
Significance. If the reported numbers hold under a fixed configuration, ITACLIP is a useful empirical advance: it combines several previously scattered ideas—self-self attention, FFN removal, multi-layer attention averaging, image augmentation ensembling, and LLM-generated auxiliary text—into one training-free pipeline and reports consistent improvements over SCLIP and NACLIP on all five benchmarks. The paper is transparent in its experimental design: the datasets are standard, each component is ablated, the baseline list is explicit, and the code is released. The main caveat is that several free parameters and the auxiliary-text type are selected on the same validation splits used for the headline numbers, so the generality of the claimed state of the art is not yet established.
major comments (3)
- [Sec. 3.4, Appendix A.1, Table 9] The auxiliary text type is chosen per dataset "based on their segmentation performance" (Sec. 3.4), and Table 9 assigns dataset-specific values for the coefficients λ and α. Because this selection and tuning are performed on the same official validation splits whose mIoU numbers appear in Table 1, the reported results are the best among a family of per-dataset configurations rather than the output of a single fixed ITACLIP system. This matters because the reported advantages over the previous best are modest in several cases (+1.3 on COCO-Stuff, +1.9 on Cityscapes, +1.5 on COCO-Object), and a fixed choice of auxiliary text type could shrink or eliminate some of these margins. Please report results for a unified configuration across all datasets—for example, always definitions, always synonyms, and both combined—and, if possible, a held-out dataset or class set.
- [Sec. 4.3, Table 4] The intermediate layers l' = {7, 8, 10} are selected using Pascal VOC performance without PAMR and are then applied to all five datasets, but no per-dataset sensitivity analysis is reported for this choice. Since the selected layers are part of the architectural contribution, the paper should show that this choice is not overfit to VOC, or it should use a selection criterion that does not rely on the validation splits of the other four benchmarks.
- [Sec. 4.1, Table 1] The baseline comparison mixes numbers that appear to be copied from prior papers, a reimplementation (TagCLIP), and methods with different post-processing (PAMR vs. Dense-CRF), but the table does not state which baseline rows were rerun under the authors' evaluation protocol and which were taken verbatim. Please document the source and evaluation protocol of each baseline row so that the claimed margins are verifiable.
minor comments (5)
- [Abstract] The abstract says the method outperforms state-of-the-art approaches "on segmentation benchmarks such as COCO-Stuff, COCO-Object, Pascal Context, and Pascal VOC," while the main text says "five popular segmentation benchmarks"; please make the dataset list consistent and name all five datasets.
- [Table 1] The column header "VOC Context" is ambiguous because Pascal VOC and Pascal Context are separate benchmarks; please separate the columns clearly and add a note explaining the missing entries for CaR and TagCLIP.
- [Appendix A.1, Fig. 4] The LLM prompt templates in Fig. 4 appear as garbled token sequences in the provided version; please render them as plain text so that the auxiliary-text generation procedure is reproducible from the paper itself.
- [Eq. (10)] The symbol K is used for the number of augmentations, which could be confused with the key matrix notation introduced in Eq. (3); please rename the augmentation count to avoid ambiguity.
- [Sec. 5] The conclusion states that the Image Engineering module and LLM-based Text Generation strategy "can be seamlessly integrated into a range of computer vision tasks," but no transfer experiment supports this claim; please soften the claim or add a small demonstration.
Circularity Check
No circularity: ITACLIP is an empirical benchmark evaluation; the reported mIoU values are measurements on external validation sets, not predictions derived from fitted equations or self-referential definitions.
full rationale
The paper's central claim is an empirical one: a set of architectural, textual, and image-augmentation modifications to CLIP yields higher mIoU on five segmentation validation sets. No equation in the paper defines one quantity in terms of the target result, and no fitted parameter is renamed as a prediction. The hyperparameters (lambda, alpha, auxiliary-text type, intermediate layers, stride) are selected by validation performance, which is a legitimate correctness/generalization concern about possible validation-set overfitting, but it is not circular reasoning: the reported numbers are the measured outcome of the chosen configuration, not quantities forced by construction to equal the selection criterion. The paper contains no author self-citations; all cited priors (SCLIP, NACLIP, CLIP-DIY, CLIPSurgery, etc.) are external works, and the improvements are not justified by an imported uniqueness theorem or by a self-citation chain. Because the claim is a benchmark measurement rather than a derivation, there is no load-bearing step that reduces to its own inputs, so the circularity score is 0.
Assumptions & free parameters
free parameters (5)
- λ (image engineering coefficient) =
0.75 (COCO-Stuff, COCO-Object), 0.7 (Pascal VOC, Cityscapes), 0.75 (Pascal Context)
- α (auxiliary text coefficient) =
0.2 (COCO-Stuff), 0.1 (COCO-Object), 0.05 (Pascal VOC, Cityscapes), 0.15 (Pascal Context)
- Intermediate layers l' =
{7, 8, 10}
- Auxiliary text type =
Definitions for COCO-Stuff, Pascal VOC, Pascal Context; synonyms for COCO-Object, Cityscapes
- Slide inference stride =
28 pixels
assumptions (5)
- domain assumption CLIP's frozen ViT features contain enough spatial and semantic information for pixel-level segmentation after the proposed architectural modifications
- domain assumption LLM-generated synonyms and definitions are compatible with CLIP's text embedding space and improve class discrimination
- domain assumption Averaging attention maps from intermediate layers with the final layer produces a relevance map that retains localization
- domain assumption First-category augmentations (blur, grayscale) preserve semantic layout in feature space, and second-category flips can be reversed to restore spatial order
- domain assumption The hand-defined background set for Pascal VOC and COCO-Object is appropriate and does not unfairly advantage the method
Cite this review
Pith. "Pith review of ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements." pith.science (2026). https://pith.science/paper/QG27UG25
@misc{pith2026241112044,
author = {Pith},
title = {Pith review of: ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements},
year = {2026},
howpublished = {\url{https://pith.science/paper/QG27UG25}},
note = {Machine review of arXiv:2411.12044}
}
read the original abstract
Recent advances in foundational Vision Language Models (VLMs) have reshaped the evaluation paradigm in computer vision tasks. These foundational models, especially CLIP, have accelerated research in open-vocabulary computer vision tasks, including Open-Vocabulary Semantic Segmentation (OVSS). Although the initial results are promising, the dense prediction capabilities of VLMs still require further improvement. In this study, we enhance the semantic segmentation performance of CLIP by introducing new modules and modifications: 1) architectural changes in the last layer of ViT and the incorporation of attention maps from the middle layers with the last layer, 2) Image Engineering: applying data augmentations to enrich input image representations, and 3) using Large Language Models (LLMs) to generate definitions and synonyms for each class name to leverage CLIP's open-vocabulary capabilities. Our training-free method, ITACLIP, outperforms current state-of-the-art approaches on segmentation benchmarks such as COCO-Stuff, COCO-Object, Pascal Context, and Pascal VOC. Our code is available at https://github.com/m-arda-aydn/ITACLIP.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Cdul: Clip-driven unsupervised learning for multi-label image classification
Rabab Abdelfattah, Qing Guo, Xiaoguang Li, Xiaofeng Wang, and Song Wang. Cdul: Clip-driven unsupervised learning for multi-label image classification. In ICCV, pages 1348–1357, 2023. 1
work page 2023
-
[2]
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. In NeurIPS,
-
[3]
Self- supervised multimodal versatile networks
Jean-Baptiste Alayrac, Adria Recasens, Rosalia Schneider, Relja Arandjelovi´c, Jason Ramapuram, Jeffrey De Fauw, Lu- cas Smaira, Sander Dieleman, and Andrew Zisserman. Self- supervised multimodal versatile networks. NeurIPS, 33:25– 37, 2020. 2
work page 2020
-
[4]
Single-stage semantic segmentation from image labels
Nikita Araslanov and Stefan Roth. Single-stage semantic segmentation from image labels. In CVPR, pages 4253– 4262, 2020. 7
2020
-
[5]
Fossil: Free open-vocabulary semantic seg- mentation through synthetic references retrieval
Luca Barsellotti, Roberto Amoroso, Lorenzo Baraldi, and Rita Cucchiara. Fossil: Free open-vocabulary semantic seg- mentation through synthetic references retrieval. In WACV, pages 1464–1473, 2024. 1, 2
work page 2024
-
[6]
Grounding everything: Emerging localiza- tion properties in vision-language transformers
Walid Bousselham, Felix Petersen, Vittorio Ferrari, and Hilde Kuehne. Grounding everything: Emerging localiza- tion properties in vision-language transformers. In CVPR, pages 3828–3837, 2024. 2, 4
work page 2024
-
[7]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Sub- biah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Lan- guage models are few-shot learners. In NeurIPS, 2020. 5
work page 2020
-
[8]
Coco- stuff: Thing and stuff classes in context
Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Coco- stuff: Thing and stuff classes in context. In CVPR, pages 1209–1218, 2018. 1, 2, 7, 12
work page 2018
Show all 71 references
-
[9]
Cambridge dictionary
Cambridge University Press. Cambridge dictionary. https : / / dictionary . cambridge . org/. Ac- cessed: 2024-08-24. 12
2024
-
[10]
Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs
Junbum Cha, Jonghwan Mun, and Byungseok Roh. Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs. In CVPR, pages 11165–11174, 2023. 3, 5, 6, 7
2023
-
[11]
Reproducible scal- ing laws for contrastive language-image learning
Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuh- mann, Ludwig Schmidt, and Jenia Jitsev. Reproducible scal- ing laws for contrastive language-image learning. In CVPR, pages 2818–2829, 2023. 2
2023
-
[12]
The cityscapes dataset for semantic urban scene understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Proceed- ings of the IEEE conference on computer vision and pattern re...
2016
-
[13]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009. 7
2009
-
[14]
Segment anything model (sam) for digital pathology: Assess zero- shot segmentation on whole slide imaging
Ruining Deng, Can Cui, Quan Liu, Tianyuan Yao, Lu- cas W Remedios, Shunxing Bao, Bennett A Landman, Lee E Wheless, Lori A Coburn, Keith T Wilson, et al. Segment anything model (sam) for digital pathology: Assess zero- shot segmentation on whole slide imaging. arXiv preprint ar...
2023 arXiv
-
[15]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 2
2018 arXiv
-
[16]
De- coupling zero-shot semantic segmentation
Jian Ding, Nan Xue, Gui-Song Xia, and Dengxin Dai. De- coupling zero-shot semantic segmentation. In CVPR, pages 11583–11592, 2022. 2
2022
-
[17]
An image is worth 16x16 words: Trans- formers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. In ICLR, 2020. 2, 3
2020
-
[18]
Learning to prompt for open-vocabulary ob- ject detection with vision-language model
Yu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi, Yue Gao, and Guoqi Li. Learning to prompt for open-vocabulary ob- ject detection with vision-language model. In CVPR, pages 14084–14093, 2022. 1
2022
-
[19]
The llama 3 herd of models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 ,
-
[20]
The pascal visual object classes (voc) challenge
Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. IJCV, 88:303–338, 2010. 2, 7, 12, 13
2010
-
[21]
Improving clip training with language rewrites
Lijie Fan, Dilip Krishnan, Phillip Isola, Dina Katabi, and Yonglong Tian. Improving clip training with language rewrites. In NeurIPS, 2023. 5
2023
-
[22]
In- terpreting clip’s image representation via text-based decom- position
Yossi Gandelsman, Alexei A Efros, and Jacob Steinhardt. In- terpreting clip’s image representation via text-based decom- position. In ICLR, 2024. 5, 8
2024
-
[23]
Scal- ing open-vocabulary image segmentation with image-level labels
Golnaz Ghiasi, Xiuye Gu, Yin Cui, and Tsung-Yi Lin. Scal- ing open-vocabulary image segmentation with image-level labels. In ECCV, pages 540–557. Springer, 2022. 3
2022
-
[24]
Open-vocabulary object detection via vision and language knowledge distillation
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui. Open-vocabulary object detection via vision and language knowledge distillation. In ICLR, 2022. 1, 5, 6
2022
-
[25]
Pay attention to your neighbours: Training-free open-vocabulary semantic segmentation
Sina Hajimiri, Ismail Ben Ayed, and Jose Dolz. Pay attention to your neighbours: Training-free open-vocabulary semantic segmentation. arXiv preprint arXiv:2404.08181, 2024. 1, 2, 4, 6, 7, 12, 13
2024 arXiv
-
[26]
Open-vocabulary semantic segmentation with decou- pled one-pass network
Cong Han, Yujie Zhong, Dengjie Li, Kai Han, and Lin Ma. Open-vocabulary semantic segmentation with decou- pled one-pass network. In ICCV, pages 1086–1096, 2023. 3
2023
-
[27]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 3
2016
-
[28]
Computer- vision benchmark segment-anything model (sam) in med- ical images: Accuracy in 12 datasets
Sheng He, Rina Bao, Jingpeng Li, Jeffrey Stout, Atle Bjornerud, P Ellen Grant, and Yangming Ou. Computer- vision benchmark segment-anything model (sam) in med- ical images: Accuracy in 12 datasets. arXiv preprint arXiv:2304.09324, 2023. 2
2023 arXiv
-
[29]
Open- clip, July 2021
Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Han- 9 naneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. Open- clip, July 2021. 1, 2
2021
-
[30]
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML, pages 4904–
-
[31]
Learning mask-aware clip representations for zero-shot segmentation
Siyu Jiao, Yunchao Wei, Yaowei Wang, Yao Zhao, and Humphrey Shi. Learning mask-aware clip representations for zero-shot segmentation. NeurIPS, 36:35631–35653,
-
[32]
Segment any- thing
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment any- thing. In ICCV, pages 4015–4026, 2023. 2
2023
-
[33]
Clearclip: Decom- posing clip representations for dense vision-language infer- ence
Mengcheng Lan, Chaofeng Chen, Yiping Ke, Xinjiang Wang, Litong Feng, and Wayne Zhang. Clearclip: Decom- posing clip representations for dense vision-language infer- ence. arXiv preprint arXiv:2407.12442, 2024. 1, 2, 4, 5, 6, 7, 12
2024 arXiv
-
[34]
Language-driven semantic seg- mentation
Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and Rene Ranftl. Language-driven semantic seg- mentation. In ICLR, 2022. 3
2022
-
[35]
Align before fuse: Vision and language representation learn- ing with momentum distillation
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learn- ing with momentum distillation. NeurIPS, 34:9694–9705,
-
[36]
Cascade-clip: Cascaded vision-language embeddings alignment for zero-shot semantic segmentation
Yunheng Li, ZhongYu Li, Quansheng Zeng, Qibin Hou, and Ming-Ming Cheng. Cascade-clip: Cascaded vision-language embeddings alignment for zero-shot semantic segmentation. In ICML, 2024. 1
2024
-
[37]
Clip surgery for better explainability with enhancement in open- vocabulary tasks
Yi Li, Hualiang Wang, Yiqun Duan, and Xiaomeng Li. Clip surgery for better explainability with enhancement in open- vocabulary tasks. arXiv preprint arXiv:2304.05653, 2023. 1, 2, 4, 5
2023 arXiv
-
[38]
Open-vocabulary object segmentation with diffusion models
Ziyi Li, Qinye Zhou, Xiaoyun Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. Open-vocabulary object segmentation with diffusion models. In ICCV, pages 7667–7676, 2023. 3
2023
-
[39]
Open-vocabulary semantic segmentation with mask-adapted clip
Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, and Diana Marculescu. Open-vocabulary semantic segmentation with mask-adapted clip. In CVPR, pages 7061–7070, 2023. 3
2023
-
[40]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV. Springer, 2014. 2, 7, 12, 13
2014
-
[41]
Clip is also an ef- ficient segmenter: A text-driven approach for weakly super- vised semantic segmentation
Yuqi Lin, Minghao Chen, Wenxiao Wang, Boxi Wu, Ke Li, Binbin Lin, Haifeng Liu, and Xiaofei He. Clip is also an ef- ficient segmenter: A text-driven approach for weakly super- vised semantic segmentation. InCVPR, pages 15305–15314,
-
[42]
Tagclip: A local-to-global framework to enhance open-vocabulary multi-label classification of clip without training
Yuqi Lin, Minghao Chen, Kaipeng Zhang, Hengjia Li, Ming- ming Li, Zheng Yang, Dongqin Lv, Binbin Lin, Haifeng Liu, and Deng Cai. Tagclip: A local-to-global framework to enhance open-vocabulary multi-label classification of clip without training. In AAAI, volume 38, pages 3513–3521,
-
[43]
Object- centric learning with slot attention
Francesco Locatello, Dirk Weissenborn, Thomas Un- terthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf. Object- centric learning with slot attention. NeurIPS, 33:11525– 11538, 2020. 3
2020
-
[44]
Segclip: Patch aggregation with learn- able centers for open-vocabulary semantic segmentation
Huaishao Luo, Junwei Bao, Youzheng Wu, Xiaodong He, and Tianrui Li. Segclip: Patch aggregation with learn- able centers for open-vocabulary semantic segmentation. In ICML, pages 23033–23044. PMLR, 2023. 3
2023
-
[45]
Segment anything model for medical image analysis: an experimental study
Maciej A Mazurowski, Haoyu Dong, Hanxue Gu, Jichen Yang, Nicholas Konz, and Yixin Zhang. Segment anything model for medical image analysis: an experimental study. Medical Image Analysis, 89:102918, 2023. 2
2023
-
[46]
MetaICL: Learning to learn in context
Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. MetaICL: Learning to learn in context. In NAACL-HLT, 2022. 5
2022
-
[47]
The role of context for object detection and se- mantic segmentation in the wild
Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, and Alan Yuille. The role of context for object detection and se- mantic segmentation in the wild. In CVPR, pages 891–898,
-
[48]
Peters, Mark Neumann, Mohit Iyyer, Matt Gard- ner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gard- ner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Hu...
2018
-
[49]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In ICML, 2021. 1, 2, 4, 7
2021
-
[50]
Improving language understanding by gen- erative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by gen- erative pre-training. Technical report, OpenAI, 2018. 2
2018
-
[51]
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67, 2020. 2
2020
-
[52]
Per- ceptual grouping in contrastive vision-language models
Kanchana Ranasinghe, Brandon McKinzie, Sachin Ravi, Yinfei Yang, Alexander Toshev, and Jonathon Shlens. Per- ceptual grouping in contrastive vision-language models. In ICCV, pages 5571–5584, 2023. 1, 3
2023
-
[53]
Laion-5b: An open large-scale dataset for train- ing next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, et al. Laion-5b: An open large-scale dataset for train- ing next generation image-text models. NeurIPS, 35:25278– 252...
2022
-
[54]
Ex- plore the potential of clip for training-free open vocabulary semantic segmentation
Tong Shao, Zhuotao Tian, Hang Zhao, and Jingyong Su. Ex- plore the potential of clip for training-free open vocabulary semantic segmentation. In ECCV. Springer, 2024. 1
2024
-
[55]
Edadet: Open-vocabulary ob- ject detection using early dense alignment
Cheng Shi and Sibei Yang. Edadet: Open-vocabulary ob- ject detection using early dense alignment. In ICCV, pages 15724–15734, 2023. 1
2023
-
[56]
Reco: Retrieve and co-segment for zero-shot transfer
Gyungin Shin, Weidi Xie, and Samuel Albanie. Reco: Retrieve and co-segment for zero-shot transfer. NeurIPS, 35:33754–33767, 2022. 1, 6, 7
2022
-
[57]
Going denser with open-vocabulary part segmentation
Peize Sun, Shoufa Chen, Chenchen Zhu, Fanyi Xiao, Ping 10 Luo, Saining Xie, and Zhicheng Yan. Going denser with open-vocabulary part segmentation. In ICCV, pages 15453– 15465, 2023. 1
2023
-
[58]
Clip as rnn: Segment countless visual concepts without training endeavor
Shuyang Sun, Runjia Li, Philip Torr, Xiuye Gu, and Siyang Li. Clip as rnn: Segment countless visual concepts without training endeavor. In CVPR, pages 13171–13182, 2024. 2, 3, 6, 7, 12
2024
-
[59]
Galip: Generative adversarial clips for text-to-image synthe- sis
Ming Tao, Bing-Kun Bao, Hao Tang, and Changsheng Xu. Galip: Generative adversarial clips for text-to-image synthe- sis. In CVPR, pages 14214–14223, 2023. 1
2023
-
[60]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. 5
2017
-
[61]
Sclip: Rethink- ing self-attention for dense vision-language inference
Feng Wang, Jieru Mei, and Alan Yuille. Sclip: Rethink- ing self-attention for dense vision-language inference. arXiv preprint arXiv:2312.01597, 2023. 1, 2, 4, 6, 7, 12, 13
2023 arXiv
-
[62]
Clip-gen: Language-free training of a text-to-image genera- tor with clip
Zihao Wang, Wei Liu, Qian He, Xinglong Wu, and Zili Yi. Clip-gen: Language-free training of a text-to-image genera- tor with clip. arXiv preprint arXiv:2203.00386, 2022. 1
2022 arXiv
-
[63]
Clip-diy: Clip dense infer- ence yields open-vocabulary semantic segmentation for-free
Monika Wysocza ´nska, Micha ¨el Ramamonjisoa, Tomasz Trzci´nski, and Oriane Sim ´eoni. Clip-diy: Clip dense infer- ence yields open-vocabulary semantic segmentation for-free. In WACV, pages 1403–1413, 2024. 2, 3, 5, 6, 7
2024
-
[64]
Demystify- ing clip data
Hu Xu, Saining Xie, Xiaoqing Ellen Tan, Po-Yao Huang, Russell Howes, Vasu Sharma, Shang-Wen Li, Gargi Ghosh, Luke Zettlemoyer, and Christoph Feichtenhofer. Demystify- ing clip data. arXiv preprint arXiv:2309.16671, 2023. 2
2023 arXiv
-
[65]
Groupvit: Semantic segmentation emerges from text supervision
Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang. Groupvit: Semantic segmentation emerges from text supervision. In CVPR, pages 18134–18144, 2022. 3, 6, 7
2022
-
[66]
Learning open-vocabulary semantic segmentation models from natural language supervision
Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng, Yi Wang, Yu Qiao, and Weidi Xie. Learning open-vocabulary semantic segmentation models from natural language supervision. In CVPR, pages 2935–2944, 2023. 3
2023
-
[67]
Open-vocabulary panop- tic segmentation with text-to-image diffusion models
Jiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon, Xiao- long Wang, and Shalini De Mello. Open-vocabulary panop- tic segmentation with text-to-image diffusion models. In CVPR, pages 2955–2966, 2023. 3
2023
-
[68]
A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- language model
Mengde Xu, Zheng Zhang, Fangyun Wei, Yutong Lin, Yue Cao, Han Hu, and Xiang Bai. A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- language model. In ECCV. Springer, 2022. 2
2022
-
[69]
Learning deep features for discrimi- native localization
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discrimi- native localization. In CVPR, pages 2921–2929, 2016. 3
2016
-
[70]
Extract free dense labels from clip
Chong Zhou, Chen Change Loy, and Bo Dai. Extract free dense labels from clip. In ECCV, pages 696–712. Springer,
-
[71]
thing” classes and one explicit background class, our method fails to distinguish foreground classes from the background when the word “background
Ziqin Zhou, Yinjie Lei, Bowen Zhang, Lingqiao Liu, and Yifan Liu. Zegclip: Towards adapting clip for zero-shot se- mantic segmentation. In CVPR, pages 11175–11185, 2023. 1, 2 11 A. Appendix A.1. Additional Implementation Details LLM-based Auxiliary Text Generation. As detailed...
2023
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.