Pith. sign in

REVIEW 59 references

A Holistically Point-guided Text Framework for Weakly-Supervised Camouflaged Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.06038 v1 pith:OCIGGZWY submitted 2025-01-10 cs.CV

classification cs.CV
keywords detectionobjecttextcamouflagedchoosemaskmethodspoint-guided
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Weakly-Supervised Camouflaged Object Detection (WSCOD) has gained popularity for its promise to train models with weak labels to segment objects that visually blend into their surroundings. Recently, some methods using sparsely-annotated supervision shown promising results through scribbling in WSCOD, while point-text supervision remains underexplored. Hence, this paper introduces a novel holistically point-guided text framework for WSCOD by decomposing into three phases: segment, choose, train. Specifically, we propose Point-guided Candidate Generation (PCG), where the point's foreground serves as a correction for the text path to explicitly correct and rejuvenate the loss detection object during the mask generation process (SEGMENT). We also introduce a Qualified Candidate Discriminator (QCD) to choose the optimal mask from a given text prompt using CLIP (CHOOSE), and employ the chosen pseudo mask for training with a self-supervised Vision Transformer (TRAIN). Additionally, we developed a new point-supervised dataset (P2C-COD) and a text-supervised dataset (T-COD). Comprehensive experiments on four benchmark datasets demonstrate our method outperforms state-of-the-art methods by a large margin, and also outperforms some existing fully-supervised camouflaged object detection methods.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

59 extracted references · 9 linked inside Pith

  1. [1]

    Frequency-tuned salient region detection

    Radhakrishna Achanta, Sheila Hemami, Francisco Estrada, and Sabine S¨ usstrunk. Frequency-tuned salient region detection. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition , number CONF, pages 1597–1604, 2009

  2. [2]

    Emerging properties in self-supervised vision transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Herv´ e J´ egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021

  3. [3]

    Sam-adapter: Adapt- ing segment anything in underperformed scenes

    Tianrun Chen, Lanyun Zhu, Chaotao Deng, Runlong Cao, Yan Wang, Shangzhan Zhang, Zejian Li, Lingyun Sun, Ying Zang, and Papa Mao. Sam-adapter: Adapt- ing segment anything in underperformed scenes. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 3367–3375, 2023

  4. [4]

    Weakly-supervised semantic segmentation with image-level labels: from traditional models to foundation models

    Zhaozheng Chen and Qianru Sun. Weakly-supervised semantic segmentation with image-level labels: from traditional models to foundation models. arXiv preprint arXiv:2310.13026, 2023

  5. [5]

    Out-of-candidate rectification for weakly supervised semantic segmentation

    Zesen Cheng, Pengchong Qiao, Kehan Li, Siheng Li, Pengxu Wei, Xiangyang Ji, Li Yuan, Chang Liu, and Jie Chen. Out-of-candidate rectification for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 23673–23684, 2023

  6. [6]

    Camouflage images

    Hung-Kuo Chu, Wei-Hsin Hsu, Niloy J Mitra, Daniel Cohen-Or, Tien-Tsin Wong, and Tong-Yee Lee. Camouflage images. ACM Trans. Graph., 29(4):51–1, 2010

  7. [7]

    A weakly supervised learning framework for salient object detec- tion via hybrid labels

    Runmin Cong, Qi Qin, Chen Zhang, Qiuping Jiang, Shiqi Wang, Yao Zhao, and Sam Kwong. A weakly supervised learning framework for salient object detec- tion via hybrid labels. IEEE Transactions on Circuits and Systems for Video Technology, 33(2):534–548, 2022

  8. [8]

    A tutorial on the cross-entropy method

    Pieter-Tjerk De Boer, Dirk P Kroese, Shie Mannor, and Reuven Y Rubinstein. A tutorial on the cross-entropy method. Annals of operations research , 134:19–67, 2005

Show all 59 references
  1. [9]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiao- hua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arX...

  2. [10]

    Structure- measure: A new way to evaluate foreground maps

    Deng-Ping Fan, Ming-Ming Cheng, Yun Liu, Tao Li, and Ali Borji. Structure- measure: A new way to evaluate foreground maps. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 4548–4557, 2017

  3. [11]

    Enhanced-alignment measure for binary foreground map evaluation

    Deng-Ping Fan, Cheng Gong, Yang Cao, Bo Ren, Ming-Ming Cheng, and Ali Borji. Enhanced-alignment measure for binary foreground map evaluation. In IJCAI, pages 698–704, 2018

  4. [12]

    Concealed 22 object detection

    Deng-Ping Fan, Ge-Peng Ji, Ming-Ming Cheng, and Ling Shao. Concealed 22 object detection. IEEE transactions on pattern analysis and machine intelligence, 44(10):6024–6042, 2021

  5. [13]

    Cognitive vision inspired object segmentation metric and loss function

    Deng-Ping Fan, Ge-Peng Ji, Xuebin Qin, and Ming-Ming Cheng. Cognitive vision inspired object segmentation metric and loss function. Scientia Sinica Informationis, 6(6), 2021

  6. [14]

    Camouflaged object detection

    Deng-Ping Fan, Ge-Peng Ji, Guolei Sun, Ming-Ming Cheng, Jianbing Shen, and Ling Shao. Camouflaged object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 2777–2787, 2020

  7. [15]

    Pranet: Parallel reverse attention network for polyp segmentation

    Deng-Ping Fan, Ge-Peng Ji, Tao Zhou, Geng Chen, Huazhu Fu, Jianbing Shen, and Ling Shao. Pranet: Parallel reverse attention network for polyp segmentation. In International conference on medical image computing and computer-assisted intervention, pages 263–273. Springer, 2020

  8. [16]

    Inf-net: Automatic covid-19 lung infection segmentation from ct images

    Deng-Ping Fan, Tao Zhou, Ge-Peng Ji, Yi Zhou, Geng Chen, Huazhu Fu, Jianbing Shen, and Ling Shao. Inf-net: Automatic covid-19 lung infection segmentation from ct images. IEEE transactions on medical imaging , 39(8):2626–2637, 2020

  9. [17]

    Weakly-supervised salient object detection using point supervision

    Shuyong Gao, Wei Zhang, Yan Wang, Qianyu Guo, Chenglong Zhang, Yangji He, and Wenqiang Zhang. Weakly-supervised salient object detection using point supervision. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 36, pages 670–678, 2022

  10. [18]

    Image editing by object-aware optimal boundary searching and mixed-domain composition

    Shiming Ge, Xin Jin, Qiting Ye, Zhao Luo, and Qiang Li. Image editing by object-aware optimal boundary searching and mixed-domain composition. Computational Visual Media , 4:71–82, 2018

  11. [19]

    Box-based refinement for weakly supervised and unsupervised localization tasks

    Eyal Gomel, Tal Shaharbany, and Lior Wolf. Box-based refinement for weakly supervised and unsupervised localization tasks. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 16044–16054, 2023

  12. [20]

    Camouflaged object detection with feature decomposition and edge reconstruction

    Chunming He, Kai Li, Yachao Zhang, Longxiang Tang, Yulun Zhang, Zhenhua Guo, and Xiu Li. Camouflaged object detection with feature decomposition and edge reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 22046–22055, 2023

  13. [21]

    Weakly-supervised concealed object segmentation with sam-based pseudo labeling and multi-scale feature grouping

    Chunming He, Kai Li, Yachao Zhang, Guoxia Xu, Longxiang Tang, Yulun Zhang, Zhenhua Guo, and Xiu Li. Weakly-supervised concealed object segmentation with sam-based pseudo labeling and multi-scale feature grouping. NeurIPS, 2023

  14. [22]

    Strategic preys make acute predators: Enhancing camouflaged object detectors by generating camouflaged objects

    Chunming He, Kai Li, Yachao Zhang, Yulun Zhang, Zhenhua Guo, Xiu Li, Mar- tin Danelljan, and Fisher Yu. Strategic preys make acute predators: Enhancing camouflaged object detectors by generating camouflaged objects. arXiv preprint arXiv:2308.03166, 2023

  15. [23]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016

  16. [24]

    Weakly-supervised camouflaged object detection with scribble annotations

    Ruozhen He, Qihua Dong, Jiaying Lin, and Rynson WH Lau. Weakly-supervised camouflaged object detection with scribble annotations. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 37, pages 781–789, 2023

  17. [25]

    Relax image-specific prompt requirement in sam: A single generic prompt for segmenting camouflaged objects

    Jian Hu, Jiayi Lin, Weitong Cai, and Shaogang Gong. Relax image-specific prompt requirement in sam: A single generic prompt for segmenting camouflaged objects. arXiv preprint arXiv:2312.07374 , 2023. 23

  18. [26]

    Scribble-based boundary-aware network for weakly supervised salient object detection in remote sensing images

    Zhou Huang, Tian-Zhu Xiang, Huai-Xin Chen, and Hang Dai. Scribble-based boundary-aware network for weakly supervised salient object detection in remote sensing images. ISPRS Journal of Photogrammetry and Remote Sensing, 191:290– 301, 2022

  19. [27]

    The distribution of the flora in the alpine zone

    Paul Jaccard. The distribution of the flora in the alpine zone. 1. New phytologist, 11(2):37–50, 1912

  20. [28]

    Sam struggles in concealed scenes–empirical study on” segment anything”

    Ge-Peng Ji, Deng-Ping Fan, Peng Xu, Ming-Ming Cheng, Bowen Zhou, and Luc Van Gool. Sam struggles in concealed scenes–empirical study on” segment anything”. arXiv preprint arXiv:2304.06022 , 2023

  21. [29]

    Segment, magnify and reiterate: Detecting camouflaged objects the hard way

    Qi Jia, Shuilian Yao, Yu Liu, Xin Fan, Risheng Liu, and Zhongxuan Luo. Segment, magnify and reiterate: Detecting camouflaged objects the hard way. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 4713–4722, 2022

  22. [30]

    Distilling self- supervised vision transformers for weakly-supervised few-shot classification & segmentation

    Dahyun Kang, Piotr Koniusz, Minsu Cho, and Naila Murray. Distilling self- supervised vision transformers for weakly-supervised few-shot classification & segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19627–19638, 2023

  23. [31]

    Segment anything

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. arXiv preprint arXiv:2304.02643 , 2023

  24. [32]

    Anabranch network for camouflaged object segmentation

    Trung-Nghia Le, Tam V Nguyen, Zhongliang Nie, Minh-Triet Tran, and Akihiro Sugimoto. Anabranch network for camouflaged object segmentation. Computer vision and image understanding , 184:45–56, 2019

  25. [33]

    Weakly-supervised salient object detection on light fields

    Zijian Liang, Pengjie Wang, Ke Xu, Pingping Zhang, and Rynson WH Lau. Weakly-supervised salient object detection on light fields. IEEE Transactions on Image Processing, 31:6295–6305, 2022

  26. [34]

    A metaheuristic-based approach to optimizing color design for military camouflage using particle swarm optimization

    Chiuhsiang Joe Lin and Yogi Tri Prasetyo. A metaheuristic-based approach to optimizing color design for military camouflage using particle swarm optimization. Color Research & Application , 44(5):740–748, 2019

  27. [35]

    Poolnet+: Exploring the potential of pooling for salient object detection

    Jiang-Jiang Liu, Qibin Hou, Zhi-Ang Liu, and Ming-Ming Cheng. Poolnet+: Exploring the potential of pooling for salient object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence , 45(1):887–904, 2022

  28. [36]

    Grounding dino: Marrying dino with grounded pre-training for open-set object detection

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chun- yuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023

  29. [37]

    Mscaf-net: a general framework for camouflaged object detection via learning multi-scale context-aware features

    Yu Liu, Haihang Li, Juan Cheng, and Xun Chen. Mscaf-net: a general framework for camouflaged object detection via learning multi-scale context-aware features. IEEE Transactions on Circuits and Systems for Video Technology , 2023

  30. [38]

    Fixing weight decay regularization in adam

    Ilya Loshchilov and Frank Hutter. Fixing weight decay regularization in adam. ArXiv, 5, 2017

  31. [39]

    Weakly-supervised contrastive learning for unsupervised object discovery

    Yunqiu Lv, Jing Zhang, Nick Barnes, and Yuchao Dai. Weakly-supervised contrastive learning for unsupervised object discovery. arXiv preprint arXiv:2307.03376, 2023

  32. [40]

    Simultaneously localize, segment and rank the camouflaged objects

    Yunqiu Lv, Jing Zhang, Yuchao Dai, Aixuan Li, Bowen Liu, Nick Barnes, and 24 Deng-Ping Fan. Simultaneously localize, segment and rank the camouflaged objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11591–11601, 2021

  33. [41]

    Boosting broader receptive fields for salient object detection

    Mingcan Ma, Changqun Xia, Chenxi Xie, Xiaowu Chen, and Jia Li. Boosting broader receptive fields for salient object detection. IEEE Transactions on Image Processing, 32:1026–1038, 2023

  34. [42]

    How to evaluate foreground maps? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 248–255, 2014

    Ran Margolin, Lihi Zelnik-Manor, and Ayellet Tal. How to evaluate foreground maps? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 248–255, 2014

  35. [43]

    Deeproadmapper: Extract- ing road topology from aerial images

    Gell´ ert M´ attyus, Wenjie Luo, and Raquel Urtasun. Deeproadmapper: Extract- ing road topology from aerial images. In Proceedings of the IEEE international conference on computer vision , pages 3438–3446, 2017

  36. [44]

    Camouflaged object segmentation with distraction mining

    Haiyang Mei, Ge-Peng Ji, Ziqi Wei, Xin Yang, Xiaopeng Wei, and Deng-Ping Fan. Camouflaged object segmentation with distraction mining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 8772–8781, 2021

  37. [45]

    Discriminative sampling of proposals in self-supervised transformers for weakly supervised object localization

    Shakeeb Murtaza, Soufiane Belharbi, Marco Pedersoli, Aydin Sarraf, and Eric Granger. Discriminative sampling of proposals in self-supervised transformers for weakly supervised object localization. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vis...

  38. [46]

    Zoom in and out: A mixed-scale triplet network for camouflaged object detec- tion

    Youwei Pang, Xiaoqi Zhao, Tian-Zhu Xiang, Lihe Zhang, and Huchuan Lu. Zoom in and out: A mixed-scale triplet network for camouflaged object detec- tion. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pages 2160–2170, 2022

  39. [47]

    Saliency filters: Contrast based filtering for salient region detection

    Federico Perazzi, Philipp Kr¨ ahenb¨ uhl, Yael Pritch, and Alexander Hornung. Saliency filters: Contrast based filtering for salient region detection. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 733–740, 2012

  40. [48]

    Early evolution and ecology of camouflage in insects

    Ricardo P´ erez-de la Fuente, Xavier Delcl` os, Enrique Pe˜ nalver, Mariela Speranza, Jacek Wierzchos, Carmen Ascaso, and Michael S Engel. Early evolution and ecology of camouflage in insects. Proceedings of the National Academy of Sciences, 109(52):21414–21419, 2012

  41. [49]

    Mfnet: Multi-filter direc- tive network for weakly supervised salient object detection

    Yongri Piao, Jian Wang, Miao Zhang, and Huchuan Lu. Mfnet: Multi-filter direc- tive network for weakly supervised salient object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 4136–4145, 2021

  42. [50]

    Basnet: Boundary-aware salient object detection

    Xuebin Qin, Zichen Zhang, Chenyang Huang, Chao Gao, Masood Dehghan, and Martin Jagersand. Basnet: Boundary-aware salient object detection. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition , pages 7479–7489, 2019

  43. [51]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning , p...

  44. [52]

    Optimizing intersection-over-union in deep 25 neural networks for image segmentation

    Md Atiqur Rahman and Yang Wang. Optimizing intersection-over-union in deep 25 neural networks for image segmentation. In International symposium on visual computing, pages 234–244. Springer, 2016

  45. [53]

    Semantic segmentation using foundation models for cultural heritage: an experimental study on notre-dame de paris

    K´ evin R´ eby, Ana ¨ ıs Guilhelm, and Livio De Luca. Semantic segmentation using foundation models for cultural heritage: an experimental study on notre-dame de paris. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1689–1697, 2023

  46. [54]

    Boundary-enhanced co- training for weakly supervised semantic segmentation

    Shenghai Rong, Bohai Tu, Zilei Wang, and Junjie Li. Boundary-enhanced co- training for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 19574–19584, 2023

  47. [55]

    Token contrast for weakly- supervised semantic segmentation

    Lixiang Ru, Heliang Zheng, Yibing Zhan, and Bo Du. Token contrast for weakly- supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 3093–3102, 2023

  48. [56]

    Imagenet large scale visual recognition challenge

    Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115:211–252, 2015

  49. [57]

    Application of an image and environmental sensor network for automated greenhouse insect pest monitoring.Journal of Asia-Pacific Entomology, 23(1):17–28, 2020

    Dan Jeric Arcega Rustia, Chien Erh Lin, Jui-Yung Chung, Yi-Ji Zhuang, Ju- Chun Hsu, and Ta-Te Lin. Application of an image and environmental sensor network for automated greenhouse insect pest monitoring.Journal of Asia-Pacific Entomology, 23(1):17–28, 2020

  50. [58]

    Similarity maps for self-training weakly- supervised phrase grounding

    Tal Shaharabany and Lior Wolf. Similarity maps for self-training weakly- supervised phrase grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 6925–6934, 2023

  51. [59]

    What does clip know about a red circle? visual prompt engineering for vlms

    Aleksandar Shtedritski, Christian Rupprecht, and Andrea Vedaldi. What does clip know about a red circle? visual prompt engineering for vlms. arXiv preprint arXiv:2304.06712, 2023

Pith tools