REVIEW 4 major objections 5 minor 99 references
Rethinking Detecting Salient and Camouflaged Objects in Unconstrained Scenes
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper claims that SOD/COD models fail in unconstrained scenes because datasets assume scenes contain either salient or camouflaged objects but never both, and that a new dataset (USC12K), a new model (USCNet), and a new metric (CSCS)…
desk verdict A genuinely useful four-scene dataset with a plausible SOTA method; the main gaps are missing annotation agreement stats and unreleased artifacts, not the SAM-annotation circularity that dominates the stress-test note. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing pieces are threefold. First, USC12K: a 12,000-image benchmark with four balanced scene types (only salient, only camouflaged, both, neither) and labels for the three attributes of saliency, camouflage, and background, built from existing DUTS, HKU-IS, COD10K, and CAMO datasets plus newly collected internet and underwater images. Second, the Attribute Relation Modeling (ARM) module inside USCNet: it generates three attribute prompts (saliency, camouflage, background) by combining Inter-SPQ (learnable cross-sample queries) and Intra-SPQ (per-sample queries computed as $[Q_{Sa}, Q_{Ca}, Q_{Ba}] = \mathrm{Linear}(\sigma(\Phi_{AH}(F)) \otimes F)$), then applies self-attention and query-to-image attention to produce the prompts that feed a frozen SAM mask decoder; this is the mechanism that explicitly models how the two attributes relate both across samples and within a single image. Third, the Camouflage-Saliency Confusion Score, $\mathrm{CSCS} = \frac{1}{2}\left(\frac{P_{CS}}{P_{BS}+P_{SS}+P_{CS}} + \frac{P_{SC}}{P_{BC}+P_{SC}+P_{CC}}\right)$, which quantifies how often camouflaged pixels are predicted as salient and vice versa, a quantity that existing foreground/background metrics such as mIoU and weighted F-measure do not isolate.
What would settle it
One concrete experiment that would settle the central claim is to re-annotate a stratified random sample of the 3,000 Scene C images and the 1,436 internet Scene D images with independent expert annotators and compute per-pixel inter-annotator agreement (e.g., Cohen's kappa) against the released labels; if agreement is low, or if the original labels are systematically biased toward 'salient' or 'camouflaged' in ambiguous cases, the USC12K ground truth is unstable and the reported CSCS ranking of USCNet over SAM2-Adapter could be an artifact of that label noise rather than a genuine modeling gain.
Extended reading notes
Core claim
On its own terms, the central discovery is that the mutual-exclusivity annotation paradigm of existing SOD/COD datasets—labeling a scene either salient or camouflaged, but never both and never neither—systematically corrupts both single-task and unified detectors, and that this can be fixed by two complementary contributions. First, a dataset, USC12K, whose four balanced scenes (A: only salient, B: only camouflaged, C: both, D: neither) carry full pixel-level masks for all three attributes and whose Scene C contains both salient and camouflaged objects in the same image. Second, a model, USCNet, that explicitly models the relationship between the two attributes through an Attribute Relation Modeling module using Inter-Sample Prompt Query (learnable, sample-invariant queries that capture generic attribute difference) and Intra-Sample Prompt Query (per-image queries derived from an attention-weighted feature map that capture within-sample context), built on a SAM/SAM2 encoder with adapters and a frozen mask decoder. The paper reports that USCNet achieves state-of-the-art performance across all four scenes and all metrics on the USC12K benchmark, and that training on USC12K sharply reduces the cross-task misinterpretation scores shown in Table 1, for example a SOD model's F_beta score on COD10K dropping from 0.6384 to 0.0146.
Load-bearing premise
The label quality of the newly internet-collected images (2,617 'both' and 1,436 'neither' images) is asserted through a voting and refinement process without reported inter-annotator agreement statistics; if those labels are noisy, systematically biased, or inconsistent with existing SOD/COD label conventions, the benchmark comparisons and the conclusion that USCNet achieves state-of-the-art results would not transfer to real unconstrained scenes.
Editorial extensions
If this is right
- If USC12K is adopted as a training set, SOD and COD models should no longer exhibit the cross-task misinterpretation shown in Table 1; the paper reports that retraining baselines on USC12K drops their false-positive F_beta scores on the opposite task from values like 0.6384 to below 0.1.
- In scenes containing both salient and camouflaged objects (Scene C), explicit inter- and intra-sample attribute modeling yields the largest gains; the ablation shows that both Intra-SPQ and Inter-SPQ improve IoU and mIoU and lower CSCS, with Intra-SPQ contributing more.
- Because USCNet attaches the ARM module to a frozen SAM2 mask decoder with only 4.04 million tunable parameters, the claimed state-of-the-art results come with parameter-efficient fine-tuning, suggesting the approach can be ported to other large vision backbones.
- The proposed CSCS metric provides a direct way to compare detectors on salient-versus-camouflaged discrimination, complementing standard foreground/background metrics and giving the community a concrete target for reducing attribute confusion in unconstrained scenes.
Reading between the lines
- Editorial inference: the mutual-exclusivity bias probably affects other binary foreground tasks beyond SOD/COD, such as defect detection or anomaly segmentation, so the four-scene design of USC12K could be transferred to those domains to test whether models similarly confuse opposing attributes.
- Editorial inference: the Scene D background-only images, which are often trivial for standard metrics, may be the most important test for real deployment; a detector that reports false positives on them would be unsafe, and the CSCS-style evaluation could be extended to measure background-foreground false alarms in addition to attribute confusion.
- Editorial inference: a minimal attribute-switch experiment could be run with USCNet—keep the encoder and decoder fixed, swap only the saliency and camouflage attribute prompts; if the prompts truly carry attribute information, the same image should produce the opposite mask, which would isolate whether the model disentangles attributes rather than learning a single foreground detector.
- Editorial inference: if the annotation-quality assumption holds, retraining existing SOD and COD models on USC12K should produce larger improvements than architecture changes alone, a prediction that the paper's Table 4 already supports for a few baselines and that could be tested on a broader set of models.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims that existing SOD/COD datasets and models enforce a mutual-exclusivity constraint on salient and camouflaged objects, which causes cross-task misinterpretation. To address this, the authors introduce USC12K, a 12,000-image dataset with four scene types covering all logical combinations of salient and camouflaged objects (salient-only, camouflaged-only, both, neither), propose a SAM2-based model USCNet with an Attribute Relation Modeling (ARM) module that uses inter-sample and intra-sample prompt queries, and design a new evaluation metric, the Camouflage-Saliency Confusion Score (CSCS). They retrain 21 SOD/COD/unified methods on USC12K, report that USCNet achieves state-of-the-art performance across all scenes and metrics, and demonstrate that training on USC12K reduces cross-task false positives.
Significance. If validated, USC12K would be a valuable new benchmark for unconstrained salient/camouflaged object detection, and the CSCS metric addresses a real gap in evaluating the confusion between these two attributes. The ARM module is a sensible approach for explicitly modeling attribute relationships, and the authors provide a fairly extensive comparison, including generalization results on six standard SOD/COD datasets. The paper also ships promising reproducibility artifacts: it promises to release code and data, and the benchmark protocol is described in detail. The main risks are that the Scene C ground truth is partly SAM-derived without quantified annotation agreement, and the scene definitions are internally inconsistent; these issues must be resolved before the benchmark and SOTA claims can be fully trusted.
major comments (4)
- [Section 3.1 vs. Figure 3 / Table 3] Section 3.1 defines Scene A as containing only salient objects and Scene B as containing only camouflaged objects, but Figure 3 captions these scenes in the opposite order, and the numbers in Table 3 (e.g., high IoUC values in Scene A and high IoUS values in Scene B) are only interpretable under the swapped definitions. This internal inconsistency must be resolved for the benchmark results to be meaningful.
- [Section 3.2 and Appendix §5] The authors state that Scene C annotations are produced using SAM for coarse annotation followed by manual refinement, and the appendix describes the ISAT tool with SAM semi-automatic labeling. No inter-annotator agreement statistics or quantitative measure of the refinement effort are provided. Because USCNet uses SAM2 as its backbone and a frozen SAM mask decoder, the Scene C benchmark may be biased in favor of SAM-based models. Please provide independent validation, such as pixel-level agreement among annotators, the distribution of manual corrections, and evaluation on a subset with fully manual masks.
- [Tables 3, 5, 7–10] All quantitative comparisons report single-run point estimates without variance, and several margins over the second-best method are small (e.g., Table 3, Scene A IoUS: USCNet 79.70 vs. SAM2-Adapter 78.75). Without multiple seeds or significance tests, the claim of state-of-the-art performance across all scenes is not robust.
- [Appendix §4] The benchmark protocol retools every conventional SOD/COD model's output layer into a three-class softmax and modifies the training procedure for unified models (using two copies of the dataset for VSCode and EVP). These adaptations may disadvantage the competitors relative to USCNet, so the authors should justify their neutrality or supplement the benchmark with baselines that use the models' native binary output heads plus a post-hoc background class.
minor comments (5)
- [Eq. (6) / Eq. (7)] The CSCS denominator can vanish when no pixels are predicted as salient or as camouflaged; please specify the convention for such degenerate cases.
- [Section 5] The text says 21 methods are compared, but the enumerated list omits PGNet, which appears in Table 3; please ensure the count and the list are consistent.
- [Figure 3] The caption of Figure 3 reverses the definitions of Scene A and Scene B relative to Section 3.1; beyond the major comment, please ensure all captions use consistent scene definitions.
- [Section 4.2] The term 'inter-sample' for the Inter-SPQ queries is potentially misleading, because the queries are fixed at inference and do not depend on other samples; consider renaming them 'global' or 'dataset-level' queries.
- [Table 3] The column labeled 'Update' is unexplained; please clarify whether it indicates that all models are retrained on USC12K or some other setting.
Circularity Check
No demonstrated circularity: the benchmark, held-out evaluation, and external generalization tests are self-contained; only a mild SAM-based annotation fairness concern remains.
full rationale
The paper's main claims are empirical rather than derived from their own definitions. USCNet is trained on the USC12K training split (8,400 images) and evaluated on the held-out test split (3,600 images), while all 21 compared methods are retrained under the same protocol, so the reported SOTA numbers are independent measured outcomes. The CSCS metric (Eq. 6) is used only for evaluation and does not appear in the loss (Eqs. 4-5); the focal-loss alpha values (1:4:6) are class-balance weights based on pixel-count ratios, not fitted parameters that encode the target result. External generalization results on DUTS, HKU-IS, NC4K, and COD10K further show that the model is not merely tuned to the new benchmark. The self-citations in Section 2.2 (refs. [29,51,52]) are survey-level mentions and are not load-bearing for the architecture, the dataset, or the SOTA claim. The only substantive concern is the Scene C annotation pipeline: Section 3.2 states 'we use SAM for coarse annotation, followed by manual refinement,' and Appendix §5 describes ISAT with SAM semi-automatic labeling plus review by 3 observers, but no inter-annotator agreement statistics are reported. Because USCNet is also SAM/SAM2-based, this raises a real question about label independence. However, the paper does not assert that the ground truth equals SAM outputs, and manual refinement and observer review are stated steps. Without evidence that refinement was trivial, claiming the labels reduce to SAM priors would be speculation rather than an exhibited reduction. Therefore no specific circular step is established; the score of 2 reflects the mild, non-load-bearing self-citation and the unresolved annotation-fairness concern, not demonstrated circularity.
Assumptions & free parameters
free parameters (4)
- Focal loss weighting factors alpha_ti =
1:4:6 (background:salient:camouflaged)
- Loss balance weights lambda_p and lambda_a =
1 and 0.5
- Number and dimensionality of attribute queries (N, C) =
not reported
- Adapter hidden dimensions and placement =
not reported
assumptions (4)
- domain assumption The four logical scenes (salient-only, camouflaged-only, both, neither) are exhaustive for unconstrained detection.
- domain assumption Manual annotation via voting and refinement yields accurate ground truth for saliency and camouflage in Scene C and D.
- domain assumption SAM/SAM2 image features, adapted with lightweight modules, provide a sufficient representation for the combined SOD and COD task.
- domain assumption The ternary output space (background, salient, camouflaged) is the appropriate formulation for unconstrained scenes.
Cite this review
Pith. "Pith review of Rethinking Detecting Salient and Camouflaged Objects in Unconstrained Scenes." pith.science (2026). https://pith.science/paper/2E35S5OG
@misc{pith2026241210943,
author = {Pith},
title = {Pith review of: Rethinking Detecting Salient and Camouflaged Objects in Unconstrained Scenes},
year = {2026},
howpublished = {\url{https://pith.science/paper/2E35S5OG}},
note = {Machine review of arXiv:2412.10943}
}
read the original abstract
While the human visual system employs distinct mechanisms to perceive salient and camouflaged objects, existing models struggle to disentangle these tasks. Specifically, salient object detection (SOD) models frequently misclassify camouflaged objects as salient, while camouflaged object detection (COD) models conversely misinterpret salient objects as camouflaged. We hypothesize that this can be attributed to two factors: (i) the specific annotation paradigm of current SOD and COD datasets, and (ii) the lack of explicit attribute relationship modeling in current models. Prevalent SOD/COD datasets enforce a mutual exclusivity constraint, assuming scenes contain either salient or camouflaged objects, which poorly aligns with the real world. Furthermore, current SOD/COD methods are primarily designed for these highly constrained datasets and lack explicit modeling of the relationship between salient and camouflaged objects. In this paper, to promote the development of unconstrained salient and camouflaged object detection, we construct a large-scale dataset, USC12K, which features comprehensive labels and four different scenes that cover all possible logical existence scenarios of both salient and camouflaged objects. To explicitly model the relationship between salient and camouflaged objects, we propose a model called USCNet, which introduces two distinct prompt query mechanisms for modeling inter-sample and intra-sample attribute relationships. Additionally, to assess the model's ability to distinguish between salient and camouflaged objects, we design an evaluation metric called CSCS. The proposed method achieves state-of-the-art performance across all scenes in various metrics. The code and dataset will be available at https://github.com/ssecv/USCNet.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
Frequency-tuned salient region detec- tion
Radhakrishna Achanta, Sheila Hemami, Francisco Estrada, and Sabine Susstrunk. Frequency-tuned salient region detec- tion. In CVPR, pages 1597–1604, 2009. 12
2009
-
[2]
Sam-adapter: Adapting segment any- thing in underperformed scenes
Tianrun Chen, Lanyun Zhu, Chaotao Ding, Runlong Cao, Yan Wang, Shangzhan Zhang, Zejian Li, Lingyun Sun, Ying Zang, and Papa Mao. Sam-adapter: Adapting segment any- thing in underperformed scenes. In ICCVW, pages 3359– 3367, 2023. 3, 5, 6, 7, 14
2023
-
[3]
Tianrun Chen, Ankang Lu, Lanyun Zhu, Chaotao Ding, Chu- nan Yu, Deyi Ji, Zejian Li, Lingyun Sun, Papa Mao, and Ying Zang. Sam2-adapter: Evaluating & adapting seg- ment anything 2 in downstream tasks: Camouflage, shadow, medical image segmentation, and more. arXiv preprint arXiv:2408.04579, 2024. 3, 6, 7, 13, 14
arXiv 2024
-
[4]
Camodiffusion: Camouflaged object detection via conditional diffusion mod- els
Zhongxi Chen, Ke Sun, and Xianming Lin. Camodiffusion: Camouflaged object detection via conditional diffusion mod- els. In AAAI, pages 1272–1280, 2024. 6, 7
2024
-
[5]
Mitra, Xiaolei Huang, Philip H
Ming-Ming Cheng, Niloy J. Mitra, Xiaolei Huang, Philip H. S. Torr, and Shi-Min Hu. Global contrast based salient region detection. IEEE TPAMI, 37(3):569–582, 2015. 4
2015
-
[6]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, pages 248–255, 2009. 13
2009
-
[7]
R3net: Recurrent residual refinement network for saliency detection
Zijun Deng, Xiaowei Hu, Lei Zhu, Xuemiao Xu, Jing Qin, Guoqiang Han, and Pheng-Ann Heng. R3net: Recurrent residual refinement network for saliency detection. In AAAI,
-
[8]
Structure-measure: A new way to evaluate foreground maps
Deng-Ping Fan, Ming-Ming Cheng, Yun Liu, Tao Li, and Ali Borji. Structure-measure: A new way to evaluate foreground maps. In ICCV, pages 4548–4557, 2017. 12
2017
Show all 99 references
-
[9]
Salient objects in clut- ter: Bringing salient object detection to the foreground
Deng-Ping Fan, Ming-Ming Cheng, Jiang-Jiang Liu, Shang- Hua Gao, Qibin Hou, and Ali Borji. Salient objects in clut- ter: Bringing salient object detection to the foreground. In ECCV, 2018. 4
2018
-
[10]
Enhanced-alignment measure for binary foreground map evaluation
Deng-Ping Fan, Cheng Gong, Yang Cao, Bo Ren, Ming- Ming Cheng, and Ali Borji. Enhanced-alignment measure for binary foreground map evaluation. InIJCAI, pages 1–10,
-
[11]
Camouflaged object detec- tion
Deng-Ping Fan, Ge-Peng Ji, Guolei Sun, Ming-Ming Cheng, Jianbing Shen, and Ling Shao. Camouflaged object detec- tion. In CVPR, pages 2777–2787, 2020. 2, 3, 4, 12, 15
2020
-
[12]
Concealed object detection
Deng-Ping Fan, Ge-Peng Ji, Ming-Ming Cheng, and Ling Shao. Concealed object detection. IEEE TPAMI, 44(10): 6024–6042, 2021. 1, 2, 3, 6, 7, 12, 13, 14
2021
-
[13]
Advances in deep concealed scene understanding
Deng-Ping Fan, Ge-Peng Ji, Peng Xu, Ming-Ming Cheng, Christos Sakaridis, and Luc Van Gool. Advances in deep concealed scene understanding. Visual Intelligence, 1(1):16,
-
[14]
Densely nested top-down flows for salient object detection.Science China Information Sciences, 65(8):182103, 2022
Chaowei Fang, Haibin Tian, Dingwen Zhang, Qiang Zhang, Jungong Han, and Junwei Han. Densely nested top-down flows for salient object detection.Science China Information Sciences, 65(8):182103, 2022. 3
2022
-
[15]
Multi-scale and detail-enhanced segment anything model for salient object detection
Shixuan Gao, Pingping Zhang, Tianyu Yan, and Huchuan Lu. Multi-scale and detail-enhanced segment anything model for salient object detection. arXiv preprint arXiv:2408.04326, 2024. 2, 3, 6, 7
2024 arXiv
-
[16]
Res2net: A new multi-scale backbone architecture
Shang-Hua Gao, Ming-Ming Cheng, Kai Zhao, Xin-Yu Zhang, Ming-Hsuan Yang, and Philip Torr. Res2net: A new multi-scale backbone architecture. IEEE TPAMI, 43(2):652– 662, 2019. 13
2019
-
[17]
Camouflaged object detection with feature decomposition and edge reconstruc- tion
Chunming He, Kai Li, Yachao Zhang, Longxiang Tang, Yu- lun Zhang, Zhenhua Guo, and Xiu Li. Camouflaged object detection with feature decomposition and edge reconstruc- tion. In CVPR, pages 22046–22055, 2023. 3, 6, 7, 13, 14
2023
-
[18]
Strategic preys make acute predators: Enhancing cam- ouflaged object detectors by generating camouflaged objects
Chunming He, Kai Li, Yachao Zhang, Yulun Zhang, Chenyu You, Zhenhua Guo, Xiu Li, Martin Danelljan, and Fisher Yu. Strategic preys make acute predators: Enhancing cam- ouflaged object detectors by generating camouflaged objects. In ICLR, 2024. 6, 7, 13, 14
2024
-
[19]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 13
2016
-
[20]
Deeply supervised salient object detection with short connections
Qibin Hou, Ming-Ming Cheng, Xiaowei Hu, Ali Borji, Zhuowen Tu, and Philip HS Torr. Deeply supervised salient object detection with short connections. In CVPR, pages 3203–3212, 2017. 3
2017
-
[21]
Efficient camouflaged object detection network based on global localization perception and local guidance refinement
Xihang Hu, Xiaoli Zhang, Fasheng Wang, Jing Sun, and Fuming Sun. Efficient camouflaged object detection network based on global localization perception and local guidance refinement. IEEE TCSVT, 2024. 6, 7, 13, 14
2024
-
[22]
ISAT with Segment Anything: An Interactive Semi-Automatic Annotation Tool,
Shuwei Ji and Hongyuan Zhang. ISAT with Segment Anything: An Interactive Semi-Automatic Annotation Tool,
-
[23]
Segment any- thing
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment any- thing. In ICCV, pages 4015–4026, 2023. 3, 4, 5, 13, 16
2023
-
[24]
Anabranch network for camouflaged object segmentation
Trung-Nghia Le, Tam V Nguyen, Zhongliang Nie, Minh- Triet Tran, and Akihiro Sugimoto. Anabranch network for camouflaged object segmentation. CVIU, 184:45–56, 2019. 3, 4, 12
2019
-
[25]
Uncertainty-aware joint salient object and camouflaged object detection
Aixuan Li, Jing Zhang, Yunqiu Lv, Bowen Liu, Tong Zhang, and Yuchao Dai. Uncertainty-aware joint salient object and camouflaged object detection. In CVPR, pages 10071– 10081, 2021. 1, 2, 3
2021
-
[26]
Size- invariance matters: Rethinking metrics and losses for imbal- anced multi-object salient object detection
Feiran Li, Qianqian Xu, Shilong Bao, Zhiyong Yang, Run- min Cong, Xiaochun Cao, and Qingming Huang. Size- invariance matters: Rethinking metrics and losses for imbal- anced multi-object salient object detection. In ICML, pages 28989–29021, 2024. 6
2024
-
[27]
Visual saliency based on multi- scale deep features
Guanbin Li and Yizhou Yu. Visual saliency based on multi- scale deep features. In CVPR, 2015. 3, 4, 12
2015
-
[28]
The secrets of salient object segmentation
Yin Li, Xiaodi Hou, Christof Koch, James M Rehg, and Alan L Yuille. The secrets of salient object segmentation. In CVPR, 2014. 4
2014
-
[29]
Evaluation of segment anything model 2: The role of sam2 in the underwater environment
Shijie Lian and Hua Li. Evaluation of segment anything model 2: The role of sam2 in the underwater environment. arXiv preprint arXiv:2408.02924, 2024. 3
2024 arXiv
-
[30]
A metaheuristic- based approach to optimizing color design for military cam- ouflage using particle swarm optimization
Chiuhsiang Joe Lin and Yogi Tri Prasetyo. A metaheuristic- based approach to optimizing color design for military cam- ouflage using particle swarm optimization. Color Research & Application, 44(5):740–748, 2019. 1 9
2019
-
[31]
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Doll´ar. Focal loss for dense object detection. In ICCV, pages 2980–2988, 2017. 6
2017
-
[32]
Scale-aware modulation meet transformer
Weifeng Lin, Ziheng Wu, Jiayu Chen, Jun Huang, and Lian- wen Jin. Scale-aware modulation meet transformer. InICCV, pages 6015–6026, 2023. 13
2023
-
[33]
Dhsnet: Deep hierarchical saliency network for salient object detection
Nian Liu and Junwei Han. Dhsnet: Deep hierarchical saliency network for salient object detection. InCVPR, pages 678–686, 2016. 3
2016
-
[34]
Picanet: Learning pixel-wise contextual attention for saliency detec- tion
Nian Liu, Junwei Han, and Ming-Hsuan Yang. Picanet: Learning pixel-wise contextual attention for saliency detec- tion. In CVPR, pages 3089–3098, 2018. 3
2018
-
[35]
Visual saliency transformer
Nian Liu, Ni Zhang, Kaiyuan Wan, Ling Shao, and Junwei Han. Visual saliency transformer. In ICCV, pages 4722– 4732, 2021. 3, 6, 7, 13, 14
2021
-
[36]
Learning to detect a salient object
Tie Liu, Zejian Yuan, Jian Sun, Jingdong Wang, Nanning Zheng, Xiaoou Tang, and Heung-Yeung Shum. Learning to detect a salient object. IEEE TPAMI, 33(2):353–367, 2011. 4
2011
-
[37]
Explicit visual prompting for low-level structure segmenta- tions
Weihuang Liu, Xi Shen, Chi-Man Pun, and Xiaodong Cun. Explicit visual prompting for low-level structure segmenta- tions. In CVPR, pages 19434–19445, 2023. 1, 3, 6, 7, 13, 14
2023
-
[38]
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, pages 10012–10022, 2021. 13
2021
-
[39]
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In CVPR, pages 3431–3440, 2015. 6, 14
2015
-
[40]
Vscode: General visual salient and camouflaged object de- tection with 2d prompt learning
Ziyang Luo, Nian Liu, Wangbo Zhao, Xuguang Yang, Ding- wen Zhang, Deng-Ping Fan, Fahad Khan, and Junwei Han. Vscode: General visual salient and camouflaged object de- tection with 2d prompt learning. In CVPR, pages 17169– 17180, 2024. 1, 2, 3, 6, 7, 13, 14
2024
-
[41]
Simultaneously localize, segment and rank the camouflaged objects
Yunqiu Lv, Jing Zhang, Yuchao Dai, Aixuan Li, Bowen Liu, Nick Barnes, and Deng-Ping Fan. Simultaneously localize, segment and rank the camouflaged objects. In CVPR, pages 11591–11601, 2021. 3, 4, 12
2021
-
[42]
Segment anything in medical images
Jun Ma, Yuting He, Feifei Li, Lin Han, Chenyu You, and Bo Wang. Segment anything in medical images. Nature Communications, 15(1):654, 2024. 3
2024
-
[43]
How to evaluate foreground maps? In CVPR, pages 248–255, 2014
Ran Margolin, Lihi Zelnik-Manor, and Ayellet Tal. How to evaluate foreground maps? In CVPR, pages 248–255, 2014. 3, 12
2014
-
[44]
Camouflaged object segmentation with distraction mining
Haiyang Mei, Ge-Peng Ji, Ziqi Wei, Xin Yang, Xiaopeng Wei, and Deng-Ping Fan. Camouflaged object segmentation with distraction mining. In CVPR, pages 8772–8781, 2021. 1, 3, 6, 7, 13, 14
2021
-
[45]
Design and perceptual validation of performance measures for salient object seg- mentation
Vida Movahedi and James H Elder. Design and perceptual validation of performance measures for salient object seg- mentation. In CVPRW, 2010. 4
2010
-
[46]
Multi-scale interactive network for salient object detection
Youwei Pang, Xiaoqi Zhao, Lihe Zhang, and Huchuan Lu. Multi-scale interactive network for salient object detection. In CVPR, pages 9413–9422, 2020. 3
2020
-
[47]
Zoom in and out: A mixed-scale triplet network for camouflaged object detection
Youwei Pang, Xiaoqi Zhao, Tian-Zhu Xiang, Lihe Zhang, and Huchuan Lu. Zoom in and out: A mixed-scale triplet network for camouflaged object detection. In CVPR, pages 2160–2170, 2022. 3, 6, 7, 13, 14
2022
-
[48]
Osformer: One-stage camouflaged instance segmentation with transformers
Jialun Pei, Tianyang Cheng, Deng-Ping Fan, He Tang, Chuanbo Chen, and Luc Van Gool. Osformer: One-stage camouflaged instance segmentation with transformers. In ECCV, pages 19–37, 2022. 3
2022
-
[49]
Transformer-based efficient salient instance segmentation networks with orientative query.IEEE TMM, 25:1964–1978,
Jialun Pei, Tianyang Cheng, He Tang, and Chuanbo Chen. Transformer-based efficient salient instance segmentation networks with orientative query.IEEE TMM, 25:1964–1978,
1964
-
[50]
Calibnet: Dual- branch cross-modal calibration for rgb-d salient instance seg- mentation
Jialun Pei, Tao Jiang, He Tang, Nian Liu, Yueming Jin, Deng-Ping Fan, and Pheng-Ann Heng. Calibnet: Dual- branch cross-modal calibration for rgb-d salient instance seg- mentation. IEEE TIP, 2024. 3
2024
-
[51]
Evaluation study on sam 2 for class-agnostic instance-level segmenta- tion
Jialun Pei, Zhangjun Zhou, and Tiantian Zhang. Evaluation study on sam 2 for class-agnostic instance-level segmenta- tion. arXiv preprint arXiv:2409.02567, 2024. 3
2024 arXiv
-
[52]
Synergistic bleeding region and point detection in laparoscopic surgical videos
Jialun Pei, Zhangjun Zhou, Diandian Guo, Zhixi Li, Jing Qin, Bo Du, and Pheng-Ann Heng. Synergistic bleeding region and point detection in laparoscopic surgical videos. arXiv preprint arXiv:2503.22174, 2025. 3
2025
-
[53]
U-shape trans- former for underwater image enhancement
Lintao Peng, Chunli Zhu, and Liheng Bian. U-shape trans- former for underwater image enhancement. IEEE TIP, 2023. 4
2023
-
[54]
Saliency filters: Contrast based filtering for salient region detection
Federico Perazzi, Philipp Kr ¨ahenb¨uhl, Yael Pritch, and Alexander Hornung. Saliency filters: Contrast based filtering for salient region detection. In CVPR, pages 733–740, 2012. 12
2012
-
[55]
Depth-induced multi-scale recurrent attention network for saliency detection
Yongri Piao, Wei Ji, Jingjing Li, Miao Zhang, and Huchuan Lu. Depth-induced multi-scale recurrent attention network for saliency detection. In ICCV, pages 7254–7263, 2019. 3
2019
-
[56]
The attention system of the human brain
Michael I Posner, Steven E Petersen, et al. The attention system of the human brain. Annual review of neuroscience, 13(1):25–42, 1990. 1
1990
-
[57]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In ICML, pages 8748–8763, 2021. 4, 15
2021
-
[58]
Sam 2: Seg- ment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chloe Rolland, Laura Gustafson, et al. Sam 2: Seg- ment anything in images and videos. In ICLR, 2025. 3, 13, 16
2025
-
[59]
Deep texture-aware features for camouflaged object detec- tion
Jingjing Ren, Xiaowei Hu, Lei Zhu, Xuemiao Xu, Yangyang Xu, Weiming Wang, Zijun Deng, and Pheng-Ann Heng. Deep texture-aware features for camouflaged object detec- tion. IEEE TCSVT, 33(3):1157–1167, 2021. 3
2021
-
[60]
Animal camouflage analysis: Chameleon database
Przemysław Skurowski, Hassan Abdulameer, Jakub Błaszczyk, Tomasz Depta, Adam Kornacki, and Przemysław Kozieł. Animal camouflage analysis: Chameleon database. Unpublished Manuscript, 2018. 4
2018
-
[61]
Animal camouflage: current issues and new perspectives
Martin Stevens and Sami Merilaita. Animal camouflage: current issues and new perspectives. Philosophical Transac- tions of the Royal Society B: Biological Sciences, 364(1516): 423–427, 2009. 1 10
2009
-
[62]
Segmenter: Transformer for semantic segmenta- tion
Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. Segmenter: Transformer for semantic segmenta- tion. In ICCV, pages 7262–7272, 2021. 14
2021
-
[63]
Boundary-guided camouflaged object detection
Yujia Sun, Shuo Wang, Chenglizhao Chen, and Tian-Zhu Xi- ang. Boundary-guided camouflaged object detection. arXiv preprint arXiv:2207.00794, 2022. 3
2022 arXiv
-
[64]
Source-free domain adaptive fundus image segmen- tation with class-balanced mean teacher
Longxiang Tang, Kai Li, Chunming He, Yulun Zhang, and Xiu Li. Source-free domain adaptive fundus image segmen- tation with class-balanced mean teacher. In MICCAI, pages 684–694, 2023. 1
2023
-
[65]
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research , 9 (11), 2008. 8
2008
-
[66]
Learning to de- tect salient objects with image-level supervision
Lijun Wang, Huchuan Lu, Yifan Wang, Mengyang Feng, Dong Wang, Baocai Yin, and Xiang Ruan. Learning to de- tect salient objects with image-level supervision. In CVPR,
-
[67]
Salient object detection with recurrent fully convolutional networks
Linzhao Wang, Lijun Wang, Huchuan Lu, Pingping Zhang, and Xiang Ruan. Salient object detection with recurrent fully convolutional networks. IEEE TPAMI, 41(7):1734–1746,
-
[68]
Camouflaged object segmentation with prior via two-stage training.CVIU, 246:104061, 2024
Rui Wang, Caijuan Shi, Changyu Duan, Weixiang Gao, Hongli Zhu, Yunchao Wei, and Meiqin Liu. Camouflaged object segmentation with prior via two-stage training.CVIU, 246:104061, 2024. 6, 7, 13, 14
2024
-
[69]
A stagewise refinement model for detecting salient objects in images
Tiantian Wang, Ali Borji, Lihe Zhang, Pingping Zhang, and Huchuan Lu. A stagewise refinement model for detecting salient objects in images. In ICCV, pages 4019–4028, 2017. 3
2017
-
[70]
Salient object detection in the deep learning era: An in-depth survey
Wenguan Wang, Qiuxia Lai, Huazhu Fu, Jianbing Shen, Haibin Ling, and Ruigang Yang. Salient object detection in the deep learning era: An in-depth survey. IEEE TPAMI, 44(6):3239–3259, 2021. 12
2021
-
[71]
Pvtv2: Improved baselines with pyramid vision transformer
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pvtv2: Improved baselines with pyramid vision transformer. Computational Visual Media, 8(3):1–10, 2022. 13
2022
-
[72]
F3net: fusion, feedback and focus for salient object detection
Jun Wei, Shuhui Wang, and Qingming Huang. F3net: fusion, feedback and focus for salient object detection. In AAAI, pages 12321–12328, 2020. 1, 6, 7, 13, 14
2020
-
[73]
Edn: Salient object detection via extremely- downsampled network
Yu-Huan Wu, Yun Liu, Le Zhang, Ming-Ming Cheng, and Bo Ren. Edn: Salient object detection via extremely- downsampled network. IEEE TIP, 31:3125–3136, 2022. 6, 7, 13, 14
2022
-
[74]
Zero-shot learning—a comprehensive eval- uation of the good, the bad and the ugly
Yongqin Xian, Christoph H Lampert, Bernt Schiele, and Zeynep Akata. Zero-shot learning—a comprehensive eval- uation of the good, the bad and the ugly. IEEE TPAMI, 41 (9):2251–2265, 2018. 4
2018
-
[75]
Pyramid grafting network for one- stage high resolution saliency detection
Chenxi Xie, Changqun Xia, Mingcan Ma, Zhirui Zhao, Xi- aowu Chen, and Jia Li. Pyramid grafting network for one- stage high resolution saliency detection. In CVPR, pages 11717–11726, 2022. 7
2022
-
[76]
Segformer: Simple and ef- ficient design for semantic segmentation with transformers
Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and ef- ficient design for semantic segmentation with transformers. NeurIPS, 34:12077–12090, 2021. 6, 13
2021
-
[77]
Efficientsam: Leveraged masked image pretraining for efficient segment anything
Yunyang Xiong, Bala Varadarajan, Lemeng Wu, Xiaoyu Xi- ang, Fanyi Xiao, Chenchen Zhu, Xiaoliang Dai, Dilin Wang, Fei Sun, Forrest Iandola, et al. Efficientsam: Leveraged masked image pretraining for efficient segment anything. arXiv preprint arXiv:2312.00863, 2023. 3
2023 arXiv
-
[78]
Hierarchical saliency detection
Qiong Yan, Li Xu, Jianping Shi, and Jiaya Jia. Hierarchical saliency detection. In CVPR, 2013. 4
2013
-
[79]
Biomedical sam 2: Segment anything in biomedical images and videos
Zhiling Yan, Weixiang Sun, Rong Zhou, Zhengqing Yuan, Kai Zhang, Yiwei Li, Tianming Liu, Quanzheng Li, Xi- ang Li, Lifang He, et al. Biomedical sam 2: Segment anything in biomedical images and videos. arXiv preprint arXiv:2408.03286, 2024. 3
2024 arXiv
-
[80]
Saliency detection via graph-based man- ifold ranking
Chuan Yang, Lihe Zhang, Huchuan Lu, Xiang Ruan, and Ming-Hsuan Yang. Saliency detection via graph-based man- ifold ranking. In CVPR, 2013. 4, 12
2013
-
[81]
Uncertainty-guided transformer reasoning for camouflaged object detection
Fan Yang, Qiang Zhai, Xin Li, Rui Huang, Ao Luo, Hong Cheng, and Deng-Ping Fan. Uncertainty-guided transformer reasoning for camouflaged object detection. In ICCV, pages 4146–4155, 2021. 3
2021
-
[82]
Camo- former: Masked separable attention for camouflaged object detection
Bowen Yin, Xuying Zhang, Deng-Ping Fan, Shaohui Jiao, Ming-Ming Cheng, Luc Van Gool, and Qibin Hou. Camo- former: Masked separable attention for camouflaged object detection. IEEE TPAMI, 2024. 6, 7, 13, 14
2024
-
[83]
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zi-Hang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. In ICCV, pages 558–567, 2021. 13
2021
-
[84]
Mutual graph learning for cam- ouflaged object detection
Qiang Zhai, Xin Li, Fan Yang, Chenglizhao Chen, Hong Cheng, and Deng-Ping Fan. Mutual graph learning for cam- ouflaged object detection. In CVPR, pages 12997–13007,
-
[85]
Ex- ploring figure-ground assignment mechanism in perceptual organization
Wei Zhai, Yang Cao, Jing Zhang, and Zheng-Jun Zha. Ex- ploring figure-ground assignment mechanism in perceptual organization. NeurIPS, 35:17030–17042, 2022. 3
2022
-
[86]
Auto-msfnet: Search multi-scale fusion net- work for salient object detection
Miao Zhang, Tingwei Liu, Yongri Piao, Shunyu Yao, and Huchuan Lu. Auto-msfnet: Search multi-scale fusion net- work for salient object detection. In ACM MM, pages 667– 676, 2021. 3, 6, 7, 13, 14
2021
-
[87]
Preynet: Preying on cam- ouflaged objects
Miao Zhang, Shuang Xu, Yongri Piao, Dongxiang Shi, Shusen Lin, and Huchuan Lu. Preynet: Preying on cam- ouflaged objects. In ACM MM, pages 5323–5332, 2022. 3
2022
-
[88]
Progressive attention guided recurrent net- work for salient object detection
Xiaoning Zhang, Tiantian Wang, Jinqing Qi, Huchuan Lu, and Gang Wang. Progressive attention guided recurrent net- work for salient object detection. In CVPR, pages 714–722,
-
[89]
Suppress and balance: A simple gated network for salient object detection
Xiaoqi Zhao, Youwei Pang, Lihe Zhang, Huchuan Lu, and Lei Zhang. Suppress and balance: A simple gated network for salient object detection. In ECCV, pages 35–51, 2020. 6, 7, 13, 14
2020
-
[90]
Spider: A unified frame- work for context-dependent concept segmentation
Xiaoqi Zhao, Youwei Pang, Wei Ji, Baicheng Sheng, Jiaming Zuo, Lihe Zhang, and Huchuan Lu. Spider: A unified frame- work for context-dependent concept segmentation. InICML,
-
[91]
Salient object detection via integrity learning
Mingchen Zhuge, Deng-Ping Fan, Nian Liu, Dingwen Zhang, Dong Xu, and Ling Shao. Salient object detection via integrity learning. IEEE TPAMI, 45(3):3738–3752, 2022. 1, 2, 3, 6, 7, 13, 14 11 Appendix We summarize the supplementary material from the follow- ing aspects: Table of ...
2022
-
[93]
CSCS Metric Contrary to the Intersection over Union (IoU) that measures accuracy for a single class, the Camouflage-Saliency Con- fusion Score (CSCS) assesses the misclassification between two distinct classes. The CSCS, designed to evaluate the confusion between camouflaged a...
-
[94]
Confusion matrix of our USCNet on the USC12K test set
Performance in Detecting Objects of Differ- ent Sizes To evaluate the model’s ability to detect objects of vary- ing sizes, we employ several metrics: AUC ↑, SI-AUC ↑, Predicted labelTrue label PBBPBSPBC PSBPSSPSC PCBPCSPCC BSC C S B Predicted labelTrue label 38777277318 27727...
-
[95]
We adopt five metrics that are widely used in COD and SOD tasks [12, 70]
Results on Popular COD and SOD Datasets To further validate the effectiveness and robustness of our method regarding generalizability, we conduct tests on popular SOD datasets (DUTS [66], HKU-IS [27], and DUT-OMRON [80]) and COD datasets (CAMO [24], COD10K [11], and NC4K [41])...
-
[96]
Horizontal flipping and random cropping are applied for data aug- mentation
More Technical Details All models are retrained using the training set of USC12K with an input image resolution of 352 ×352. Horizontal flipping and random cropping are applied for data aug- mentation. The experiments are conducted in PyTorch on one NVIDIA L40 GPU. For our mod...
-
[97]
We obtain an initial coarse classification using CLIP [57], followed by manual verifi- cation and refinement
More USC12K Dataset Detail and Examples Object category distribution. We obtain an initial coarse classification using CLIP [57], followed by manual verifi- cation and refinement. Except for images collected from COD10K [11], which already include camouflage object category la...
-
[98]
As illustrated in Figure 11, our model outperforms its competitors
Additional Qualitative Results We present additional predictive results of our USC- Net model compared to other COD and SOD models in the USC12K test set. As illustrated in Figure 11, our model outperforms its competitors. Specifically, across four dif- ferent scenes, our mode...
-
[99]
We conducted ablation experiments to evaluate the performance of differ- ent base models, as presented in Table 12
Additional Ablation Study Performance of Different Base Models. We conducted ablation experiments to evaluate the performance of differ- ent base models, as presented in Table 12. First, as shown in the first two and last two rows of the table, our model demonstrates significa...
-
[2023]
Updated on 2023-06-03. 15
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.