Pith. sign in

REVIEW 77 references

Pseudolabel guided pixels contrast for domain adaptive semantic segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.09040 v1 pith:QQ2R2EQ4 submitted 2025-01-15 cs.CV cs.LG

classification cs.CVcs.LG
keywords learningsegmentationsemanticwithoutworksannotationscityscapesclass
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Semantic segmentation is essential for comprehending images, but the process necessitates a substantial amount of detailed annotations at the pixel level. Acquiring such annotations can be costly in the real-world. Unsupervised domain adaptation (UDA) for semantic segmentation is a technique that uses virtual data with labels to train a model and adapts it to real data without labels. Some recent works use contrastive learning, which is a powerful method for self-supervised learning, to help with this technique. However, these works do not take into account the diversity of features within each class when using contrastive learning, which leads to errors in class prediction. We analyze the limitations of these works and propose a novel framework called Pseudo-label Guided Pixel Contrast (PGPC), which overcomes the disadvantages of previous methods. We also investigate how to use more information from target images without adding noise from pseudo-labels. We test our method on two standard UDA benchmarks and show that it outperforms existing methods. Specifically, we achieve relative improvements of 5.1% mIoU and 4.6% mIoU on the Grand Theft Auto V (GTA5) to Cityscapes and SYNTHIA to Cityscapes tasks based on DAFormer, respectively. Furthermore, our approach can enhance the performance of other UDA approaches without increasing model complexity. Code is available at https://github.com/embar111/pgpc

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

77 extracted references · 17 canonical work pages

  1. [1]

    Ronneberger, P

    O. Ronneberger, P. Fischer, T. Brox, U-net: Convolutional networks for biomedical image segmentation, in: Medical Image Computing and Computer- Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, Springer, 2015, pp. 234–241

  2. [2]

    Zurbr ¨ugg, H

    R. Zurbr ¨ugg, H. Blum, C. Cadena, R. Siegwart, L. Schmid, Embodied active domain adaptation for semantic segmentation via informative path planning, IEEE Robotics and Automation Letters 7 (4) (2022) 8691–8698

  3. [3]

    Yurtsever, J

    E. Yurtsever, J. Lambert, A. Carballo, K. Takeda, A survey of autonomous driving: Common practices and emerging technologies, IEEE access 8 (2020) 58443–58469

  4. [4]

    Chen, Liang-Chieh and Papandreou, George and Kokkinos, Iasonas and Murphy, Kevin and Yuille, Alan L, Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs, IEEE transactions on pattern analysis and machine intelligence 40 (4) (2017) 834–848

  5. [5]

    L.-C. Chen, G. Papandreou, F. Schroff, H. Adam, Rethinking atrous convolution for semantic image segmentation, arXiv preprint arXiv:1706.05587 (2017)

  6. [6]

    Y . Li, L. Song, Y . Chen, Z. Li, X. Zhang, X. Wang, J. Sun, Learning dynamic routing for semantic segmentation, in: Proceedings of the IEEE /CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 8553–8562

  7. [7]

    J. Fan, F. Wang, H. Chu, X. Hu, Y . Cheng, B. Gao, Mlfnet: Multi-level fusion network for real-time semantic segmentation of autonomous driving, IEEE Transactions on Intelligent Vehicles 8 (1) (2023) 756–767. doi:10.1109/TIV.2022.3176860

  8. [8]

    D. Sun, G. Gao, L. Huang, Y . Liu, D. Liu, Extraction of water bodies from high-resolution remote sensing imagery based on a deep semantic segmentation network, Scientific Reports 14 (1) (2024) 14604

Show all 77 references
  1. [9]

    L. Lu, Y . Xiao, X. Chang, X. Wang, P. Ren, Z. Ren, Deformable attention-oriented feature pyramid network for semantic segmentation, Knowledge-Based Systems 254 (2022) 109623

  2. [10]

    Zheng, J

    S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y . Wang, Y . Fu, J. Feng, T. Xiang, P. H. Torr, et al., Rethinking semantic segmentation from a sequence-to- sequence perspective with transformers, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition...

  3. [11]

    E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, P. Luo, Segformer: Simple and e fficient design for semantic segmentation with transformers, Advances in Neural Information Processing Systems 34 (2021) 12077–12090

  4. [12]

    Y . Miao, Y . Sun, Y . Zhang, J. Wang, X. Zhang, An efficient point cloud semantic segmentation network with multiscale super-patch transformer, Scientific Reports 14 (1) (2024) 14581

  5. [13]

    S. R. Richter, V . Vineet, S. Roth, V . Koltun, Playing for data: Ground truth from computer games, in: European conference on computer vision, Springer, 2016, pp. 102–118

  6. [14]

    G. Ros, L. Sellart, J. Materzynska, D. Vazquez, A. M. Lopez, The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 3234–3243

  7. [15]

    Z. Yan, X. Yu, Y . Qin, Y . Wu, X. Han, S. Cui, Pixel-level intra-domain adaptation for semantic segmentation, in: Proceedings of the 29th ACM International Conference on Multimedia, 2021, pp. 404–413

  8. [16]

    X. Huo, L. Xie, H. Hu, W. Zhou, H. Li, Q. Tian, Domain-agnostic prior for transfer semantic segmentation, in: Proceedings of the IEEE /CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 7075–7085

  9. [17]

    Y . Wang, J. Peng, Z. Zhang, Uncertainty-aware pseudo label refinery for domain adaptive semantic segmentation, in: 2021 IEEE /CVF International Conference on Computer Vision (ICCV), 2021, pp. 9072–9081. doi:10.1109/ICCV48922.2021.00896

  10. [18]

    M. Liao, S. Tian, Y . Zhang, G. Hua, W. Zou, X. Li, Pda: Progressive domain adaptation for semantic segmentation, Knowledge-Based Systems 284 (2024) 111179

  11. [19]

    Zhang, M

    Y . Zhang, M. Ye, Y . Gan, W. Zhang, Knowledge based domain adaptation for semantic segmentation, Knowledge-Based Systems 193 (2020) 105444

  12. [20]

    Ren, Y .-H

    C.-X. Ren, Y .-H. Liu, X.-W. Zhang, K.-K. Huang, Multi-source unsupervised domain adaptation via pseudo target domain, IEEE Transactions on Image Processing 31 (2022) 2122–2135

  13. [21]

    H. Lin, Y . Zhang, Z. Qiu, S. Niu, C. Gan, Y . Liu, M. Tan, Prototype-guided continual adaptation for class-incremental unsupervised domain adaptation, in: European Conference on Computer Vision, Springer, 2022, pp. 351–368

  14. [22]

    Y . Yang, D. Lao, G. Sundaramoorthi, S. Soatto, Phase consistent ecological domain adaptation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 9011–9020

  15. [23]

    Corbi `ere, N

    C. Corbi `ere, N. Thome, A. Saporta, T.-H. Vu, M. Cord, P. P´erez, Confidence estimation via auxiliary models, IEEE Transactions on Pattern Analysis and Machine Intelligence 44 (10) (2021) 6043–6055

  16. [25]

    H. Xu, M. Yang, L. Deng, Y . Qian, C. Wang, Neutral cross-entropy loss based unsupervised domain adaptation for semantic segmentation, IEEE Transac- tions on Image Processing 30 (2021) 4516–4525

  17. [26]

    Vayyat, J

    M. Vayyat, J. Kasi, A. Bhattacharya, S. Ahmed, R. Tallamraju, Cluda: Contrastive learning in unsupervised domain adaptation for semantic segmentation, arXiv preprint arXiv:2208.14227 (2022)

  18. [27]

    Chopra, R

    S. Chopra, R. Hadsell, Y . LeCun, Learning a similarity metric discriminatively, with application to face verification, in: 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), V ol. 1, IEEE, 2005, pp. 539–546

  19. [28]

    B. Xie, S. Li, M. Li, C. H. Liu, G. Huang, G. Wang, Sepico: Semantic-guided pixel contrast for domain adaptive semantic segmentation, IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)

  20. [29]

    Jiang, Y

    Z. Jiang, Y . Li, C. Yang, P. Gao, Y . Wang, Y . Tai, C. Wang, Prototypical contrast adaptation for domain adaptive semantic segmentation, in: Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXIV , Springer, 2022, ...

  21. [30]

    Huang, D

    J. Huang, D. Guan, A. Xiao, S. Lu, L. Shao, Category contrast for unsupervised domain adaptation in visual tasks, in: Proceedings of the IEEE /CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 1203–1214

  22. [31]

    G. Lee, C. Eom, W. Lee, H. Park, B. Ham, Bi-directional contrastive learning for domain adaptive semantic segmentation, in: Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXX, Springer, 2022, pp. 38–55

  23. [32]

    Arazo, D

    E. Arazo, D. Ortego, P. Albert, N. E. O’Connor, K. McGuinness, Pseudo-labeling and confirmation bias in deep semi-supervised learning, in: 2020 International Joint Conference on Neural Networks (IJCNN), IEEE, 2020, pp. 1–8

  24. [33]

    Shelhamer, J

    E. Shelhamer, J. Long, T. Darrell, Fully convolutional networks for semantic segmentation, IEEE Transactions on Pattern Analysis and Machine Intelligence 39 (4) (2017) 640–651. doi:10.1109/TPAMI.2016.2572683

  25. [34]

    H. Zhao, J. Shi, X. Qi, X. Wang, J. Jia, Pyramid scene parsing network, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2881–2890

  26. [35]

    F. Yu, V . Koltun, Multi-scale context aggregation by dilated convolutions, arXiv preprint arXiv:1511.07122 (2015)

  27. [36]

    Y . Yuan, X. Chen, J. Wang, Object-contextual representations for semantic segmentation, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16, Springer, 2020, pp. 173–190

  28. [37]

    J. Fu, J. Liu, H. Tian, Y . Li, Y . Bao, Z. Fang, H. Lu, Dual attention network for scene segmentation, in: Proceedings of the IEEE /CVF conference on computer vision and pattern recognition, 2019, pp. 3146–3154

  29. [38]

    Huang, Y

    L. Huang, Y . Yuan, J. Guo, C. Zhang, X. Chen, J. Wang, Interlaced sparse self-attention for semantic segmentation, arXiv preprint arXiv:1907.12273 (2019)

  30. [39]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, Attention is all you need, Advances in neural information processing systems 30 (2017)

  31. [40]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al., An image is worth 16x16 words: Transformers for image recognition at scale, arXiv preprint arXiv:2010.11929 (2020)

  32. [41]

    Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, B. Guo, Swin transformer: Hierarchical vision transformer using shifted windows, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10012–10022

  33. [42]

    Ho ffman, D

    J. Ho ffman, D. Wang, F. Yu, T. Darrell, Fcns in the wild: Pixel-level adversarial and constraint-based adaptation, arXiv preprint arXiv:1612.02649 (2016)

  34. [43]

    Ho ffman, E

    J. Ho ffman, E. Tzeng, T. Park, J.-Y . Zhu, P. Isola, K. Saenko, A. Efros, T. Darrell, Cycada: Cycle-consistent adversarial domain adaptation, in: International conference on machine learning, PMLR, 2018, pp. 1989–1998

  35. [44]

    M. Kim, H. Byun, Learning texture invariant representation for domain adaptation of semantic segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 12975–12984

  36. [45]

    K. Mei, C. Zhu, J. Zou, S. Zhang, Instance adaptive self-training for unsupervised domain adaptation, in: European conference on computer vision, Springer, 2020, pp. 415–430

  37. [46]

    Y . Zou, Z. Yu, X. Liu, B. Kumar, J. Wang, Confidence regularized self-training, in: Proceedings of the IEEE /CVF International Conference on Computer Vision, 2019, pp. 5982–5991

  38. [47]

    K. Sohn, D. Berthelot, N. Carlini, Z. Zhang, H. Zhang, C. A. Ra ffel, E. D. Cubuk, A. Kurakin, C.-L. Li, Fixmatch: Simplifying semi-supervised learning with consistency and confidence, Advances in neural information processing systems 33 (2020) 596–608

  39. [48]

    L. Gao, J. Zhang, L. Zhang, D. Tao, Dsp: Dual soft-paste for unsupervised domain adaptive semantic segmentation, in: Proceedings of the 29th ACM International Conference on Multimedia, 2021, pp. 2825–2833

  40. [49]

    Hoyer, D

    L. Hoyer, D. Dai, Q. Wang, Y . Chen, L. Van Gool, Improving semi-supervised and domain-adaptive semantic segmentation with self-supervised depth estimation, arXiv preprint arXiv:2108.12545 (2021)

  41. [50]

    R. Gong, Q. Wang, M. Danelljan, D. Dai, L. Van Gool, Continuous pseudo-label rectified domain adaptive semantic segmentation with implicit neural representations, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 7225–7235

  42. [51]

    Hoyer, D

    L. Hoyer, D. Dai, L. Van Gool, Daformer: Improving network architectures and training strategies for domain-adaptive semantic segmentation, in: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 9924–9935

  43. [52]

    Hoyer, D

    L. Hoyer, D. Dai, L. Van Gool, Hrda: Context-aware high-resolution domain-adaptive semantic segmentation, in: Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXX, Springer, 2022, pp. 372–391

  44. [53]

    T. Chen, S. Kornblith, K. Swersky, M. Norouzi, G. E. Hinton, Big self-supervised models are strong semi-supervised learners, Advances in neural infor- mation processing systems 33 (2020) 22243–22255

  45. [54]

    X. Chen, S. Xie, K. He, An empirical study of training self-supervised vision transformers, in: Proceedings of the IEEE /CVF International Conference on Computer Vision, 2021, pp. 9640–9649

  46. [55]

    Grill, F

    J.-B. Grill, F. Strub, F. Altch ´e, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, et al., Bootstrap your own latent-a new approach to self-supervised learning, Advances in neural information processing systems 33 (2020) 21271–21284

  47. [56]

    H. Hu, J. Cui, L. Wang, Region-aware contrastive learning for semantic segmentation, in: Proceedings of the IEEE /CVF International Conference on Computer Vision, 2021, pp. 16291–16301

  48. [57]

    Zhong, B

    Y . Zhong, B. Yuan, H. Wu, Z. Yuan, J. Peng, Y .-X. Wang, Pixel contrastive-consistent semi-supervised semantic segmentation, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 7273–7282

  49. [58]

    X. Lai, Z. Tian, L. Jiang, S. Liu, H. Zhao, L. Wang, J. Jia, Semi-supervised semantic segmentation with directional context-aware consistency, in: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 1205–1214

  50. [59]

    Y . Wang, H. Wang, Y . Shen, J. Fei, W. Li, G. Jin, L. Wu, R. Zhao, X. Le, Semi-supervised semantic segmentation using unreliable pseudo-labels, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 4248–4257

  51. [60]

    W. Wang, T. Zhou, F. Yu, J. Dai, E. Konukoglu, L. Van Gool, Exploring cross-image pixel contrast for semantic segmentation, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 7303–7313

  52. [61]

    A. v. d. Oord, Y . Li, O. Vinyals, Representation learning with contrastive predictive coding, arXiv preprint arXiv:1807.03748 (2018)

  53. [62]

    X. Chen, H. Fan, R. Girshick, K. He, Improved baselines with momentum contrastive learning, arXiv preprint arXiv:2003.04297 (2020)

  54. [63]

    Cordts, M

    M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, B. Schiele, The cityscapes dataset for semantic urban scene understanding, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 3213–3223

  55. [64]

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large-scale hierarchical image database, in: 2009 IEEE conference on computer vision and pattern recognition, Ieee, 2009, pp. 248–255

  56. [65]

    Contributors, Mmsegmentation: Openmmlab semantic segmentation toolbox and benchmark, https://github.com/openmmlab/mmsegmentation (2020)

    M. Contributors, Mmsegmentation: Openmmlab semantic segmentation toolbox and benchmark, https://github.com/openmmlab/mmsegmentation (2020)

  57. [66]

    Loshchilov, F

    I. Loshchilov, F. Hutter, Decoupled weight decay regularization, arXiv preprint arXiv:1711.05101 (2017)

  58. [67]

    Zhang, B

    P. Zhang, B. Zhang, T. Zhang, D. Chen, Y . Wang, F. Wen, Prototypical pseudo label denoising and target structure learning for domain adaptive semantic segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 12414–12424

  59. [68]

    Hoyer, D

    L. Hoyer, D. Dai, H. Wang, L. Van Gool, Mic: Masked image consistency for context-enhanced domain adaptation, in: Proceedings of the IEEE /CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 11721–11732

  60. [69]

    Tranheden, V

    W. Tranheden, V . Olsson, J. Pinto, L. Svensson, Dacs: Domain adaptation via cross-domain mixed sampling, in: Proceedings of the IEEE /CVF Winter Conference on Applications of Computer Vision, 2021, pp. 1379–1389

  61. [70]

    Araslanov, S

    N. Araslanov, S. Roth, Self-supervised augmentation consistency for adapting semantic segmentation, in: Proceedings of the IEEE /CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 15384–15394

  62. [71]

    T.-H. Vu, H. Jain, M. Bucher, M. Cord, P. P ´erez, Advent: Adversarial entropy minimization for domain adaptation in semantic segmentation, Proceedings / CVPR, IEEE Computer Society Conference on Computer Vision and Pattern Recognition (2019) 2517–2526

  63. [72]

    Sakaridis, D

    C. Sakaridis, D. Dai, L. V . Gool, Guided curriculum model adaptation and uncertainty-aware evaluation for semantic nighttime image segmentation, in: Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 7374–7383

  64. [73]

    Sakaridis, D

    C. Sakaridis, D. Dai, L. Van Gool, Map-guided curriculum domain adaptation and uncertainty-aware evaluation for semantic nighttime image segmentation, IEEE Transactions on Pattern Analysis and Machine Intelligence 44 (6) (2020) 3139–3153

  65. [74]

    X. Wu, Z. Wu, H. Guo, L. Ju, S. Wang, Dannet: A one-stage domain adaptation network for unsupervised nighttime semantic segmentation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 15769–15778

  66. [75]

    Van der Maaten, G

    L. Van der Maaten, G. Hinton, Visualizing data using t-sne., Journal of machine learning research 9 (11) (2008)

  67. [76]

    Y . Li, L. Yuan, N. Vasconcelos, Bidirectional learning for domain adaptation of semantic segmentation, in: Proceedings of the IEEE /CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 6936–6945

  68. [77]

    Y . Zou, Z. Yu, B. Kumar, J. Wang, Unsupervised domain adaptation for semantic segmentation via class-balanced self-training, in: Proceedings of the European conference on computer vision (ECCV), 2018, pp. 289–305

  69. [78]

    Saporta, T.-H

    A. Saporta, T.-H. Vu, M. Cord, P. P ´erez, Esl: Entropy-guided self-supervised learning for domain adaptation in semantic segmentation, arXiv preprint arXiv:2006.08658 (2020)

Pith tools