Pith. sign in

REVIEW 92 references

SleeperMark: Towards Robust Watermark against Fine-Tuning Text-to-image Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.04852 v2 pith:CPKZSS35 submitted 2024-12-06 cs.CV

classification cs.CV
keywords modelsdiffusionmodelsleepermarkwatermarkfine-tuningdownstreamapplications
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in large-scale text-to-image (T2I) diffusion models have enabled a variety of downstream applications, including style customization, subject-driven personalization, and conditional generation. As T2I models require extensive data and computational resources for training, they constitute highly valued intellectual property (IP) for their legitimate owners, yet making them incentive targets for unauthorized fine-tuning by adversaries seeking to leverage these models for customized, usually profitable applications. Existing IP protection methods for diffusion models generally involve embedding watermark patterns and then verifying ownership through generated outputs examination, or inspecting the model's feature space. However, these techniques are inherently ineffective in practical scenarios when the watermarked model undergoes fine-tuning, and the feature space is inaccessible during verification ((i.e., black-box setting). The model is prone to forgetting the previously learned watermark knowledge when it adapts to a new task. To address this challenge, we propose SleeperMark, a novel framework designed to embed resilient watermarks into T2I diffusion models. SleeperMark explicitly guides the model to disentangle the watermark information from the semantic concepts it learns, allowing the model to retain the embedded watermark while continuing to be adapted to new downstream tasks. Our extensive experiments demonstrate the effectiveness of SleeperMark across various types of diffusion models, including latent diffusion models (e.g., Stable Diffusion) and pixel diffusion models (e.g., DeepFloyd-IF), showing robustness against downstream fine-tuning and various attacks at both the image and model levels, with minimal impact on the model's generative capability. The code is available at https://github.com/taco-group/SleeperMark.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

92 extracted references · 39 canonical work pages

  1. [1]

    Waves: Bench- marking the robustness of image watermarks

    Bang An, Mucong Ding, Tahseen Rabbani, Aakriti Agrawal, Yuancheng Xu, Chenghao Deng, Sicheng Zhu, Abdirisak Mohamed, Yuxin Wen, Tom Goldstein, et al. Waves: Bench- marking the robustness of image watermarks. In Forty-first International Conference on Machine Learning. 1

  2. [2]

    Wasserstein generative adversarial networks

    Martin Arjovsky, Soumith Chintala, and L ´eon Bottou. Wasserstein generative adversarial networks. In Interna- tional conference on machine learning , pages 214–223. PMLR, 2017. 2

  3. [3]

    DeepFloyd IF: a novel state-of- the-art open-source text-to-image model with a high degree of photorealism and language understanding

    DeepFloyd Lab at Stability. DeepFloyd IF: a novel state-of- the-art open-source text-to-image model with a high degree of photorealism and language understanding. https:// github.com/deep-floyd/IF, 2023. 1, 2, 5

  4. [4]

    ediff-i: Text-to-image dif- fusion models with an ensemble of expert denoisers

    Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Ji- aming Song, Qinsheng Zhang, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, et al. ediff-i: Text-to-image dif- fusion models with an ensemble of expert denoisers. arXiv preprint arXiv:2211.01324, 2022. 1

  5. [5]

    Improving image generation with better captions

    James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al. Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3):8, 2023. 7, 8

  6. [6]

    Betker, G

    J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y . Guo, et al. Improving image generation with better captions. https://cdn.openai. com/papers/dall- e- 3.pdf , 2023. Computer Sci- ence. 1

  7. [7]

    A computational approach to edge detection

    John Canny. A computational approach to edge detection. IEEE Transactions on pattern analysis and machine intelli- gence, pages 679–698, 1986. 7

  8. [8]

    Naruto blip captions

    Eole Cervenka. Naruto blip captions. https : / / huggingface . co / datasets / lambdalabs / naruto-blip-captions/, 2022. 2, 6, 5

Show all 92 references
  1. [9]

    Wmadapter: Adding watermark control to latent dif- fusion models

    Hai Ci, Yiren Song, Pei Yang, Jinheng Xie, and Mike Zheng Shou. Wmadapter: Adding watermark control to latent dif- fusion models. arXiv preprint arXiv:2406.08337, 2024. 1, 3

  2. [10]

    Insider threats to cloud computing: Directions for new research challenges

    William R Claycomb and Alex Nicoll. Insider threats to cloud computing: Directions for new research challenges. In 2012 IEEE 36th annual computer software and applications conference, pages 387–394. IEEE, 2012. 3

  3. [11]

    https://civitai.com/models/22354/clearvae

    ClearV AE. https://civitai.com/models/22354/clearvae. 7, 8

  4. [12]

    Diffusionshield: A wa- termark for copyright protection against generative diffusion models

    Yingqian Cui, Jie Ren, Han Xu, Pengfei He, Hui Liu, Lichao Sun, Yue Xing, and Jiliang Tang. Diffusionshield: A wa- termark for copyright protection against generative diffusion models. arXiv preprint arXiv:2306.04642, 2023. 3

  5. [13]

    Pregip: Watermarking the pretraining of graph neural networks for deep intellectual property protection

    Enyan Dai, Minhua Lin, and Suhang Wang. Pregip: Watermarking the pretraining of graph neural networks for deep intellectual property protection. arXiv preprint arXiv:2402.04435, 2024. 3

  6. [14]

    Diffusion models beat gans on image synthesis

    Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in neural informa- tion processing systems, 34:8780–8794, 2021. 1

  7. [15]

    Scaling recti- fied flow transformers for high-resolution image synthesis

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling recti- fied flow transformers for high-resolution image synthesis. In Forty-first International Conference on Mach...

  8. [16]

    train text to image lora.py

    Hugging Face. train text to image lora.py. https : //github.com/huggingface/diffusers/blob/ main/examples/text_to_image/train_text_ to_image_lora.py, 2024. 5

  9. [17]

    Diffusers dreambooth example

    Hugging Face. Diffusers dreambooth example. https: //github.com/huggingface/diffusers/tree/ main/examples/dreambooth, 2024. 6

  10. [18]

    Wide flat minimum watermarking for robust ownership ver- ification of gans

    Jianwei Fei, Zhihua Xia, Benedetta Tondi, and Mauro Barni. Wide flat minimum watermarking for robust ownership ver- ification of gans. IEEE Transactions on Information Foren- sics and Security, 2024. 3

  11. [19]

    Aqualora: Toward white-box protection for customized stable diffusion models via watermark lora

    Weitao Feng, Wenbo Zhou, Jiyan He, Jie Zhang, Tianyi Wei, Guanlin Li, Tianwei Zhang, Weiming Zhang, and Nenghai Yu. Aqualora: Toward white-box protection for customized stable diffusion models via watermark lora. arXiv preprint arXiv:2405.11135, 2024. 1, 2, 3, 4, 5

  12. [20]

    The stable signature: Rooting watermarks in latent diffusion models

    Pierre Fernandez, Guillaume Couairon, Herv ´e J ´egou, Matthijs Douze, and Teddy Furon. The stable signature: Rooting watermarks in latent diffusion models. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 22466–22477, 2023. 1, 3, 5

  13. [21]

    Dream- sim: Learning new dimensions of human visual similar- ity using synthetic data

    Stephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. Dream- sim: Learning new dimensions of human visual similar- ity using synthetic data. arXiv preprint arXiv:2306.09344 ,

  14. [22]

    Watermarking plms on classification tasks by combining contrastive learning with weight perturbation

    Chenxi Gu, Xiaoqing Zheng, Jianhan Xu, Muling Wu, Cenyuan Zhang, Chengsong Huang, Hua Cai, and Xuan- Jing Huang. Watermarking plms on classification tasks by combining contrastive learning with weight perturbation. In Findings of the Association for Computational Linguistics: ...

  15. [23]

    Vec- tor quantized diffusion model for text-to-image synthesis

    Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. Vec- tor quantized diffusion model for text-to-image synthesis. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 10696–10706, 2022. 1, 2

  16. [24]

    Stable diffusion prompts dataset

    Gustavosta. Stable diffusion prompts dataset. https: / / huggingface . co / datasets / Gustavosta / Stable-Diffusion-Prompts, 2023. 5

  17. [25]

    Classifier-free diffusion guidance

    Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022. 8

  18. [26]

    Denoising dif- fusion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 1, 3, 6

  19. [27]

    Lora: Low-rank adaptation of large language models

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021. 1, 2

  20. [28]

    On the trustworthiness of generative founda- 9 tion models: Guideline, assessment, and perspective

    Yue Huang, Chujie Gao, Siyuan Wu, Haoran Wang, Xiangqi Wang, Yujun Zhou, Yanbo Wang, Jiayi Ye, Jiawen Shi, Qihui Zhang, et al. On the trustworthiness of generative founda- 9 tion models: Guideline, assessment, and perspective. arXiv preprint arXiv:2502.14296, 2025. 3

  21. [29]

    Security issues in cloud com- puting and countermeasures

    Danish Jamil and Hassan Zaki. Security issues in cloud com- puting and countermeasures. International Journal of Engi- neering Science and Technology (IJEST) , 3(4):2672–2676,

  22. [30]

    Elucidating the design space of diffusion-based generative models

    Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. Advances in neural information processing systems, 35:26565–26577, 2022. 6

  23. [31]

    Wouaf: Weight modulation for user attri- bution and fingerprinting in text-to-image diffusion models

    Changhoon Kim, Kyle Min, Maitreya Patel, Sheng Cheng, and Yezhou Yang. Wouaf: Weight modulation for user attri- bution and fingerprinting in text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8974–8983, 2...

  24. [32]

    Multi-concept customization of text-to-image diffusion

    Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. Multi-concept customization of text-to-image diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 1931–1941, 2023. 1, 2

  25. [33]

    Evaluating the robustness of trigger set-based watermarks embedded in deep neural networks.IEEE Trans- actions on Dependable and Secure Computing, 20(4):3434– 3448, 2022

    Suyoung Lee, Wonho Song, Suman Jana, Meeyoung Cha, and Sooel Son. Evaluating the robustness of trigger set-based watermarks embedded in deep neural networks.IEEE Trans- actions on Dependable and Secure Computing, 20(4):3434– 3448, 2022. 3

  26. [34]

    Dif- fusetrace: A transparent and flexible watermarking scheme for latent diffusion model

    Liangqi Lei, Keke Gai, Jing Yu, and Liehuang Zhu. Dif- fusetrace: A transparent and flexible watermarking scheme for latent diffusion model. arXiv preprint arXiv:2405.02696,

  27. [35]

    Light the night: A multi-condition diffusion framework for unpaired low-light enhancement in autonomous driving

    Jinlong Li, Baolu Li, Zhengzhong Tu, Xinyu Liu, Qing Guo, Felix Juefei-Xu, Runsheng Xu, and Hongkai Yu. Light the night: A multi-condition diffusion framework for unpaired low-light enhancement in autonomous driving. In Proceed- ings of the IEEE/CVF Conference on Computer Visi...

  28. [36]

    Plmmark: A secure and robust black-box watermarking framework for pre-trained language models

    Peixuan Li, Pengzhou Cheng, Fangqi Li, Wei Du, Haodong Zhao, and Gongshen Liu. Plmmark: A secure and robust black-box watermarking framework for pre-trained language models. Proceedings of the AAAI Conference on Artificial Intelligence, 37(12):14991–14999, 2023. 3

  29. [37]

    A survey of deep neural network watermarking techniques

    Yue Li, Hongxia Wang, and Mauro Barni. A survey of deep neural network watermarking techniques. Neurocomputing, 461:171–193, 2021. 2

  30. [38]

    Gligen: Open-set grounded text-to-image generation

    Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jian- wei Yang, Jianfeng Gao, Chunyuan Li, and Yong Jae Lee. Gligen: Open-set grounded text-to-image generation. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22511–22521, 2023. 1, 2

  31. [39]

    Hunyuan-dit: A powerful multi-resolution diffusion trans- former with fine-grained chinese understanding, 2024

    Zhimin Li, Jianwei Zhang, Qin Lin, Jiangfeng Xiong, Yanxin Long, Xinchi Deng, Yingfang Zhang, Xingchao Liu, Minbin Huang, Zedong Xiao, Dayou Chen, Jiajun He, Jiahao Li, Wenyue Li, Chen Zhang, Rongwei Quan, Jianxiang Lu, Jiabin Huang, Xiaoyan Yuan, Xiaoxiao Zheng, Yixuan Li, Ji...

  32. [40]

    Microsoft coco: Common objects in context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceeding...

  33. [41]

    Lazy layers to make fine- tuned diffusion models more traceable

    Haozhe Liu, Wentian Zhang, Bing Li, Bernard Ghanem, and J ¨urgen Schmidhuber. Lazy layers to make fine- tuned diffusion models more traceable. arXiv preprint arXiv:2405.00466, 2024. 2, 3

  34. [42]

    Pseudo numerical methods for diffusion models on manifolds.arXiv preprint arXiv:2202.09778, 2022

    Luping Liu, Yi Ren, Zhijie Lin, and Zhou Zhao. Pseudo numerical methods for diffusion models on manifolds.arXiv preprint arXiv:2202.09778, 2022. 6

  35. [43]

    Watermarking diffusion model

    Yugeng Liu, Zheng Li, Michael Backes, Yun Shen, and Yang Zhang. Watermarking diffusion model. arXiv preprint arXiv:2305.12502, 2023. 2, 3

  36. [44]

    Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps

    Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. Advances in Neural Information Processing Systems , 35:5775–5787,

  37. [45]

    Countering language drift with seeded iterated learning

    Yuchen Lu, Soumye Singhal, Florian Strub, Aaron Courville, and Olivier Pietquin. Countering language drift with seeded iterated learning. In International Conference on Machine Learning, pages 6437–6447. PMLR, 2020. 4

  38. [46]

    Understanding diffusion models: A unified per- spective, 2022

    Calvin Luo. Understanding diffusion models: A unified per- spective, 2022. 3

  39. [47]

    A robustness- assured white-box watermark in neural networks

    Peizhuo Lv, Pan Li, Shengzhi Zhang, Kai Chen, Ruigang Liang, Hualong Ma, Yue Zhao, and Yingjiu Li. A robustness- assured white-box watermark in neural networks. IEEE Transactions on Dependable and Secure Computing , 20(6): 5214–5229, 2023. 3

  40. [48]

    Ssl-wm: A black-box watermarking approach for encoders pre-trained by self-supervised learning, 2024

    Peizhuo Lv, Pan Li, Shenchen Zhu, Shengzhi Zhang, Kai Chen, Ruigang Liang, Chang Yue, Fan Xiang, Yuling Cai, Hualong Ma, Yingjun Zhang, and Guozhu Meng. Ssl-wm: A black-box watermarking approach for encoders pre-trained by self-supervised learning, 2024. 3

  41. [49]

    Latent watermark: Inject and detect watermarks in latent diffusion space

    Zheling Meng, Bo Peng, and Jing Dong. Latent watermark: Inject and detect watermarks in latent diffusion space. arXiv preprint arXiv:2404.00230, 2024. 4

  42. [50]

    Midjourney home page

    midjourney. Midjourney home page. https://www. midjourney.com/home, 2024. Accessed: 2024-10-23. 1

  43. [51]

    A watermark-conditioned diffusion model for ip protection

    Rui Min, Sen Li, Hongyang Chen, and Minhao Cheng. A watermark-conditioned diffusion model for ip protection. arXiv preprint arXiv:2403.10893, 2024. 1, 3, 5

  44. [52]

    T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

    Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, and Ying Shan. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 4296–4304, 2024. 1, 2

  45. [53]

    Glide: Towards photorealistic image generation 10 and editing with text-guided diffusion models.arXiv preprint arXiv:2112.10741, 2021

    Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation 10 and editing with text-guided diffusion models.arXiv preprint arXiv:2112.10741, 2021. 1, 2

  46. [54]

    Improved denoising diffusion probabilistic models

    Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In International conference on machine learning, pages 8162–8171. PMLR,

  47. [55]

    Protecting intellectual property of genera- tive adversarial networks from ambiguity attacks

    Ding Sheng Ong, Chee Seng Chan, Kam Woh Ng, Lixin Fan, and Qiang Yang. Protecting intellectual property of genera- tive adversarial networks from ambiguity attacks. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3630–3639, 2021. 3

  48. [56]

    On aliased resizing and surprising subtleties in gan evaluation

    Gaurav Parmar, Richard Zhang, and Jun-Yan Zhu. On aliased resizing and surprising subtleties in gan evaluation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11410–11420, 2022. 5

  49. [57]

    In- tellectual property protection of diffusion models via the wa- termark diffusion process

    Sen Peng, Yufei Chen, Cong Wang, and Xiaohua Jia. In- tellectual property protection of diffusion models via the wa- termark diffusion process. arXiv preprint arXiv:2306.03436,

  50. [58]

    Sdxl: Improving latent diffusion mod- els for high-resolution image synthesis

    Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M ¨uller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion mod- els for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023. 1, 2

  51. [59]

    Spire: Semantic prompt-driven image restoration

    Chenyang Qi, Zhengzhong Tu, Keren Ye, Mauricio Delbra- cio, Peyman Milanfar, Qifeng Chen, and Hossein Talebi. Spire: Semantic prompt-driven image restoration. In Eu- ropean Conference on Computer Vision , pages 446–464. Springer, 2024. 1

  52. [60]

    Learning transferable visual models from natural language supervi- sion

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...

  53. [61]

    A dwt, dct and svd based wa- termarking technique to protect the image piracy

    Md Maklachur Rahman. A dwt, dct and svd based wa- termarking technique to protect the image piracy. arXiv preprint arXiv:1307.3294, 2013. 3, 5

  54. [62]

    Hierarchical text-conditional image gener- ation with clip latents

    Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image gener- ation with clip latents. arXiv preprint arXiv:2204.06125, 1 (2):3, 2022. 1, 2

  55. [63]

    Copyright protection in generative ai: A technical perspective

    Jie Ren, Han Xu, Pengfei He, Yingqian Cui, Shenglai Zeng, Jiankun Zhang, Hongzhi Wen, Jiayuan Ding, Pei Huang, Lingjuan Lyu, et al. Copyright protection in generative ai: A technical perspective. arXiv preprint arXiv:2402.02333 ,

  56. [64]

    Stable signature

    Facebook Research. Stable signature. https : / / github . com / facebookresearch / stable _ signature, 2024. 3

  57. [65]

    Lawa: Using latent space for in-generation image watermarking

    Ahmad Rezaei, Mohammad Akbari, Saeed Ranjbar Alvar, Arezou Fatemi, and Yong Zhang. Lawa: Using latent space for in-generation image watermarking. arXiv preprint arXiv:2408.05868, 2024. 1

  58. [66]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 1, 2

  59. [67]

    Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

    Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 2250...

  60. [68]

    Photorealistic text-to-image diffusion models with deep language understanding

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information...

  61. [69]

    Image super- resolution via iterative refinement

    Chitwan Saharia, Jonathan Ho, William Chan, Tim Sali- mans, David J Fleet, and Mohammad Norouzi. Image super- resolution via iterative refinement. IEEE transactions on pattern analysis and machine intelligence, 45(4):4713–4726,

  62. [70]

    https://huggingface.co/stabilityai/sd-vae-ft- mse, 2024

    sd-vae-ft mse. https://huggingface.co/stabilityai/sd-vae-ft- mse, 2024. 7, 8

  63. [71]

    Invisible watermark

    ShieldMnt. Invisible watermark. https://github. com/ShieldMnt/invisible-watermark, 2024. 3

  64. [72]

    Jpeg-resistant adversarial im- ages

    Richard Shin and Dawn Song. Jpeg-resistant adversarial im- ages. In NIPS 2017 workshop on machine learning and com- puter security, page 8, 2017. 3

  65. [73]

    Denoising diffusion implicit models

    Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. 5, 6

  66. [74]

    Generative modeling by esti- mating gradients of the data distribution

    Yang Song and Stefano Ermon. Generative modeling by esti- mating gradients of the data distribution. Advances in neural information processing systems, 32, 2019. 3

  67. [75]

    Score-based generative modeling through stochastic differential equa- tions

    Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equa- tions. arXiv preprint arXiv:2011.13456, 2020. 1

  68. [76]

    Stegastamp: Invisible hyperlinks in physical photographs

    Matthew Tancik, Ben Mildenhall, and Ren Ng. Stegastamp: Invisible hyperlinks in physical photographs. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2117–2126, 2020. 4, 1, 2, 3

  69. [77]

    Common sense guide to mitigating insider threats

    Michael Theis, Randall F Trzeciak, Daniel L Costa, An- drew P Moore, Sarah Miller, Tracy Cassidy, and William R Claycomb. Common sense guide to mitigating insider threats. 2019. 3

  70. [78]

    Bovik, H.R

    Zhou Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing , 13(4): 600–612, 2004. 6

  71. [79]

    Malware detection in cloud comput- ing infrastructures

    Michael R Watson, Angelos K Marnerides, Andreas Mauthe, David Hutchison, et al. Malware detection in cloud comput- ing infrastructures. IEEE Transactions on Dependable and Secure Computing, 13(2):192–205, 2015. 3

  72. [80]

    Tree-rings watermarks: Invisible fingerprints for diffusion images

    Yuxin Wen, John Kirchenbauer, Jonas Geiping, and Tom Goldstein. Tree-rings watermarks: Invisible fingerprints for diffusion images. Advances in Neural Information Process- ing Systems, 36, 2024. 3 11

  73. [81]

    Flexible and secure watermarking for latent diffusion model

    Cheng Xiong, Chuan Qin, Guorui Feng, and Xinpeng Zhang. Flexible and secure watermarking for latent diffusion model. In Proceedings of the 31st ACM International Conference on Multimedia, pages 1668–1676, 2023. 1, 3

  74. [82]

    Diffusion models: A comprehensive survey of methods and applications, 2024

    Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Run- sheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming- Hsuan Yang. Diffusion models: A comprehensive survey of methods and applications, 2024. 3

  75. [83]

    Reco: Region-controlled text-to-image genera- tion

    Zhengyuan Yang, Jianfeng Wang, Zhe Gan, Linjie Li, Kevin Lin, Chenfei Wu, Nan Duan, Zicheng Liu, Ce Liu, Michael Zeng, et al. Reco: Region-controlled text-to-image genera- tion. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 14246–14255,

  76. [84]

    Text-to-image diffusion models can be easily backdoored through multimodal data poisoning

    Shengfang Zhai, Yinpeng Dong, Qingni Shen, Shi Pu, Yue- jian Fang, and Hang Su. Text-to-image diffusion models can be easily backdoored through multimodal data poisoning. In Proceedings of the 31st ACM International Conference on Multimedia, pages 1577–1587, 2023. 2, 3

  77. [85]

    Robust invisible video watermark- ing with attention

    Kevin Alex Zhang, Lei Xu, Alfredo Cuesta-Infante, and Kalyan Veeramachaneni. Robust invisible video watermark- ing with attention. arXiv preprint arXiv:1909.01285, 2019. 3

  78. [86]

    Adding conditional control to text-to-image diffusion models

    Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847, 2023. 1, 2, 6

  79. [87]

    The unreasonable effectiveness of deep features as a perceptual metric

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 586–595, 2018. 4

  80. [88]

    Uni-controlnet: All-in-one control to text-to-image diffusion models

    Shihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao, Shaozhe Hao, Lu Yuan, and Kwan-Yee K Wong. Uni-controlnet: All-in-one control to text-to-image diffusion models. Advances in Neural Information Processing Sys- tems, 36, 2024. 1, 2

  81. [89]

    Unipc: A unified predictor-corrector framework for fast sampling of diffusion models

    Wenliang Zhao, Lujia Bai, Yongming Rao, Jie Zhou, and Jiwen Lu. Unipc: A unified predictor-corrector framework for fast sampling of diffusion models. Advances in Neural Information Processing Systems, 36, 2024. 6

  82. [90]

    A recipe for watermarking dif- fusion models

    Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Ngai- Man Cheung, and Min Lin. A recipe for watermarking dif- fusion models. arXiv preprint arXiv:2303.10137, 2023. 2, 3, 5, 1

  83. [91]

    *[Z]& A dog⋯

    Jiren Zhu, Russell Kaplan, Justin Johnson, and Li Fei-Fei. Hidden: Hiding data with deep networks, 2018. 1 12 SleeperMark: Towards Robust Watermark against Fine-Tuning Text-to-image Diffusion Models Supplementary Material A. Intuition and Post-hoc Explanation The training loss...

  84. [200]

    *[Z]&%#{@}Aˆ˜$

    The model requires a substantial number of iterations (up to 10,000 steps) to adapt to the new condition. Nev- ertheless, we find that integrating this additional condition has minimal impact on the effectiveness of our watermark- ing method, which has been demonstrated in the...

Pith tools