REVIEW 5 major objections 4 minor 57 references
Automating Evaluation of Diffusion Model Unlearning with (Vision-) Language Model World Knowledge
T0 review · 5 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper claims that (vision-) language models can automate evaluation of diffusion-model unlearning, using their world knowledge to rank nearby concepts by expected damage and to generate adversarial prompts that slip past erasure.
desk verdict A useful automated evaluation tool for diffusion unlearning, but the headline correlations and circumvention rates are under-supported by small, unreplicated measurements and a CLIP-only detector. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the (vision-) language model used as a source of structured concept knowledge: given a target, it outputs nearby concepts, ranks them by similarity through repeated re-ranking, and writes both misspellings and 'evocative' prompts that describe the target obliquely. Those outputs feed a metric script that computes KID between original and unlearned image distributions for each nearby concept and tracks the CLIP target-prediction rate for each adversarial prompt. The Spearman rank correlation between V-LM similarity order and KID damage is the quantitative link claimed between semantic closeness and erasure side effects.
What would settle it
Re-run the locality experiments with, say, 200 to 500 generated images per concept, compute KID per concept with bootstrap confidence intervals, and recompute the Spearman correlations; if the negative correlations largely disappear or flip sign, the claimed alignment between V-LM semantic order and unlearning damage does not survive a more stable damage estimate. Similarly, recompute CLIP target-prediction rates over a larger sample; if adversarial prompts no longer exceed the direct-prompt rate, the circumvention claim fails.
Extended reading notes
Core claim
The central claim is that a (vision-) language model's internal world knowledge contains a usable ordering of concepts around any erased target—an ordering that tracks the collateral damage unlearning inflicts—and that the same knowledge can write prompts which evoke the erased concept without naming it, defeating popular erasure methods. In the paper's experiments, Spearman correlations between V-LM similarity rank and KID-based damage were negative across ESD and Receler (roughly -0.13 to -0.67 for ESD, -0.36 to -0.64 for REC), meaning closer concepts bore more damage, and larger models gave stronger correlations. For 'Mickey Mouse', every adversarial prompt produced more CLIP predictions of the target than the direct prompt did, often about 30% more, and for 'Van Gogh style' some prompts added 60–80%. On a FLUX.1-dev LoRA ablation of Ablating Concepts, adversarial prompts recovered about 20% of target predictions where the direct prompt got 0%.
Load-bearing premise
All damage and circumvention conclusions are read out of KID computed from only 30 generated images per concept and CLIP prediction rates, with no confidence intervals or significance tests reported, so the measured correlations and prompt-success rates could partly reflect sampling noise.
Editorial extensions
If this is right
- Unlearning audits should include oblique, paraphrased, and misspelled prompts, not just the target phrase, because direct queries systematically understate residual knowledge.
- Semantic rankings from a V-LM can tell evaluators which nearby concepts to check first, converting a broad search for collateral damage into a ranked list.
- Larger vision-language models are likely to produce more informative audits, since their similarity rankings correlated more strongly with measured damage.
- Erasure claims made with direct-prompt evaluations alone are probably optimistic; the paper shows substantially higher target recovery with adversarial prompts.
- The approach transfers to at least one other architecture (FLUX.1-dev with LoRA), so the tool is not tied to Stable Diffusion-specific unlearning.
Reading between the lines
- A natural extension would be to use the ranked nearby concepts as a formal test set for unlearning certification, with pass/fail thresholds tied to acceptable KID shift, something the paper does not propose.
- Adversarial prompt success might also indicate memorization or lexical leakage in the text encoder rather than failure of the diffusion weights; distinguishing these would matter for choosing a fix.
- The negative rank correlations suggest a hard trade-off: erasing a concept in semantic space tends to push nearby concepts along with it, implying that 'clean' erasure may be impossible without conditioning on a boundary in prompt space.
- A testable extension is to see whether the same V-LM ranking can predict damage from finetuning or unlearning in other generative modalities such as text or video, not just image generators.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces autoeval-dmun, an automated pipeline that uses (vision-)language models to evaluate concept unlearning in diffusion models. Given a target concept, the tool asks a V-LM to generate nearby concepts, ranks them by semantic similarity, and measures unlearning damage via KID between original and unlearned model outputs. It also generates adversarial prompts (misspellings and oblique references) and measures the rate at which CLIP classifies generated images as the target concept. The paper reports that V-LM semantic orderings correlate with unlearning damage and that adversarial prompts partially circumvent ESD, REC, and an AC implementation. Experiments are conducted on Stable Diffusion v1.4 with three Llama-based assistant models, using two main target concepts for the correlation analysis and two targets for the adversarial-prompt analysis.
Significance. If the central claims are established, the tool would be a practically useful, low-cost way to automate red-teaming and damage assessment for diffusion model unlearning, an area where current evaluation is often manual and concept-specific. The paper's strengths include a concrete end-to-end pipeline, evaluation against multiple unlearning methods, and the use of an external image-distribution metric (KID) rather than only text-side checks. The authors also make their generated prompt lists available in the appendix, which aids reproducibility. However, the empirical evidence is currently preliminary: the headline correlations rest on 10 ranked concepts per target, the KID estimates use 30 images per concept, the CLIP-based circumvention rates have no independent validation, and the reported results cover very few target concepts. The central claims therefore need substantially stronger statistical and validation support before they can be taken as established.
major comments (5)
- [Section 4.1, Fig. 3] The Spearman correlations in Fig. 3 are computed from only 10 ranked concepts per target, and the reported range (-0.126 to -0.672) includes values that are not statistically distinguishable from zero at conventional significance levels: for n=10, a two-tailed 5% test requires |rho| around 0.65. No p-values, confidence intervals, permutation tests, or repeated-seed estimates are provided. The sentence in Section 4.1 that these correlations show semantic ordering 'correlates well' with unlearning damage is therefore not supported by the data in several of the six panels, especially the -0.126 ESD/Llama-3.1-8B case. The authors should report uncertainty bounds and test the null hypothesis that the ranking and the damage metric are unrelated.
- [Appendix C, Fig. 5] The claim in Section 4.1 that 'more capable V-LMs produce semantically similar concepts that are more highly correlated' is not internally consistent with the additional results in Appendix C. For REC on 'Van Gogh style', the Spearman correlations are -0.664 for Llama-3.1-8B, -0.753 for Llama-3.3-70B, and -0.717 for Llama-3.2-90B-Vision-Instruct. The 90B vision model underperforms the 70B model, so the monotone capability trend asserted in the main text does not hold. The authors should either soften the capability claim, explain the discrepancy, or provide more evidence across multiple targets and model families.
- [Section 4.2, Fig. 4] The circumvention rate is measured solely by the fraction of CLIP predictions matching the target concept, with no human evaluation or independent detector. Several adversarial prompts explicitly describe generic features of the target: for 'Mickey Mouse' they mention red shorts, white gloves, round ears, and yellow shoes, and for 'Van Gogh style' they mention thick brushstrokes, cypress trees, and swirling skies. High CLIP target rates could therefore reflect CLIP's sensitivity to these feature-level cues rather than successful generation of the erased concept itself. The headline numbers, including the 60-80% additional predictions for 'Van Gogh style', need validation on human-labeled or independently detected images before they can be interpreted as genuine unlearning circumvention.
- [Section 4.1 and Appendix B] The damage metric KID is estimated from only 30 generated images per concept, and no confidence intervals, multiple seeds, or significance tests are reported. With 30 images, KID has substantial sampling variance, and differences between concepts or between unlearning methods could easily arise from noise. Because all locality conclusions are read out of these KID values, the authors should provide variance estimates, for example by bootstrapping over generated images or sampling multiple independent image sets, and show that the reported correlations are stable under this uncertainty.
- [Generalization across targets] The main correlational evidence is limited to one target ('Formula 1 car') in Section 4.1, one additional target ('Van Gogh style') in Appendix C, and two targets for the adversarial-prompt analysis ('Mickey Mouse' and 'Van Gogh style'). This is a very narrow slice of the concept space (a vehicle, a cartoon character, and an art style). The abstract and conclusion claim a general-purpose evaluation tool, but the current evidence does not establish that the ranking-quality or circumvention results generalize to other categories such as objects, scenes, animals, or abstract concepts. The authors should either add more target concepts or explicitly scope the claims to the tested domains.
minor comments (4)
- [Appendix D] The raw JSON examples contain trailing commas, missing commas between string elements, and stray '.,,->' artifacts, which makes the appendix harder to use for reproduction. The authors should format the examples as valid JSON or clearly mark line breaks in the prompt strings.
- [Fig. 2] Figure 2 shows a CLIP confusion matrix, but the axes are not labeled and the color scale is not defined, making it difficult to interpret the row-wise distributions the text refers to.
- [Section 4.2] The number of generated images per adversarial prompt is not stated, even though the CLIP prediction rates are proportions. Without the denominator, the precision of the reported rates cannot be assessed.
- [Section 2] The related work section mentions existing jailbreaking tools and benchmarks but does not provide a quantitative or qualitative comparison of autoeval-dmun against them; adding a comparison or at least a discussion of how the tool complements existing benchmarks such as UnlearnCanvas would strengthen the positioning.
Circularity Check
No significant circularity: the evaluation pipeline measures external damage metrics on independently generated images; no claim reduces to its own inputs.
full rationale
The paper is an empirical evaluation tool, not a derivation from first principles. In Section 4.1, the Spearman correlation of Eq. 1 compares a V-LM-assigned similarity rank R[s_i] against a damage rank R[m_i] where m_i is KID computed between images generated by the original and unlearned diffusion models. KID is an external, separately computed distributional metric, not a value fitted from the V-LM ranking, so the correlation is not forced by construction. The concept list is generated and ranked by the same V-LM, which can restrict the range of sampled similarity, but this is a sampling and generalizability concern rather than circularity: the measured correlation is still an empirical quantity that could have been zero or positive. In Section 4.2, adversarial prompts are produced by a V-LM and then evaluated by the CLIP target prediction rate on images generated by the unlearned diffusion model; the success metric is read from model outputs and is not an input to prompt generation. There is no self-citation chain that bears the load of the central claims, no fitted parameter renamed as a prediction, and no uniqueness theorem imported from the authors' prior work. The reported instability of 30-image KID estimates and the lack of confidence intervals are measurement-validity concerns, not evidence that any result is equivalent to its assumptions by definition. Therefore, no specific circular step can be exhibited, and the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (4)
- k (number of top-ranked concepts) =
10
- n (initial concepts per V-LM call) =
10
- ranking repetitions =
3
- images per concept for KID =
30
assumptions (3)
- domain assumption Stable Diffusion v1.4 with ESD and Receler default hyperparameters is a representative testbed for diffusion unlearning evaluation.
- domain assumption CLIP prediction rate is a valid measure of whether the target concept was generated.
- domain assumption KID between base and unlearned distributions is a valid unlearning damage metric at 30 samples per concept.
Cite this review
Pith. "Pith review of Automating Evaluation of Diffusion Model Unlearning with (Vision-) Language Model World Knowledge." pith.science (2026). https://pith.science/paper/HDZQYS5M
@misc{pith2026250707137,
author = {Pith},
title = {Pith review of: Automating Evaluation of Diffusion Model Unlearning with (Vision-) Language Model World Knowledge},
year = {2026},
howpublished = {\url{https://pith.science/paper/HDZQYS5M}},
note = {Machine review of arXiv:2507.07137}
}
read the original abstract
Machine unlearning (MU) is a promising cost-effective method to cleanse undesired information (generated concepts, biases, or patterns) from foundational diffusion models. While MU is orders of magnitude less costly than retraining a diffusion model without the undesired information, it can be challenging and labor-intensive to prove that the information has been fully removed from the model. Moreover, MU can damage diffusion model performance on surrounding concepts that one would like to retain, making it unclear if the diffusion model is still fit for deployment. We introduce autoeval-dmun, an automated tool which leverages (vision-) language models to thoroughly assess unlearning in diffusion models. Given a target concept, autoeval-dmun extracts structured, relevant world knowledge from the language model to identify nearby concepts which are likely damaged by unlearning and to circumvent unlearning with adversarial prompts. We use our automated tool to evaluate popular diffusion model unlearning methods, revealing that language models (1) impose semantic orderings of nearby concepts which correlate well with unlearning damage and (2) effectively circumvent unlearning with synthetic adversarial prompts.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
Legal issues associated with artificial intelligence (ai)
Tolulope Awoyomi. Legal issues associated with artificial intelligence (ai). Available at SSRN, 2024
work page 2024
-
[3]
Miko aj Bi \'n kowski, Danica J Sutherland, Michael Arbel, and Arthur Gretton. Demystifying mmd gans. arXiv preprint arXiv:1801.01401, 2018
arXiv 2018
-
[4]
Lucas Bourtoule, Varun Chandrasekaran, Christopher A Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot. Machine unlearning. In 2021 IEEE symposium on security and privacy (SP), pages 141--159. IEEE, 2021
work page 2021
-
[5]
Fantastic targets for concept erasure in diffusion models and where to find them
Anh Bui, Trang Vu, Long Vuong, Trung Le, Paul Montague, Tamas Abraham, Junae Kim, and Dinh Phung. Fantastic targets for concept erasure in diffusion models and where to find them. arXiv preprint arXiv:2501.18950, 2025
arXiv 2025
-
[6]
Towards making systems forget with machine unlearning
Yinzhi Cao and Junfeng Yang. Towards making systems forget with machine unlearning. In 2015 IEEE symposium on security and privacy, pages 463--480. IEEE, 2015
work page 2015
-
[7]
Extracting training data from diffusion models
Nicolas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramer, Borja Balle, Daphne Ippolito, and Eric Wallace. Extracting training data from diffusion models. In 32nd USENIX Security Symposium (USENIX Security 23), pages 5253--5270, 2023
2023
-
[8]
Arantxa Casanova, Marlene Careil, Jakob Verbeek, Michal Drozdzal, and Adriana Romero Soriano. Instance-conditioned gan. Advances in Neural Information Processing Systems, 34: 0 27517--27529, 2021
work page 2021
Show all 57 references
-
[9]
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3558--3568, 2021
2021
-
[10]
Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models
Jaemin Cho, Abhay Zala, and Mohit Bansal. Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models. In Proceedings of the IEEE/CVF international conference on computer vision, pages 3043--3054, 2023
2023
-
[11]
Training data attribution for diffusion models
Zheng Dai and David K Gifford. Training data attribution for diffusion models. arXiv preprint arXiv:2306.02174, 2023
2023 arXiv
-
[12]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248--255. Ieee, 2009
2009
-
[13]
Genie: Higher-order denoising diffusion solvers
Tim Dockhorn, Arash Vahdat, and Karsten Kreis. Genie: Higher-order denoising diffusion solvers. Advances in Neural Information Processing Systems, 35: 0 30150--30166, 2022
2022
-
[14]
Jailbreaking text-to-image models with llm-based agents
Yingkai Dong, Zheng Li, Xiangtao Meng, Ning Yu, and Shanqing Guo. Jailbreaking text-to-image models with llm-based agents. arXiv preprint arXiv:2408.00523, 2024 a
2024 arXiv
-
[15]
Unmemorization in large language models via self-distillation and deliberate imagination
Yijiang River Dong, Hongzhou Lin, Mikhail Belkin, Ramon Huerta, and Ivan Vulic. Unmemorization in large language models via self-distillation and deliberate imagination. arXiv preprint arXiv:2402.10052, 2024 b
2024 arXiv
-
[16]
Scalable detection of offensive and non-compliant content/logo in product images
Shreyansh Gandhi, Samrat Kokkula, Abon Chaudhuri, Alessandro Magnani, Theban Stanley, Behzad Ahmadi, Venkatesh Kandaswamy, Omer Ovenc, and Shie Mannor. Scalable detection of offensive and non-compliant content/logo in product images. In Proceedings of the IEEE/CVF winter confe...
2020
-
[17]
Erasing concepts from diffusion models
Rohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, and David Bau. Erasing concepts from diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2426--2436, 2023
2023
-
[18]
Second-order information matters: Revisiting machine unlearning for large language models
Kang Gu, Md Rafi Ur Rashid, Najrin Sultana, and Shagufta Mehnaz. Second-order information matters: Revisiting machine unlearning for large language models. arXiv preprint arXiv:2403.10557, 2024
2024 arXiv
-
[19]
Selective amnesia: A continual learning approach to forgetting in deep generative models
Alvin Heng and Harold Soh. Selective amnesia: A continual learning approach to forgetting in deep generative models. Advances in Neural Information Processing Systems, 36: 0 17170--17194, 2023
2023
-
[20]
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33: 0 6840--6851, 2020
2020
-
[21]
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models. ICLR, 1 0 (2): 0 3, 2022
2022
-
[22]
Receler: Reliable concept erasing of text-to-image diffusion models via lightweight erasers
Chi-Pin Huang, Kai-Po Chang, Chung-Ting Tsai, Yung-Hsuan Lai, Fu-En Yang, and Yu-Chiang Frank Wang. Receler: Reliable concept erasing of text-to-image diffusion models via lightweight erasers. In European Conference on Computer Vision, pages 360--376. Springer, 2024 a
2024
-
[23]
Offset unlearning for large language models
James Y Huang, Wenxuan Zhou, Fei Wang, Fred Morstatter, Sheng Zhang, Hoifung Poon, and Muhao Chen. Offset unlearning for large language models. arXiv preprint arXiv:2404.11045, 2024 b
2024 arXiv
-
[24]
Knowledge unlearning for mitigating privacy risks in language models
Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, and Minjoon Seo. Knowledge unlearning for mitigating privacy risks in language models. arXiv preprint arXiv:2210.01504, 2022
2022 arXiv
-
[25]
Fairsisa: Ensemble post-processing to improve fairness of unlearning in llms
Swanand Ravindra Kadhe, Anisa Halimi, Ambrish Rawat, and Nathalie Baracaldo. Fairsisa: Ensemble post-processing to improve fairness of unlearning in llms. arXiv preprint arXiv:2312.07420, 2023
2023 arXiv
-
[26]
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4401--4410, 2019
2019
-
[27]
Alias-free generative adversarial networks
Tero Karras, Miika Aittala, Samuli Laine, Erik H \"a rk \"o nen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Alias-free generative adversarial networks. Advances in neural information processing systems, 34: 0 852--863, 2021
2021
-
[28]
Ablating concepts in text-to-image diffusion models
Nupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman, Richard Zhang, and Jun-Yan Zhu. Ablating concepts in text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 22691--22702, 2023
2023
-
[29]
Rethinking machine unlearning for large language models
Sijia Liu, Yuanshun Yao, Jinghan Jia, Stephen Casper, Nathalie Baracaldo, Peter Hase, Yuguang Yao, Chris Yuhao Liu, Xiaojun Xu, Hang Li, et al. Rethinking machine unlearning for large language models. Nature Machine Intelligence, pages 1--14, 2025
2025
-
[30]
Machine unlearning in generative ai: A survey
Zheyuan Liu, Guangyao Dou, Zhaoxuan Tan, Yijun Tian, and Meng Jiang. Machine unlearning in generative ai: A survey. arXiv preprint arXiv:2407.20516, 2024
2024 arXiv
-
[31]
Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. Advances in Neural Information Processing Systems, 35: 0 5775--5787, 2022
2022
-
[32]
Mace: Mass concept erasure in diffusion models
Shilin Lu, Zilan Wang, Leyang Li, Yanzhu Liu, and Adams Wai-Kin Kong. Mace: Mass concept erasure in diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6430--6440, 2024 a
2024
-
[33]
Eraser: Jailbreaking defense in large language models via unlearning harmful knowledge
Weikai Lu, Ziqian Zeng, Jianwei Wang, Zhengdong Lu, Zelin Chen, Huiping Zhuang, and Cen Chen. Eraser: Jailbreaking defense in large language models via unlearning harmful knowledge. arXiv preprint arXiv:2404.05880, 2024 b
2024 arXiv
-
[34]
Stable bias: Analyzing societal representations in diffusion models
Alexandra Sasha Luccioni, Christopher Akiki, Margaret Mitchell, and Yacine Jernite. Stable bias: Analyzing societal representations in diffusion models. arXiv preprint arXiv:2303.11408, 2023
2023 arXiv
-
[35]
A dataset and benchmark for copyright infringement unlearning from text-to-image diffusion models
Rui Ma, Qiang Zhou, Yizhu Jin, Daquan Zhou, Bangjun Xiao, Xiuyu Li, Yi Qu, Aishani Singh, Kurt Keutzer, Jingtong Hu, et al. A dataset and benchmark for copyright infringement unlearning from text-to-image diffusion models. arXiv preprint arXiv:2403.12052, 2024
2024 arXiv
-
[36]
Holistic unlearning benchmark: A multi-faceted evaluation for text-to-image diffusion model unlearning
Saemi Moon, Minjong Lee, Sangdon Park, and Dongwoo Kim. Holistic unlearning benchmark: A multi-faceted evaluation for text-to-image diffusion model unlearning. arXiv preprint arXiv:2410.05664, 2024
2024
-
[37]
Improved denoising diffusion probabilistic models
Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In International conference on machine learning, pages 8162--8171. PMLR, 2021
2021
-
[38]
Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models
Yiting Qu, Xinyue Shen, Xinlei He, Michael Backes, Savvas Zannettou, and Yang Zhang. Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models. In Proceedings of the 2023 ACM SIGSAC conference on computer and communications security, page...
2023
-
[39]
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International conference on machine learning, pages 8821--8831. Pmlr, 2021
2021
-
[40]
Red-teaming the stable diffusion safety filter
Javier Rando, Daniel Paleka, David Lindner, Lennart Heim, and Florian Tram \`e r. Red-teaming the stable diffusion safety filter. arXiv preprint arXiv:2210.04610, 2022
2022 arXiv
-
[41]
Generative adversarial text to image synthesis
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. Generative adversarial text to image synthesis. In International conference on machine learning, pages 1060--1069. PMLR, 2016
2016
-
[42]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684--10695, 2022
2022
-
[43]
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information...
2022
-
[44]
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in neural informa...
2022
-
[45]
Singan: Learning a generative model from a single natural image
Tamar Rott Shaham, Tali Dekel, and Tomer Michaeli. Singan: Learning a generative model from a single natural image. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4570--4580, 2019
2019
-
[46]
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International conference on machine learning, pages 2256--2265. pmlr, 2015
2015
-
[47]
Diffusion art or digital forgery? investigating data replication in diffusion models
Gowthami Somepalli, Vasu Singla, Micah Goldblum, Jonas Geiping, and Tom Goldstein. Diffusion art or digital forgery? investigating data replication in diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6048--6058, 2023
2023
-
[48]
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020
2011 arXiv
-
[49]
Unrolling sgd: Understanding factors influencing machine unlearning
Anvith Thudi, Gabriel Deza, Varun Chandrasekaran, and Nicolas Papernot. Unrolling sgd: Understanding factors influencing machine unlearning. In 2022 IEEE 7th European Symposium on Security and Privacy (EuroS&P), pages 303--319. IEEE, 2022
2022
-
[50]
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023
2023 arXiv
-
[51]
Kga: A general machine unlearning framework based on knowledge gap alignment
Lingzhi Wang, Tong Chen, Wei Yuan, Xingshan Zeng, Kam-Fai Wong, and Hongzhi Yin. Kga: A general machine unlearning framework based on knowledge gap alignment. arXiv preprint arXiv:2305.06535, 2023
2023 arXiv
-
[52]
Sneakyprompt: Jailbreaking text-to-image generative models
Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong, and Yinzhi Cao. Sneakyprompt: Jailbreaking text-to-image generative models. In 2024 IEEE symposium on security and privacy (SP), pages 897--912. IEEE, 2024
2024
-
[53]
Large language model unlearning
Yuanshun Yao, Xiaojun Xu, and Yang Liu. Large language model unlearning. Advances in Neural Information Processing Systems, 37: 0 105425--105475, 2025
2025
-
[54]
Unlearning bias in language models by partitioning gradients
Charles Yu, Sullam Jeoung, Anish Kasi, Pengfei Yu, and Heng Ji. Unlearning bias in language models by partitioning gradients. In Findings of the Association for Computational Linguistics: ACL 2023, pages 6032--6048, 2023
2023
-
[55]
Scaling autoregressive models for content-rich text-to-image generation
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al. Scaling autoregressive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789, 2 0 (3): 0 5, 2022
2022 arXiv
-
[56]
Forget-me-not: Learning to forget in text-to-image diffusion models
Gong Zhang, Kai Wang, Xingqian Xu, Zhangyang Wang, and Humphrey Shi. Forget-me-not: Learning to forget in text-to-image diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1755--1764, 2024 a
2024
-
[57]
Unlearncanvas: Stylized image dataset for enhanced machine unlearning evaluation in diffusion models
Yihua Zhang, Chongyu Fan, Yimeng Zhang, Yuguang Yao, Jinghan Jia, Jiancheng Liu, Gaoyuan Zhang, Gaowen Liu, Ramana Rao Kompella, Xiaoming Liu, et al. Unlearncanvas: Stylized image dataset for enhanced machine unlearning evaluation in diffusion models. arXiv preprint arXiv:2402...
2024 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.