Pith. sign in

REVIEW 9 cited by

Towards Safe Self-Distillation of Internet-Scale Text-to-Image Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.05977 v1 pith:3QLVNAAB submitted 2023-07-12 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords modelscontentdiffusionmethodconceptgenerationharmfulimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale image generation models, with impressive quality made possible by the vast amount of data available on the Internet, raise social concerns that these models may generate harmful or copyrighted content. The biases and harmfulness arise throughout the entire training process and are hard to completely remove, which have become significant hurdles to the safe deployment of these models. In this paper, we propose a method called SDD to prevent problematic content generation in text-to-image diffusion models. We self-distill the diffusion model to guide the noise estimate conditioned on the target removal concept to match the unconditional one. Compared to the previous methods, our method eliminates a much greater proportion of harmful content from the generated images without degrading the overall image quality. Furthermore, our method allows the removal of multiple concepts at once, whereas previous works are limited to removing a single concept at a time.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LU-500: A Logo Benchmark for Concept Unlearning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A new 500-company benchmark shows current concept-erasure methods cannot remove small logos from generated images without also changing unrelated content.

  2. SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing

    cs.CV 2025-06 reject novelty 6.0 of 10

    SAGE erases concepts from diffusion models by optimizing attack prompts against the model's own text encoder and then fine-tuning that encoder with a global-local retention loss.

  3. Model Immunization from a Condition Number Perspective

    cs.LG 2025-05 reject novelty 6.0 of 10

    A new regularizer increases the condition number of the linear-probing Hessian on harmful tasks, making gradient-descent fine-tuning slower, but the theoretical analysis contains a false claim.

  4. AdvAnchor: Enhancing Diffusion Model Unlearning with Adversarial Anchors

    cs.LG 2024-12 conditional novelty 6.0 of 10

    AdvAnchor generates adversarial anchors, embeddings perturbed to be dissimilar from the target concept, and fine-tunes the model toward them, improving the erasure-preservation trade-off in diffusion model unlearning.

  5. Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models

    cs.AI 2024-11 conditional novelty 6.0 of 10

    Fine-tuning text-to-image diffusion models on benign data can reactivate suppressed unsafe concepts, and training the task adapter separately from a frozen safety LoRA prevents this.

  6. Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A system-first taxonomy and literature synthesis of multimodal unlearning across vision, language, video, and audio, with datasets, benchmarks, metrics, applications, and open challenges.

  7. DuMo: Dual Encoder Modulation Network for Precise Concept Erasure

    cs.CV 2025-01 conditional novelty 5.0 of 10

    DuMo erases target concepts from text-to-image models by adding a frozen-backbone skip-connection eraser with learned timestep and layer modulation, reporting the best trade-off on three concept erasure benchmarks.

  8. SecureT2I: No More Unauthorized Manipulation on AI Generated Images from Prompts

    cs.CR 2025-07 reject novelty 4.0 of 10

    A diffusion editing model is fine-tuned with a blur target for forbidden images and the original output for permitted images, claiming selective suppression of unauthorized edits.

  9. Open Foundation Models in Healthcare: Challenges, Paradoxes, and Opportunities with GenAI Driven Personalized Prescription

    cs.CL 2025-02 conditional novelty 4.0 of 10

    Open LLMs (LLaMA-2, LLaMA-3, Mistral, Meditron) roughly match GPT-4 on a 25-patient prescription-suitability check when given SmPC context via RAG, though some interaction classes degrade with RAG.

Pith tools