REVIEW 7 cited by
Reliable and Efficient Concept Erasure of Text-to-Image Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text-to-image models encounter safety issues, including concerns related to copyright and Not-Safe-For-Work (NSFW) content. Despite several methods have been proposed for erasing inappropriate concepts from diffusion models, they often exhibit incomplete erasure, consume a lot of computing resources, and inadvertently damage generation ability. In this work, we introduce Reliable and Efficient Concept Erasure (RECE), a novel approach that modifies the model in 3 seconds without necessitating additional fine-tuning. Specifically, RECE efficiently leverages a closed-form solution to derive new target embeddings, which are capable of regenerating erased concepts within the unlearned model. To mitigate inappropriate content potentially represented by derived embeddings, RECE further aligns them with harmless concepts in cross-attention layers. The derivation and erasure of new representation embeddings are conducted iteratively to achieve a thorough erasure of inappropriate concepts. Besides, to preserve the model's generation ability, RECE introduces an additional regularization term during the derivation process, resulting in minimizing the impact on unrelated concepts during the erasure process. All the processes above are in closed-form, guaranteeing extremely efficient erasure in only 3 seconds. Benchmarking against previous approaches, our method achieves more efficient and thorough erasure with minor damage to original generation ability and demonstrates enhanced robustness against red-teaming tools. Code is available at \url{https://github.com/CharlesGong12/RECE}.
Forward citations
Cited by 7 Pith papers
-
Concept Pinpoint Eraser for Text-to-image Diffusion Models via Residual Attention Gate
CPE uses nonlinear residual attention gates with anchoring and adversarial training to erase target concepts from text-to-image diffusion models while preserving remaining concepts better than prior fine-tuning methods.
-
ACE: Anti-Editing Concept Erasure in Text-to-Image Models
ACE trains a LoRA adapter on both conditional and unconditional noise predictions so that erased concepts are suppressed during both generation and text-guided editing.
-
SafeCFG: Controlling Harmful Features with Dynamic Safe Guidance for Safe Generation
SafeCFG adapts classifier-free guidance with a learned feature controller so that clean prompts generate normally while harmful prompts are pushed away from unsafe content.
-
TraSCE: Trajectory Steering for Concept Erasure
TraSCE steers diffusion trajectories with a modified negative-prompt formulation and a Gaussian loss to erase concepts at inference time without training or weight updates.
-
Moderating the Generalization of Score-based Generative Model
MSGM is a score-adjustment unlearning method for score-based generative models that suppresses targeted content generation without full retraining.
-
DuMo: Dual Encoder Modulation Network for Precise Concept Erasure
DuMo erases target concepts from text-to-image models by adding a frozen-backbone skip-connection eraser with learned timestep and layer modulation, reporting the best trade-off on three concept erasure benchmarks.
-
FameBias: Embedding Manipulation Bias Attack in Text-to-Image Models
FameBias linearly combines a famous person's embedding with a trigger word's embedding to make text-to-image models generate that person, reaching 53% bias success without training.
Discussion (0). Continue with ORCID to comment.