REVIEW 7 cited by
Universal Prompt Optimizer for Safe Text-to-Image Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Text-to-Image (T2I) models have shown great performance in generating images based on textual prompts. However, these models are vulnerable to unsafe input to generate unsafe content like sexual, harassment and illegal-activity images. Existing studies based on image checker, model fine-tuning and embedding blocking are impractical in real-world applications. Hence, we propose the first universal prompt optimizer for safe T2I (POSI) generation in black-box scenario. We first construct a dataset consisting of toxic-clean prompt pairs by GPT-3.5 Turbo. To guide the optimizer to have the ability of converting toxic prompt to clean prompt while preserving semantic information, we design a novel reward function measuring toxicity and text alignment of generated images and train the optimizer through Proximal Policy Optimization. Experiments show that our approach can effectively reduce the likelihood of various T2I models in generating inappropriate images, with no significant impact on text alignment. It is also flexible to be combined with methods to achieve better performance. Our code is available at https://github.com/wu-zongyu/POSI.
Forward citations
Cited by 7 Pith papers
-
Safe Text-to-Image Generation: Simply Sanitize the Prompt Embedding
Embedding Sanitizer removes inappropriate concepts directly from text prompt embeddings with per-token scores, claiming state-of-the-art robustness against adversarial prompts.
-
SafeCFG: Controlling Harmful Features with Dynamic Safe Guidance for Safe Generation
SafeCFG adapts classifier-free guidance with a learned feature controller so that clean prompts generate normally while harmful prompts are pushed away from unsafe content.
-
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
Prompt-A-Video refines text prompts for video diffusion models via evolutionary search and DPO alignment, improving generated video quality on Open-Sora and CogVideoX.
-
Finding a Wolf in Sheep's Clothing: Combating Adversarial Text-To-Image Prompts with Text Summarization
Summarizing LLM-obfuscated text-to-image prompts before classification improves content-moderation F1 scores on the new ATTIP dataset.
-
Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise Optimization
Prompt-Noise Optimization jointly tunes the prompt embedding and diffusion noise at inference time to suppress unsafe images while keeping outputs close to the prompt.
-
Text to Image Generation and Editing: A Survey
A broad survey of text-to-image generation and editing research from 2021 to 2024, organized by architecture and comparison tables.
-
Text-to-Image Synthesis: A Decade Survey
A decade-spanning survey categorizes over 440 text-to-image papers by architecture, research problem, dataset, and evaluation metric.
Discussion (0). Continue with ORCID to comment.