Pith. sign in

REVIEW 3 cited by

On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.16613 v2 pith:M7CFCSUE submitted 2023-10-25 cs.CR

classification cs.CR
keywords modelspromptstext-to-imageimagesunsafeattackbenigneffects
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Malicious or manipulated prompts are known to exploit text-to-image models to generate unsafe images. Existing studies, however, focus on the passive exploitation of such harmful capabilities. In this paper, we investigate the proactive generation of unsafe images from benign prompts (e.g., a photo of a cat) through maliciously modified text-to-image models. Our preliminary investigation demonstrates that poisoning attacks are a viable method to achieve this goal but uncovers significant side effects, where unintended spread to non-targeted prompts compromises attack stealthiness. Root cause analysis identifies conceptual similarity as an important contributing factor to these side effects. To address this, we propose a stealthy poisoning attack method that balances covertness and performance. Our findings highlight the potential risks of adopting text-to-image models in real-world scenarios, thereby calling for future research and safety measures in this space.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Attacks on Approximate Caches in Text-to-Image Diffusion Models

    cs.CR 2025-08 conditional novelty 7.0 of 10

    Remote attackers can abuse approximate caches in text-to-image diffusion services to transmit hidden messages, steal other users' cached prompts, and inject logos into other users' generated images.

  2. Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models

    cs.AI 2024-11 conditional novelty 6.0 of 10

    Fine-tuning text-to-image diffusion models on benign data can reactivate suppressed unsafe concepts, and training the task adapter separately from a frozen safety LoRA prevents this.

  3. Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models

    cs.CL 2024-11 conditional novelty 6.0 of 10

    An automated agent can jailbreak GPT-4o and other vision-language models using only individually safe images and benign-sounding prompts, escalating responses to harmful content.

Pith tools