Pith. sign in

REVIEW 19 cited by

Red-Teaming the Stable Diffusion Safety Filter

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.04610 v5 pith:IEGNBUD7 submitted 2022-10-03 cs.AI cs.CRcs.CVcs.CYcs.LG

classification cs.AIcs.CRcs.CVcs.CYcs.LG
keywords filtersafetycontentdiffusionpreventstableaimsdisturbing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Stable Diffusion is a recent open-source image generation model comparable to proprietary models such as DALLE, Imagen, or Parti. Stable Diffusion comes with a safety filter that aims to prevent generating explicit images. Unfortunately, the filter is obfuscated and poorly documented. This makes it hard for users to prevent misuse in their applications, and to understand the filter's limitations and improve it. We first show that it is easy to generate disturbing content that bypasses the safety filter. We then reverse-engineer the filter and find that while it aims to prevent sexual content, it ignores violence, gore, and other similarly disturbing content. Based on our analysis, we argue safety measures in future model releases should strive to be fully open and properly documented to stimulate security contributions from the community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Models

    cs.CR 2026-07 conditional novelty 6.5 of 10

    Safety alignment fails to transfer from text replies to text-in-image, and TYPO’s dual-channel combinatorial search jailbreaks four commercial image models at >90% ASR for ~$0.04.

  2. Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    MIND learns a 'defense profile' of a T2I model from fine-grained feedback, then uses it to guide an evolutionary search, achieving 95.62% ASR across six defenses and 91.58% on Wan-2.5.

  3. Evaluating Intellectual Property Guardrails of Generative Image Models: A Technical Report

    cs.CV 2026-07 conditional novelty 6.0 of 10

    All 14 tested text-to-image models readily generate recognizable IP; private models refuse at highly uneven rates, with commercial logos refused least and generated most.

  4. UnHype: CLIP-Guided Hypernetworks for Dynamic LoRA Unlearning

    cs.CV 2026-02 conditional novelty 6.0 of 10

    UnHype generates concept-specific LoRA unlearning weights on the fly from CLIP text embeddings by training a hypernetwork to follow the gradient of an unlearning loss, enabling single- and multi-concept erasure in dif...

  5. Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns

    cs.CY 2025-11 conditional novelty 6.0 of 10

    A few open-weight video models and distribution platforms dominate the creation and spread of NSFW AI video, making developer and platform choices the main intervention points for reducing non-consensual deepfake abuse.

  6. LoReUn: Data Itself Implicitly Provides Cues to Improve Machine Unlearning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    LoReUn, a plug-in loss-based reweighting strategy, improves approximate machine unlearning by focusing updates on hard-to-forget low-loss data points.

  7. PLA: Prompt Learning Attack against Text-to-Image Generative Models

    cs.CR 2025-07 conditional novelty 6.0 of 10

    PLA trains adversarial prompts with a zero-order gradient method and multimodal CLIP losses to bypass safety filters and post-hoc checkers in black-box text-to-image models, outperforming earlier word-substitution attacks.

  8. "I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI Products

    cs.HC 2025-06 conditional novelty 6.0 of 10

    GAI tools' content moderation policies are comprehensive in scope but thin on user reporting and appeals, and Reddit users report frequent frustration with opaque moderation decisions.

  9. GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models

    cs.CR 2025-06 conditional novelty 6.0 of 10

    A reinforcement-learning red-team LLM generates stealthy prompts that bypass text-to-image safety filters and produce toxic images, with reported transfer success against commercial APIs.

  10. SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing

    cs.CV 2025-06 reject novelty 6.0 of 10

    SAGE erases concepts from diffusion models by optimizing attack prompts against the model's own text encoder and then fine-tuning that encoder with a global-local retention loss.

  11. CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CuRe scores text-to-image systems by how much their output changes as prompts add cultural details, and reports better agreement with human ratings than existing proxies.

  12. Red-Teaming Text-to-Image Systems by Rule-based Preference Modeling

    cs.LG 2025-05 conditional novelty 6.0 of 10

    RPG-RT iteratively fine-tunes an LLM with rule-based preferences from a decoupled CLIP scoring model, letting it rewrite prompts that bypass unknown safety defenses in black-box text-to-image systems.

  13. DECAF: De-Clustering for Adaptive Representational Unlearning

    cs.LG 2026-07 conditional novelty 5.0 of 10

    DECAF is a forget-only unlearning method that adds input noise, suppresses the forget-class probability, and diversifies outputs, achieving 0.10% forget accuracy and 79.4% retain accuracy on CIFAR-10/ResNet-18 while d...

  14. Introspective Attention Modulation for Safe Text-to-Image Generation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Inference-time attention modulation suppresses unsafe content in diffusion-transformer T2I models without retraining and beats concept-erasure baselines in the paper's benchmarks.

  15. GIFT: Gradient-aware Immunization of diffusion models against malicious Fine-Tuning with safe concepts retention

    cs.CR 2025-07 conditional novelty 5.0 of 10

    GIFT immunizes diffusion models against malicious fine-tuning by combining loss maximization and representation noising, preserving safe concept generation.

  16. Automating Evaluation of Diffusion Model Unlearning with (Vision-) Language Model World Knowledge

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A new evaluation tool uses (vision-)language model world knowledge to rank nearby concepts and craft adversarial prompts, showing that diffusion unlearning is incomplete and that semantic similarity correlates with co...

  17. A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination

    cs.CR 2026-08 conditional novelty 4.0 of 10

    A new jailbreak framework, HACA, combines atomic text and image attack strategies selected by a cross-modal planner and generates attacks with LLMs and text-to-image models, reaching 95.48% average attack success acro...

  18. Zero-Residual Concept Erasure via Progressive Alignment in Text-to-Image Model

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A closed-form concept erasure method that enforces zero alignment residual in the optimization objective and applies updates progressively across layers to better preserve generation quality.

  19. Memory Enhanced Fractional-Order Dung Beetle Optimization for Photovoltaic Parameter Identification

    cs.NE 2025-08 reject novelty 3.0 of 10

    The claimed MFO-DBO algorithm and its CEC2017/PV results are absent from the manuscript, which instead contains an unrelated prompt-stealing attack paper.

Pith tools