REVIEW 4 cited by
DiffGuard: Text-Based Safety Checker for Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent advances in Diffusion Models have enabled the generation of images from text, with powerful closed-source models like DALL-E and Midjourney leading the way. However, open-source alternatives, such as StabilityAI's Stable Diffusion, offer comparable capabilities. These open-source models, hosted on Hugging Face, come equipped with ethical filter protections designed to prevent the generation of explicit images. This paper reveals first their limitations and then presents a novel text-based safety filter that outperforms existing solutions. Our research is driven by the critical need to address the misuse of AI-generated content, especially in the context of information warfare. DiffGuard enhances filtering efficacy, achieving a performance that surpasses the best existing filters by over 14%.
Forward citations
Cited by 4 Pith papers
-
Mind the Gap: Zero-Query Jailbreaks via Filter-Generator Discrepancy in Text-to-Image Systems
Zero-query jailbreak attacks on text-to-image systems that exploit filter-generator discrepancy reach 29-33% average success and beat baselines on six pipelines and GPT-image-2.
-
Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering
Trigger-gated additive steering vectors embedded in VLM architecture definitions create dormant backdoors that work across VQA, text-to-image, retrieval, and brand/safety biasing without data poisoning.
-
Between Safe Boundaries: Exploiting Temporal Consistency for Jailbreaking Text-To-Video Generation Models
Two benign boundary prompts can make T2V models render harmful intermediate frames via temporal interpolation; BSB's tree-search version beats prior jailbreaks by an average 18.6% relative ASR across four commercial systems.
-
AEGIS: A Mechanism-Guided Defense against Visual Synonym Jailbreaks in Text-to-Image Models
AEGIS localizes sparse semantic-injecting attention heads in diffusion models and applies similarity-aware repulsion at those heads to block visual synonym jailbreaks while preserving benign generation.
Discussion (0). Sign in to comment.