Pith. sign in

REVIEW 11 cited by

Harnessing LLM to Attack LLM-Guarded Text-to-Image Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.07130 v4 pith:PB4Y5OQL submitted 2023-12-12 cs.AI

classification cs.AI
keywords adversarialattackfilterspromptsdacadrawingimagesintended
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To prevent Text-to-Image (T2I) models from generating unethical images, people deploy safety filters to block inappropriate drawing prompts. Previous works have employed token replacement to search adversarial prompts that attempt to bypass these filters, but they have become ineffective as nonsensical tokens fail semantic logic checks. In this paper, we approach adversarial prompts from a different perspective. We demonstrate that rephrasing a drawing intent into multiple benign descriptions of individual visual components can obtain an effective adversarial prompt. We propose a LLM-piloted multi-agent method named DACA to automatically complete intended rephrasing. Our method successfully bypasses the safety filters of DALL-E 3 and Midjourney to generate the intended images, achieving success rates of up to 76.7% and 64% in the one-time attack, and 98% and 84% in the re-use attack, respectively. We open-source our code and dataset on [this link](https://github.com/researchcode003/DACA).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Models

    cs.CR 2026-07 conditional novelty 6.5 of 10

    Safety alignment fails to transfer from text replies to text-in-image, and TYPO’s dual-channel combinatorial search jailbreaks four commercial image models at >90% ASR for ~$0.04.

  2. Mind the Gap: Zero-Query Jailbreaks via Filter-Generator Discrepancy in Text-to-Image Systems

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Zero-query jailbreak attacks on text-to-image systems that exploit filter-generator discrepancy reach 29-33% average success and beat baselines on six pipelines and GPT-image-2.

  3. Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    MIND learns a 'defense profile' of a T2I model from fine-grained feedback, then uses it to guide an evolutionary search, achieving 95.62% ASR across six defenses and 91.58% on Wan-2.5.

  4. Between Safe Boundaries: Exploiting Temporal Consistency for Jailbreaking Text-To-Video Generation Models

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Two benign boundary prompts can make T2V models render harmful intermediate frames via temporal interpolation; BSB's tree-search version beats prior jailbreaks by an average 18.6% relative ASR across four commercial systems.

  5. $PC^2$: Politically Controversial Content Generation via Jailbreaking Attacks on GPT-based Text-to-Image Models

    cs.CR 2026-01 conditional novelty 6.0 of 10

    PC2, a multilingual descriptive-rewriting attack, makes GPT-based text-to-image models generate politically controversial images of real public figures despite safety filters.

  6. GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models

    cs.CR 2025-06 conditional novelty 6.0 of 10

    A reinforcement-learning red-team LLM generates stealthy prompts that bypass text-to-image safety filters and produce toxic images, with reported transfer success against commercial APIs.

  7. Localizing Persona Representations in LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Persona information is most separable in the final third of LLM layers, and in Llama3's last layer ethical personas share 17.6% of salient activations while political personas have 2.1% to 5.5% unique activations.

  8. Red-Teaming Text-to-Image Systems by Rule-based Preference Modeling

    cs.LG 2025-05 conditional novelty 6.0 of 10

    RPG-RT iteratively fine-tunes an LLM with rule-based preferences from a decoupled CLIP scoring model, letting it rewrite prompts that bypass unknown safety defenses in black-box text-to-image systems.

  9. Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation

    cs.CL 2025-08 reject novelty 5.0 of 10

    Sparse autoencoder activation perturbation (SFPF) applied on top of existing jailbreak prompts raises attack success rate on Qwen3-32B, but with no defense evaluation and weak reproducibility.

  10. Dynamic Optimization and Safety Indicator Injection for Jailbreaking Text-to-Image Models with Multimodal Safety Filters

    cs.LG 2025-05 conditional novelty 5.0 of 10

    The OptJail framework (called GhostPrompt in the body) uses dynamic LLM rewriting with filter and CLIP feedback, plus bandit-selected visual indicators, to bypass text and image safety filters in T2I models, reporting...

  11. PRJ: Perception-Retrieval-Judgement for Generated Images

    cs.CV 2025-06 reject novelty 4.0 of 10

    A new safety checker for AI-generated images, built from a vision-language model, retrieval-augmented knowledge lookup, and an LLM judge, reports higher detection rates and category-level toxicity scores than three ex...

Pith tools