REVIEW 11 cited by
Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce new jailbreak attacks on vision language models (VLMs), which use aligned LLMs and are resilient to text-only jailbreak attacks. Specifically, we develop cross-modality attacks on alignment where we pair adversarial images going through the vision encoder with textual prompts to break the alignment of the language model. Our attacks employ a novel compositional strategy that combines an image, adversarially targeted towards toxic embeddings, with generic prompts to accomplish the jailbreak. Thus, the LLM draws the context to answer the generic prompt from the adversarial image. The generation of benign-appearing adversarial images leverages a novel embedding-space-based methodology, operating with no access to the LLM model. Instead, the attacks require access only to the vision encoder and utilize one of our four embedding space targeting strategies. By not requiring access to the LLM, the attacks lower the entry barrier for attackers, particularly when vision encoders such as CLIP are embedded in closed-source LLMs. The attacks achieve a high success rate across different VLMs, highlighting the risk of cross-modality alignment vulnerabilities, and the need for new alignment approaches for multi-modal models.
Forward citations
Cited by 11 Pith papers
-
On Surjectivity of Neural Networks: Can you elicit any behavior from your model?
Pre-LayerNorm transformers and linear attention are almost always surjective, so any target output has an input that produces it in the continuous embedding space.
-
One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models
Cross-modal unlearning transfer in vision-language models is asymmetric, architecture-dependent, and shallow under typographic attacks; influence-guided block selection reduces the measured gap.
-
VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models
A single adversarially optimized image can reproduce activation-steering behavior in multiple VLMs and partially transfer to unseen models.
-
MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation
MGC, a two-stage compiler framework, generates functional malware by decomposing malicious intents into benign-appearing MDIR components that strong aligned LLMs will implement, bypassing safety alignment.
-
When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
Picture-only instructions can jailbreak large image editing models with up to 80.9% attack success; a simple appended safety trigger mitigates the threat.
-
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
Sparse autoencoder activation perturbation (SFPF) applied on top of existing jailbreak prompts raises attack success rate on Qwen3-32B, but with no defense evaluation and weak reproducibility.
-
Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models
The paper reports higher harmful-output rates in three open-source VLMs from detailed image descriptions, in-context examples, and positive openings, and from a skip connection between internal layers, with memes riva...
-
Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack
A two-stage evaluation framework and token-projection analysis show that LVLMs encode harmful semantic cues from images even without OCR, while remaining vulnerable to cross-modal attacks.
-
Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques
Multimodal LLMs can be organized by fusion mechanism, fusion level, representation paradigm, and training paradigm, with 125 models classified accordingly.
-
System Prompt Extraction Attacks and Defenses in Large Language Models
A benchmarking study shows that chain-of-thought, few-shot, and modified sandwich queries can recover LLM system prompts with high similarity-based success, and output filtering is the most reliable tested defense.
-
Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025
The ATLAS 2025 competition demonstrates that vision-language models remain highly vulnerable to flowchart-based and cross-modal jailbreak attacks, with top scores exceeding 93%.
Discussion (0). Sign in to comment.