Pith. sign in

REVIEW 16 cited by

Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.04031 v2 pith:R7DSUVG7 submitted 2024-06-06 cs.CV cs.CR

classification cs.CVcs.CR
keywords lvlmsprompttextualvisualadversarialattacksharmfulimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the realm of large vision language models (LVLMs), jailbreak attacks serve as a red-teaming approach to bypass guardrails and uncover safety implications. Existing jailbreaks predominantly focus on the visual modality, perturbing solely visual inputs in the prompt for attacks. However, they fall short when confronted with aligned models that fuse visual and textual features simultaneously for generation. To address this limitation, this paper introduces the Bi-Modal Adversarial Prompt Attack (BAP), which executes jailbreaks by optimizing textual and visual prompts cohesively. Initially, we adversarially embed universally harmful perturbations in an image, guided by a few-shot query-agnostic corpus (e.g., affirmative prefixes and negative inhibitions). This process ensures that image prompt LVLMs to respond positively to any harmful queries. Subsequently, leveraging the adversarial image, we optimize textual prompts with specific harmful intent. In particular, we utilize a large language model to analyze jailbreak failures and employ chain-of-thought reasoning to refine textual prompts through a feedback-iteration manner. To validate the efficacy of our approach, we conducted extensive evaluations on various datasets and LVLMs, demonstrating that our method significantly outperforms other methods by large margins (+29.03% in attack success rate on average). Additionally, we showcase the potential of our attacks on black-box commercial LVLMs, such as Gemini and ChatGLM.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On Surjectivity of Neural Networks: Can you elicit any behavior from your model?

    cs.LG 2025-08 conditional novelty 7.0 of 10

    Pre-LayerNorm transformers and linear attention are almost always surjective, so any target output has an input that produces it in the continuous embedding space.

  2. Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts

    cs.CV 2025-05 conditional novelty 7.0 of 10

    Adversarial images optimized to map harmless text prefixes to toxic tokens jailbreak vision-language models more effectively than continuing toxic text.

  3. GhostPrompt: Cross-Image Adversarial Prompt for Vision-Language Models

    cs.CR 2026-07 conditional novelty 6.0 of 10

    GhostPrompt is a universal adversarial text suffix that, after one optimization, steers VLMs to attacker-chosen outputs across diverse unseen images, reporting >30% ASR gains over prior prompt attacks.

  4. SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses

    cs.CR 2025-10 conditional novelty 6.0 of 10

    A systemization of LLM jailbreak security that adds linked taxonomies, an evaluation platform, and JailbreakDB, while its main attack–defense comparison results remain deferred.

  5. SafeMobile: Chain-level Jailbreak Detection and Automated Evaluation for Multimodal Mobile Agents

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A history-aware guard model with an LLM judge is reported to cut jailbreak success on mobile agent tasks from 86.1% to 8.4% while keeping task completion unchanged at 77.8%.

  6. Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space

    cs.CR 2025-05 conditional novelty 6.0 of 10

    A component-based, genetically optimized jailbreak framework reports over 90% success on Claude-3.5 and strong cross-model transferability.

  7. Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    HSR restores safety in pruned vision-language models by selectively restoring safety-critical neurons inside the attention heads that matter most for safety.

  8. Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A single universal adversarial image perturbation can route different input semantics to different attacker-defined outputs in multimodal LLMs, with up to 66% success over five targets.

  9. Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models

    cs.CV 2025-09 conditional novelty 5.0 of 10

    MSEA+ARC, a multi-scale and ranking-based residualization method, claims consistent F1-IoU gains over TAM for token-level MLLM visual attribution.

  10. Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation

    cs.CL 2025-08 reject novelty 5.0 of 10

    Sparse autoencoder activation perturbation (SFPF) applied on top of existing jailbreak prompts raises attack success rate on Qwen3-32B, but with no defense evaluation and weak reproducibility.

  11. VSF-Med:A Vulnerability Scoring Framework for Medical Vision-Language Models

    cs.CV 2025-06 reject novelty 5.0 of 10

    VSF-Med introduces an eight-dimension, judge-scored vulnerability score for medical VLMs and reports that all five tested models are most vulnerable to persistent attack effects, with Llama-3.2 showing the largest drop.

  12. Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A stacked-cipher jailbreak with adaptive code selection achieves 80.8% to 100% attack success on commercial large reasoning models.

  13. PRJ: Perception-Retrieval-Judgement for Generated Images

    cs.CV 2025-06 reject novelty 4.0 of 10

    A new safety checker for AI-generated images, built from a vision-language model, retrieval-augmented knowledge lookup, and an LLM judge, reports higher detection rates and category-level toxicity scores than three ex...

  14. Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models

    cs.CR 2025-06 conditional novelty 4.0 of 10

    An alternating image-text optimization produces a universal adversarial suffix and image that transfer across open multimodal LLMs more effectively than single-modality jailbreaks.

  15. A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations

    cs.CR 2025-02 conditional novelty 4.0 of 10

    A survey of LVLM safety that adds a lifecycle taxonomy and new benchmark results showing Janus-Pro-7B has weaker safety than several open-source LVLMs.

  16. Empowering Multimodal LLMs with External Tools: A Comprehensive Survey

    cs.CV 2025-08 unverdicted novelty 2.0 of 10

    A survey paper maps how external tools are used to augment multimodal large language models across data, tasks, evaluation, and future directions.

Pith tools