Pith. sign in

REVIEW 15 cited by

Image Hijacks: Adversarial Images can Control Generative Models at Runtime

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.00236 v4 pith:C4B2FMJV submitted 2023-09-01 cs.LG cs.CLcs.CR

classification cs.LGcs.CLcs.CR
keywords hijacksimagebehaviourmatchingpromptadversarialattackattacks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Are foundation models secure against malicious actors? In this work, we focus on the image input to a vision-language model (VLM). We discover image hijacks, adversarial images that control the behaviour of VLMs at inference time, and introduce the general Behaviour Matching algorithm for training image hijacks. From this, we derive the Prompt Matching method, allowing us to train hijacks matching the behaviour of an arbitrary user-defined text prompt (e.g. 'the Eiffel Tower is now located in Rome') using a generic, off-the-shelf dataset unrelated to our choice of prompt. We use Behaviour Matching to craft hijacks for four types of attack, forcing VLMs to generate outputs of the adversary's choice, leak information from their context window, override their safety training, and believe false statements. We study these attacks against LLaVA, a state-of-the-art VLM based on CLIP and LLaMA-2, and find that all attack types achieve a success rate of over 80%. Moreover, our attacks are automated and require only small image perturbations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On Surjectivity of Neural Networks: Can you elicit any behavior from your model?

    cs.LG 2025-08 conditional novelty 7.0 of 10

    Pre-LayerNorm transformers and linear attention are almost always surjective, so any target output has an input that produces it in the continuous embedding space.

  2. On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Aligning adversarial perturbations with the near-null singular directions of intermediate linear layers in transformer VLMs yields stronger attacks than existing feature- and output-space methods.

  3. VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A single adversarially optimized image can reproduce activation-steering behavior in multiple VLMs and partially transfer to unseen models.

  4. Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security

    cs.CR 2025-07 conditional novelty 6.0 of 10

    An iterative attacker-defender reinforcement learning method that makes a multimodal LLM refuse more jailbreak prompts without over-refusing ordinary queries.

  5. One Object, Multiple Lies: A Benchmark for Cross-task Adversarial Attack on Unified Vision-Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single adversarial image can make a unified vision-language model misclassify the same object across captioning, detection, region classification, and localization, and the new CrossVLAD benchmark and CRAFT attack m...

  6. MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation

    cs.CR 2025-07 conditional novelty 6.0 of 10

    MGC, a two-stage compiler framework, generates functional malware by decomposing malicious intents into benign-appearing MDIR components that strong aligned LLMs will implement, bypassing safety alignment.

  7. Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities

    cs.CR 2025-05 conditional novelty 6.0 of 10

    Con Instruction embeds harmful textual instructions into adversarial images or audio by aligning their representations, achieving successful jailbreaks on several vision- and audio-language models.

  8. AdInject: Real-World Black-Box Attacks on Web Agents via Advertising Delivery

    cs.CR 2025-05 conditional novelty 6.0 of 10

    Fake 'Close AD' ads make VLM web agents click them over 60% of the time, and near 100% in some settings.

  9. Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A single universal adversarial image perturbation can route different input semantics to different attacker-defined outputs in multimodal LLMs, with up to 66% success over five targets.

  10. Adversarial-Guided Diffusion for Multimodal LLM Attacks

    cs.CV 2025-07 conditional novelty 5.0 of 10

    AGD steers the final denoising steps of Stable Diffusion with CLIP-based target gradients and momentum, producing targeted MLLM attacks with high image fidelity and better survival under defenses.

  11. Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment

    cs.CV 2025-05 conditional novelty 5.0 of 10

    FOA-Attack aligns global and clustered local features via optimal transport with dynamic ensemble weighting to create targeted adversarial images that transfer to closed-source multimodal LLMs.

  12. A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination

    cs.CR 2026-08 conditional novelty 4.0 of 10

    A new jailbreak framework, HACA, combines atomic text and image attack strategies selected by a cross-modal planner and generates attacks with LLMs and text-to-image models, reaching 95.48% average attack success acro...

  13. Bridging Symbolic Control and Neural Reasoning in LLM Agents -- The Structured Cognitive Loop

    cs.AI 2025-11 reject novelty 4.0 of 10

    A five-module LLM agent loop (retrieval, cognition, control, action, memory) is claimed to eliminate policy violations and redundant calls, though validation does not compare against real baselines.

  14. PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training

    cs.CR 2025-07 reject novelty 3.0 of 10

    A PRM-free alignment pipeline combining genetic algorithm red teaming and multi-objective adversarial training is claimed to beat PRM-based methods at 61% lower cost, but the experiments are unverifiable.

  15. Empowering Multimodal LLMs with External Tools: A Comprehensive Survey

    cs.CV 2025-08 unverdicted novelty 2.0 of 10

    A survey paper maps how external tools are used to augment multimodal large language models across data, tasks, evaluation, and future directions.

Pith tools