Pith. sign in

REVIEW 14 cited by

Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.08567 v2 pith:YYRA2LNN submitted 2024-02-13 cs.CL cs.CRcs.CVcs.LGcs.MA

classification cs.CLcs.CRcs.CVcs.LGcs.MA
keywords jailbreakinfectiousagentagentsmulti-agentadversarialadversarybehaviors
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A multimodal large language model (MLLM) agent can receive instructions, capture images, retrieve histories from memory, and decide which tools to use. Nonetheless, red-teaming efforts have revealed that adversarial images/prompts can jailbreak an MLLM and cause unaligned behaviors. In this work, we report an even more severe safety issue in multi-agent environments, referred to as infectious jailbreak. It entails the adversary simply jailbreaking a single agent, and without any further intervention from the adversary, (almost) all agents will become infected exponentially fast and exhibit harmful behaviors. To validate the feasibility of infectious jailbreak, we simulate multi-agent environments containing up to one million LLaVA-1.5 agents, and employ randomized pair-wise chat as a proof-of-concept instantiation for multi-agent interaction. Our results show that feeding an (infectious) adversarial image into the memory of any randomly chosen agent is sufficient to achieve infectious jailbreak. Finally, we derive a simple principle for determining whether a defense mechanism can provably restrain the spread of infectious jailbreak, but how to design a practical defense that meets this principle remains an open question to investigate. Our project page is available at https://sail-sg.github.io/Agent-Smith/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems

    cs.RO 2026-08 conditional novelty 6.0 of 10

    Prompt injection can hijack multi-agent LLM robot planners, spread from an injected agent to clean teammates through shared prompts, and partially survives a per-agent separation defense via shared memory.

  2. Unifying Adversarially Robust Model Experts in Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    CARE collaboratively fine-tunes two CLIP experts (image-text alignment and image-invariance) with embedding harmonization and EMA merging, yielding a single model with better clean and adversarial accuracy than either expert.

  3. Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Changing one LLM agent's secret objective in Werewolf lowers its team's win rate and changes its reasoning, while its public chat stays deceptively normal.

  4. Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks

    cs.CR 2025-10 conditional novelty 6.0 of 10

    Coding agents jailbroken with simple prompts produced executable malicious code in 27–32% of attempts, and single/multi-file scaffolds drove compliance to roughly 100% for frontier models.

  5. Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous

    cs.CR 2025-08 conditional novelty 6.0 of 10

    Malicious calendar invites and emails can poison Gemini's context, enabling data exfiltration, app control, and physical-world actions.

  6. AdInject: Real-World Black-Box Attacks on Web Agents via Advertising Delivery

    cs.CR 2025-05 conditional novelty 6.0 of 10

    Fake 'Close AD' ads make VLM web agents click them over 60% of the time, and near 100% in some settings.

  7. Agents at Risk: How Users Unwittingly Undermine LLM Safety

    cs.CR 2026-01 conditional novelty 5.0 of 10

    Commercial AI agents routinely trust user-relayed unverified content and execute risky actions unless the user explicitly demands a safety check.

  8. Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation

    cs.CL 2025-08 reject novelty 5.0 of 10

    Sparse autoencoder activation perturbation (SFPF) applied on top of existing jailbreak prompts raises attack success rate on Qwen3-32B, but with no defense evaluation and weak reproducibility.

  9. Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message

    cs.AI 2025-07 reject novelty 5.0 of 10

    Trojan Horse Prompting injects malicious instructions into a fabricated assistant message in the API chat history, aiming to bypass Gemini's safety filters, but no quantitative evidence is provided.

  10. Goal-Aware Identification and Rectification of Misinformation in Multi-Agent Systems

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A training-free, graph-aware defense (ARGUS) with a new dataset (MisinfoTask) reduces misinformation impact in LLM multi-agent systems by about 28% in toxicity and 10% in success rate.

  11. MASTER: Multi-Agent Security Through Exploration of Roles and Topological Structures -- A Comprehensive Framework

    cs.MA 2025-05 conditional novelty 5.0 of 10

    A role- and topology-aware attack framework for LLM multi-agent systems, showing large ASR gains over plain jailbreak baselines and defenses that reduce ASR below 20 percent.

  12. Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A Context Reasoner pipeline that cold-starts LLMs on distilled legal reasoning and applies PPO with a rule-based compliance reward improves performance on CI-based legal compliance benchmarks and transfers to general ...

  13. SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation

    cs.CR 2025-06 conditional novelty 3.0 of 10

    A systematization-of-knowledge survey that categorizes LLM privacy risks into training data, prompts, outputs, and agents, and reviews limitations of current mitigations.

  14. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools