Pith. sign in

REVIEW 9 cited by

Multi-Agent Systems Execute Arbitrary Malicious Code

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.12188 v2 pith:GTBKDWMZ submitted 2025-03-15 cs.CR cs.LG

classification cs.CRcs.LG
keywords multi-agentagentsmalicioussystemsarbitrarycodeattackscontent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-agent systems coordinate LLM-based agents to perform tasks on users' behalf. In real-world applications, multi-agent systems will inevitably interact with untrusted inputs, such as malicious Web content, files, email attachments, and more. Using several recently proposed multi-agent frameworks as concrete examples, we demonstrate that adversarial content can hijack control and communication within the system to invoke unsafe agents and functionalities. This results in a complete security breach, up to execution of arbitrary malicious code on the user's device or exfiltration of sensitive data from the user's containerized environment. For example, when agents are instantiated with GPT-4o, Web-based attacks successfully cause the multi-agent system execute arbitrary malicious code in 58-90\% of trials (depending on the orchestrator). In some model-orchestrator configurations, the attack success rate is 100\%. We also demonstrate that these attacks succeed even if individual agents are not susceptible to direct or indirect prompt injection, and even if they refuse to perform harmful actions. We hope that these results will motivate development of trust and security models for multi-agent systems before they are widely deployed.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Monoliths to Swarms: A Study of Attack Surface Evolution in the Transition to Multi-Agent Web Systems

    cs.CR 2026-07 conditional novelty 7.0 of 10

    A new 'Telephone Loop' attack stalls multi-agent web systems in delegation cycles, succeeding in about 80% of baseline runs for three frontier models while failing against single-agent systems.

  2. Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Decomposing multi-agent LLM pipeline safety into operational reframing, planner behavior, and approval-framed delegation reveals that raw-direct model rankings mispredict deployed behavior.

  3. Attacking and Defending Multi-Agent Collaborative Filtering Systems Through Connectivity

    cs.IR 2026-08 conditional novelty 6.0 of 10

    In agent-based collaborative filtering, attack spread and privacy leakage grow with interaction connectivity, but the effect is asymmetric between user and item agents and differs between early and steady-state phases.

  4. When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems

    cs.RO 2026-08 conditional novelty 6.0 of 10

    Prompt injection can hijack multi-agent LLM robot planners, spread from an injected agent to clean teammates through shared prompts, and partially survives a per-agent separation defense via shared memory.

  5. Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Changing one LLM agent's secret objective in Werewolf lowers its team's win rate and changes its reasoning, while its public chat stays deceptively normal.

  6. Cross-Agent Campaign Attribution: Linking Asynchronous Attacks Across LLM Agents

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Cross-agent asynchronous attack sessions can be linked at 0.82 pairwise AUC from proxy-visible tool-use and prompt-style residue in the authors' synthetic SCD-v1 benchmark, far above adapted per-session detectors and ...

  7. Topology Matters: Measuring Memory Leakage in Multi-Agent LLMs

    cs.CR 2025-12 conditional novelty 6.0 of 10

    In multi-agent LLM systems, denser communication topologies, shorter attacker-target distances, and higher target centrality increase the leakage of private information, with most leakage occurring in early interactio...

  8. Detailed analysis of possible new-physics effects in the semileptonic decay $B_s \to D_s^{(*)}\tau\bar{\nu}$

    hep-ph 2026-03 unverdicted novelty 4.0 of 10

    Constraints on beyond-SM Wilson coefficients in B_s → D_s(*) τ ν̄ are derived from data using covariant-quark-model form factors, with full observable predictions for future experiments.

  9. A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents

    cs.AI 2025-06 conditional novelty 4.0 of 10

    The paper surveys security risks of LLM agents, organizes them into a five-level autonomy taxonomy, and proposes an untested CMDP-based architecture called R2A2.

Pith tools