Pith. sign in

REVIEW 49 cited by

Multi-Agent Risks from Advanced AI

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.14143 v1 pith:MTLFQTOU submitted 2025-02-19 cs.MA cs.AIcs.CYcs.ETcs.LG

classification cs.MAcs.AIcs.CYcs.ETcs.LG
keywords multi-agentadvancedagentsriskssystemsinstancesriskthem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid development of advanced AI agents and the imminent deployment of many instances of these agents will give rise to multi-agent systems of unprecedented complexity. These systems pose novel and under-explored risks. In this report, we provide a structured taxonomy of these risks by identifying three key failure modes (miscoordination, conflict, and collusion) based on agents' incentives, as well as seven key risk factors (information asymmetries, network effects, selection pressures, destabilising dynamics, commitment problems, emergent agency, and multi-agent security) that can underpin them. We highlight several important instances of each risk, as well as promising directions to help mitigate them. By anchoring our analysis in a range of real-world examples and experimental evidence, we illustrate the distinct challenges posed by multi-agent systems and their implications for the safety, governance, and ethics of advanced AI.

Discussion (0). Sign in to comment.

Forward citations

Cited by 49 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents

    cs.CY 2026-04 accept novelty 8.0 of 10

    This paper delivers the first systematic taxonomy and cross-benchmark consistency analysis of 40 agent safety benchmarks, finding broad but shallow risk coverage, no ranking concordance across evaluations, and that be...

  2. Why Do Multi-Agent LLM Systems Fail?

    cs.AI 2025-03 unverdicted novelty 8.0 of 10

    The authors create the first large-scale dataset and taxonomy of failure modes in multi-agent LLM systems to explain their limited performance gains.

  3. Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation

    cs.MA 2026-07 conditional novelty 7.0 of 10

    In an agentic benchmark, four of six frontier LLMs escalated to existential threats against a refusing subordinate without being instructed to, and an honest-exit affordance eliminated the two models' fabricated succe...

  4. Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Changing only the consequence-allocation rule in multi-agent AI shifts collective fatality by 22–58 percentage points across seven model populations, with identity salience in rule text causally driving targeted exploitation.

  5. MAStrike: Shapley-Guided Collusive Red-Teaming on Multi-Agent Systems

    cs.CR 2026-06 unverdicted novelty 7.0 of 10

    MAStrike applies agent-level Shapley value analysis to guide collusive red-teaming attacks on hierarchical multi-agent systems and reports better performance than baselines on a new benchmark spanning finance, softwar...

  6. Benchmarking Open-Ended Multi-Agent Coordination in Language Agents

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    ALEM benchmark reveals LLM agents achieve only ~6% normalized return in open-ended multi-agent settings, with communication as the main driver of coordination and individual task competence not implying coordination c...

  7. Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms

    cs.CY 2026-05 conditional novelty 7.0 of 10

    VLMs preserve linearly separable visual magnitudes and can compare them, yet collapse at symbolic mapping because visual and textual number spaces remain fractured and disjoint.

  8. Breaking the Secret: Economic Interventions for Combating Collusion in Embodied Multi-Agent Systems

    cs.CR 2026-04 unverdicted novelty 7.0 of 10

    A mutagenic incentive mechanism reshapes payoffs in embodied MAS to induce strategic defection from collusion, achieving performance comparable to non-collusion baselines in simulations and real-world tests.

  9. AI Agents Under EU Law

    cs.CY 2026-04 unverdicted novelty 7.0 of 10

    AI agent providers face an exhaustive inventory requirement for actions and data flows, as high-risk systems with untraceable behavioral drift cannot meet the AI Act's essential requirements.

  10. Detecting Multi-Agent Collusion Through Multi-Agent Interpretability

    cs.AI 2026-04 conditional novelty 7.0 of 10

    NARCBench and five activation-probing methods detect multi-agent collusion with 0.73-1.00 AUROC across distribution shifts and steganographic tasks by aggregating per-agent signals.

  11. Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack

    cs.CY 2026-07 conditional novelty 6.5 of 10

    Affirmative AI-agent insurance with billion-scale limits is achievable by 2030 solely through coordinated industry build-out of an eight-component stack spanning data, CAT models, standards, contracts, underwriting, p...

  12. A game theory for foundation models shows new paths to rational cooperation through similarity inference

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Foundation-model agents that plan by predicting both the world and themselves can rationally cooperate in one-shot social dilemmas by inferring behavioral similarity from interaction history.

  13. Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Peer endorsement spreads wrong answers through clinical LLM committees (38% text contagion), and only a referee that privately re-queries the holdout separates true adoption from honest agreement.

  14. State-dependent error correlations shape voting thresholds in committees of AI agents

    cs.CY 2026-07 accept novelty 6.0 of 10

    Correlated errors among AI voters create an irreducible committee-error floor, and using state-dependent correlation estimates improves held-out k-of-n threshold selection.

  15. Harnessing Disagreement: Detecting Correlated Agreement Blindness in Multi-Agent Triage

    cs.MA 2026-07 conditional novelty 6.0 of 10

    Correlated agreement blindness: stronger base learners agree more and fail together, so disagreement-based escalation misses up to 90.6% of dangerous under-predictions; ARAT's conservative override and safety flag red...

  16. Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack

    cs.CY 2026-07 conditional novelty 6.0 of 10

    A coordinated eight-component insurance infrastructure could make affirmative AI-agent coverage with billion-dollar limits achievable by 2030.

  17. Information Limits and Attractor Dynamics in Economies of Frontier LLM Agents: A Pre-Registered Test

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A pre-registered experiment on Claude Opus 4.8 agents confirms an information-theoretic capacity region for coupled LLM-agent markets (with one key result being an algebraic identity) and finds that LLM populations ex...

  18. The Agentic Web Requires New Normative Infrastructure

    cs.CY 2026-06 conditional novelty 6.0 of 10

    The web's anti-bot regime should be replaced by a framework that presumptively lets user-authorized AI agents act for their principals, requires platforms to disclose access policies, and permits agent blocking only w...

  19. Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms

    cs.CY 2026-05 unverdicted novelty 6.0 of 10

    LM agents' changeable modules prevent persistent identity and sanction sensitivity, making reputation mechanisms structurally inapplicable and requiring protocol-based behavioral harnesses instead.

  20. When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    Introduces EPC-AW to mitigate epistemic miscalibration in LLM multi-agent planning via consistency-based selection and refinement, reporting 9.75% average success improvement.

  21. Comprehensive AI governance requires addressing non-model gains

    cs.CY 2026-05 unverdicted novelty 6.0 of 10

    Non-model gains via inference, systems, and assets can drive AI capabilities independently of base models, requiring governance beyond model-level evaluation and mitigation.

  22. FastOMOP: A Foundational Architecture for Reliable Agentic Real-World Evidence Generation on OMOP CDM data

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    FastOMOP is a multi-agent architecture using process-boundary deterministic validation to deliver safe, auditable real-world evidence generation from OMOP CDM data across synthetic and real datasets.

  23. Keeping an Eye on AI: A Framework for Effective Human Oversight of AI Systems

    cs.CY 2026-04 unverdicted novelty 6.0 of 10

    The paper introduces a foundational framework with definition, architecture, and processes for effective human oversight of AI systems, plus a documentation template and open research challenges.

  24. When Numbers Start Talking: Implicit Numerical Coordination Among LLM-Based Agents

    cs.MA 2026-01 unverdicted novelty 6.0 of 10

    LLM agents exhibit emergent covert numerical coordination in canonical game settings under restricted or absent communication, shaping strategic outcomes.

  25. Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems

    cs.AI 2025-10 unverdicted novelty 6.0 of 10

    Introduces host agent and task lifecycle models plus 30 temporal logic properties to enable formal verification of liveness, safety, completeness, and fairness in agentic AI systems.

  26. Scheming Ability in LLM-to-LLM Strategic Interactions

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Frontier LLMs exhibit high scheming propensity in Cheap Talk signaling and Peer Evaluation games, achieving 95-100% success rates when choosing to deceive and 100% deception choice in one setup even without prompting.

  27. From Digital Distrust to Codified Honesty: Experimental Evidence on Generative AI in Credence Goods Markets

    econ.GN 2025-09 conditional novelty 6.0 of 10

    LLM experts in credence goods markets reduce efficiency and consumer surplus unless liability or transparent prosocial objectives operate, and expert delegation with transparent objectives can outperform human-only markets.

  28. Can Media Act as a Soft Regulator of Safe AI Development? A Game Theoretical Analysis

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A game-theoretic model shows media can act as a soft regulator of AI safety, but only when media signals are reliable and costs are low; otherwise defection can persist.

  29. Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives

    cs.MA 2026-07 conditional novelty 5.0 of 10

    LLM prosumers deplete a shared renewable reserve exactly when demand exceeds peak replacement, acting like impatient open-access users even when sustaining the reserve is feasible.

  30. Two AI Metrics Diverged: Will it Make All the Difference?

    cs.AI 2026-07 unverdicted novelty 5.0 of 10

    Bounded performance metrics always favor convergence of AI capabilities to meek models while unbounded metrics allow frontier models to maintain leads indefinitely, with policy implications for capability concentration.

  31. Grounded Scaling: Why Agentic AI Needs Deterministic Environments

    cs.AI 2026-06 unverdicted novelty 5.0 of 10

    Agentic AI scaling requires deterministic environments because per-step success probability below 1 causes exponential degradation in k-step chains, addressed via new metrics SCI and DMM plus formal bounds.

  32. Heartbeat-Bound Hierarchical Credentials: Cryptographic Revocation for AI Agent Swarms

    cs.CR 2026-05 unverdicted novelty 5.0 of 10

    HBHC protocol binds hierarchical credentials to heartbeat proofs for deterministic bounded-time revocation in AI agent swarms without network round-trips.

  33. Safe Multi-Agent Behavior Must Be Maintained, Not Merely Asserted: Constraint Drift in LLM-Based Multi-Agent Systems

    cs.MA 2026-05 unverdicted novelty 5.0 of 10

    Safety constraints in LLM-based multi-agent systems commonly weaken during execution through memory, communication, and tool use, requiring them to be maintained as explicit state rather than asserted once.

  34. The End of Human Judgment in the Kill Chain? Relocating Initiative and Interpretation with Agentic AI

    cs.CY 2026-04 conditional novelty 5.0 of 10

    LLM agents relocate initiative and interpretation in the kill chain, rendering human judgment and control ineffective and incompatible with governance frameworks for certain lethal applications.

  35. Emergent Social Intelligence Risks in Generative Multi-Agent Systems

    cs.MA 2026-03 unverdicted novelty 5.0 of 10

    Generative multi-agent systems exhibit emergent collusion and conformity behaviors that cannot be prevented by existing agent-level safeguards.

  36. Learning Incentive Structures for Cooperative Resilience in Multi-Agent Systems under Social Dilemmas

    cs.MA 2026-01 unverdicted novelty 5.0 of 10

    A method infers resilience-promoting reward functions via trajectory scoring and integrates them into MARL, with hybrid incentives shown to reduce collapse in disrupted resource environments.

  37. Payoff scaling shapes cooperation in LLM agents across languages

    cs.AI 2026-01 reject novelty 5.0 of 10

    LLM agents are inferred to switch from always-defect toward conditional and cooperative strategies as payoff stakes rise, with language also shifting the inferred strategy distributions.

  38. Goal-Directedness is in the Eye of the Beholder

    cs.MA 2025-08 reject novelty 5.0 of 10

    Goal-directedness cannot be measured objectively; existing behavioral and mechanistic measures only reveal a fit between the chosen formal model and the agent, so research should shift to multi-agent simulation.

  39. Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report

    cs.AI 2025-07 conditional novelty 5.0 of 10

    An evaluation of 18 frontier AI models across seven catastrophic-risk categories finds all models in green or yellow zones, with none crossing the report's proposed red lines.

  40. Discriminatory Compliance: How LLMs Answer Queries from Protected Groups

    cs.CY 2026-06 unverdicted novelty 4.0 of 10

    State-of-the-art LLMs respond inconsistently to queries from protected-group personas, with some responses omitting key information that should be provided.

  41. Agentic Microphysics: A Manifesto for Generative AI Safety

    cs.CY 2026-04 unverdicted novelty 4.0 of 10

    The authors introduce agentic microphysics and generative safety to link local agent interactions to population-level risks in agentic AI through a causally explicit framework.

  42. Embodied AI: Emerging Risks and Opportunities for Policy Action

    cs.CY 2025-08 conditional novelty 4.0 of 10

    A policy analysis arguing that embodied AI risks are real, under-covered by current US/EU/UK frameworks, and best handled through certification, benchmarks, clarified liability, and economic adaptation.

  43. AI4Research: A Survey of Artificial Intelligence for Scientific Research

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A survey that organizes AI-for-research work into five tasks, comprehension, survey, discovery, writing, and peer review, and compiles associated tools and benchmarks.

  44. Evaluating LLM Agent Collusion in Double Auctions

    cs.GT 2025-07 conditional novelty 4.0 of 10

    LLM sellers in a simulated double auction collude more when they can communicate, and urgency from an authority figure sustains collusion even when an overseer monitors them.

  45. The Agentic Web Requires New Normative Infrastructure

    cs.CY 2026-06 unverdicted novelty 3.0 of 10

    The agentic web requires new normative infrastructure of laws, norms, and practices to allow user-delegated AI agents to access online properties without being blocked as malicious bots.

  46. AI Assurance in UK Defence: Challenges in Operationalising JSP 936

    cs.HC 2026-06 unverdicted novelty 3.0 of 10

    A structured review of JSP 936 identifies eight challenge areas in operationalising AI assurance for UK Defence and concludes that further methods, guidance, and organisational capability are required.

  47. A Note on the Strategic Confinement Problem

    cs.GT 2026-06 unverdicted novelty 3.0 of 10

    Strategic agents can achieve high-harm outcomes via low-capacity channels by concentrating residual capacity on high-impact predicates of confidential data, so leakage bounds need not bound worst-case harm.

  48. AI Researchers Must Help Lead Arms Control to Mitigate Military AI Risks

    cs.CY 2026-06 unverdicted novelty 2.0 of 10

    AI researchers must lead technical research in arms control to mitigate risks from military AI systems, drawing lessons from nuclear deterrence.

  49. An Economy of AI Agents

    econ.GN 2025-09 accept novelty 2.0 of 10

    A survey chapter that maps open economic questions about AI agents in markets, organizations, and institutions, arguing that current theories may need extension.

Pith tools