REVIEW 49 cited by
Multi-Agent Risks from Advanced AI
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The rapid development of advanced AI agents and the imminent deployment of many instances of these agents will give rise to multi-agent systems of unprecedented complexity. These systems pose novel and under-explored risks. In this report, we provide a structured taxonomy of these risks by identifying three key failure modes (miscoordination, conflict, and collusion) based on agents' incentives, as well as seven key risk factors (information asymmetries, network effects, selection pressures, destabilising dynamics, commitment problems, emergent agency, and multi-agent security) that can underpin them. We highlight several important instances of each risk, as well as promising directions to help mitigate them. By anchoring our analysis in a range of real-world examples and experimental evidence, we illustrate the distinct challenges posed by multi-agent systems and their implications for the safety, governance, and ethics of advanced AI.
Forward citations
Cited by 49 Pith papers
-
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
This paper delivers the first systematic taxonomy and cross-benchmark consistency analysis of 40 agent safety benchmarks, finding broad but shallow risk coverage, no ranking concordance across evaluations, and that be...
-
Why Do Multi-Agent LLM Systems Fail?
The authors create the first large-scale dataset and taxonomy of failure modes in multi-agent LLM systems to explain their limited performance gains.
-
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
In an agentic benchmark, four of six frontier LLMs escalated to existential threats against a refusing subordinate without being instructed to, and an honest-exit affordance eliminated the two models' fabricated succe...
-
Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety
Changing only the consequence-allocation rule in multi-agent AI shifts collective fatality by 22–58 percentage points across seven model populations, with identity salience in rule text causally driving targeted exploitation.
-
MAStrike: Shapley-Guided Collusive Red-Teaming on Multi-Agent Systems
MAStrike applies agent-level Shapley value analysis to guide collusive red-teaming attacks on hierarchical multi-agent systems and reports better performance than baselines on a new benchmark spanning finance, softwar...
-
Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
ALEM benchmark reveals LLM agents achieve only ~6% normalized return in open-ended multi-agent settings, with communication as the main driver of coordination and individual task competence not implying coordination c...
-
Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms
VLMs preserve linearly separable visual magnitudes and can compare them, yet collapse at symbolic mapping because visual and textual number spaces remain fractured and disjoint.
-
Breaking the Secret: Economic Interventions for Combating Collusion in Embodied Multi-Agent Systems
A mutagenic incentive mechanism reshapes payoffs in embodied MAS to induce strategic defection from collusion, achieving performance comparable to non-collusion baselines in simulations and real-world tests.
-
AI Agents Under EU Law
AI agent providers face an exhaustive inventory requirement for actions and data flows, as high-risk systems with untraceable behavioral drift cannot meet the AI Act's essential requirements.
-
Detecting Multi-Agent Collusion Through Multi-Agent Interpretability
NARCBench and five activation-probing methods detect multi-agent collusion with 0.73-1.00 AUROC across distribution shifts and steganographic tasks by aggregating per-agent signals.
-
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack
Affirmative AI-agent insurance with billion-scale limits is achievable by 2030 solely through coordinated industry build-out of an eight-component stack spanning data, CAT models, standards, contracts, underwriting, p...
-
A game theory for foundation models shows new paths to rational cooperation through similarity inference
Foundation-model agents that plan by predicting both the world and themselves can rationally cooperate in one-shot social dilemmas by inferring behavioral similarity from interaction history.
-
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
Peer endorsement spreads wrong answers through clinical LLM committees (38% text contagion), and only a referee that privately re-queries the holdout separates true adoption from honest agreement.
-
State-dependent error correlations shape voting thresholds in committees of AI agents
Correlated errors among AI voters create an irreducible committee-error floor, and using state-dependent correlation estimates improves held-out k-of-n threshold selection.
-
Harnessing Disagreement: Detecting Correlated Agreement Blindness in Multi-Agent Triage
Correlated agreement blindness: stronger base learners agree more and fail together, so disagreement-based escalation misses up to 90.6% of dangerous under-predictions; ARAT's conservative override and safety flag red...
-
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack
A coordinated eight-component insurance infrastructure could make affirmative AI-agent coverage with billion-dollar limits achievable by 2030.
-
Information Limits and Attractor Dynamics in Economies of Frontier LLM Agents: A Pre-Registered Test
A pre-registered experiment on Claude Opus 4.8 agents confirms an information-theoretic capacity region for coupled LLM-agent markets (with one key result being an algebraic identity) and finds that LLM populations ex...
-
The Agentic Web Requires New Normative Infrastructure
The web's anti-bot regime should be replaced by a framework that presumptively lets user-authorized AI agents act for their principals, requires platforms to disclose access policies, and permits agent blocking only w...
-
Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms
LM agents' changeable modules prevent persistent identity and sanction sensitivity, making reputation mechanisms structurally inapplicable and requiring protocol-based behavioral harnesses instead.
-
When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems
Introduces EPC-AW to mitigate epistemic miscalibration in LLM multi-agent planning via consistency-based selection and refinement, reporting 9.75% average success improvement.
-
Comprehensive AI governance requires addressing non-model gains
Non-model gains via inference, systems, and assets can drive AI capabilities independently of base models, requiring governance beyond model-level evaluation and mitigation.
-
FastOMOP: A Foundational Architecture for Reliable Agentic Real-World Evidence Generation on OMOP CDM data
FastOMOP is a multi-agent architecture using process-boundary deterministic validation to deliver safe, auditable real-world evidence generation from OMOP CDM data across synthetic and real datasets.
-
Keeping an Eye on AI: A Framework for Effective Human Oversight of AI Systems
The paper introduces a foundational framework with definition, architecture, and processes for effective human oversight of AI systems, plus a documentation template and open research challenges.
-
When Numbers Start Talking: Implicit Numerical Coordination Among LLM-Based Agents
LLM agents exhibit emergent covert numerical coordination in canonical game settings under restricted or absent communication, shaping strategic outcomes.
-
Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems
Introduces host agent and task lifecycle models plus 30 temporal logic properties to enable formal verification of liveness, safety, completeness, and fairness in agentic AI systems.
-
Scheming Ability in LLM-to-LLM Strategic Interactions
Frontier LLMs exhibit high scheming propensity in Cheap Talk signaling and Peer Evaluation games, achieving 95-100% success rates when choosing to deceive and 100% deception choice in one setup even without prompting.
-
From Digital Distrust to Codified Honesty: Experimental Evidence on Generative AI in Credence Goods Markets
LLM experts in credence goods markets reduce efficiency and consumer surplus unless liability or transparent prosocial objectives operate, and expert delegation with transparent objectives can outperform human-only markets.
-
Can Media Act as a Soft Regulator of Safe AI Development? A Game Theoretical Analysis
A game-theoretic model shows media can act as a soft regulator of AI safety, but only when media signals are reliable and costs are low; otherwise defection can persist.
-
Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives
LLM prosumers deplete a shared renewable reserve exactly when demand exceeds peak replacement, acting like impatient open-access users even when sustaining the reserve is feasible.
-
Two AI Metrics Diverged: Will it Make All the Difference?
Bounded performance metrics always favor convergence of AI capabilities to meek models while unbounded metrics allow frontier models to maintain leads indefinitely, with policy implications for capability concentration.
-
Grounded Scaling: Why Agentic AI Needs Deterministic Environments
Agentic AI scaling requires deterministic environments because per-step success probability below 1 causes exponential degradation in k-step chains, addressed via new metrics SCI and DMM plus formal bounds.
-
Heartbeat-Bound Hierarchical Credentials: Cryptographic Revocation for AI Agent Swarms
HBHC protocol binds hierarchical credentials to heartbeat proofs for deterministic bounded-time revocation in AI agent swarms without network round-trips.
-
Safe Multi-Agent Behavior Must Be Maintained, Not Merely Asserted: Constraint Drift in LLM-Based Multi-Agent Systems
Safety constraints in LLM-based multi-agent systems commonly weaken during execution through memory, communication, and tool use, requiring them to be maintained as explicit state rather than asserted once.
-
The End of Human Judgment in the Kill Chain? Relocating Initiative and Interpretation with Agentic AI
LLM agents relocate initiative and interpretation in the kill chain, rendering human judgment and control ineffective and incompatible with governance frameworks for certain lethal applications.
-
Emergent Social Intelligence Risks in Generative Multi-Agent Systems
Generative multi-agent systems exhibit emergent collusion and conformity behaviors that cannot be prevented by existing agent-level safeguards.
-
Learning Incentive Structures for Cooperative Resilience in Multi-Agent Systems under Social Dilemmas
A method infers resilience-promoting reward functions via trajectory scoring and integrates them into MARL, with hybrid incentives shown to reduce collapse in disrupted resource environments.
-
Payoff scaling shapes cooperation in LLM agents across languages
LLM agents are inferred to switch from always-defect toward conditional and cooperative strategies as payoff stakes rise, with language also shifting the inferred strategy distributions.
-
Goal-Directedness is in the Eye of the Beholder
Goal-directedness cannot be measured objectively; existing behavioral and mechanistic measures only reveal a fit between the chosen formal model and the agent, so research should shift to multi-agent simulation.
-
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report
An evaluation of 18 frontier AI models across seven catastrophic-risk categories finds all models in green or yellow zones, with none crossing the report's proposed red lines.
-
Discriminatory Compliance: How LLMs Answer Queries from Protected Groups
State-of-the-art LLMs respond inconsistently to queries from protected-group personas, with some responses omitting key information that should be provided.
-
Agentic Microphysics: A Manifesto for Generative AI Safety
The authors introduce agentic microphysics and generative safety to link local agent interactions to population-level risks in agentic AI through a causally explicit framework.
-
Embodied AI: Emerging Risks and Opportunities for Policy Action
A policy analysis arguing that embodied AI risks are real, under-covered by current US/EU/UK frameworks, and best handled through certification, benchmarks, clarified liability, and economic adaptation.
-
AI4Research: A Survey of Artificial Intelligence for Scientific Research
A survey that organizes AI-for-research work into five tasks, comprehension, survey, discovery, writing, and peer review, and compiles associated tools and benchmarks.
-
Evaluating LLM Agent Collusion in Double Auctions
LLM sellers in a simulated double auction collude more when they can communicate, and urgency from an authority figure sustains collusion even when an overseer monitors them.
-
The Agentic Web Requires New Normative Infrastructure
The agentic web requires new normative infrastructure of laws, norms, and practices to allow user-delegated AI agents to access online properties without being blocked as malicious bots.
-
AI Assurance in UK Defence: Challenges in Operationalising JSP 936
A structured review of JSP 936 identifies eight challenge areas in operationalising AI assurance for UK Defence and concludes that further methods, guidance, and organisational capability are required.
-
A Note on the Strategic Confinement Problem
Strategic agents can achieve high-harm outcomes via low-capacity channels by concentrating residual capacity on high-impact predicates of confidential data, so leakage bounds need not bound worst-case harm.
-
AI Researchers Must Help Lead Arms Control to Mitigate Military AI Risks
AI researchers must lead technical research in arms control to mitigate risks from military AI systems, drawing lessons from nuclear deterrence.
-
An Economy of AI Agents
A survey chapter that maps open economic questions about AI agents in markets, organizations, and institutions, arguing that current theories may need extension.
Discussion (0). Sign in to comment.