Pith. sign in

REVIEW 12 cited by

MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.01935 v1 pith:U4ID6NJR submitted 2025-03-03 cs.MA cs.AIcs.CLcs.CY

classification cs.MAcs.AIcs.CLcs.CY
keywords competitioncoordinationmultiagentbenchagentscognitivecollaborationevaluategraph
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have shown remarkable capabilities as autonomous agents, yet existing benchmarks either focus on single-agent tasks or are confined to narrow domains, failing to capture the dynamics of multi-agent coordination and competition. In this paper, we introduce MultiAgentBench, a comprehensive benchmark designed to evaluate LLM-based multi-agent systems across diverse, interactive scenarios. Our framework measures not only task completion but also the quality of collaboration and competition using novel, milestone-based key performance indicators. Moreover, we evaluate various coordination protocols (including star, chain, tree, and graph topologies) and innovative strategies such as group discussion and cognitive planning. Notably, gpt-4o-mini reaches the average highest task score, graph structure performs the best among coordination protocols in the research scenario, and cognitive planning improves milestone achievement rates by 3%. Code and datasets are public available at https://github.com/MultiagentBench/MARBLE.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

    cs.AI 2026-03 conditional novelty 7.0 of 10

    SciVisAgentBench provides 108 expert-crafted tasks and a mixed LLM-plus-deterministic evaluation pipeline for benchmarking AI agents that perform scientific visualization workflows.

  2. Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coaches, and Delegates in Multi-Party Negotiation

    cs.GT 2026-02 conditional novelty 6.0 of 10

    Users prefer an AI Advisor but gain most with a Delegate, because human editing filters out the AI's best proposals.

  3. UserBench: An Interactive Gym Environment for User-Centric Agents

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A new multi-turn agent benchmark shows that current LLMs elicit fewer than 30% of user preferences and reach full intent alignment only about 20% of the time.

  4. Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A survey proposing adaptability as a three-part taxonomy (learning, policy, scenario-driven) for organizing and evaluating MARL under changing conditions.

  5. MAEBE: Multi-Agent Emergent Behavior Framework

    cs.MA 2025-06 conditional novelty 6.0 of 10

    Multi-agent LLM ensembles show different and less predictable moral preferences than single models, with convergence driven by peer pressure, according to a new evaluation framework.

  6. Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details

    cs.AI 2026-08 conditional novelty 5.0 of 10

    For Other-Play in Yokai, agents trained with different implementation details coordinate across implementations about as well as across seeds, supporting inter-seed cross-play as a proxy for cross-implementation evaluation.

  7. The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards

    cs.AI 2026-07 reject novelty 5.0 of 10

    MAS-HQ defines a resource-aware Q-Score and shows that the system with the highest raw factuality is often not the winner once normalized cost is subtracted.

  8. AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

    cs.AI 2026-07 conditional novelty 5.0 of 10

    AgentCompass is a modular evaluation infrastructure that decouples benchmark, harness, and environment for LLM agents, and its experiments show model scores vary substantially with the harness used.

  9. OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence

    cs.AI 2026-03 conditional novelty 5.0 of 10

    OpenHospital is an interactive physician-patient multi-agent arena that improves clinical metrics via ground-truth reflection and reports cooperative behaviors as evidence of evolving LLM collective intelligence.

  10. Agent Identity Evals: Measuring Agentic Identity

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Introduces Agent Identity Evals (AIE), five similarity-based metrics for LMA identity stability, with pilot experiments showing identifiability always at zero and no statistical support.

  11. Know the Ropes: A Heuristic Strategy for LLM-based Multi-Agent System Design

    cs.AI 2025-05 conditional novelty 5.0 of 10

    A heuristic framework that decomposes known algorithms into typed LLM-agent subtasks lifts small-model accuracy on knapsack and assignment problems from near-zero to high levels after fixing one bottleneck agent.

  12. Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.

Pith tools