REVIEW 8 cited by
LLM-Coordination: Evaluating and Analyzing Multi-agent Coordination Abilities in Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) have demonstrated emergent common-sense reasoning and Theory of Mind (ToM) capabilities, making them promising candidates for developing coordination agents. This study introduces the LLM-Coordination Benchmark, a novel benchmark for analyzing LLMs in the context of Pure Coordination Settings, where agents must cooperate to maximize gains. Our benchmark evaluates LLMs through two distinct tasks. The first is Agentic Coordination, where LLMs act as proactive participants in four pure coordination games. The second is Coordination Question Answering (CoordQA), which tests LLMs on 198 multiple-choice questions across these games to evaluate three key abilities: Environment Comprehension, ToM Reasoning, and Joint Planning. Results from Agentic Coordination experiments reveal that LLM-Agents excel in multi-agent coordination settings where decision-making primarily relies on environmental variables but face challenges in scenarios requiring active consideration of partners' beliefs and intentions. The CoordQA experiments further highlight significant room for improvement in LLMs' Theory of Mind reasoning and joint planning capabilities. Zero-Shot Coordination (ZSC) experiments in the Agentic Coordination setting demonstrate that LLM agents, unlike RL methods, exhibit robustness to unseen partners. These findings indicate the potential of LLMs as Agents in pure coordination setups and underscore areas for improvement. Code Available at https://github.com/eric-ai-lab/llm_coordination.
Forward citations
Cited by 8 Pith papers
-
Learning social norms enhances compatibility in dynamic human-AI coordination
Encoding three extracted social-norm principles into LLMs enables near-4x better human-AI coordination in a dynamic pedestrian-vehicle game, surpassing human-human baselines.
-
Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coaches, and Delegates in Multi-Party Negotiation
Users prefer an AI Advisor but gain most with a Delegate, because human editing filters out the AI's best proposals.
-
Tacit Coordination of Large Language Models
Across 20+ open-source LLMs, tacit coordination in focal-point games is often at or above human levels, with systematic failures on cultural and numerical salience that culture prompts partially fix.
-
When Identity Overrides Incentives: Representational Choices as Governance Decisions in Multi-Agent LLM Systems
Role-based personas in multi-agent LLM systems suppress payoff-aligned behavior, shifting equilibrium selection by up to 90 percentage points in Tragedy of the Commons versus Green Transition scenarios even with full ...
-
PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments
A new open Minecraft benchmark for 2v2 LLM-agent competition, and a system, TactiCrafter, that beats its baselines on points and win rate.
-
AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs
A benchmark built from five distributed computing problems shows that frontier LLM agent networks solve small coordination tasks but break down as the network scales to 100 agents.
-
TextAtari: 100K Frames Game Playing with Language Agents
TextAtari is a text-based Atari benchmark for language agents; 7-8B LLMs stay below 10% of human scores in over 90% of tested conditions, and knowledge injection helps more than chain-of-thought.
-
Structuring the Unstructured: A Multi-Agent System for Extracting and Querying Financial KPIs and Guidance
A multi-agent LLM system with hand-crafted rule validation reports about 95% extraction accuracy and 91% correct query answers, but only on a private, unreleased dataset.
Discussion (0). Continue with ORCID to comment.