REVIEW 10 cited by
The Surprising Effectiveness of PPO in Cooperative, Multi-Agent Games
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Proximal Policy Optimization (PPO) is a ubiquitous on-policy reinforcement learning algorithm but is significantly less utilized than off-policy learning algorithms in multi-agent settings. This is often due to the belief that PPO is significantly less sample efficient than off-policy methods in multi-agent systems. In this work, we carefully study the performance of PPO in cooperative multi-agent settings. We show that PPO-based multi-agent algorithms achieve surprisingly strong performance in four popular multi-agent testbeds: the particle-world environments, the StarCraft multi-agent challenge, Google Research Football, and the Hanabi challenge, with minimal hyperparameter tuning and without any domain-specific algorithmic modifications or architectures. Importantly, compared to competitive off-policy methods, PPO often achieves competitive or superior results in both final returns and sample efficiency. Finally, through ablation studies, we analyze implementation and hyperparameter factors that are critical to PPO's empirical performance, and give concrete practical suggestions regarding these factors. Our results show that when using these practices, simple PPO-based methods can be a strong baseline in cooperative multi-agent reinforcement learning. Source code is released at \url{https://github.com/marlbenchmark/on-policy}.
Forward citations
Cited by 10 Pith papers
-
Multi-Agent Reinforcement Learning for SLA-Aware Network Slicing in UAV-Enabled MEC
Predictive MAPPO with D3QL mobility forecasts reduces slice SLA violation probability and duration in UAV-MEC simulations versus greedy, GA, and random baselines, nearing an informed oracle when forecasts are accurate.
-
Offline Nash Solvers Meet Online Tree Search in Multi-Agent Games on Graphs
Primitive-Guided Tree Search combines offline exact Nash solutions of 1v1/2v1 subgames with online SM-MCTS to produce coordinated multi-agent pursuit policies that outperform PSRO and neural MCTS baselines.
-
Multi-Agent Reinforcement Learning for V2X Resource Allocation: Disentangling MARL Challenges Through Benchmarking
Robustness to varied vehicle topologies, not partial observability or coordination, is the dominant obstacle for MARL in C-V2X spectrum sharing.
-
Towards Microgrid Resilience Enhancement via Mobile Power Sources and Repair Crews: A Multi-Agent Reinforcement Learning Approach
A hierarchical multi-agent reinforcement learning method coordinates mobile power sources and repair crews to restore loads after extreme events in coupled power-transport networks.
-
Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details
For Other-Play in Yokai, agents trained with different implementation details coordinate across implementations about as well as across seeds, supporting inter-seed cross-play as a proxy for cross-implementation evaluation.
-
COLMAR: Cooperative View Policy Learning for Multi-Agent Active 3D Reconstruction
A shared PPO policy with overlap-aware rewards improves multi-agent active 3D reconstruction coverage and accuracy in simulated indoor scenes.
-
Self-Supervised Goal-Reaching Results in Multi-Agent Cooperation and Exploration
Self-supervised multi-agent goal-reaching, where each agent independently learns a contrastive critic of its own observations, achieves cooperation and exploration in sparse-reward MARL tasks where standard baselines fail.
-
Improving monotonic optimization in heterogeneous multi-agent reinforcement learning with optimal marginal deterministic policy gradient
OMDPG combines optimal marginal Q-values with pessimistic Q-critics to reconcile monotonic improvement with partial parameter sharing in heterogeneous multi-agent RL.
-
Multi-agent Reinforcement Learning for Robotized Coral Reef Sample Collection
An RL controller for coral sample collection, trained in Unity, is transferred zero-shot to a physical BlueROV2 using real-time underwater motion capture to drive the digital twin.
-
Multi-Agent Reinforcement Learning in Cybersecurity: From Fundamentals to Applications
A narrative survey of multi-agent reinforcement learning for cyber defense, reviewing game-theoretic models, cyber gyms, and applications, concluding MARL is promising but faces scalability and simulation-to-real tran...
Discussion (0). Continue with ORCID to comment.