Pith. sign in

REVIEW 10 cited by

The Surprising Effectiveness of PPO in Cooperative, Multi-Agent Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.01955 v4 pith:I65B5MFL submitted 2021-03-02 cs.LG cs.AIcs.MA

classification cs.LGcs.AIcs.MA
keywords multi-agentcooperativelearningmethodsoff-policyperformancealgorithmschallenge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Proximal Policy Optimization (PPO) is a ubiquitous on-policy reinforcement learning algorithm but is significantly less utilized than off-policy learning algorithms in multi-agent settings. This is often due to the belief that PPO is significantly less sample efficient than off-policy methods in multi-agent systems. In this work, we carefully study the performance of PPO in cooperative multi-agent settings. We show that PPO-based multi-agent algorithms achieve surprisingly strong performance in four popular multi-agent testbeds: the particle-world environments, the StarCraft multi-agent challenge, Google Research Football, and the Hanabi challenge, with minimal hyperparameter tuning and without any domain-specific algorithmic modifications or architectures. Importantly, compared to competitive off-policy methods, PPO often achieves competitive or superior results in both final returns and sample efficiency. Finally, through ablation studies, we analyze implementation and hyperparameter factors that are critical to PPO's empirical performance, and give concrete practical suggestions regarding these factors. Our results show that when using these practices, simple PPO-based methods can be a strong baseline in cooperative multi-agent reinforcement learning. Source code is released at \url{https://github.com/marlbenchmark/on-policy}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 594 citations worldwide. Full citation record

  1. Multi-Agent Reinforcement Learning for SLA-Aware Network Slicing in UAV-Enabled MEC

    cs.NI 2026-07 conditional novelty 6.0 of 10

    Predictive MAPPO with D3QL mobility forecasts reduces slice SLA violation probability and duration in UAV-MEC simulations versus greedy, GA, and random baselines, nearing an informed oracle when forecasts are accurate.

  2. Offline Nash Solvers Meet Online Tree Search in Multi-Agent Games on Graphs

    cs.GT 2026-07 conditional novelty 6.0 of 10

    Primitive-Guided Tree Search combines offline exact Nash solutions of 1v1/2v1 subgames with online SM-MCTS to produce coordinated multi-agent pursuit policies that outperform PSRO and neural MCTS baselines.

  3. Multi-Agent Reinforcement Learning for V2X Resource Allocation: Disentangling MARL Challenges Through Benchmarking

    cs.MA 2026-02 conditional novelty 6.0 of 10

    Robustness to varied vehicle topologies, not partial observability or coordination, is the dominant obstacle for MARL in C-V2X spectrum sharing.

  4. Towards Microgrid Resilience Enhancement via Mobile Power Sources and Repair Crews: A Multi-Agent Reinforcement Learning Approach

    eess.SY 2025-07 conditional novelty 6.0 of 10

    A hierarchical multi-agent reinforcement learning method coordinates mobile power sources and repair crews to restore loads after extreme events in coupled power-transport networks.

  5. Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details

    cs.AI 2026-08 conditional novelty 5.0 of 10

    For Other-Play in Yokai, agents trained with different implementation details coordinate across implementations about as well as across seeds, supporting inter-seed cross-play as a proxy for cross-implementation evaluation.

  6. COLMAR: Cooperative View Policy Learning for Multi-Agent Active 3D Reconstruction

    cs.RO 2026-07 conditional novelty 5.0 of 10

    A shared PPO policy with overlap-aware rewards improves multi-agent active 3D reconstruction coverage and accuracy in simulated indoor scenes.

  7. Self-Supervised Goal-Reaching Results in Multi-Agent Cooperation and Exploration

    cs.LG 2025-09 conditional novelty 5.0 of 10

    Self-supervised multi-agent goal-reaching, where each agent independently learns a contrastive critic of its own observations, achieves cooperation and exploration in sparse-reward MARL tasks where standard baselines fail.

  8. Improving monotonic optimization in heterogeneous multi-agent reinforcement learning with optimal marginal deterministic policy gradient

    cs.AI 2025-07 reject novelty 5.0 of 10

    OMDPG combines optimal marginal Q-values with pessimistic Q-critics to reconcile monotonic improvement with partial parameter sharing in heterogeneous multi-agent RL.

  9. Multi-agent Reinforcement Learning for Robotized Coral Reef Sample Collection

    cs.RO 2025-07 conditional novelty 4.0 of 10

    An RL controller for coral sample collection, trained in Unity, is transferred zero-shot to a physical BlueROV2 using real-time underwater motion capture to drive the digital twin.

  10. Multi-Agent Reinforcement Learning in Cybersecurity: From Fundamentals to Applications

    cs.MA 2025-05 conditional novelty 3.0 of 10

    A narrative survey of multi-agent reinforcement learning for cyber defense, reviewing game-theoretic models, cyber gyms, and applications, concluding MARL is promising but faces scalability and simulation-to-real tran...

Pith tools