Pith. sign in

REVIEW 31 cited by

Is Independent Learning All You Need in the StarCraft Multi-Agent Challenge?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.09533 v1 pith:P7TTZ2I5 submitted 2020-11-18 cs.AI

classification cs.AI
keywords learningindependentippomulti-agentapproachescentralizedfunctionjoint
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Most recently developed approaches to cooperative multi-agent reinforcement learning in the \emph{centralized training with decentralized execution} setting involve estimating a centralized, joint value function. In this paper, we demonstrate that, despite its various theoretical shortcomings, Independent PPO (IPPO), a form of independent learning in which each agent simply estimates its local value function, can perform just as well as or better than state-of-the-art joint learning approaches on popular multi-agent benchmark suite SMAC with little hyperparameter tuning. We also compare IPPO to several variants; the results suggest that IPPO's strong performance may be due to its robustness to some forms of environment non-stationarity.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 185 citations worldwide. Full citation record

  1. Aggregate in the Advantage, Not the Ratio: A Canonical-Form Analysis of Cooperative Multi-Agent Policy Optimization

    cs.MA 2026-07 conditional novelty 7.0 of 10

    For cooperative PPO, the expected gradient at the on-policy point depends on advantage and ratio aggregation supports only through their matrix product, and the variance-optimal design keeps the ratio per-agent and ag...

  2. Dreamer-CPC: Message Learning with World Models for Decentralized Multi-agent Reinforcement Learning

    cs.MA 2026-07 conditional novelty 6.0 of 10

    Dreamer-CPC has each decentralized agent send messages drawn from its learned world-model memory, outperforming current-observation messaging baselines, especially when key observations are temporarily missing.

  3. Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A tabular UCB controller trained on task success improves LLM-agent memory use over fixed heuristics, without extra LLM calls.

  4. Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Online action-space factorization via Kalman-refined cross-capacitance lets shared multi-agent policies zero-shot tune larger quantum-dot arrays with near-constant steps.

  5. Contract-Based Compositional Shielding for Safe Multi-Agent Reinforcement Learning

    cs.LG 2026-06 conditional novelty 6.0 of 10

    Certified local LTL_safe contracts, jointly fixed-point checked and selected by a bandit, recover coordinated safe team policies under decentralised multi-agent RL execution.

  6. TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning

    cs.MA 2026-02 unverdicted novelty 6.0 of 10

    TABX is a JAX-based, GPU-accelerated, configurable multi-agent battle simulator that lets researchers vary units, terrain, and physics to benchmark cooperative MARL algorithms.

  7. Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic

    cs.AI 2026-01 unverdicted novelty 6.0 of 10

    Multi-agent actor-critic methods with a centralized critic improve decentralized LLM collaboration over Monte Carlo baselines in long-horizon and sparse-reward settings.

  8. Automata-Conditioned Cooperative Multi-Agent Reinforcement Learning

    cs.MA 2025-11 conditional novelty 6.0 of 10

    ACC-MARL trains decentralized multi-agent policies that solve many automaton-specified cooperative tasks at once, with a proof of optimality for the Markovian reformulation and value-based task assignment.

  9. Hierarchical Message-Passing Policies for Multi-Agent Reinforcement Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A feudal hierarchical MARL method where lower-level policies are rewarded with the upper level's advantage function, with theoretical alignment guarantees and strong benchmark results.

  10. Hierarchical Multi-Agent Reinforcement Learning with Control Barrier Functions for Safety-Critical Autonomous Systems

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A hierarchical multi-agent RL method that selects skills at a high level and enforces pointwise safety with learned CBF-QP policies achieves about 99 percent success in simulated traffic scenarios.

  11. Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A survey proposing adaptability as a three-part taxonomy (learning, policy, scenario-driven) for organizing and evaluating MARL under changing conditions.

  12. Modeling Latent Partner Strategies for Adaptive Zero-Shot Human-Agent Collaboration

    cs.AI 2025-07 conditional novelty 6.0 of 10

    An agent that represents teammate styles as latent clusters and updates its belief with fixed-share regret minimization outperformed baselines with unfamiliar human partners in Overcooked.

  13. Multi-Agent Reinforcement Learning for Inverse Design in Photonic Integrated Circuits

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A multi-agent bandit RL formulation for binary photonic topology optimization produces designs that outperform gradient-based optimization on eight simulated 2D and 3D tasks.

  14. Learned Controllers for Agile Quadrotors in Pursuit-Evasion Games

    cs.RO 2025-06 conditional novelty 6.0 of 10

    Population-based opponent sampling during drone pursuit-evasion training improves robustness against older and unseen strategies, with rate-based control outperforming velocity-based control in simulation.

  15. Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details

    cs.AI 2026-08 conditional novelty 5.0 of 10

    For Other-Play in Yokai, agents trained with different implementation details coordinate across implementations about as well as across seeds, supporting inter-seed cross-play as a proxy for cross-implementation evaluation.

  16. Diffusing to Coordinate: Efficient Online Multi-Agent Diffusion Policies

    cs.AI 2026-02 conditional novelty 5.0 of 10

    A diffusion-policy multi-agent RL framework substitutes an ELBO for intractable joint entropy and reports 2.5–5× sample-efficiency gains on 10 MPE/MAMuJoCo tasks.

  17. CoRL-MPPI: Enhancing MPPI With Learnable Behaviours For Efficient And Provably-Safe Multi-Robot Collision Avoidance

    cs.RO 2025-11 conditional novelty 5.0 of 10

    CoRL-MPPI injects a learned cooperative policy into MPPI's sampling distribution to speed up and make safer multi-robot navigation in dense simulations.

  18. Self-Supervised Goal-Reaching Results in Multi-Agent Cooperation and Exploration

    cs.LG 2025-09 conditional novelty 5.0 of 10

    Self-supervised multi-agent goal-reaching, where each agent independently learns a contrastive critic of its own observations, achieves cooperation and exploration in sparse-reward MARL tasks where standard baselines fail.

  19. cMALC-D: Contextual Multi-Agent LLM-Guided Curriculum Learning with Diversity-Based Context Blending

    cs.LG 2025-08 reject novelty 5.0 of 10

    cMALC-D uses an LLM to generate training contexts for multi-agent RL and a diversity-blending mechanism to avoid mode collapse, claiming improved generalization on traffic signal control.

  20. Multi-Agent Trust Region Policy Optimisation: A Joint Constraint Approach

    cs.AI 2025-08 unverdicted novelty 5.0 of 10

    Adaptive per-agent KL-threshold allocation via KKT (HATRPO-W) and greedy (HATRPO-G) improves HATRPO's final reward by over 22.5% in MARL benchmarks.

  21. From General Relation Patterns to Task-Specific Decision-Making in Continual Multi-Agent Coordination

    cs.MA 2025-07 conditional novelty 5.0 of 10

    RPG uses a relation capturer plus a task-conditioned hypernetwork to reduce catastrophic forgetting and enable zero-shot transfer in continual multi-agent coordination.

  22. Light Aircraft Game : Basic Implementation and training results analysis

    cs.LG 2025-06 reject novelty 5.0 of 10

    In the new LAG air-combat environment, HASAC scores higher than HAPPO in no-weapon coordination tasks while HAPPO scores higher in missile combat, but the results come from single runs without error bars.

  23. Shapley Machine: A Game-Theoretic Framework for N-Agent Ad Hoc Teamwork

    cs.MA 2025-06 conditional novelty 5.0 of 10

    Shapley Machine reshapes rewards and TD targets so per-agent value functions approximately satisfy the Shapley axioms, and it beats the prior POAM baseline in several NAHT test environments.

  24. Ego-centric Learning of Communicative World Models for Autonomous Driving

    cs.RO 2025-06 reject novelty 5.0 of 10

    Sharing compressed latent states and planned waypoints between agents, triggered by prediction errors, improves multi-agent driving performance in CARLA while cutting communication bandwidth by roughly 50x.

  25. Reward-Independent Messaging for Decentralized Multi-Agent Reinforcement Learning

    cs.MA 2025-05 conditional novelty 5.0 of 10

    MARL-CPC lets decentralized agents learn to send informative messages through a self-supervised reconstruction objective, and outperforms message-as-action baselines in non-cooperative multi-agent tasks.

  26. SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control

    cs.AI 2025-08 conditional novelty 4.0 of 10

    A multi-agent RL workflow that interleaves single-agent updates, applied to mobile GUI control, achieves SOTA zero-shot performance and a +14.8 MATH500 gain.

  27. Multi-agent Reinforcement Learning for Robotized Coral Reef Sample Collection

    cs.RO 2025-07 conditional novelty 4.0 of 10

    An RL controller for coral sample collection, trained in Unity, is transferred zero-shot to a physical BlueROV2 using real-time underwater motion capture to drive the digital twin.

  28. PILOC: A Pheromone Inverse Guidance Mechanism and Local-Communication Framework for Dynamic Target Search of Multi-Agent in Unknown Environments

    cs.RO 2025-07 conditional novelty 4.0 of 10

    A decentralized reinforcement-learning framework using virtual pheromone marks and local map sharing outperforms MARL baselines for finding dynamic targets in simulated unknown grid environments.

  29. GenAI-based Multi-Agent Reinforcement Learning towards Distributed Agent Intelligence: A Generative-RL Agent Perspective

    cs.AI 2025-07 unverdicted novelty 3.0 of 10

    A position paper claiming that generative-AI agents that model and predict multi-agent dynamics will replace today's reactive MARL approaches.

  30. Multi-Agent Reinforcement Learning in Cybersecurity: From Fundamentals to Applications

    cs.MA 2025-05 conditional novelty 3.0 of 10

    A narrative survey of multi-agent reinforcement learning for cyber defense, reviewing game-theoretic models, cyber gyms, and applications, concluding MARL is promising but faces scalability and simulation-to-real tran...

  31. MEAL: A Benchmark for Continual Multi-Agent Reinforcement Learning

    cs.AI 2025-06

Pith tools