REVIEW 31 cited by
Is Independent Learning All You Need in the StarCraft Multi-Agent Challenge?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Most recently developed approaches to cooperative multi-agent reinforcement learning in the \emph{centralized training with decentralized execution} setting involve estimating a centralized, joint value function. In this paper, we demonstrate that, despite its various theoretical shortcomings, Independent PPO (IPPO), a form of independent learning in which each agent simply estimates its local value function, can perform just as well as or better than state-of-the-art joint learning approaches on popular multi-agent benchmark suite SMAC with little hyperparameter tuning. We also compare IPPO to several variants; the results suggest that IPPO's strong performance may be due to its robustness to some forms of environment non-stationarity.
Forward citations
Cited by 31 Pith papers
-
Aggregate in the Advantage, Not the Ratio: A Canonical-Form Analysis of Cooperative Multi-Agent Policy Optimization
For cooperative PPO, the expected gradient at the on-policy point depends on advantage and ratio aggregation supports only through their matrix product, and the variance-optimal design keeps the ratio per-agent and ag...
-
Dreamer-CPC: Message Learning with World Models for Decentralized Multi-agent Reinforcement Learning
Dreamer-CPC has each decentralized agent send messages drawn from its learned world-model memory, outperforming current-observation messaging baselines, especially when key observations are temporarily missing.
-
Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents
A tabular UCB controller trained on task success improves LLM-agent memory use over fixed heuristics, without extra LLM calls.
-
Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning
Online action-space factorization via Kalman-refined cross-capacitance lets shared multi-agent policies zero-shot tune larger quantum-dot arrays with near-constant steps.
-
Contract-Based Compositional Shielding for Safe Multi-Agent Reinforcement Learning
Certified local LTL_safe contracts, jointly fixed-point checked and selected by a bandit, recover coordinated safe team policies under decentralised multi-agent RL execution.
-
TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning
TABX is a JAX-based, GPU-accelerated, configurable multi-agent battle simulator that lets researchers vary units, terrain, and physics to benchmark cooperative MARL algorithms.
-
Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic
Multi-agent actor-critic methods with a centralized critic improve decentralized LLM collaboration over Monte Carlo baselines in long-horizon and sparse-reward settings.
-
Automata-Conditioned Cooperative Multi-Agent Reinforcement Learning
ACC-MARL trains decentralized multi-agent policies that solve many automaton-specified cooperative tasks at once, with a proof of optimality for the Markovian reformulation and value-based task assignment.
-
Hierarchical Message-Passing Policies for Multi-Agent Reinforcement Learning
A feudal hierarchical MARL method where lower-level policies are rewarded with the upper level's advantage function, with theoretical alignment guarantees and strong benchmark results.
-
Hierarchical Multi-Agent Reinforcement Learning with Control Barrier Functions for Safety-Critical Autonomous Systems
A hierarchical multi-agent RL method that selects skills at a high level and enforces pointwise safety with learned CBF-QP policies achieves about 99 percent success in simulated traffic scenarios.
-
Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review
A survey proposing adaptability as a three-part taxonomy (learning, policy, scenario-driven) for organizing and evaluating MARL under changing conditions.
-
Modeling Latent Partner Strategies for Adaptive Zero-Shot Human-Agent Collaboration
An agent that represents teammate styles as latent clusters and updates its belief with fixed-share regret minimization outperformed baselines with unfamiliar human partners in Overcooked.
-
Multi-Agent Reinforcement Learning for Inverse Design in Photonic Integrated Circuits
A multi-agent bandit RL formulation for binary photonic topology optimization produces designs that outperform gradient-based optimization on eight simulated 2D and 3D tasks.
-
Learned Controllers for Agile Quadrotors in Pursuit-Evasion Games
Population-based opponent sampling during drone pursuit-evasion training improves robustness against older and unseen strategies, with rate-based control outperforming velocity-based control in simulation.
-
Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details
For Other-Play in Yokai, agents trained with different implementation details coordinate across implementations about as well as across seeds, supporting inter-seed cross-play as a proxy for cross-implementation evaluation.
-
Diffusing to Coordinate: Efficient Online Multi-Agent Diffusion Policies
A diffusion-policy multi-agent RL framework substitutes an ELBO for intractable joint entropy and reports 2.5–5× sample-efficiency gains on 10 MPE/MAMuJoCo tasks.
-
CoRL-MPPI: Enhancing MPPI With Learnable Behaviours For Efficient And Provably-Safe Multi-Robot Collision Avoidance
CoRL-MPPI injects a learned cooperative policy into MPPI's sampling distribution to speed up and make safer multi-robot navigation in dense simulations.
-
Self-Supervised Goal-Reaching Results in Multi-Agent Cooperation and Exploration
Self-supervised multi-agent goal-reaching, where each agent independently learns a contrastive critic of its own observations, achieves cooperation and exploration in sparse-reward MARL tasks where standard baselines fail.
-
cMALC-D: Contextual Multi-Agent LLM-Guided Curriculum Learning with Diversity-Based Context Blending
cMALC-D uses an LLM to generate training contexts for multi-agent RL and a diversity-blending mechanism to avoid mode collapse, claiming improved generalization on traffic signal control.
-
Multi-Agent Trust Region Policy Optimisation: A Joint Constraint Approach
Adaptive per-agent KL-threshold allocation via KKT (HATRPO-W) and greedy (HATRPO-G) improves HATRPO's final reward by over 22.5% in MARL benchmarks.
-
From General Relation Patterns to Task-Specific Decision-Making in Continual Multi-Agent Coordination
RPG uses a relation capturer plus a task-conditioned hypernetwork to reduce catastrophic forgetting and enable zero-shot transfer in continual multi-agent coordination.
-
Light Aircraft Game : Basic Implementation and training results analysis
In the new LAG air-combat environment, HASAC scores higher than HAPPO in no-weapon coordination tasks while HAPPO scores higher in missile combat, but the results come from single runs without error bars.
-
Shapley Machine: A Game-Theoretic Framework for N-Agent Ad Hoc Teamwork
Shapley Machine reshapes rewards and TD targets so per-agent value functions approximately satisfy the Shapley axioms, and it beats the prior POAM baseline in several NAHT test environments.
-
Ego-centric Learning of Communicative World Models for Autonomous Driving
Sharing compressed latent states and planned waypoints between agents, triggered by prediction errors, improves multi-agent driving performance in CARLA while cutting communication bandwidth by roughly 50x.
-
Reward-Independent Messaging for Decentralized Multi-Agent Reinforcement Learning
MARL-CPC lets decentralized agents learn to send informative messages through a self-supervised reconstruction objective, and outperforms message-as-action baselines in non-cooperative multi-agent tasks.
-
SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control
A multi-agent RL workflow that interleaves single-agent updates, applied to mobile GUI control, achieves SOTA zero-shot performance and a +14.8 MATH500 gain.
-
Multi-agent Reinforcement Learning for Robotized Coral Reef Sample Collection
An RL controller for coral sample collection, trained in Unity, is transferred zero-shot to a physical BlueROV2 using real-time underwater motion capture to drive the digital twin.
-
PILOC: A Pheromone Inverse Guidance Mechanism and Local-Communication Framework for Dynamic Target Search of Multi-Agent in Unknown Environments
A decentralized reinforcement-learning framework using virtual pheromone marks and local map sharing outperforms MARL baselines for finding dynamic targets in simulated unknown grid environments.
-
GenAI-based Multi-Agent Reinforcement Learning towards Distributed Agent Intelligence: A Generative-RL Agent Perspective
A position paper claiming that generative-AI agents that model and predict multi-agent dynamics will replace today's reactive MARL approaches.
-
Multi-Agent Reinforcement Learning in Cybersecurity: From Fundamentals to Applications
A narrative survey of multi-agent reinforcement learning for cyber defense, reviewing game-theoretic models, cyber gyms, and applications, concluding MARL is promising but faces scalability and simulation-to-real tran...
- MEAL: A Benchmark for Continual Multi-Agent Reinforcement Learning
Discussion (0). Continue with ORCID to comment.