REVIEW 12 cited by
Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We explore deep reinforcement learning methods for multi-agent domains. We begin by analyzing the difficulty of traditional algorithms in the multi-agent case: Q-learning is challenged by an inherent non-stationarity of the environment, while policy gradient suffers from a variance that increases as the number of agents grows. We then present an adaptation of actor-critic methods that considers action policies of other agents and is able to successfully learn policies that require complex multi-agent coordination. Additionally, we introduce a training regimen utilizing an ensemble of policies for each agent that leads to more robust multi-agent policies. We show the strength of our approach compared to existing methods in cooperative as well as competitive scenarios, where agent populations are able to discover various physical and informational coordination strategies.
Forward citations
Cited by 12 Pith papers
-
Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning
A happiness-regression contrast from the SoDec dataset is used as a reward-shaping weight in a two-agent Social Lottery, yielding a safe rate of 0.459 versus a human 0.484, but the contrast is statistically indistingu...
-
Coupling Smoothed Particle Hydrodynamics with Multi-Agent Deep Reinforcement Learning for Cooperative Control of Point Absorbers
A GPU-coupled SPH and multi-agent reinforcement learning platform learns cooperative PTO damping policies that increase simulated wave-energy capture by up to 23.8% over fixed damping.
-
A Learning Framework For Cooperative Collision Avoidance of UAV Swarms Leveraging Domain Knowledge
A MARL framework that uses an active-contour-inspired reward to train UAV swarms for collision avoidance without credit assignment or observation sharing.
-
Hold My Beer: Learning Gentle Humanoid Locomotion and End-Effector Stabilization Control
A slow-fast two-agent reinforcement learning architecture with separate upper- and lower-body policies reduces end-effector shaking during humanoid locomotion.
-
Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details
For Other-Play in Yokai, agents trained with different implementation details coordinate across implementations about as well as across seeds, supporting inter-seed cross-play as a proxy for cross-implementation evaluation.
-
Integrated Altruistic and Fairness Preference Induces Advanced Mutual Cooperation in Sequential Social Dilemmas
Integrating forward-looking altruistic and fairness preferences into agent utilities produces mutual cooperation and higher collective returns than egoistic or inequity-aversion baselines in sequential social dilemmas.
-
Multi-agent DRL-based Lane Change Decision Model for Cooperative Platooning in Mixed Traffic
A CNN-QMIX lane-change controller lifts simulated cooperative-platoon formation over MOBIL and greedy baselines and keeps working as the number of connected agents varies.
-
Agent-Based Decentralized Energy Management of EV Charging Station with Solar Photovoltaics via Multi-Agent Reinforcement Learning
An LSTM-based multi-agent RL approach with decentralized execution and a dense reward reduces charging cost and unfinished charging demand in a simulated EV charging station under partial charger faults.
-
Conflict Resolution under Degraded Surveillance in Air Corridors Using Multi-Agent Reinforcement Learning
Separate DQN policies for UAV and eVTOL agents are trained to keep separation in a simulated structured corridor with degraded surveillance, and their behavior is summarized as action shares and Pareto-optimal safety/...
-
Adversarial Agent Behavior Learning in Autonomous Driving Using Deep Reinforcement Learning
An adversarial car trained with a collision-based reward reliably decreases the reward of a PPO-trained ego vehicle in Highway-Env, and a robust PPO policy trained against it recovers performance.
-
A Survey of Multi Agent Reinforcement Learning: Federated Learning and Cooperative and Noncooperative Decentralized Regimes
A review of multi-agent reinforcement learning that catalogues federated, decentralized cooperative, and noncooperative regimes from the existing literature.
-
Multi-Agent Reinforcement Learning in Cybersecurity: From Fundamentals to Applications
A narrative survey of multi-agent reinforcement learning for cyber defense, reviewing game-theoretic models, cyber gyms, and applications, concluding MARL is promising but faces scalability and simulation-to-real tran...
Discussion (0). Sign in to comment.