Pith. sign in

REVIEW 10 cited by

Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1706.02275 v4 pith:6RD3BRGA submitted 2017-06-07 cs.LG cs.AIcs.NE

classification cs.LGcs.AIcs.NE
keywords multi-agentpoliciesmethodsableactor-criticagentagentscoordination
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We explore deep reinforcement learning methods for multi-agent domains. We begin by analyzing the difficulty of traditional algorithms in the multi-agent case: Q-learning is challenged by an inherent non-stationarity of the environment, while policy gradient suffers from a variance that increases as the number of agents grows. We then present an adaptation of actor-critic methods that considers action policies of other agents and is able to successfully learn policies that require complex multi-agent coordination. Additionally, we introduce a training regimen utilizing an ensemble of policies for each agent that leads to more robust multi-agent policies. We show the strength of our approach compared to existing methods in cooperative as well as competitive scenarios, where agent populations are able to discover various physical and informational coordination strategies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 1,015 citations worldwide. Full citation record

  1. Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning

    cs.AI 2026-08 reject novelty 6.0 of 10

    A happiness-regression contrast from the SoDec dataset is used as a reward-shaping weight in a two-agent Social Lottery, yielding a safe rate of 0.459 versus a human 0.484, but the contrast is statistically indistingu...

  2. Coupling Smoothed Particle Hydrodynamics with Multi-Agent Deep Reinforcement Learning for Cooperative Control of Point Absorbers

    eess.SY 2026-01 conditional novelty 6.0 of 10

    A GPU-coupled SPH and multi-agent reinforcement learning platform learns cooperative PTO damping policies that increase simulated wave-energy capture by up to 23.8% over fixed damping.

  3. A Learning Framework For Cooperative Collision Avoidance of UAV Swarms Leveraging Domain Knowledge

    cs.MA 2025-07 reject novelty 6.0 of 10

    A MARL framework that uses an active-contour-inspired reward to train UAV swarms for collision avoidance without credit assignment or observation sharing.

  4. Hold My Beer: Learning Gentle Humanoid Locomotion and End-Effector Stabilization Control

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A slow-fast two-agent reinforcement learning architecture with separate upper- and lower-body policies reduces end-effector shaking during humanoid locomotion.

  5. Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details

    cs.AI 2026-08 conditional novelty 5.0 of 10

    For Other-Play in Yokai, agents trained with different implementation details coordinate across implementations about as well as across seeds, supporting inter-seed cross-play as a proxy for cross-implementation evaluation.

  6. Integrated Altruistic and Fairness Preference Induces Advanced Mutual Cooperation in Sequential Social Dilemmas

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Integrating forward-looking altruistic and fairness preferences into agent utilities produces mutual cooperation and higher collective returns than egoistic or inequity-aversion baselines in sequential social dilemmas.

  7. Multi-agent DRL-based Lane Change Decision Model for Cooperative Platooning in Mixed Traffic

    cs.AI 2026-01 conditional novelty 5.0 of 10

    A CNN-QMIX lane-change controller lifts simulated cooperative-platoon formation over MOBIL and greedy baselines and keeps working as the number of connected agents varies.

  8. Conflict Resolution under Degraded Surveillance in Air Corridors Using Multi-Agent Reinforcement Learning

    cs.LG 2026-07 reject novelty 4.0 of 10

    Separate DQN policies for UAV and eVTOL agents are trained to keep separation in a simulated structured corridor with degraded surveillance, and their behavior is summarized as action shares and Pareto-optimal safety/...

  9. Adversarial Agent Behavior Learning in Autonomous Driving Using Deep Reinforcement Learning

    cs.CV 2025-08 reject novelty 4.0 of 10

    An adversarial car trained with a collision-based reward reliably decreases the reward of a PPO-trained ego vehicle in Highway-Env, and a robust PPO policy trained against it recovers performance.

  10. A Survey of Multi Agent Reinforcement Learning: Federated Learning and Cooperative and Noncooperative Decentralized Regimes

    cs.MA 2025-07 reject novelty 3.0 of 10

    A review of multi-agent reinforcement learning that catalogues federated, decentralized cooperative, and noncooperative regimes from the existing literature.

Pith tools