Pith. sign in

REVIEW 9 cited by

Benchmarking Multi-Agent Deep Reinforcement Learning Algorithms in Cooperative Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.07869 v4 pith:ABIRKBBT submitted 2020-06-14 cs.LG cs.AIcs.MAstat.ML

classification cs.LGcs.AIcs.MAstat.ML
keywords learningmulti-agentalgorithmstasksdifferentapproachescooperativedeep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-agent deep reinforcement learning (MARL) suffers from a lack of commonly-used evaluation tasks and criteria, making comparisons between approaches difficult. In this work, we provide a systematic evaluation and comparison of three different classes of MARL algorithms (independent learning, centralised multi-agent policy gradient, value decomposition) in a diverse range of cooperative multi-agent learning tasks. Our experiments serve as a reference for the expected performance of algorithms across different learning tasks, and we provide insights regarding the effectiveness of different learning approaches. We open-source EPyMARL, which extends the PyMARL codebase to include additional algorithms and allow for flexible configuration of algorithm implementation details such as parameter sharing. Finally, we open-source two environments for multi-agent research which focus on coordination under sparse rewards.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 14 citations worldwide. Full citation record

  1. MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Credit assignment via LMM pairwise comparisons plus Bradley–Terry rank aggregation and potential-based shaping improves cooperative MARL under sparse rewards and dynamic agent counts.

  2. Contract-Based Compositional Shielding for Safe Multi-Agent Reinforcement Learning

    cs.LG 2026-06 conditional novelty 6.0 of 10

    Certified local LTL_safe contracts, jointly fixed-point checked and selected by a bandit, recover coordinated safe team policies under decentralised multi-agent RL execution.

  3. Multi-Agent Reinforcement Learning for V2X Resource Allocation: Disentangling MARL Challenges Through Benchmarking

    cs.MA 2026-02 conditional novelty 6.0 of 10

    Robustness to varied vehicle topologies, not partial observability or coordination, is the dominant obstacle for MARL in C-V2X spectrum sharing.

  4. Transformer World Model for Sample Efficient Multi-Agent Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A new transformer-based multi-agent world model with teammate prediction and prioritized replay achieves near-optimal performance on cooperative benchmarks in as few as 50,000 environment steps.

  5. Integrating Counterfactual Simulations with Language Models for Explaining Multi-Agent Behaviour

    cs.AI 2025-05 conditional novelty 6.0 of 10

    AXIS couples an LLM with a multi-agent simulator to produce counterfactual action explanations, and reports higher judged correctness and goal-prediction accuracy than baselines on ten autonomous-driving scenarios.

  6. Explaining Reinforcement Learning Agents via Inductive Logic Programming

    cs.AI 2026-07 conditional novelty 5.0 of 10

    This paper introduces activation rate, feature coverage, syntactic distance, and semantic distance as objective metrics for logic-rule explanations of RL and multi-agent RL policies.

  7. cMALC-D: Contextual Multi-Agent LLM-Guided Curriculum Learning with Diversity-Based Context Blending

    cs.LG 2025-08 reject novelty 5.0 of 10

    cMALC-D uses an LLM to generate training contexts for multi-agent RL and a diversity-blending mechanism to avoid mode collapse, claiming improved generalization on traffic signal control.

  8. From General Relation Patterns to Task-Specific Decision-Making in Continual Multi-Agent Coordination

    cs.MA 2025-07 conditional novelty 5.0 of 10

    RPG uses a relation capturer plus a task-conditioned hypernetwork to reduce catastrophic forgetting and enable zero-shot transfer in continual multi-agent coordination.

  9. Feedback Attribution and Representation Geometry: Metrics for Comparing Individual and Shared Rewards in MARL

    cs.LG 2026-07 conditional novelty 4.0 of 10

    On SMACv2, role geometry in shared-encoder MARL is set by whether unit type is observed, not by individual vs shared reward; reward attribution affects behavior, mainly action diversity.

Pith tools