Pith. sign in

REVIEW 2 cited by

Assigning Credit with Partial Reward Decoupling in Multi-Agent Proximal Policy Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.04295 v3 pith:6C4CZR43 submitted 2024-08-08 cs.MA cs.AIcs.LGcs.RO

classification cs.MAcs.AIcs.LGcs.RO
keywords creditmappomulti-agentrewardagentsassignmentlearningperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-agent proximal policy optimization (MAPPO) has recently demonstrated state-of-the-art performance on challenging multi-agent reinforcement learning tasks. However, MAPPO still struggles with the credit assignment problem, wherein the sheer difficulty in ascribing credit to individual agents' actions scales poorly with team size. In this paper, we propose a multi-agent reinforcement learning algorithm that adapts recent developments in credit assignment to improve upon MAPPO. Our approach leverages partial reward decoupling (PRD), which uses a learned attention mechanism to estimate which of a particular agent's teammates are relevant to its learning updates. We use this estimate to dynamically decompose large groups of agents into smaller, more manageable subgroups. We empirically demonstrate that our approach, PRD-MAPPO, decouples agents from teammates that do not influence their expected future reward, thereby streamlining credit assignment. We additionally show that PRD-MAPPO yields significantly higher data efficiency and asymptotic performance compared to both MAPPO and other state-of-the-art methods across several multi-agent tasks, including StarCraft II. Finally, we propose a version of PRD-MAPPO that is applicable to \textit{shared} reward settings, where PRD was previously not applicable, and empirically show that this also leads to performance improvements over MAPPO.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning

    cs.MA 2025-02 conditional novelty 5.0 of 10

    LLM-generated, agent-specific potential-based rewards accelerate sparse-reward MARL training in grid world and pistonball benchmarks.

  2. Decentralized Multi-agent Reinforcement Learning for Resilient Critical Infrastructures

    cs.MA 2026-07 conditional novelty 4.0 of 10

    Decentralized MARL is argued to be structurally aligned with resilient critical infrastructure, with credit assignment and communication identified as the key open challenges.

Pith tools