Pith. sign in

REVIEW 5 cited by

Revisiting Parameter Sharing in Multi-Agent Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.13625 v8 pith:XT7ZR7UN submitted 2020-05-27 cs.LG cs.AIcs.MAstat.ML

classification cs.LGcs.AIcs.MAstat.ML
keywords agentindicationparameterpoliciessharingspaceslearningmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Parameter sharing, where each agent independently learns a policy with fully shared parameters between all policies, is a popular baseline method for multi-agent deep reinforcement learning. Unfortunately, since all agents share the same policy network, they cannot learn different policies or tasks. This issue has been circumvented experimentally by adding an agent-specific indicator signal to observations, which we term "agent indication". Agent indication is limited, however, in that without modification it does not allow parameter sharing to be applied to environments where the action spaces and/or observation spaces are heterogeneous. This work formalizes the notion of agent indication and proves that it enables convergence to optimal policies for the first time. Next, we formally introduce methods to extend parameter sharing to learning in heterogeneous observation and action spaces, and prove that these methods allow for convergence to optimal policies. Finally, we experimentally confirm that the methods we introduce function empirically, and conduct a wide array of experiments studying the empirical efficacy of many different agent indication schemes for image based observation spaces.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hierarchical Message-Passing Policies for Multi-Agent Reinforcement Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A feudal hierarchical MARL method where lower-level policies are rewarded with the upper level's advantage function, with theoretical alignment guarantees and strong benchmark results.

  2. Low-Rank Agent-Specific Adaptation (LoRASA) for Multi-Agent Policy Learning

    cs.MA 2025-02 conditional novelty 6.0 of 10

    Per-agent low-rank adapters on a shared backbone let multi-agent policies specialize at a fraction of the memory cost of separate networks, with competitive benchmark performance.

  3. An Extended Benchmarking of Multi-Agent Reinforcement Learning Algorithms in Complex Fully Cooperative Tasks

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Algorithms that set state of the art on SMAC and GRF often underperform standard baselines on fully cooperative benchmarks, including image-based tasks.

  4. Multi-Agent Reinforcement Learning for Dynamic Mobility Resource Allocation with Hierarchical Adaptive Grouping

    cs.AI 2025-07 conditional novelty 5.0 of 10

    HAG-PS adds hierarchical adaptive grouping and identity embeddings to MARL parameter sharing, reaching 77.21% fulfilled service ratio on simulated January 2024 Manhattan bike sharing.

  5. Wasserstein-Barycenter Consensus for Cooperative Multi-Agent Reinforcement Learning

    eess.SY 2025-06 reject novelty 5.0 of 10

    A cooperative MARL algorithm that regularizes each agent's policy toward the Sinkhorn barycenter of the team's visitation distributions, with a claimed but insufficiently proven geometric convergence guarantee.

Pith tools