Pith. sign in

REVIEW 3 cited by

Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1810.08647 v4 pith:MIDNW3MV submitted 2018-10-19 cs.LG cs.AIcs.MAstat.ML

classification cs.LGcs.AIcs.MAstat.ML
keywords agentsinfluenceactionscommunicationotherdeeplearningbehavior
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We propose a unified mechanism for achieving coordination and communication in Multi-Agent Reinforcement Learning (MARL), through rewarding agents for having causal influence over other agents' actions. Causal influence is assessed using counterfactual reasoning. At each timestep, an agent simulates alternate actions that it could have taken, and computes their effect on the behavior of other agents. Actions that lead to bigger changes in other agents' behavior are considered influential and are rewarded. We show that this is equivalent to rewarding agents for having high mutual information between their actions. Empirical results demonstrate that influence leads to enhanced coordination and communication in challenging social dilemma environments, dramatically increasing the learning curves of the deep RL agents, and leading to more meaningful learned communication protocols. The influence rewards for all agents can be computed in a decentralized way by enabling agents to learn a model of other agents using deep neural networks. In contrast, key previous works on emergent communication in the MARL setting were unable to learn diverse policies in a decentralized manner and had to resort to centralized training. Consequently, the influence reward opens up a window of new opportunities for research in this area.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Emergence of Fair Leaders via Mediators in Multi-Agent Reinforcement Learning

    cs.MA 2025-08 conditional novelty 6.0 of 10

    A mediator that dynamically selects leaders in Stackelberg MARL can induce self-interested agents to adopt fair policies, improving fairness of returns.

  2. InvestESG: A multi-agent reinforcement learning benchmark for studying climate investment as a social dilemma

    cs.LG 2024-11 conditional novelty 6.0 of 10

    InvestESG is a MARL benchmark showing that ESG-conscious investors, not the disclosure mandate itself, drive corporate mitigation in long-run simulated markets.

  3. No Press Diplomacy: Modeling Multi-Agent Gameplay

    cs.AI 2019-09 accept novelty 6.0 of 10

    A neural policy trained on 150,000 human Diplomacy games, then refined by self-play, beats rule-based bots in No Press Diplomacy.

Pith tools