Pith. sign in

REVIEW 5 cited by

FACMAC: Factored Multi-Agent Centralised Policy Gradients

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.06709 v5 pith:X72JH4KJ submitted 2020-03-14 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords multi-agentfacmaccentralisedpolicyfactoredactioncriticgradients
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose FACtored Multi-Agent Centralised policy gradients (FACMAC), a new method for cooperative multi-agent reinforcement learning in both discrete and continuous action spaces. Like MADDPG, a popular multi-agent actor-critic method, our approach uses deep deterministic policy gradients to learn policies. However, FACMAC learns a centralised but factored critic, which combines per-agent utilities into the joint action-value function via a non-linear monotonic function, as in QMIX, a popular multi-agent Q-learning algorithm. However, unlike QMIX, there are no inherent constraints on factoring the critic. We thus also employ a nonmonotonic factorisation and empirically demonstrate that its increased representational capacity allows it to solve some tasks that cannot be solved with monolithic, or monotonically factored critics. In addition, FACMAC uses a centralised policy gradient estimator that optimises over the entire joint action space, rather than optimising over each agent's action space separately as in MADDPG. This allows for more coordinated policy changes and fully reaps the benefits of a centralised critic. We evaluate FACMAC on variants of the multi-agent particle environments, a novel multi-agent MuJoCo benchmark, and a challenging set of StarCraft II micromanagement tasks. Empirical results demonstrate FACMAC's superior performance over MADDPG and other baselines on all three domains.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 106 citations worldwide. Full citation record

  1. Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Online action-space factorization via Kalman-refined cross-capacitance lets shared multi-agent policies zero-shot tune larger quantum-dot arrays with near-constant steps.

  2. Bridging MARL to SARL: An Order-Independent Multi-Agent Transformer via Latent Consensus

    cs.LG 2026-04 conditional novelty 6.0 of 10

    CMAT uses a transformer decoder to produce a high-level consensus vector in latent space, enabling simultaneous order-independent actions by all agents and optimization via single-agent PPO, with superior results on S...

  3. A Self-Evolving Default Action for Cooperative Tasks with Continuous Action Space

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Sampling a 'default action' from an agent's replay buffer as a counterfactual baseline gives unbiased policy gradients and strong empirical results in continuous cooperative control.

  4. Self-Supervised Goal-Reaching Results in Multi-Agent Cooperation and Exploration

    cs.LG 2025-09 conditional novelty 5.0 of 10

    Self-supervised multi-agent goal-reaching, where each agent independently learns a contrastive critic of its own observations, achieves cooperation and exploration in sparse-reward MARL tasks where standard baselines fail.

  5. Light Aircraft Game : Basic Implementation and training results analysis

    cs.LG 2025-06 reject novelty 5.0 of 10

    In the new LAG air-combat environment, HASAC scores higher than HAPPO in no-weapon coordination tasks while HAPPO scores higher in missile combat, but the results come from single runs without error bars.

Pith tools