REVIEW 5 cited by
Qatten: A General Framework for Cooperative Multiagent Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
In many real-world tasks, multiple agents must learn to coordinate with each other given their private observations and limited communication ability. Deep multiagent reinforcement learning (Deep-MARL) algorithms have shown superior performance in such challenging settings. One representative class of work is multiagent value decomposition, which decomposes the global shared multiagent Q-value $Q_{tot}$ into individual Q-values $Q^{i}$ to guide individuals' behaviors, i.e. VDN imposing an additive formation and QMIX adopting a monotonic assumption using an implicit mixing method. However, most of the previous efforts impose certain assumptions between $Q_{tot}$ and $Q^{i}$ and lack theoretical groundings. Besides, they do not explicitly consider the agent-level impact of individuals to the whole system when transforming individual $Q^{i}$s into $Q_{tot}$. In this paper, we theoretically derive a general formula of $Q_{tot}$ in terms of $Q^{i}$, based on which we can naturally implement a multi-head attention formation to approximate $Q_{tot}$, resulting in not only a refined representation of $Q_{tot}$ with an agent-level attention mechanism, but also a tractable maximization algorithm of decentralized policies. Extensive experiments demonstrate that our method outperforms state-of-the-art MARL methods on the widely adopted StarCraft benchmark across different scenarios, and attention analysis is further conducted with valuable insights.
Forward citations
Cited by 5 Pith papers
-
Heterogeneous Value Decomposition Policy Fusion for Multi-Agent Cooperation
Adaptively fusing two value-decomposition policies with a KL constraint improves cooperative MARL performance on matrix game, SMAC, and predator-prey tasks.
-
Dual Ensembled Multiagent Q-Learning with Hypernet Regularizer
DEMAR combines dual ensembled Q-value targets with an L1 regularizer on mixing-network weights to curb multiagent Q-value overestimation in value-mixing Q-learning.
-
Multi-Year Maintenance Planning for Large-Scale Infrastructure Systems: A Novel Network Deep Q-Learning Approach
A deep Q-learning framework with cost-normalized rewards and an annual budget-allocation linear program outperforms LP and GA baselines on a 68,800-segment pavement network.
-
ToMacVF : Temporal Macro-action Value Factorization for Asynchronous Multi-Agent Reinforcement Learning
A temporal macro-action value factorization method with a segmented replay buffer improves asynchronous multi-agent RL performance, but the claimed proof that it generalizes standard IGM is invalid.
-
Double Distillation Network for Multi-Agent Reinforcement Learning
DDN, a double distillation network for cooperative MARL, claims to eliminate CTDE inherent error via global-to-local knowledge distillation and an RND-style intrinsic reward, with better SMAC and Predator-Prey win rates.
Discussion (0). Continue with ORCID to comment.