Pith. sign in

REVIEW 1 cited by

Multi-Agent Coordination via Multi-Level Communication

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.12713 v2 pith:GFJE73GP submitted 2022-09-26 cs.MA cs.LG

classification cs.MAcs.LG
keywords agentscommunicationseqcommmulti-agentactionscommunicatecoordinationdecisions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The partial observability and stochasticity in multi-agent settings can be mitigated by accessing more information about others via communication. However, the coordination problem still exists since agents cannot communicate actual actions with each other at the same time due to the circular dependencies. In this paper, we propose a novel multi-level communication scheme, Sequential Communication (SeqComm). SeqComm treats agents asynchronously (the upper-level agents make decisions before the lower-level ones) and has two communication phases. In the negotiation phase, agents determine the priority of decision-making by communicating hidden states of observations and comparing the value of intention, obtained by modeling the environment dynamics. In the launching phase, the upper-level agents take the lead in making decisions and then communicate their actions with the lower-level agents. Theoretically, we prove the policies learned by SeqComm are guaranteed to improve monotonically and converge. Empirically, we show that SeqComm outperforms existing methods in various cooperative multi-agent tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy Optimization

    cs.AI 2024-12 conditional novelty 6.0 of 10

    InSPO updates agents sequentially with in-sample objectives and entropy regularization, avoiding out-of-distribution joint actions and converging to a quantal response equilibrium in offline MARL.

Pith tools