Pith. sign in

REVIEW

Options as responses: Grounding behavioural hierarchies in multi-agent RL

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.01470 v3 pith:OUVDML6T submitted 2019-06-04 cs.LG cs.AIcs.MAcs.NEstat.ML

classification cs.LGcs.AIcs.MAcs.NEstat.ML
keywords opponentsagentfailgamesgeneralisationgeneralisegroundinghierarchical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper investigates generalisation in multi-agent games, where the generality of the agent can be evaluated by playing against opponents it hasn't seen during training. We propose two new games with concealed information and complex, non-transitive reward structure (think rock/paper/scissors). It turns out that most current deep reinforcement learning methods fail to efficiently explore the strategy space, thus learning policies that generalise poorly to unseen opponents. We then propose a novel hierarchical agent architecture, where the hierarchy is grounded in the game-theoretic structure of the game -- the top level chooses strategic responses to opponents, while the low level implements them into policy over primitive actions. This grounding facilitates credit assignment across the levels of hierarchy. Our experiments show that the proposed hierarchical agent is capable of generalisation to unseen opponents, while conventional baselines fail to generalise whatsoever.

Discussion (0). Continue with ORCID to comment.

Pith tools