Pith. sign in

REVIEW 1 cited by

Population-based Evaluation in Repeated Rock-Paper-Scissors as a Benchmark for Multiagent Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.03196 v2 pith:BXWSXS6P submitted 2023-03-02 cs.GT cs.AIcs.LGcs.MA

classification cs.GTcs.AIcs.LGcs.MA
keywords learningbenchmarkmultiagentevaluationrepeatedsomeadversarialagent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Progress in fields of machine learning and adversarial planning has benefited significantly from benchmark domains, from checkers and the classic UCI data sets to Go and Diplomacy. In sequential decision-making, agent evaluation has largely been restricted to few interactions against experts, with the aim to reach some desired level of performance (e.g. beating a human professional player). We propose a benchmark for multiagent learning based on repeated play of the simple game Rock, Paper, Scissors along with a population of forty-three tournament entries, some of which are intentionally sub-optimal. We describe metrics to measure the quality of agents based both on average returns and exploitability. We then show that several RL, online learning, and language model approaches can learn good counter-strategies and generalize well, but ultimately lose to the top-performing bots, creating an opportunity for research in multiagent learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. IFlowNets: Extending Generative Samplers to Learn Strategies in Incomplete Information Games

    cs.LG 2026-08 conditional novelty 6.0 of 10

    IFlowNets make generative flow networks work for imperfect-information games by adding an information-set aggregation constraint that restores valid flow matching.

Pith tools