Pith. sign in

REVIEW 2 cited by

Multi-Player Bandits Revisited

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1711.02317 v3 pith:BRZ3JN2K submitted 2017-11-07 stat.ML cs.LG

classification stat.MLcs.LG
keywords algorithmsapplicationsmulti-playeralgorithmbanditsexistinginformationintroduce
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Multi-player Multi-Armed Bandits (MAB) have been extensively studied in the literature, motivated by applications to Cognitive Radio systems. Driven by such applications as well, we motivate the introduction of several levels of feedback for multi-player MAB algorithms. Most existing work assume that sensing information is available to the algorithm. Under this assumption, we improve the state-of-the-art lower bound for the regret of any decentralized algorithms and introduce two algorithms, RandTopM and MCTopM, that are shown to empirically outperform existing algorithms. Moreover, we provide strong theoretical guarantees for these algorithms, including a notion of asymptotic optimality in terms of the number of selections of bad arms. We then introduce a promising heuristic, called Selfish, that can operate without sensing information, which is crucial for emerging applications to Internet of Things networks. We investigate the empirical performance of this algorithm and provide some first theoretical elements for the understanding of its behavior.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Heterogeneous Multi-Player Multi-Armed Bandits Robust To Adversarial Attacks

    stat.ML 2025-01 reject novelty 6.0 of 10

    A decentralized policy for heterogeneous multiplayer bandits is claimed to achieve O(log^{1+δ}T + W) regret under adversarial zero-reward attacks using one-bit communication.

  2. Accelerated learning from recommender systems using multi-armed bandit

    cs.IR 2019-08 conditional novelty 4.0 of 10

    A Vrbo team used daily Thompson sampling to rank four recommendation models by click-through rate, but the A/B validation they report is for a previous campaign's winner, not the current one.

Pith tools