REVIEW 2 cited by
On Lai's Upper Confidence Bound in Multi-Armed Bandits
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this memorial paper, we honor Tze Leung Lai's seminal contributions to the topic of multi-armed bandits, with a specific focus on his pioneering work on the upper confidence bound. We establish sharp non-asymptotic regret bounds for an upper confidence bound index with a constant level of exploration for Gaussian rewards. Furthermore, we establish a non-asymptotic regret bound for the upper confidence bound index of Lai (1987) which employs an exploration function that decreases with the sample size of the corresponding arm. The regret bounds have leading constants that match the Lai-Robbins lower bound. Our results highlight an aspect of Lai's seminal works that deserves more attention in the machine learning literature.
Forward citations
Cited by 2 Pith papers
-
UCB algorithms for multi-armed bandits: Precise regret and adaptive inference
For Gaussian bandits, UCB arm-pull counts concentrate around a deterministic fixed-point schedule, yielding a precise regret formula and quantitative CLTs for adaptive inference.
-
Selective Reviews of Bandit Problems in AI via a Statistical View
A statistical survey of multi-armed, contextual, and continuum-armed bandits that restates known minimax and regret results, adds an alternative UCB proof, and reports small simulation comparisons.
Discussion (0). Continue with ORCID to comment.