REVIEW 1 cited by
A Note on KL-UCB+ Policy for the Stochastic Bandit
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
A classic setting of the stochastic K-armed bandit problem is considered in this note. In this problem it has been known that KL-UCB policy achieves the asymptotically optimal regret bound and KL-UCB+ policy empirically performs better than the KL-UCB policy although the regret bound for the original form of the KL-UCB+ policy has been unknown. This note demonstrates that a simple proof of the asymptotic optimality of the KL-UCB+ policy can be given by the same technique as those used for analyses of other known policies.
Forward citations
Cited by 1 Pith paper
-
Optimal Regret of Bernoulli Bandits under Global Differential Privacy
For epsilon-global-DP Bernoulli bandits, the paper proves a tighter lower bound and matching upper bounds (up to a factor alpha that can approach 1) using a new quantity d_epsilon and a new DP-Chernoff concentration i...
Discussion (0). Continue with ORCID to comment.