Pith. sign in

REVIEW 2 cited by

On Lai's Upper Confidence Bound in Multi-Armed Bandits

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.02279 v2 pith:Z7SEQS77 submitted 2024-10-03 stat.ML cs.LGmath.STstat.TH

classification stat.MLcs.LGmath.STstat.TH
keywords boundconfidenceupperregretbanditsboundsestablishexploration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this memorial paper, we honor Tze Leung Lai's seminal contributions to the topic of multi-armed bandits, with a specific focus on his pioneering work on the upper confidence bound. We establish sharp non-asymptotic regret bounds for an upper confidence bound index with a constant level of exploration for Gaussian rewards. Furthermore, we establish a non-asymptotic regret bound for the upper confidence bound index of Lai (1987) which employs an exploration function that decreases with the sample size of the corresponding arm. The regret bounds have leading constants that match the Lai-Robbins lower bound. Our results highlight an aspect of Lai's seminal works that deserves more attention in the machine learning literature.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UCB algorithms for multi-armed bandits: Precise regret and adaptive inference

    math.ST 2024-12 conditional novelty 7.0 of 10

    For Gaussian bandits, UCB arm-pull counts concentrate around a deterministic fixed-point schedule, yielding a precise regret formula and quantitative CLTs for adaptive inference.

  2. Selective Reviews of Bandit Problems in AI via a Statistical View

    stat.ML 2024-12 unverdicted novelty 1.0 of 10

    A statistical survey of multi-armed, contextual, and continuum-armed bandits that restates known minimax and regret results, adds an alternative UCB proof, and reports small simulation comparisons.

Pith tools