Pith. sign in

REVIEW 2 cited by

Ensemble sampling for linear bandits: small ensembles suffice

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.08376 v4 pith:UIYE6NYC submitted 2023-11-14 stat.ML cs.LG

classification stat.MLcs.LG
keywords ensemblesamplingfirstlinearorderbanditregretresult
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We provide the first useful and rigorous analysis of ensemble sampling for the stochastic linear bandit setting. In particular, we show that, under standard assumptions, for a $d$-dimensional stochastic linear bandit with an interaction horizon $T$, ensemble sampling with an ensemble of size of order $d \log T$ incurs regret at most of the order $(d \log T)^{5/2} \sqrt{T}$. Ours is the first result in any structured setting not to require the size of the ensemble to scale linearly with $T$ -- which defeats the purpose of ensemble sampling -- while obtaining near $\smash{\sqrt{T}}$ order regret. Our result is also the first to allow for infinite action sets.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Practical and Optimal Algorithm for Linear Contextual Bandits with Rare Parameter Updates

    stat.ML 2026-05 unverdicted novelty 7.0 of 10

    BLCE-G and BLCE achieve minimax-optimal regret for linear contextual bandits with only O(log log T) parameter updates and reduced computational cost by avoiding near G-optimal design.

  2. Quantile of Means: A Bonus-Free Ensemble Method for Minimax Optimal Reinforcement Learning

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    A quantile-of-means ensemble method achieves minimax optimal variance-dependent regret bounds for finite-horizon MDPs without count-based uncertainty estimates.

Pith tools