REVIEW 3 minor 227 references
Practical and Optimal Algorithm for Linear Contextual Bandits with Rare Parameter Updates
T0 review · 0 major / 3 minor · reviewed 2026-06-28 · grok-4.3
Pith's one-line read Two algorithms achieve minimax-optimal regret for linear contextual bandits using only O(log log T) parameter updates.
desk verdict The paper delivers two algorithms that hit minimax-optimal regret for linear contextual bandits with only O(log log T) static updates and drops the G-optimal design step in one of them. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
A static grid of O(log log T) update times that divides the horizon into intervals inside which the policy can still adapt to the realized sequence of contexts and actions.
What would settle it
A concrete lower-bound construction or numerical simulation in which any algorithm limited to o(log log T) updates incurs regret strictly larger than the known minimax rate by more than polylog factors.
Extended reading notes
Core claim
Under a static schedule of O(log log T) update times, the BLCE-G algorithm attains minimax-optimal regret up to polylogarithmic factors in T simultaneously for small and large action sets, while the BLCE algorithm removes the near G-optimal design computation, preserves the same regret guarantee, and records the lowest known runtime among optimal algorithms for the setting.
Load-bearing premise
A fixed non-adaptive schedule of O(log log T) updates is enough to recover minimax-optimal regret when action selection inside each interval can still depend on the realized contexts.
Editorial extensions
If this is right
- Minimax-optimal regret holds simultaneously across small-K and large-K regimes under the same static schedule.
- The dominant computational cost of near G-optimal design can be removed without sacrificing the optimal regret rate.
- The rare-update construction extends directly to generalized linear contextual bandits.
- The resulting runtime is the lowest known among all algorithms that achieve the minimax rate.
Reading between the lines
- If the static schedule works, then full online parameter updates are unnecessary for optimality, which could reduce communication overhead in distributed or federated deployments.
- The same interval structure might transfer to other online decision problems where model refits are expensive.
- Empirical checks on large-scale recommendation data could test whether the polylog factors remain small in practice.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript studies linear contextual bandits where parameter updates are restricted to a small number of times (O(log log T)) under a static schedule, while still permitting context-dependent action selection within each interval. It introduces BLCE-G, which attains minimax-optimal regret (up to polylog factors in T) simultaneously for small-K and large-K regimes, and BLCE, which eliminates the near G-optimal design computation while preserving the same regret guarantee and achieving the lowest known runtime among optimal methods. The approach is extended to generalized linear contextual bandits.
Significance. If the central claims hold, the work is significant because it delivers statistically optimal algorithms that are also practical under rare updates, explicitly separating the setting from strictly batched baselines by allowing within-interval adaptivity. This addresses a common practical constraint in deployment while matching known minimax bounds, and the runtime improvement in BLCE is a concrete computational contribution.
minor comments (3)
- [Abstract] Abstract: the phrase 'lowest known runtime complexity among optimal algorithms' would be strengthened by an explicit big-O expression for the per-round or total runtime of BLCE.
- [Introduction] The distinction between the proposed static schedule with within-interval adaptivity and strictly batched methods is central; a short clarifying paragraph or table contrasting the two (e.g., what information is available for action selection inside an interval) would improve readability.
- [Preliminaries] Notation for the update times and the intervals between them is used repeatedly; ensuring a single, consistently referenced definition (perhaps in a preliminary section) would reduce ambiguity for readers.
Simulated Author's Rebuttal
We thank the referee for the positive summary, recognition of the significance of separating rare-update linear contextual bandits from strictly batched settings, and the recommendation for minor revision. No specific major comments were provided in the report.
Circularity Check
No significant circularity; derivation self-contained
full rationale
The paper introduces BLCE-G and BLCE algorithms for linear contextual bandits under rare (O(log log T)) parameter updates, claiming minimax-optimal regret (up to polylog factors) in both small-K and large-K regimes. The abstract and description separate the method from strictly batched baselines by allowing within-interval context-adaptive action selection, but the optimality claim rests on standard regret analysis rather than any reduction of the target bound to a fitted quantity, self-definition, or self-citation chain. No equations or steps in the provided text exhibit self-definitional equivalence, fitted inputs renamed as predictions, or load-bearing self-citations. The result is presented as an extension of known minimax bounds to the rare-update setting without internal circularity.
Assumptions & free parameters
assumptions (1)
- domain assumption Linear reward model with sub-Gaussian noise
Cite this review
Pith. "Pith review of Practical and Optimal Algorithm for Linear Contextual Bandits with Rare Parameter Updates." pith.science (2026). https://pith.science/paper/UU6PCMBY
@misc{pith2026260600984,
author = {Pith},
title = {Pith review of: Practical and Optimal Algorithm for Linear Contextual Bandits with Rare Parameter Updates},
year = {2026},
howpublished = {\url{https://pith.science/paper/UU6PCMBY}},
note = {Machine review of arXiv:2606.00984}
}
abstract
We study linear contextual bandits under rare parameter updates: the learner may incorporate reward feedback into its parameter estimate only at a small number of update times, while still observing contexts online and selecting actions sequentially. This viewpoint clarifies a practical distinction that is often blurred in the literature: many "strictly batched" methods additionally restrict within-interval context adaptivity, meaning that the action rule inside an interval cannot depend on the sequence of realized contexts/actions in that interval (beyond the current round's context). For linear contextual bandits, we propose two practical algorithms with only $O(\log\log T)$ parameter updates. Our first algorithm BLCE-G attains minimax-optimal regret (up to polylogarithmic factors in $T$) simultaneously in both the small-$K$ and large-$K$ regimes under a static schedule. Our second algorithm BLCE removes the near G-optimal design step -- a dominant computational bottleneck in prior strictly batched static-grid methods -- yet preserves minimax-optimal regret and achieves the lowest known runtime complexity among optimal algorithms. We further extend these rare-update and computational principles to generalized linear contextual bandits. Overall, our results yield statistically optimal algorithms under $O(\log\log T)$ parameter updates that are also computationally efficient in practice.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Advances in Neural Information Processing Systems , volume=
An analysis of ensemble sampling , author=. Advances in Neural Information Processing Systems , volume=
-
[2]
Advances in Neural Information Processing Systems , volume=
Ensemble sampling , author=. Advances in Neural Information Processing Systems , volume=
-
[3]
Proceedings of the 17th ACM Conference on Recommender Systems , pages=
Deep exploration for recommendation systems , author=. Proceedings of the 17th ACM Conference on Recommender Systems , pages=
-
[4]
arXiv preprint arXiv:2311.08376 , year=
Ensemble sampling for linear bandits: small ensembles suffice , author=. arXiv preprint arXiv:2311.08376 , year=
-
[5]
Journal of the American Statistical Association , pages=
Stochastic low-rank tensor bandits for multi-dimensional online decision making , author=. Journal of the American Statistical Association , pages=. 2024 , publisher=
2024
-
[6]
International Conference on Machine Learning , pages=
Multiplier bootstrap-based exploration , author=. International Conference on Machine Learning , pages=. 2023 , organization=
2023
-
[7]
Abeille, Marc and Lazaric, Alessandro , year = 2017, booktitle =
2017
-
[8]
Annals of statistics , pages=
Adaptive estimation of a quadratic functional by model selection , author=. Annals of statistics , pages=. 2000 , publisher=
2000
Show all 227 references
-
[9]
Advances in Neural Information Processing Systems , volume=
Deep exploration via bootstrapped DQN , author=. Advances in Neural Information Processing Systems , volume=
-
[10]
Advances in Neural Information Processing Systems , volume=
Randomized prior functions for deep reinforcement learning , author=. Advances in Neural Information Processing Systems , volume=
-
[11]
Proceedings of the 12th ACM Conference on Recommender Systems , pages=
Efficient online recommendation via low-rank ensemble sampling , author=. Proceedings of the 12th ACM Conference on Recommender Systems , pages=
-
[12]
International Conference on Artificial Intelligence and Statistics , pages=
Exploration via linearly perturbed loss minimisation , author=. International Conference on Artificial Intelligence and Statistics , pages=. 2024 , organization=
2024
-
[13]
Advances in Neural Information Processing Systems , volume=
Scalable representation learning in linear contextual bandits with constant regret guarantees , author=. Advances in Neural Information Processing Systems , volume=
-
[14]
Advances in Neural Information Processing Systems , volume=
Model selection for contextual bandits , author=. Advances in Neural Information Processing Systems , volume=
-
[15]
International Conference on Artificial Intelligence and Statistics , pages=
Osom: A simultaneously optimal algorithm for multi-armed and linear contextual bandits , author=. International Conference on Artificial Intelligence and Statistics , pages=. 2020 , organization=
2020
-
[16]
International Conference on Artificial Intelligence and Statistics , pages=
Stochastic linear contextual bandits with diverse contexts , author=. International Conference on Artificial Intelligence and Statistics , pages=. 2020 , organization=
2020
-
[17]
2017 International Conference on Sampling Theory and Applications (SampTA) , pages=
Sparse linear contextual bandits via relevance vector machines , author=. 2017 International Conference on Sampling Theory and Applications (SampTA) , pages=. 2017 , organization=
2017
-
[18]
Artificial Intelligence and Statistics , pages=
Online-to-confidence-set conversions and application to sparse stochastic bandits , author=. Artificial Intelligence and Statistics , pages=. 2012 , organization=
2012
-
[19]
Mobile health: sensors, analytic methods, and applications , pages=
From ads to interventions: Contextual bandits in mobile health , author=. Mobile health: sensors, analytic methods, and applications , pages=. 2017 , publisher=
2017
-
[20]
arXiv preprint arXiv:1904.10040 , year=
A survey on practical applications of multi-armed and contextual bandits , author=. arXiv preprint arXiv:1904.10040 , year=
1904 arXiv
-
[21]
2020 , publisher=
Bandit algorithms , author=. 2020 , publisher=
2020
-
[22]
Proceedings of the 22nd ACM SIGKDD international conference on Knowledge discovery and data mining , pages=
Online context-aware recommendation with time varying multi-armed bandit , author=. Proceedings of the 22nd ACM SIGKDD international conference on Knowledge discovery and data mining , pages=
-
[23]
Proceedings of the 39th International ACM SIGIR conference on Research and Development in Information Retrieval , pages=
Collaborative filtering bandits , author=. Proceedings of the 39th International ACM SIGIR conference on Research and Development in Information Retrieval , pages=
-
[24]
Some aspects of the sequential design of experiments , author=
-
[25]
Journal of the American Statistical Association , pages=
Nearly dimension-independent sparse linear bandit over small action spaces via best subset selection , author=. Journal of the American Statistical Association , pages=. 2022 , publisher=
2022
-
[26]
Bayesian linear regression with sparse priors , author=
-
[27]
arXiv preprint arXiv:1002.1583 , year=
Thresholded Lasso for high dimensional variable selection and statistical estimation , author=. arXiv preprint arXiv:1002.1583 , year=
-
[28]
Biometrics , volume=
Doubly robust estimation in missing data and causal inference models , author=. Biometrics , volume=. 2005 , publisher=
2005
-
[29]
Nearly unbiased variable selection under minimax concave penalty , author=
-
[30]
Probability Theory and Related Fields , volume=
The lower tail of random quadratic forms with applications to ordinary least squares , author=. Probability Theory and Related Fields , volume=. 2016 , publisher=
2016
-
[31]
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=
Regression shrinkage and selection via the lasso , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 1996 , publisher=
1996
-
[32]
Management Science , volume=
Mostly exploration-free algorithms for contextual bandits , author=. Management Science , volume=. 2021 , publisher=
2021
-
[33]
International Conference on Machine Learning , pages=
Leveraging good representations in linear contextual bandits , author=. International Conference on Machine Learning , pages=. 2021 , organization=
2021
-
[34]
Stochastic Systems , volume=
A linear response bandit problem , author=. Stochastic Systems , volume=. 2013 , publisher=
2013
-
[35]
International Conference on Machine Learning , pages=
A simple unified framework for high dimensional bandit problems , author=. International Conference on Machine Learning , pages=. 2022 , organization=
2022
-
[36]
International Conference on Artificial Intelligence and Statistics , pages=
Adaptive exploration in linear contextual bandit , author=. International Conference on Artificial Intelligence and Statistics , pages=. 2020 , organization=
2020
-
[37]
International Conference on Machine Learning , pages=
Minimax concave penalized multi-armed bandit model with high-dimensional covariates , author=. International Conference on Machine Learning , pages=. 2018 , organization=
2018
-
[38]
International Conference on Machine Learning , pages=
Thresholded lasso bandit , author=. International Conference on Machine Learning , pages=. 2022 , organization=
2022
-
[39]
Advances in Neural Information Processing Systems , volume=
High-dimensional sparse linear bandits , author=. Advances in Neural Information Processing Systems , volume=
-
[40]
Electronic Journal of Statistics , volume=
Regret lower bound and optimal algorithm for high-dimensional contextual linear bandit , author=. Electronic Journal of Statistics , volume=. 2021 , publisher=
2021
-
[41]
International Conference on Machine Learning , pages=
Thompson sampling for high-dimensional sparse linear contextual bandits , author=. International Conference on Machine Learning , pages=. 2023 , organization=
2023
-
[42]
2011 , publisher=
Statistics for high-dimensional data: methods, theory and applications , author=. 2011 , publisher=
2011
-
[43]
Finite-time analysis of kernelised contextual bandits , author =
-
[44]
, author =
X-Armed Bandits. , author =
-
[45]
Conference on Learning Theory , pages=
Towards minimax policies for online linear optimization with bandit feedback , author=. Conference on Learning Theory , pages=. 2012 , organization=
2012
-
[46]
Doubly-robust lasso bandit , author =
-
[47]
Operations Research , publisher =
Online decision making with high-dimensional covariates , author =. Operations Research , publisher =
-
[48]
International Joint Conference on Artificial Intelligence , url =
Perturbed-History Exploration in Stochastic Multi-Armed Bandits , author =. International Joint Conference on Artificial Intelligence , url =
-
[49]
Conference on learning theory , pages =
Analysis of thompson sampling for the multi-armed bandit problem , author =. Conference on learning theory , pages =
-
[50]
Proceedings of the 24th annual conference on learning theory , pages =
The KL-UCB algorithm for bounded stochastic bandits and beyond , author =. Proceedings of the 24th annual conference on learning theory , pages =
-
[51]
2013 IEEE Information Theory Workshop (ITW) , pages=
Informational confidence bounds for self-normalized averages and applications , author=. 2013 IEEE Information Theory Workshop (ITW) , pages=. 2013 , organization=
2013
-
[52]
Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages=
Contextual bandit algorithms with supervised learning guarantees , author=. Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages=. 2011 , organization=
2011
-
[53]
, author =
Minimax Policies for Adversarial and Stochastic Bandits. , author =. COLT , volume = 7, pages =
-
[54]
, author =
Action Elimination and Stopping Conditions for the Multi-Armed Bandit and Reinforcement Learning Problems. , author =
-
[55]
Proceedings of the 40th International Conference on Machine Learning , publisher =
Combinatorial Neural Bandits , author =. Proceedings of the 40th International Conference on Machine Learning , publisher =
-
[56]
Composite convex minimization involving self-concordant-like cost functions , author =
-
[57]
European Journal of Operational Research , publisher =
A tractable online learning algorithm for the multinomial logit contextual bandit , author =. European Journal of Operational Research , publisher =
-
[58]
Proceedings of the 2014 SIAM International Conference on Data Mining , pages =
Contextual combinatorial bandit and its application on diversified online recommendation , author =. Proceedings of the 2014 SIAM International Conference on Data Mining , pages =
2014
-
[59]
Self-concordant analysis for logistic regression , author =
-
[60]
Dynamic assortment optimization with changing contextual information , author =
-
[61]
International conference on machine learning , pages =
Why is posterior sampling better than optimism for reinforcement learning? , author =. International conference on machine learning , pages =
-
[62]
An empirical evaluation of thompson sampling , author =
-
[63]
The International Journal of Robotics Research , publisher =
Reinforcement learning in robotics: A survey , author =. The International Journal of Robotics Research , publisher =
-
[64]
Nature , publisher =
Discovering faster matrix multiplication algorithms with reinforcement learning , author =. Nature , publisher =
-
[65]
Science , publisher =
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play , author =. Science , publisher =
-
[66]
nature , publisher =
Mastering the game of go without human knowledge , author =. nature , publisher =
-
[67]
nature , publisher =
Human-level control through deep reinforcement learning , author =. nature , publisher =
-
[68]
9th International Conference on Learning Representations,
Optimism in Reinforcement Learning with Generalized Linear Function Approximation , author =. 9th International Conference on Learning Representations,
-
[69]
Advances in Neural Information Processing Systems , volume = 31, pages =
On Oracle-Efficient PAC RL with Rich Observations , author =. Advances in Neural Information Processing Systems , volume = 31, pages =
-
[70]
Advances in Neural Information Processing Systems , volume = 29, pages =
PAC Reinforcement Learning with Rich Observations , author =. Advances in Neural Information Processing Systems , volume = 29, pages =
-
[71]
International Conference on Machine Learning , pages =
Provably efficient RL with rich observations via latent state decoding , author =. International Conference on Machine Learning , pages =
-
[72]
, author =
Eluder Dimension and the Sample Complexity of Optimistic Exploration. , author =. Advances in Neural Information Processing Systems , pages =
-
[73]
Learning for Dynamics and Control , pages =
Model-based reinforcement learning with value-targeted regression , author =. Learning for Dynamics and Control , pages =
-
[74]
Machine learning , publisher =
Linear least-squares algorithms for temporal difference learning , author =. Machine learning , publisher =
-
[75]
International Conference on Machine Learning , pages =
Provably efficient reinforcement learning for discounted mdps with feature mapping , author =. International Conference on Machine Learning , pages =
-
[76]
International Conference on Machine Learning , pages =
Logarithmic regret for reinforcement learning with linear function approximation , author =. International Conference on Machine Learning , pages =
-
[77]
Algorithmic Learning Theory , pages =
Exponential lower bounds for planning in mdps with linearly-realizable optimal action-value functions , author =. Algorithmic Learning Theory , pages =
-
[78]
Reinforcement Learning with General Value Function Approximation: Provably Efficient Approach via Bounded Eluder Dimension , author =
-
[79]
International Conference on Machine Learning , pages =
Provably efficient exploration in policy optimization , author =. International Conference on Machine Learning , pages =
-
[80]
8th International Conference on Learning Representations,
Is a Good Representation Sufficient for Sample Efficient Reinforcement Learning? , author =. 8th International Conference on Learning Representations,
-
[81]
International Conference on Artificial Intelligence and Statistics , pages =
Sample complexity of reinforcement learning using linearly combined model ensembles , author =. International Conference on Artificial Intelligence and Statistics , pages =
-
[82]
International Conference on Machine Learning , pages =
Contextual decision processes with low Bellman rank are PAC-learnable , author =. International Conference on Machine Learning , pages =
-
[83]
International Conference on Machine Learning , pages =
Model-free reinforcement learning: from clipped pseudo-regret to sample complexity , author =. International Conference on Machine Learning , pages =
-
[84]
Advances in Neural Information Processing Systems , volume = 33, pages =
Almost Optimal Model-Free Reinforcement Learningvia Reference-Advantage Decomposition , author =. Advances in Neural Information Processing Systems , volume = 33, pages =
-
[85]
, author =
Deep Exploration via Randomized Value Functions. , author =. Journal of Machine Learning Research , volume = 20, number = 124, pages =
-
[86]
Advances in Neural Information Processing Systems , volume = 31, pages =
Is Q-Learning Provably Efficient? , author =. Advances in Neural Information Processing Systems , volume = 31, pages =
-
[87]
Learning Unknown Markov Decision Processes:
Yi Ouyang and Mukul Gagrani and Ashutosh Nayyar and Rahul Jain , year = 2017, booktitle =. Learning Unknown Markov Decision Processes:
2017
-
[88]
Advances in Neural Information Processing Systems , pages =
Posterior sampling for reinforcement learning: worst-case regret bounds , author =. Advances in Neural Information Processing Systems , pages =
-
[89]
Advances in Neural Information Processing Systems , volume = 30, pages =
Unifying PAC and Regret: Uniform PAC Bounds for Episodic Reinforcement Learning , author =. Advances in Neural Information Processing Systems , volume = 30, pages =
-
[90]
International Conference on Machine Learning , pages =
Minimax regret bounds for reinforcement learning , author =. International Conference on Machine Learning , pages =
-
[91]
Advances in Neural Information Processing Systems , pages =
Model-based Reinforcement Learning and the Eluder Dimension , author =. Advances in Neural Information Processing Systems , pages =
-
[92]
, author =
Near-optimal Regret Bounds for Reinforcement Learning. , author =
-
[93]
(More) efficient reinforcement learning via posterior sampling , author =
-
[94]
Proceedings of the AAAI Conference on Artificial Intelligence , volume = 35, number = 10, pages =
Multinomial logit contextual bandits: Provable optimality and practicality , author =. Proceedings of the AAAI Conference on Artificial Intelligence , volume = 35, number = 10, pages =
-
[95]
International Conference on Machine Learning , pages =
Provably optimal algorithms for generalized linear contextual bandits , author =. International Conference on Machine Learning , pages =
-
[96]
Proceedings of the 23rd International Conference on Neural Information Processing Systems - Volume 1 , location =
Parametric Bandits: The Generalized Linear Case , author =. Proceedings of the 23rd International Conference on Neural Information Processing Systems - Volume 1 , location =
-
[97]
Conference on Learning Theory , pages =
Provably efficient reinforcement learning with linear function approximation , author =. Conference on Learning Theory , pages =
-
[98]
International Conference on Machine Learning , pages =
Generalization and exploration via randomized value functions , author =. International Conference on Machine Learning , pages =
-
[99]
Worst-case regret bounds for exploration via randomized value functions , author =
-
[100]
International Conference on Machine Learning , pages =
Model-based reinforcement learning with value-targeted regression , author =. International Conference on Machine Learning , pages =
-
[101]
International Conference on Machine Learning , pages =
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound , author =. International Conference on Machine Learning , pages =
-
[102]
International Conference on Machine Learning , pages =
Sample-optimal parametric Q-learning using linearly additive features , author =. International Conference on Machine Learning , pages =
-
[103]
Dynamic pricing and assortment under a contextual MNL demand , author =
-
[104]
Advances in Neural Information Processing Systems , volume = 35, pages =
Dynamic pricing and assortment under a contextual MNL demand , author =. Advances in Neural Information Processing Systems , volume = 35, pages =
-
[105]
International Conference on Machine Learning , pages =
Improved optimistic algorithms for logistic bandits , author =. International Conference on Machine Learning , pages =
-
[106]
International Conference on Artificial Intelligence and Statistics , pages =
Instance-wise minimax-optimal algorithms for logistic bandits , author =. International Conference on Artificial Intelligence and Statistics , pages =
-
[107]
International Conference on Artificial Intelligence and Statistics , pages =
Jointly Efficient and Optimal Algorithms for Logistic Bandits , author =. International Conference on Artificial Intelligence and Statistics , pages =
-
[108]
Advances in Neural Information Processing Systems , booktitle =
Improved algorithms for linear stochastic bandits , author =. Advances in Neural Information Processing Systems , booktitle =
-
[109]
Concentration inequalities , author =
-
[110]
Proceedings of the AAAI conference on artificial intelligence , volume = 37, number = 7, pages =
Model-based reinforcement learning with multinomial logistic function approximation , author =. Proceedings of the AAAI conference on artificial intelligence , volume = 37, number = 7, pages =
-
[111]
International Conference on Artificial Intelligence and Statistics , pages =
Frequentist regret bounds for randomized least-squares value iteration , author =. International Conference on Artificial Intelligence and Statistics , pages =
-
[112]
Advances in Neural Information Processing Systems , booktitle =
Thompson sampling for multinomial logit contextual bandits , author =. Advances in Neural Information Processing Systems , booktitle =
-
[113]
International Conference on Machine Learning , publisher =
Randomized Exploration in Reinforcement Learning with General Value Function Approximation , author =. International Conference on Machine Learning , publisher =
-
[114]
International Conference on Machine Learning , pages =
Neural contextual bandits with ucb-based exploration , author =. International Conference on Machine Learning , pages =
-
[115]
Conference on Learning Theory , pages =
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes , author =. Conference on Learning Theory , pages =
-
[116]
Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages =
Contextual bandits with linear payoff functions , author =. Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages =
-
[117]
International conference on machine learning , pages =
Thompson sampling for contextual bandits with linear payoffs , author =. International conference on machine learning , pages =
-
[118]
Uncertainty in Artificial Intelligence , pages =
Perturbed-History Exploration in Stochastic Linear Bandits , author =. Uncertainty in Artificial Intelligence , pages =
-
[119]
Optimism in reinforcement learning with generalized linear function approximation , author =
-
[120]
International Conference on Artificial Intelligence and Statistics , pages =
Randomized exploration in generalized linear bandits , author =. International Conference on Artificial Intelligence and Statistics , pages =
-
[121]
Advances in Neural Information Processing Systems , booktitle =
Scalable generalized linear bandits: Online computation and hashing , author =. Advances in Neural Information Processing Systems , booktitle =
-
[122]
Foundations and Trends
A tutorial on thompson sampling , author =. Foundations and Trends
-
[123]
Linear algebra and geometry , author =
-
[124]
Journal of Computer and System Sciences , publisher =
An analysis of model-based interval estimation for Markov decision processes , author =. Journal of Computer and System Sciences , publisher =
-
[125]
21st Annual Conference on Learning Theory , number=
Stochastic linear optimization under bandit feedback , author=. 21st Annual Conference on Learning Theory , number=
-
[126]
Machine learning , publisher =
Finite-time analysis of the multiarmed bandit problem , author =. Machine learning , publisher =
-
[127]
Journal of Machine Learning Research , volume = 3, number =
Using confidence bounds for exploitation-exploration trade-offs , author =. Journal of Machine Learning Research , volume = 3, number =
-
[128]
Available at SSRN , publisher =
Near-optimal algorithms for capacity constrained assortment optimization , author =. Available at SSRN , publisher =
-
[129]
Artificial intelligence and statistics , pages =
Further optimal regret bounds for thompson sampling , author =. Artificial intelligence and statistics , pages =
-
[130]
Algorithms for non-stationary generalized linear bandits , author =
-
[131]
the Annals of Probability , publisher =
On tail probabilities for martingales , author =. the Annals of Probability , publisher =
-
[132]
Advances in Neural Information Processing Systems , volume = 13, pages =
Finite-sample convergence rates for Q-learning and indirect algorithms , author =. Advances in Neural Information Processing Systems , volume = 13, pages =
-
[133]
On the sample complexity of reinforcement learning , author =
-
[134]
International Conference on Machine Learning , pages =
Sparsity-agnostic lasso bandit , author =. International Conference on Machine Learning , pages =
-
[135]
Playing atari with deep reinforcement learning , author =
-
[136]
Contextual combinatorial multi-armed bandits with volatile arms and submodular reward , author =
-
[137]
International Conference on Artificial Intelligence and Statistics , pages =
Contextual combinatorial volatile multi-armed bandit with adaptive discretization , author =. International Conference on Artificial Intelligence and Statistics , pages =
-
[138]
Mathematics of Operations Research , publisher =
Regret in online combinatorial optimization , author =. Mathematics of Operations Research , publisher =
-
[139]
International Conference on Learning Representations , url =
Learning Neural Contextual Bandits through Perturbed Rewards , author =. International Conference on Learning Representations , url =
-
[140]
Advances in neural information processing systems , booktitle =
Learning and generalization in overparameterized neural networks, going beyond two layers , author =. Advances in neural information processing systems , booktitle =
-
[141]
Mathematics of Operations Research , publisher =
Assortment optimization under variants of the nested logit model , author =. Mathematics of Operations Research , publisher =
-
[142]
1303.6746 , archiveprefix =
Exploiting correlation and budget constraints in Bayesian multi-armed bandit optimization , author =. 1303.6746 , archiveprefix =
-
[143]
Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables , year = 1964, publisher =
1964
-
[144]
1811.03962 , archiveprefix =
A Convergence Theory for Deep Learning via Over-Parameterization , author =. 1811.03962 , archiveprefix =
-
[145]
Neural Thompson Sampling , author =
-
[146]
Proceedings of the 2014 SIAM International Conference on Data Mining , publisher =
Contextual Combinatorial Bandit and its Application on Diversified Online Recommendation , author =. Proceedings of the 2014 SIAM International Conference on Data Mining , publisher =
2014
-
[147]
Proceedings of the 17th International Conference on Machine Learning (ICML 2000) , publisher =
Crafting Papers on Machine Learning , author =. Proceedings of the 17th International Conference on Machine Learning (ICML 2000) , publisher =
2000
-
[148]
The Need for Biases in Learning Generalizations , author =
-
[149]
Computational Complexity of Machine Learning , author =
-
[150]
I , year = 1983, publisher =
Machine Learning: An Artificial Intelligence Approach, Vol. I , year = 1983, publisher =
1983
-
[151]
Pattern Classification , author =
-
[152]
Bandit Algorithms , author =
-
[153]
Suppressed for Anonymity , author =
-
[154]
Cognitive Skills and Their Acquisition , publisher =
Mechanisms of Skill Acquisition and the Law of Practice , author =. Cognitive Skills and Their Acquisition , publisher =
-
[155]
IBM Journal of Research and Development , volume = 3, number = 3, pages =
Some Studies in Machine Learning Using the Game of Checkers , author =. IBM Journal of Research and Development , volume = 3, number = 3, pages =
-
[156]
International Conference on Machine Learning , pages =
Efficient learning in large-scale combinatorial semi-bandits , author =. International Conference on Machine Learning , pages =
-
[157]
International Conference on Machine Learning , pages =
Cascading bandits: Learning to rank in the cascade model , author =. International Conference on Machine Learning , pages =
-
[158]
Proceedings of the Thirty-Second Conference on Uncertainty in Artificial Intelligence , series =
Cascading Bandits for Large-scale Recommendation Problems , author =. Proceedings of the Thirty-Second Conference on Uncertainty in Artificial Intelligence , series =
-
[159]
International conference on machine learning , pages =
Contextual combinatorial cascading bandits , author =. International conference on machine learning , pages =
-
[160]
International Conference on Machine Learning , pages =
Online learning to rank with features , author =. International Conference on Machine Learning , pages =
-
[161]
Mathematics of Operations Research , publisher =
Linearly parameterized bandits , author =. Mathematics of Operations Research , publisher =
-
[162]
International Conference on Machine Learning , pages =
Associative reinforcement learning using linear probabilistic concepts , author =. International Conference on Machine Learning , pages =
-
[163]
Operations research , publisher =
Dynamic assortment optimization with a multinomial logit choice model and capacity constraint , author =. Operations research , publisher =
-
[164]
Neural Contextual Bandits with Deep Representation and Shallow Exploration , author =
-
[165]
nature , publisher =
Deep learning , author =. nature , publisher =
-
[166]
nature , publisher =
Mastering the game of Go with deep neural networks and tree search , author =. nature , publisher =
-
[167]
Biometrika , publisher =
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples , author =. Biometrika , publisher =
-
[168]
Advances in Neural Information Processing Systems , volume = 32, pages =
Generalization Bounds of Stochastic Gradient Descent for Wide and Deep Neural Networks , author =. Advances in Neural Information Processing Systems , volume = 32, pages =
-
[169]
Deep learning , author =
-
[170]
Advances in Neural Information Processing Systems , volume = 31, pages =
Neural Tangent Kernel: Convergence and Generalization in Neural Networks , author =. Advances in Neural Information Processing Systems , volume = 31, pages =
-
[171]
International Conference on Learning Representations , url =
Gradient Descent Provably Optimizes Over-parameterized Neural Networks , author =. International Conference on Learning Representations , url =
-
[172]
International Conference on Machine Learning , pages =
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks , author =. International Conference on Machine Learning , pages =
-
[173]
Advances in applied mathematics , publisher =
Asymptotically efficient adaptive allocation rules , author =. Advances in applied mathematics , publisher =
-
[174]
Online clustering of contextual cascading bandits , author =
-
[175]
What doubling tricks can and can't do for multi-armed bandits , author =
-
[176]
International Conference on Machine Learning , pages =
Gradient descent finds global minima of deep neural networks , author =. International Conference on Machine Learning , pages =
-
[177]
Advances in Neural Information Processing Systems , publisher =
An Improved Analysis of Training Over-parameterized Deep Neural Networks , author =. Advances in Neural Information Processing Systems , publisher =
-
[178]
On the conditions used to prove oracle results for the lasso , author=
-
[179]
arXiv preprint arXiv:2006.06790 , year=
On frequentist regret of linear thompson sampling , author=. arXiv preprint arXiv:2006.06790 , year=
2006
-
[180]
Mathematics of Operations Research , volume=
Learning to optimize via posterior sampling , author=. Mathematics of Operations Research , volume=. 2014 , publisher=
2014
-
[181]
Proceedings of the 53rd Annual ACM SIGACT Symposium on Theory of Computing , pages=
Linear bandits with limited adaptivity and learning distributional optimal design , author=. Proceedings of the 53rd Annual ACM SIGACT Symposium on Theory of Computing , pages=
-
[182]
Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Regret bounds for batched bandits , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=
-
[183]
Asian Journal of Mathematics & Statistics , volume=
Concise formulas for the area and volume of a hyperspherical cap , author=. Asian Journal of Mathematics & Statistics , volume=. 2010 , publisher=
2010
-
[184]
International Conference on Machine Learning , pages=
Optimal Batched Linear Bandits , author=. International Conference on Machine Learning , pages=. 2024 , organization=
2024
-
[185]
Conference on Learning Theory , pages=
Asymptotically optimal information-directed sampling , author=. Conference on Learning Theory , pages=. 2021 , organization=
2021
-
[186]
The Thirty Sixth Annual Conference on Learning Theory , pages=
Contexts can be cheap: Solving stochastic contextual bandits with linear bandit algorithms , author=. The Thirty Sixth Annual Conference on Learning Theory , pages=. 2023 , organization=
2023
-
[187]
Advances in Neural Information Processing Systems , volume=
Batched multi-armed bandits problem , author=. Advances in Neural Information Processing Systems , volume=
-
[188]
Foundations of computational mathematics , volume=
User-friendly tail bounds for sums of random matrices , author=. Foundations of computational mathematics , volume=. 2012 , publisher=
2012
-
[189]
Conference on Learning Theory , pages=
Nearly minimax-optimal regret for linearly parameterized bandits , author=. Conference on Learning Theory , pages=. 2019 , organization=
2019
-
[190]
2019 , note =
Yuan Zhou , title =. 2019 , note =
2019
-
[191]
The Thirteenth International Conference on Learning Representations , year=
Almost Optimal Batch-Regret Tradeoff for Batch Linear Contextual Bandits , author=. The Thirteenth International Conference on Learning Representations , year=
-
[192]
ACM Computing Surveys (CSUR) , volume=
Reinforcement learning in healthcare: A survey , author=. ACM Computing Surveys (CSUR) , volume=. 2021 , publisher=
2021
-
[193]
Proceedings of the 19th international conference on World wide web , pages=
A contextual-bandit approach to personalized news article recommendation , author=. Proceedings of the 19th international conference on World wide web , pages=
-
[194]
Advances in Neural Information Processing Systems , volume=
Online models for content optimization , author=. Advances in Neural Information Processing Systems , volume=
-
[195]
Operations Research , volume=
A learning approach for interactive marketing to a customer segment , author=. Operations Research , volume=. 2007 , publisher=
2007
-
[196]
Life , volume=
A contextual-bandit-based approach for informed decision-making in clinical trials , author=. Life , volume=. 2022 , publisher=
2022
-
[197]
Algorithmica , volume=
Reinforcement learning with immediate rewards and linear hypotheses , author=. Algorithmica , volume=. 2003 , publisher=
2003
-
[198]
Advances in neural information processing systems , volume=
The epoch-greedy algorithm for contextual multi-armed bandits , author=. Advances in neural information processing systems , volume=. 2007 , publisher=
2007
-
[199]
arXiv preprint arXiv:2004.06321 , year=
Sequential batch learning in finite-action linear contextual bandits , author=. arXiv preprint arXiv:2004.06321 , year=
2004
-
[200]
Canadian Journal of Mathematics , volume=
The equivalence of two extremum problems , author=. Canadian Journal of Mathematics , volume=. 1960 , publisher=
1960
-
[201]
International conference on machine learning , pages=
Learning with good feature representations in bandits and in rl with a generative model , author=. International conference on machine learning , pages=. 2020 , organization=
2020
-
[202]
Conference on Learning Theory , pages=
Batched Bandit Problems , author=. Conference on Learning Theory , pages=. 2015 , organization=
2015
-
[203]
Management Science , volume=
Shrinking the upper confidence bound: A dynamic product selection problem for urban warehouses , author=. Management Science , volume=. 2021 , publisher=
2021
-
[204]
Handbook of Statistical Methods for Precision Medicine , pages=
Bandit algorithms for precision medicine , author=. Handbook of Statistical Methods for Precision Medicine , pages=. 2021 , publisher=
2021
-
[205]
The Annals of Statistics , volume=
Batched bandit problems , author=. The Annals of Statistics , volume=
-
[206]
International Conference on Machine Learning , pages=
Almost optimal anytime algorithm for batched multi-armed bandits , author=. International Conference on Machine Learning , pages=. 2021 , organization=
2021
-
[207]
Conference on Learning Theory , pages=
Double explore-then-commit: Asymptotic optimality and beyond , author=. Conference on Learning Theory , pages=. 2021 , organization=
2021
-
[208]
Advances in Neural Information Processing Systems , volume=
Optimal batched best arm identification , author=. Advances in Neural Information Processing Systems , volume=
-
[209]
Monographies des Probabilit
Etude critique de la notion de collectif, Gauthier-Villars, Paris, 1939 , author=. Monographies des Probabilit
1939
-
[210]
Les rencontres physiciens-math
Convex trace functions and the Wigner-Yanase-Dyson conjecture , author=. Les rencontres physiciens-math
-
[211]
Journal of Mathematical Physics , volume=
Inequality with applications in statistical mechanics , author=. Journal of Mathematical Physics , volume=. 1965 , publisher=
1965
-
[212]
Advances in Neural Information Processing Systems , volume=
Generalized linear bandits with limited adaptivity , author=. Advances in Neural Information Processing Systems , volume=
-
[213]
Mathematical Programming , volume=
Near-optimal discrete optimization for experimental design: A regret minimization approach , author=. Mathematical Programming , volume=. 2021 , publisher=
2021
-
[214]
2012 , publisher=
Geometric algorithms and combinatorial optimization , author=. 2012 , publisher=
2012
-
[215]
Discrete Applied Mathematics , volume=
On Khachiyan's algorithm for the computation of minimum-volume enclosing ellipsoids , author=. Discrete Applied Mathematics , volume=. 2007 , publisher=
2007
-
[216]
arXiv e-prints , pages=
Positive Semidefinite Matrix Supermartingales , author=. arXiv e-prints , pages=
-
[217]
Advances in Neural Information Processing Systems , volume=
Efficient batched algorithm for contextual linear bandits with large action space via soft elimination , author=. Advances in Neural Information Processing Systems , volume=
-
[218]
International Conference on Machine Learning , pages=
A reduction from linear contextual bandits lower bounds to estimations lower bounds , author=. International Conference on Machine Learning , pages=. 2022 , organization=
2022
-
[219]
IEEE Control Systems Letters , volume=
Batched learning in generalized linear contextual bandits with general decision sets , author=. IEEE Control Systems Letters , volume=. 2020 , publisher=
2020
-
[220]
IEEE Transactions on Information Theory , year=
Batched nonparametric contextual bandits , author=. IEEE Transactions on Information Theory , year=
-
[221]
Advances in Neural Information Processing Systems , volume=
Batched thompson sampling , author=. Advances in Neural Information Processing Systems , volume=
-
[222]
Management Science , volume=
Dynamic batch learning in high-dimensional sparse linear contextual bandits , author=. Management Science , volume=. 2024 , publisher=
2024
-
[223]
Advances in Neural Information Processing Systems , volume=
Parallelizing thompson sampling , author=. Advances in Neural Information Processing Systems , volume=
-
[224]
arXiv preprint arXiv:2511.03708 , year=
The Adaptivity Barrier in Batched Nonparametric Bandits: Sharp Characterization of the Price of Unknown Margin , author=. arXiv preprint arXiv:2511.03708 , year=
-
[225]
The Thirty-Ninth Annual Conference on Neural Information Processing Systems , year=
Generalized linear bandits: Almost optimal regret with one-pass update , author=. The Thirty-Ninth Annual Conference on Neural Information Processing Systems , year=
-
[226]
ACM/JMS Journal of Data Science , volume=
Batched neural bandits , author=. ACM/JMS Journal of Data Science , volume=. 2024 , publisher=
2024
-
[227]
Forty-Second International Conference on Machine Learning , year=
Optimal and Practical Batched Linear Bandit Algorithm , author=. Forty-Second International Conference on Machine Learning , year=
Reviewed June 28, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.