REVIEW 1 major objections 5 minor 151 references
Selective Reviews of Bandit Problems in AI via a Statistical View
T0 review · 1 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper argues that stochastic bandit theory is best read as non-asymptotic statistical inference, where concentration inequalities yield the confidence intervals, regret decompositions, and minimax benchmarks that organize multi-armed…
desk verdict A useful but flawed survey: the alternative UCB proof in Appendix B has a load-bearing gap, so the paper's one original contribution is not established as written. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The object that carries the argument is the regret-decomposition identity for optimistic algorithms: for UCB, $\mathrm{Reg}_T(\pi, v) \le \mathbb{E}\sum_{t=1}^T [\mu^* - \mathrm{UCB}_*(t-1,\delta) + \mathrm{UCB}_{A_t}(t-1,\delta) - \mu_{A_t}]$, where the selected arm maximizes the index so the optimal arm's index term contributes a negative concentration error. This reduces regret to the sum of concentration errors of the chosen and optimal arms; feeding in a sub-Gaussian tail bound plus a union bound yields the $O(K + \sqrt{KT\log T})$ rate. The companion structural objects are the minimax lower bound constructed from Gaussian instances, which fixes the $\sqrt{KT}$ benchmark, and the information gain $\gamma_T$, which quantifies the effective dimensionality a GP-UCB learner must explore over a continuous action space.
What would settle it
Fix a K-armed bandit with near-equal means (small gaps $\Delta_k$) and rewards drawn from a mixture such as $0.4\,N(0,1) + 0.6\,N(0,9)$, so the true sub-Gaussian proxy exceeds 1; run Algorithm 3 with proxy 1 for $T=10^4$ and compare the observed cumulative regret to $3\sum_k \Delta_k + 8\sqrt{TK\log T}$, since a violation at the claimed confidence level would show that the known-proxy assumption is carrying the central theorem.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the many bandit settings share a single statistical skeleton: every algorithm's exploration is a confidence interval, every regret bound is a concentration calculation, and optimality is judged by minimax rate. The review demonstrates this by reproving the UCB regret theorem—for subG(1) rewards with $\delta = 1/T^2$, $\mathrm{Reg}_T(\pi, v) \le 3\sum_k \Delta_k + 8\sqrt{TK\log T}$, with an alternative proof giving $\mathrm{Reg}_T \le 4|\mu^*|K + 12\sqrt{\log T}\sqrt{KT}$—through a regret decomposition that mirrors excess risk in statistical learning. It then applies the same reading to linear contextual bandits, where LinUCB and linear Thompson sampling attain $\widetilde{O}(d\sqrt{T})$ regret independent of the number of arms, and to continuum-armed bandits, where the GP-UCB regret is controlled by the maximum information gain $\gamma_T$. The review also connects continuum-armed bandits to functional data analysis, treating the unknown reward as a function to be estimated over a continuous domain.
Load-bearing premise
The main theorems assume the reward noise has a known sub-Gaussian variance proxy, normally set to 1, together with a known horizon $T$ and confidence level $\delta$; if that proxy is unknown, misspecified, or the noise is heavy-tailed, the concentration bound in display (19) and Theorem 5 do not hold.
Editorial extensions
If this is right
- UCB with subG(1) rewards has problem-independent regret $O(\sqrt{TK\log T})$, and the gap to the $\sqrt{KT}$ minimax lower bound is closed by MOSS and the minimax-optimal Thompson-sampling variant MOTS.
- Contextual linear bandits, both LinUCB and linear Thompson sampling, achieve regret $\widetilde{O}(d\sqrt{T})$ independent of the number of arms, so shared features pay off whenever the feature dimension is manageable.
- For continuum-armed bandits with a Gaussian-process prior, the cumulative regret is bounded by $\sqrt{C_1 T \beta_T \gamma_T}$, where $\gamma_T$ is the maximum information gain, which gives sublinear regret for common kernels.
- The causal treatment-effect problem can be modeled as a two-armed bandit, yielding non-asymptotic confidence intervals for the treatment effect when the potential outcomes are sub-Gaussian.
- The doubling trick converts the horizon-dependent algorithms described in the review into anytime algorithms without changing their regret rates.
Reading between the lines
- Read as a template, the regret-decomposition lemma likely extends to optimistic algorithms beyond UCB: LinUCB and GP-UCB can be viewed as the same concentration-error-plus-estimation-error split with different confidence radii, though the review does not state this unification explicitly.
- The advertised connection between continuum-armed bandits and functional data analysis suggests transfers the authors leave implicit, such as using functional principal components or basis smoothing inside a continuum bandit when the reward function has low-rank structure, which could reduce the effective action dimension.
- The simulations in the unknown-variance-proxy section hint at a practical rule the paper does not state: when the variance proxy is unknown, estimating the sub-Gaussian norm online and plugging it into the confidence radius can outperform both asymptotic variance estimators and wrongly applied bounded-reward concentration bounds; this could be tested systematically across reward families.
- If the statistical reading is correct, then designing a new bandit algorithm reduces to finding the tightest valid concentration inequality for the reward family, meaning any new tail bound would directly yield a new UCB-style algorithm with a corresponding regret bound.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript is a survey of stochastic bandit problems (finite-armed, contextual, and continuum-armed) viewed through the lens of non-asymptotic statistics. It introduces concentration inequalities for sub-Gaussian and sub-exponential rewards, reviews ETC, UCB, MOSS, Thompson sampling, MOTS, LinUCB, LinTS, GP-UCB, and GP-TS, and discusses connections to functional data analysis and recent topics such as unknown variance proxies. A central feature is an 'alternative proof' of the UCB regret bound based on an excess-risk decomposition from statistical learning theory, presented in Appendix B.
Significance. If correct, the survey would provide a compact bridge between bandit theory and statistical learning theory, with a coherent narrative that many bandit guarantees reduce to concentration inequalities and risk decompositions. The paper is strongest as a compilation: it gathers standard algorithms and bounds, attributes them clearly, and includes simulations that illustrate qualitative differences among algorithms. It also honestly acknowledges the known-variance-proxy limitation in Section 6.4. However, the advertised alternative UCB proof in Appendix B contains a load-bearing gap, and the sub-exponential example in Section 2.2 contains a false inequality. These issues prevent the present version from fully delivering its methodological message, although the standard results quoted are mostly correct.
major comments (1)
- [Appendix B, Eqs. (A3)–(A4)] The alternative proof of Theorem 5 drops the term μ* − UCB*(t−1,δ) from Lemma 7 unconditionally. This term is non-positive only on the good event G; on G^c it can be positive and of order sqrt(2 log(1/δ)) even when μ* = 0, and it can persist over many rounds. The displayed bound E{1_{G^c} Σ (UCB_{A_t} − μ_{A_t})} ≤ 2|μ*| T P(G^c) is also invalid, since UCB_{A_t} − μ_{A_t} is not bounded by 2|μ*| on G^c. The martingale-difference term in the preceding display is not handled correctly either: the expectation of an adaptively weighted sum is not bounded by max_t E[Y_{A_t}(τ) − μ_{A_t}]. As a result, the claimed regret bound 4|μ*|K + 12 sqrt(log T) sqrt(KT) does not follow from the written argument. Since this proof is advertised in Section 3.2 as the main pedagogical support for the UCB analysis, it needs to be repaired or replaced; the theorem itself is standard and Appendix A provides a valid proof.
minor comments (5)
- [Section 2.2, Example 9] Example 9 asserts e^{-sμ}(1−sμ)^{-1} ≤ e^{s^2 μ^2/2} for |s| ≤ (2μ)^{-1}, but this is false: at s = 0.4/μ the left-hand side is approximately 1.117 and the right-hand side approximately 1.083. The centered exponential is indeed sub-exponential, but the displayed inequality and the implied parameterization (μ, 2μ) should be corrected.
- [Appendix B] In the display after Eq. (A3), the quantity S_{k*} should be S_{A_t}: the confidence radius for the selected arm A_t depends on the number of pulls of that arm, not on the number of pulls of the optimal arm.
- [Sections 3.2 and 6.4] Theorems 5 and 7 assume a known variance proxy (usually 1), while Section 6.4 discusses the unknown-proxy problem only later; adding a forward reference in Section 3.2 would make the scope of the main bounds clearer to readers.
- [Section 2.1, Example 4] The phrase '∑_{i=1}^n X_{ij} i.i.d.∼ N(0, nσ^2)' should read '∑_{i=1}^n X_{ij} ∼ N(0, nσ^2)', since the sum is a single random variable rather than an i.i.d. collection.
- [Various] There are minor typos, e.g., 'genalization' in Section 6.1 should be 'generalization', and the caption of Figure 5 has 'Culumative' instead of 'Cumulative'.
Circularity Check
No significant circularity: the UCB derivation reduces to standard sub-Gaussian concentration and regret decomposition, not to its own assumptions.
full rationale
The paper's central derivations are self-contained reductions to textbook material. Theorem 5 is explicitly attributed to Lattimore and Szepesvári's Theorem 7.2, and Appendix A supplies a standard proof from the regret decomposition lemma and sub-Gaussian concentration inequality (14). Lemma 7 decomposes regret into UCB terms, and Appendix B then applies concentration plus integral estimates; this is a genuine derivation chain, not an identity or a fitted-parameter prediction. Self-citations such as [24], [34], and [47] are used as references for background inequalities and for a simulation variant, not as the load-bearing justification for the main regret bounds. The Appendix B proof has a potential gap in bounding the bad-event contribution, but a proof gap is a correctness issue, not circularity. No fitted constants are renamed as predictions, no uniqueness theorem is imported from the authors, and no target result is assumed in defining its own inputs.
Assumptions & free parameters
free parameters (3)
- GP lengthscale l =
0.2575
- GP-UCB exploration parameter beta =
2.0
- MOTS truncation parameters rho and alpha =
rho = 0.8, alpha = 1.5
assumptions (6)
- domain assumption Rewards are sub-Gaussian with known variance proxy, often subG(1), for the MAB and linear bandit bounds.
- standard math The minimax regret lower bound for K-armed Gaussian bandits is Reg*_T(E) >= (1/27) sqrt((K-1)T).
- standard math The MOSS, MOTS, LinUCB, LinTS, GP-UCB, and GP-TS regret bounds in Theorems 7, 9, 10, 11, 12, and 13 are correct as quoted.
- domain assumption For SCAB, the reward function f is a sample from a Gaussian process with known kernel, or lies in the RKHS of that kernel.
- domain assumption For the Neyman-Rubin example, potential outcomes are subG(sigma^2) and treatment assignment is independent of potential outcomes.
- ad hoc to paper The centered exponential distribution is claimed to satisfy the subexponential MGF bound with parameters (mu, 2mu) in Example 9.
Cite this review
Pith. "Pith review of Selective Reviews of Bandit Problems in AI via a Statistical View." pith.science (2026). https://pith.science/paper/MTETA5FT
@misc{pith2026241202251,
author = {Pith},
title = {Pith review of: Selective Reviews of Bandit Problems in AI via a Statistical View},
year = {2026},
howpublished = {\url{https://pith.science/paper/MTETA5FT}},
note = {Machine review of arXiv:2412.02251}
}
read the original abstract
Reinforcement Learning (RL) is a widely researched area in artificial intelligence that focuses on teaching agents decision-making through interactions with their environment. A key subset includes stochastic multi-armed bandit (MAB) and continuum-armed bandit (SCAB) problems, which model sequential decision-making under uncertainty. This review outlines the foundational models and assumptions of bandit problems, explores non-asymptotic theoretical tools like concentration inequalities and minimax regret bounds, and compares frequentist and Bayesian algorithms for managing exploration-exploitation trade-offs. Additionally, we explore K-armed contextual bandits and SCAB, focusing on their methodologies and regret analyses. We also examine the connections between SCAB problems and functional data analysis. Finally, we highlight recent advances and ongoing challenges in the field.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Reinforcement Learning: An Introduction, 2012.; MIT Press: Cambridge, USA
Richard, S.; Sutton, S.; Andrew, G. Reinforcement Learning: An Introduction, 2012.; MIT Press: Cambridge, USA
2012
-
[2]
Statistical Reinforcement Learning: Modern Machine Learning Approaches; Chapman and Hall/CRC: New York, NY, USA, 2015; pp
Sugiyama, M. Statistical Reinforcement Learning: Modern Machine Learning Approaches; Chapman and Hall/CRC: New York, NY, USA, 2015; pp. 1-206
2015
-
[3]
Autonomous Driving: Technical, Legal and Social Aspects; Springer: Berlin, Germany, 2016; pp
Maurer, M.; Gerdes, J.C.; Lenz, B.; Winner, H. Autonomous Driving: Technical, Legal and Social Aspects; Springer: Berlin, Germany, 2016; pp. 1-706
2016
-
[4]
Large-scale bandit approaches for recommender systems
Zhou, Q.; Zhang, X.; Xu, J.; Liang, B. Large-scale bandit approaches for recommender systems. In Proceedings of the International Conference on Neural Information Processing (ICONIP 2017), Liu, D.; Xie, S.; Li, Y.; Zhao, D.; El-Alfy, E.S., Eds.; Springer: Cham, 2017; pp. 811–821
2017
-
[5]
Unmanned aerial vehicles in agriculture: A survey
Del Cerro, J.; Cruz Ulloa, C.; Barrientos, A.; de León Rivas, J. Unmanned aerial vehicles in agriculture: A survey. Agronomy 2021, 11, 203
2021
-
[6]
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; Xu, R.; Zhu, Q.; Ma, S.; Wang, P .; Bi, X.; et al. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv 2025, arXiv:2501.12948
arXiv 2025
-
[7]
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P .; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. Training language models to follow instructions with human feedback. Adv. Neural Inf. Process. Syst. 2022, 35, 27730–27744
2022
-
[8]
Portfolio choices with orthogonal bandit learning
Shen, W.; Wang, J.; Jiang, Y.G.; Zha, H. Portfolio choices with orthogonal bandit learning. In Proceedings of the 24th International Conference on Artificial Intelligence (IJCAI 2015), Buenos Aires, Argentina, 2015; AAAI Press: 974–980
2015
Show all 151 references
-
[9]
Causal Inference: A Statistical Learning Approach
Wager, S. Causal Inference: A Statistical Learning Approach. 2024. Available online: https://web.stanford.edu/~swager/causal_ inf_book.pdf (accessed on 14th Feb., 2025)
2024
-
[10]
Enhancing digital twins through reinforcement learning
Cronrath, C.; Aderiani, A.R.; Lennartson, B. Enhancing digital twins through reinforcement learning. In Proceedings of the 2019 IEEE 15th International Conference on Automation Science and Engineering (CASE), Vancouver, BC, Canada, 2019; IEEE: 293–298. DOI: https://doi.org/10....
2019
-
[11]
A digital twin to train deep reinforcement learning agent for smart manufacturing plants: Environment, interfaces and intelligence
Xia, K.; Sacco, C.; Kirkpatrick, M.; Saidy, C.; Nguyen, L.; Kircaliali, A.; Harik, R. A digital twin to train deep reinforcement learning agent for smart manufacturing plants: Environment, interfaces and intelligence. J. Manuf. Syst. 2021, 58, 210–230
2021
-
[12]
Contextual bandits for adapting treatment in a mouse model of de novo carcinogenesis
Durand, A.; Achilleos, C.; Iacovides, D.; Strati, K.; Mitsis, G.D.; Pineau, J. Contextual bandits for adapting treatment in a mouse model of de novo carcinogenesis. In Proceedings of the 3rd Machine Learning for Healthcare Conference, 17–18 August 2018, PMLR: 67–82. URL: https...
2018
-
[13]
Bandit algorithms for precision medicine
Lu, Y.; Xu, Z.; Tewari, A. Bandit algorithms for precision medicine. In Handbooks of Modern Statistical Methods; Bühlmann, P ., Drineas, P ., Kane, M., van der Laan, M., Eds.; 2024; Chapter 13
2024
-
[14]
A survey on practical applications of multi-armed and contextual bandits
Bouneffouf, D.; Rish, I. A survey on practical applications of multi-armed and contextual bandits. arXiv 2019, arXiv:1904.10040
2019 arXiv
-
[15]
Introduction to Multi-Armed Bandits
Slivkins, A. Introduction to Multi-Armed Bandits. Found. Trends® Mach. Learn. 2019, 12, 1–286
2019
-
[16]
Survey of multiarmed bandit algorithms applied to recommendation systems
Elena, G.; Milos, K.; Eugene, I. Survey of multiarmed bandit algorithms applied to recommendation systems. Int. J. Open Inf. Technol. 2021, 9, 12–27
2021
-
[17]
A map of bandits for e-commerce
Liu, Y.; Li, L. A map of bandits for e-commerce. In Proceedings of the KDD 2021 Workshop on Multi-Armed Bandits and Reinforcement Learning (MARBLE), 2021. Version February 20, 2025 submitted to arxiv 49 of 52
2021
-
[18]
Bandit algorithms: A comprehensive review and their dynamic selection from a portfolio for multicriteria top-k recommendation
Letard, A.; Gutowski, N.; Camp, O.; Amghar, T. Bandit algorithms: A comprehensive review and their dynamic selection from a portfolio for multicriteria top-k recommendation. Expert Syst. Appl. 2024, 26, 123151
2024
-
[19]
A survey of online experiment design with the stochastic multi-armed bandit
Burtini, G.; Loeppky, J.; Lawrence, R. A survey of online experiment design with the stochastic multi-armed bandit. arXiv 2015, arXiv:1510.00757
2015 arXiv
-
[20]
A survey on Multi-Armed, Contextual and Causal bandit algorithms for online learning Preprint
Shah, N. A survey on Multi-Armed, Contextual and Causal bandit algorithms for online learning Preprint. 2020. Available online: https://nihaarshah.github.io/project_pdfs/ML_Theory_Project_Report.pdf (accessed on 14th Feb., 2025)
2020
-
[21]
Bandit Algorithms; Cambridge University Press: Cambridge, UK, 2020
Lattimore, T.; Szepesvári, C. Bandit Algorithms; Cambridge University Press: Cambridge, UK, 2020
2020
-
[23]
Zero-Inflated Bandits
Wei, H.; Wan, R.; Shi, L.; Song, R. Zero-Inflated Bandits. arXiv 2023, arXiv:2312.15595
2023 arXiv
-
[24]
Sharper sub-weibull concentrations
Zhang, H.; Wei, H. Sharper sub-weibull concentrations. Mathematics 2022, 10, 2252
2022
-
[25]
Non-asymptotic guarantees for robust statistical learning under infinite variance assumption
Xu, L.; Yao, F.; Yao, Q.; Zhang, H. Non-asymptotic guarantees for robust statistical learning under infinite variance assumption. J. Mach. Learn. Res. 2023, 24, 1–46
2023
-
[26]
Time-uniform, nonparametric, nonasymptotic confidence sequences
Howard, S.R.; Ramdas, A.; McAuliffe, J.; Sekhon, J. Time-uniform, nonparametric, nonasymptotic confidence sequences. Ann. Stat. 2021, 49, 1055–1080
2021
-
[27]
Some aspects of the sequential design of experiments
Robbins, H. Some aspects of the sequential design of experiments. Bull. Am. Math. Soc. 1952, 58, 527–535
1952
-
[28]
Regret Analysis of Stochastic and Nonstochastic Multi-Armed Bandit Problems
Bubeck, S.; Cesa-Bianchi, N. Regret Analysis of Stochastic and Nonstochastic Multi-Armed Bandit Problems. In Foundations and Trends® in Machine Learning; Now Publishers, Inc.: 2012; Volume 5, pp. 1–122
2012
-
[29]
Asymptotically efficient adaptive allocation rules
Lai, T.L.; Robbins, H. Asymptotically efficient adaptive allocation rules. Adv. Appl. Math. 1985, 6, 4–22
1985
-
[30]
Nearly dimension-independent sparse linear bandit over small action spaces via best subset selection
Chen, Y.; Wang, Y.; Fang, E.X.; Wang, Z.; Li, R. Nearly dimension-independent sparse linear bandit over small action spaces via best subset selection. J. Am. Stat. Assoc. 2024, 119, 246–258
2024
-
[31]
High-dimensional sparse linear bandits
Hao, B.; Lattimore, T.; Wang, M. High-dimensional sparse linear bandits. Adv. Neural Inf. Process. Syst. 2020, 33, 10753–10763
2020
-
[32]
Efficient sparse linear bandits under high dimensional data
Wang, X.; Wei, M.M.; Yao, T. Efficient sparse linear bandits under high dimensional data. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Long Beach, CA, USA, 2023; pp. 2431–2443
2023
-
[33]
Provably Efficient High-Dimensional Bandit Learning with Batched Feedbacks
Fan, J.; Wang, Z.; Yang, Z.; Ye, C. Provably Efficient High-Dimensional Bandit Learning with Batched Feedbacks. arXiv 2023, arXiv:2311.13180
2023 arXiv
-
[34]
Tight non-asymptotic inference via sub-Gaussian intrinsic moment norm
Zhang, H.; Wei, H.; Cheng, G. Tight non-asymptotic inference via sub-Gaussian intrinsic moment norm. arXiv 2023, arXiv:2303.07287
2023
-
[35]
A perspective on off-policy evaluation in reinforcement learning
Li, L. A perspective on off-policy evaluation in reinforcement learning. Front. Comput. Sci. 2019, 13, 911–912
2019
-
[36]
The continuum-armed bandit problem
Agrawal, R. The continuum-armed bandit problem. SIAM J. Control Optim. 1995, 33, 1926–1951
1995
-
[37]
Optimal Design of Experiments; SIAM, Philadelphia, USA, 2006
Pukelsheim, F. Optimal Design of Experiments; SIAM, Philadelphia, USA, 2006
2006
-
[38]
Bayesian Optimization; Cambridge University Press: Cambridge, UK, 2023
Garnett, R. Bayesian Optimization; Cambridge University Press: Cambridge, UK, 2023
2023
-
[39]
Gaussian processes for regression
Williams, C.; Rasmussen, C. Gaussian processes for regression. Adv. Neural Inf. Process. Syst. 1995, 8, 514–520
1995
-
[40]
On the mathematical foundations of theoretical statistics
Fisher, R.A. On the mathematical foundations of theoretical statistics. Philos. Trans. R. Soc. Lond. Ser. A Contain. Pap. Math. Phys. Character 1922, 222, 309–368
1922
-
[41]
Propriétés locales des fonctions à séries de Fourier aléatoires
Kahane, J.P . Propriétés locales des fonctions à séries de Fourier aléatoires. Stud. Math. 1960, 19, 1–25
1960
-
[42]
Sur un nouveau théoreme-limite de la théorie des probabilités
Cramér, H. Sur un nouveau théoreme-limite de la théorie des probabilités. Actual. Sci. Ind. 1938, 736, 5–23
1938
-
[43]
Probability: A Graduate Course, 2nd ed.; Springer: New York, USA, 2013
Gut, A. Probability: A Graduate Course, 2nd ed.; Springer: New York, USA, 2013
2013
-
[44]
Introduction to High-Dimensional Statistics, 2nd ed.; Chapman and Hall/CRC: Boca Raton, USA, 2021
Giraud, C. Introduction to High-Dimensional Statistics, 2nd ed.; Chapman and Hall/CRC: Boca Raton, USA, 2021
2021
-
[45]
Probability Inequalities for Sums of Bounded Random Variables
Hoeffding, W. Probability Inequalities for Sums of Bounded Random Variables. J. Am. Stat. Assoc. 1963, 58, 13–30
1963
-
[46]
Asymptotic minimax character of the sample distribution function and of the classical multinomial estimator
Dvoretzky, A.; Kiefer, J.; Wolfowitz, J. Asymptotic minimax character of the sample distribution function and of the classical multinomial estimator. Ann. Math. Stat. 1956, 27, 642–669
1956
-
[47]
Concentration inequalities for statistical inference
Zhang, H.; Chen, S.X. Concentration inequalities for statistical inference. Commun. Math. Res. 2021, 37, 1–85
2021
-
[48]
Statistics and Information Theory
Duchi, J. Statistics and Information Theory. 2024. Available online: https://web.stanford.edu/class/stats311/lecture-notes.pdf (accessed on 14th Feb., 2025)
2024
-
[49]
High-Dimensional Statistics: A Non-Asymptotic Viewpoint; Cambridge University Press: Cambridge, UK, 2019
Wainwright, M.J. High-Dimensional Statistics: A Non-Asymptotic Viewpoint; Cambridge University Press: Cambridge, UK, 2019
2019
-
[50]
Petrov, V .V .Limit Theorems of Probability Theory; Sequences of Independent Random Variable ; Oxford University Press: Oxford, UK, 1995
1995
-
[51]
Provably optimal algorithms for generalized linear contextual bandits
Li, L.; Lu, Y.; Zhou, D. Provably optimal algorithms for generalized linear contextual bandits. InProceedings of the 34th International Conference on Machine Learning, 06–11 Aug 2017; Volume 70, pp. 2071–2080
2017
-
[52]
Towards practical mean bounds for small samples
Phan, M.; Thomas, P .; Learned-Miller, E. Towards practical mean bounds for small samples. InProceedings of the 38th International Conference on Machine Learning, 18–24 Jul 2021; Volume 139, pp. 8567–8576
2021
-
[53]
Estimating means of bounded random variables by betting
Waudby-Smith, I.; Ramdas, A. Estimating means of bounded random variables by betting. J. R. Stat. Soc. Ser. B Stat. Methodol. 2024, 86, 1–27
2024
-
[54]
Policy optimization using semiparametric models for dynamic pricing
Fan, J.; Guo, Y.; Yu, M. Policy optimization using semiparametric models for dynamic pricing. J. Am. Stat. Assoc. 2024, 119, 552–564
2024
-
[55]
Sums of Independent Random Variables; Springer: Berlin, German, 1975
Petrov, V .V . Sums of Independent Random Variables; Springer: Berlin, German, 1975
1975
-
[56]
Computer Age Statistical Inference, Student Edition: Algorithms, Evidence, and Data Science; Cambridge University Press: Cambridge, UK, 2021
Efron, B.; Hastie, T. Computer Age Statistical Inference, Student Edition: Algorithms, Evidence, and Data Science; Cambridge University Press: Cambridge, UK, 2021. Version February 20, 2025 submitted to arxiv 50 of 52
2021
-
[57]
A Probabilistic Theory of Pattern Recognition; Springer: New York, USA, 1997
Devroye, L.; Györfi, L.; Lugosi, G. A Probabilistic Theory of Pattern Recognition; Springer: New York, USA, 1997
1997
-
[58]
Large deviation methods for approximate probabilistic inference
Kearns, M.; Saul, L. Large deviation methods for approximate probabilistic inference. In Proceedings of the Fourteenth Conference on Uncertainty in Artificial Intelligence, 1998; pp. 311–319. Morgan Kaufmann Publishers Inc.: San Francisco, CA, USA, 1998
1998
-
[59]
Some nonasymptotic results on resampling in high dimension, I: confidence regions
Arlot, S.; Blanchard, G.; Roquain, E. Some nonasymptotic results on resampling in high dimension, I: confidence regions. Ann. Stat. 2010, 38, 51–82
2010
-
[60]
Inference in a class of optimization problems: Confidence regions and finite sample bounds on errors in coverage probabilities
Horowitz, J.L.; Lee, S. Inference in a class of optimization problems: Confidence regions and finite sample bounds on errors in coverage probabilities. J. Bus. Econ. Stat. 2023, 41, 927–938
2023
-
[61]
Mathematical Statistics: A Non-Asymptotic Approach
Rakhlin, A. Mathematical Statistics: A Non-Asymptotic Approach. Lecture Notes. 2020. Available online: https://web.stanford. edu/class/stats311/lecture-notes.pdf (accessed on 14th Feb., 2025)
2020
-
[62]
Finite-time analysis of vector autoregressive models under linear restrictions
Zheng, Y.; Cheng, G. Finite-time analysis of vector autoregressive models under linear restrictions. Biometrika 2021, 108, 469–489
2021
-
[63]
Fast nonasymptotic testing and support recovery for large sparse Toeplitz covariance matrices
Bettache, N.; Butucea, C.; Sorba, M. Fast nonasymptotic testing and support recovery for large sparse Toeplitz covariance matrices. J. Multivar. Anal. 2021, 190, 104883
2021
-
[64]
Valid and Approximately Valid Confidence Intervals for Current Status Data.J
Kim, S.; Fay, M.P .; Proschan, M.A. Valid and Approximately Valid Confidence Intervals for Current Status Data.J. R. Stat. Soc. Ser. B (Stat. Methodol.) 2021, 83, 438–452
2021
-
[65]
Finite sample change point inference and identification for high-dimensional mean vectors
Yu, M.; Chen, X. Finite sample change point inference and identification for high-dimensional mean vectors. J. R. Stat. Soc. Ser. B (Stat. Methodol.) 2021, 83, 247–270
2021
-
[66]
What doubling tricks can and can’t do for multi-armed bandits
Besson, L.; Kaufmann, E. What doubling tricks can and can’t do for multi-armed bandits. arXiv 2018, arXiv:1803.06971
2018 arXiv
-
[67]
Adaptive treatment allocation and the multi-armed bandit problem
Lai, T.L. Adaptive treatment allocation and the multi-armed bandit problem. Ann. Stat. 1987, 15, 1091–1114
1987
-
[68]
On Lai’s Upper Confidence Bound in Multi-Armed Bandits
Ren, H.; Zhang, C.H. On Lai’s Upper Confidence Bound in Multi-Armed Bandits. arXiv 2024, arXiv:2410.02279
2024 arXiv
-
[69]
Introduction to Nonparametric Estimation; Springer: New York, USA, 2009
Alexandre, T. Introduction to Nonparametric Estimation; Springer: New York, USA, 2009
2009
-
[70]
Minimax policies for adversarial and stochastic bandits
Audibert, J.Y.; Bubeck, S. Minimax policies for adversarial and stochastic bandits. In Proceedings of the COLT, 2009; pp. 217–226
2009
-
[71]
A tutorial on thompson sampling
Russo, D.J.; Van Roy, B.; Kazerouni, A.; Osband, I.; Wen, Z. A tutorial on thompson sampling. Found. Trends® Mach. Learn. 2018, 11, 1–96
2018
-
[72]
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W.R. On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika 1933, 25, 285–294
1933
-
[73]
MOTS: Minimax Optimal Thompson Sampling
Jin, T.; Xu, P .; Shi, J.; Xiao, X.; Gu, Q. MOTS: Minimax Optimal Thompson Sampling. In Proceedings of the 38th International Conference on Machine Learning, 2021; pp. 5074–5083. PMLR: 2021
2021
-
[74]
A contextual-bandit approach to personalized news article recommendation
Li, L.; Chu, W.; Langford, J.; Schapire, R.E. A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th International Conference on World Wide Web (WWW ’10), 2010; pp. 661–670
2010
-
[75]
Mathematical Analysis of Machine Learning Algorithms; Cambridge University Press: Cambridge, UK, 2023
Zhang, T. Mathematical Analysis of Machine Learning Algorithms; Cambridge University Press: Cambridge, UK, 2023
2023
-
[76]
Stochastic Linear Optimization under Bandit Feedback
Dani, V .; Hayes, T.P .; Kakade, S.M. Stochastic Linear Optimization under Bandit Feedback. In Proceedings of the COLT, 2008; Volume 2, p. 3
2008
-
[77]
Thompson sampling for contextual bandits with linear payoffs
Agrawal, S.; Goyal, N. Thompson sampling for contextual bandits with linear payoffs. In Proceedings of the 30th International Conference on Machine Learning, 2013; pp. 127–135
2013
-
[78]
Learning to optimize via posterior sampling
Russo, D.; Van Roy, B. Learning to optimize via posterior sampling. Math. Oper. Res. 2014, 39, 1221–1243
2014
-
[79]
Neural contextual bandits with UCB-based exploration
Zhou, D.; Li, L.; Gu, Q. Neural contextual bandits with UCB-based exploration. In Proceedings of the 37th International Conference on Machine Learning, 2020; pp. 11492–11502
2020
-
[80]
Neural Thompson Sampling
Zhang, W.; Zhou, D.; Li, L.; Gu, Q. Neural Thompson Sampling. In Proceedings of the International Conference on Learning Representation (ICLR), 2021
2021
-
[81]
Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design
Srinivas, N.; Krause, A.; Kakade, S.; Seeger, M. Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design. In Proceedings of the 27th International Conference on Machine Learning; Omnipress: 2010; pp. 1015–1022
2010
-
[82]
A tutorial on Bayesian optimization of expensive cost functions, with application to active user modeling and hierarchical reinforcement learning
Brochu, E.; Cora, V .M.; De Freitas, N. A tutorial on Bayesian optimization of expensive cost functions, with application to active user modeling and hierarchical reinforcement learning. arXiv 2010, arXiv:1012.2599
2010 arXiv
-
[83]
A tutorial on Gaussian process regression: Modelling, exploring, and exploiting functions
Schulz, E.; Speekenbrink, M.; Krause, A. A tutorial on Gaussian process regression: Modelling, exploring, and exploiting functions. J. Math. Psychol. 2018, 85, 1–16
2018
-
[84]
Elements of Information Theory; Wiley-Interscience, Hoboken, USA, 2006
Thomas, M.; Joy, A.T. Elements of Information Theory; Wiley-Interscience, Hoboken, USA, 2006
2006
-
[85]
On kernelized multi-armed bandits
Chowdhury, S.R.; Gopalan, A. On kernelized multi-armed bandits. In Proceedings of the International Conference on Machine Learning; PMLR: 2017; pp. 844–853
2017
-
[86]
Contextual Gaussian Process Bandit Optimization
Krause, A.; Ong, C. Contextual Gaussian Process Bandit Optimization. Advances in Neural Information Processing Systems 2011, 24, 2447–2455
2011
-
[87]
Active learning for level set estimation
Gotovos, A.; Casati, N.; Hitz, G.; Krause, A. Active learning for level set estimation. In Proceedings of the Twenty-Third International Joint Conference on Artificial Intelligence, 2013; pp. 1344–1350
2013
-
[88]
Gaussian process classification bandits
Hayashi, T.; Ito, N.; Tabata, K.; Nakamura, A.; Fujita, K.; Harada, Y.; Komatsuzaki, T. Gaussian process classification bandits. Pattern Recognit. 2024, 149, 110224
2024
-
[89]
Gaussian Process Bandits Preprint
Mathieu, E. Gaussian Process Bandits Preprint. 2016. Available online: https://emilemathieu.fr/files/gpbanditsreport.pdf (accessed on 14th Feb., 2025)
2016
-
[90]
Uncertainty Quantification Using Martingales for Misspecified Gaussian Processes
Neiswanger, W.; Ramdas, A. Uncertainty Quantification Using Martingales for Misspecified Gaussian Processes. Proceedings of the 32nd International Conference on Algorithmic Learning Theory 2021, 132, 963–982. Version February 20, 2025 submitted to arxiv 51 of 52
2021
-
[91]
Confidence Bound Minimization for Bayesian Optimization with Student’s-t Processes
Clare, C.; Hawe, G.; Lin, Z.; McClean, S. Confidence Bound Minimization for Bayesian Optimization with Student’s-t Processes. Proceedings of the 3rd International Conference on Applications of Intelligent Systems 2020, Article 9, 1–5
2020
-
[92]
Gaussian Processes for Machine Learning; MIT Press: Cambridge, MA, USA, 2006; Volume 2
Williams, C.K.; Rasmussen, C.E. Gaussian Processes for Machine Learning; MIT Press: Cambridge, MA, USA, 2006; Volume 2
2006
-
[93]
Continuum-Armed Bandits: A Function Space Perspective
Singh, S. Continuum-Armed Bandits: A Function Space Perspective. Proceedings of The 24th International Conference on Artificial Intelligence and Statistics 2021, 130, 2620–2628
2021
-
[94]
Theoretical Foundations of Functional Data Analysis, with an Introduction to Linear Operators; John Wiley & Sons, West Sussex, 2015
Hsing, T.; Eubank, R. Theoretical Foundations of Functional Data Analysis, with an Introduction to Linear Operators; John Wiley & Sons, West Sussex, 2015
2015
-
[95]
Gaussian Process Regression Analysis for Functional Data; CRC Press: 2011
Shi, J.Q.; Choi, T. Gaussian Process Regression Analysis for Functional Data; CRC Press: 2011
2011
-
[96]
Growing-dimensional partially functional linear models: non-asymptotic optimal prediction error
Zhang, H.; Lei, X. Growing-dimensional partially functional linear models: non-asymptotic optimal prediction error. Phys. Scr. 2023, 98, 095216
2023
-
[97]
Functional data analysis for sparse longitudinal data
Yao, F.; Müller, H.G.; Wang, J.L. Functional data analysis for sparse longitudinal data. J. Am. Stat. Assoc. 2005, 100, 577–590
2005
-
[98]
Functional linear regression for discretely observed data: From ideal to reality
Zhou, H.; Yao, F.; Zhang, H. Functional linear regression for discretely observed data: From ideal to reality. Biometrika 2023, 110, 381–393
2023
-
[99]
Stochastic continuum-armed bandits with additive models: Minimax regrets and adaptive algorithm
Cai, T.T.; Pu, H. Stochastic continuum-armed bandits with additive models: Minimax regrets and adaptive algorithm. Ann. Stat. 2022, 50, 2179–2204
2022
-
[100]
Optimal designs for longitudinal and functional data
Ji, H.; Müller, H.G. Optimal designs for longitudinal and functional data. J. R. Stat. Soc. Ser. B Stat. Methodol. 2017, 79, 859–876
2017
-
[101]
One-armed bandit problems with covariates
Sarkar, J. One-armed bandit problems with covariates. Ann. Stat. 1991, 19, 1978–2002
1991
-
[102]
Randomized allocation with nonparametric estimation for a multi-armed bandit problem with covariates
Yang, Y.; Zhu, D. Randomized allocation with nonparametric estimation for a multi-armed bandit problem with covariates. Ann. Stat. 2002, 30, 100–121
2002
-
[103]
The multi-armed bandit problem with covariates
Perchet, V .; Rigollet, P . The multi-armed bandit problem with covariates. Ann. Stat. 2013, 41, 693–721
2013
-
[104]
Sub-sampling for multi-armed bandits
Baransi, A.; Maillard, O.A.; Mannor, S. Sub-sampling for multi-armed bandits. In Proceedings of the Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2014, Nancy, France, 15–19 September 2014; Proceedings, Part I 14; Springer: 2014; pp. 115–131
2014
-
[105]
The Multi-armrd bandit Problem
Chan, H.P . The Multi-armrd bandit Problem. Ann. Stat. 2020, 48, 346–373
2020
-
[106]
A non-parametric solution to the multi-armed bandit problem with covariates
Ai, M.; Huang, Y.; Yu, J. A non-parametric solution to the multi-armed bandit problem with covariates. J. Stat. Plan. Inference 2021, 211, 402–413
2021
-
[107]
Transfer learning for contextual multi-armed bandits
Cai, C.; Cai, T.T.; Li, H. Transfer learning for contextual multi-armed bandits. Ann. Stat. 2024, 52, 207–232
2024
-
[108]
Adaptive algorithm for multi-armed bandit problem with high-dimensional covariates
Qian, W.; Ing, C.K.; Liu, J. Adaptive algorithm for multi-armed bandit problem with high-dimensional covariates. J. Am. Stat. Assoc. 2024, 119, 970–982
2024
-
[110]
Multi-armed bandit for species discovery: A Bayesian nonparametric approach
Battiston, M.; Favaro, S.; Teh, Y.W. Multi-armed bandit for species discovery: A Bayesian nonparametric approach. J. Am. Stat. Assoc. 2018, 113, 455–466
2018
-
[111]
Statistical inference for online decision making: In a contextual bandit setting
Chen, H.; Lu, W.; Song, R. Statistical inference for online decision making: In a contextual bandit setting. J. Am. Stat. Assoc. 2021, 116, 240–255
2021
-
[112]
Stochastic low-rank tensor bandits for multi-dimensional online decision making
Zhou, J.; Hao, B.; Wen, Z.; Zhang, J.; Sun, W.W. Stochastic low-rank tensor bandits for multi-dimensional online decision making. J. Am. Stat. Assoc. 2024, 1–14. https://doi.org/10.1080/01621459.2024.2311364
2024
-
[113]
Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons
Zhu, B.; Jordan, M.; Jiao, J. Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons. Proceedings of the 40th International Conference on Machine Learning 2023, 202, 43037–43067
2023
-
[114]
Optimal Design for Reward Modeling in RLHF
Scheid, A.; Boursier, E.; Durmus, A.; Jordan, M.I.; Ménard, P .; Moulines, E.; Valko, M. Optimal Design for Reward Modeling in RLHF. arXiv 2024, arXiv:2410.17055
2024 arXiv
-
[115]
A review of off-policy evaluation in reinforcement learning
Uehara, M.; Shi, C.; Kallus, N. A review of off-policy evaluation in reinforcement learning. arXiv 2022, arXiv:2212.06355
2022 arXiv
-
[116]
Bandit problems with infinitely many arms
Berry, D.A.; Chen, R.W.; Zame, A.; Heath, D.C.; Shepp, L.A. Bandit problems with infinitely many arms. Ann. Stat. 1997, 25, 2103–2116
1997
-
[117]
Bayesian nonparametric bandits
Clayton, M.K.; Berry, D.A. Bayesian nonparametric bandits. Ann. Stat. 1985, 13, 1523–1534
1985
-
[118]
Multi-armed bandits with discount factor near one: The Bernoulli case
Kelly, F. Multi-armed bandits with discount factor near one: The Bernoulli case. Ann. Stat. 1981, 9, 987–1001
1981
-
[119]
Bandit processes and dynamic allocation indices
Gittins, J.C. Bandit processes and dynamic allocation indices. J. R. Stat. Soc. Ser. B Stat. Methodol. 1979, 41, 148–164
1979
-
[120]
Multi-armed bandits and the Gittins index
Whittle, P . Multi-armed bandits and the Gittins index. J. R. Stat. Soc. Ser. B (Methodol.) 1980, 42, 143–149
1980
-
[121]
On randomized dynamic allocation indices for the sequential design of experiments
Glazebrook, K. On randomized dynamic allocation indices for the sequential design of experiments. J. R. Stat. Soc. Ser. B Stat. Methodol. 1980, 42, 342–346
1980
-
[122]
Asymptotically efficient strategies for a stochastic scheduling problem with order constraints
Fuh, C.D.; Hu, I. Asymptotically efficient strategies for a stochastic scheduling problem with order constraints. Ann. Stat. 2000, 28, 1670–1695
2000
-
[123]
The learning component of dynamic allocation indices
Gittins, J.; Wang, Y.G. The learning component of dynamic allocation indices. Ann. Stat. 1992, 20, 1625–1636
1992
-
[124]
Strategic two-sample test via the two-armed bandit process
Chen, Z.; Yan, X.; Zhang, G. Strategic two-sample test via the two-armed bandit process. J. R. Stat. Soc. Ser. B Stat. Methodol. 2023, 85, 1271–1298
2023
-
[125]
UCB algorithms for multi-armed bandits: Precise regret and adaptive inference
Han, Q.; Khamaru, K.; Zhang, C.H. UCB algorithms for multi-armed bandits: Precise regret and adaptive inference. arXiv 2024, arXiv:2412.06126
2024 arXiv
-
[126]
Finite horizon behavior of policies for two-arm bandits
Fox, B.L. Finite horizon behavior of policies for two-arm bandits. J. Am. Stat. Assoc. 1974, 69, 963–965. Version February 20, 2025 submitted to arxiv 52 of 52
1974
-
[127]
Asymptotically efficient allocation rules for two Bernoulli populations
Li, Z.; Zhang, C.H. Asymptotically efficient allocation rules for two Bernoulli populations. J. R. Stat. Soc. Ser. B Stat. Methodol. 1992, 54, 609–616
1992
-
[128]
Kullback-Leibler upper confidence bounds for optimal sequential allocation
Cappé, O.; Garivier, A.; Maillard, O.A.; Munos, R.; Stoltz, G. Kullback-Leibler upper confidence bounds for optimal sequential allocation. Ann. Stat. 2013, 41, 1516–1541
2013
-
[129]
On Bayesian index policies for sequential resource allocation
Kaufmann, E. On Bayesian index policies for sequential resource allocation. Ann. Stat. 2018, 46, 842–865
2018
-
[130]
Response surface bandits
Ginebra, J.; Clayton, M.K. Response surface bandits. J. R. Stat. Soc. Ser. B Stat. Methodol. 1995, 57, 771–784
1995
-
[131]
Pre-trained Gaussian processes for Bayesian optimization
Wang, Z.; Dahl, G.E.; Swersky, K.; Lee, C.; Nado, Z.; Gilmer, J.; Snoek, J.; Ghahramani, Z. Pre-trained Gaussian processes for Bayesian optimization. J. Mach. Learn. Res. 2024, 25, 1–83
2024
-
[132]
Modified two-armed bandit strategies for certain clinical trials
Berry, D.A. Modified two-armed bandit strategies for certain clinical trials. J. Am. Stat. Assoc. 1978, 73, 339–345
1978
-
[133]
Randomized allocation of treatments in sequential experiments
Bather, J. Randomized allocation of treatments in sequential experiments. J. R. Stat. Soc. Ser. B (Methodol.) 1981, 43, 265–283
1981
-
[134]
Optimal adaptive randomized designs for clinical trials
Cheng, Y.; Berry, D.A. Optimal adaptive randomized designs for clinical trials. Biometrika 2007, 94, 673–689
2007
-
[135]
An information theoretic approach for selecting arms in clinical trials
Mozgunov, P .; Jaki, T. An information theoretic approach for selecting arms in clinical trials. J. R. Stat. Soc. Ser. B Stat. Methodol. 2020, 82, 1223–1247
2020
-
[136]
False discovery rate control with e-values
Wang, R.; Ramdas, A. False discovery rate control with e-values. J. R. Stat. Soc. Ser. B Stat. Methodol. 2022, 84, 822–852
2022
-
[137]
Variable selection via Thompson sampling.J
Liu, Y.; Roˇ cková, V . Variable selection via Thompson sampling.J. Am. Stat. Assoc. 2023, 118, 287–304
2023
-
[138]
Dynamic online pricing with incomplete information using multiarmed bandit experiments
Misra, K.; Schwartz, E.M.; Abernethy, J. Dynamic online pricing with incomplete information using multiarmed bandit experiments. Mark. Sci. 2019, 38, 226–252
2019
-
[139]
Effective Adaptive Exploration of Prices and Promotions in Choice-Based Demand Models
Jain, L.; Li, Z.; Loghmani, E.; Mason, B.; Yoganarasimhan, H. Effective Adaptive Exploration of Prices and Promotions in Choice-Based Demand Models. Mark. Sci. 2024, 43, 925–1151
2024
-
[140]
Bandit Interpretability of Deep Models via Confidence Selection
Duan, X.; Li, H.; Wang, P .; Wang, T.; Liu, B.; Zhang, B. Bandit Interpretability of Deep Models via Confidence Selection. Neurocomputing 2023, 544, 126250
2023
-
[141]
Replication or exploration? Sequential design for stochastic simulation experiments
Binois, M.; Huang, J.; Gramacy, R.B.; Ludkovski, M. Replication or exploration? Sequential design for stochastic simulation experiments. Technometrics 2019, 61, 7–23
2019
-
[142]
Risk-averse heteroscedastic bayesian optimization
Makarova, A.; Usmanova, I.; Bogunovic, I.; Krause, A. Risk-averse heteroscedastic bayesian optimization. Adv. Neural Inf. Process. Syst. 2021, 34, 17235–17245
2021
-
[143]
Digital Triplet: A Sequential Methodology for Digital Twin Learning
Zhang, X.; Lin, D.K.; Wang, L. Digital Triplet: A Sequential Methodology for Digital Twin Learning. Mathematics 2023, 11, 2661
2023
-
[144]
Towards a digital twin framework in additive manufacturing: Machine learning and bayesian optimization for time series process optimization
Karkaria, V .; Goeckner, A.; Zha, R.; Chen, J.; Zhang, J.; Zhu, Q.; Cao, J.; Gao, R.X.; Chen, W. Towards a digital twin framework in additive manufacturing: Machine learning and bayesian optimization for time series process optimization. J. Manuf. Syst. 2024, 75, 322–332
2024
-
[145]
A survey on contextual multi-armed bandits
Zhou, L. A survey on contextual multi-armed bandits. arXiv 2015, arXiv:1508.03326
2015 arXiv
-
[146]
Conservative Bandits
Wu, Y.; Shariff, R.; Lattimore, T.; Szepesvári, C. Conservative Bandits. Proceedings of the 33rd International Conference on Machine Learning 2016, 48, 1254–1262
2016
-
[147]
Residual Bootstrap Exploration for Stochastic Linear Bandit
Wu, S.; Wang, C.H.; Li, Y.; Cheng, G. Residual Bootstrap Exploration for Stochastic Linear Bandit. Proceedings of the Thirty-Eighth Conference on Uncertainty in Artificial Intelligence 2022, 180, 2117–2127
2022
-
[148]
Estimating concentration parameters for bandit algorithms
Lieber, J. Estimating concentration parameters for bandit algorithms. Job Market Paper 2022. Available online: https://jonaslieber. com/research.html (accessed on 14th Feb., 2025)
2022
-
[149]
Multiplier Bootstrap-Based Exploration
Wan, R.; Wei, H.; Kveton, B.; Song, R. Multiplier Bootstrap-Based Exploration. Proceedings of the 40th International Conference on Machine Learning 2023, 1476, 35444–35490
2023
-
[150]
Perturbed-history exploration in stochastic multi-armed bandits
Kveton, B.; Szepesvári, C.; Ghavamzadeh, M.; Boutilier, C. Perturbed-history exploration in stochastic multi-armed bandits. In Proceedings of the 28th International Joint Conference on Artificial Intelligence, 2019; pp. 2786–2793
2019
-
[151]
Perturbed-History Exploration in Stochastic Linear Bandits.Proceedings of the 35th Uncertainty in Artificial Intelligence Conference 2020, 530–540
Kveton, B.; Szepesvári, C.; Ghavamzadeh, M.; Boutilier, C. Perturbed-History Exploration in Stochastic Linear Bandits.Proceedings of the 35th Uncertainty in Artificial Intelligence Conference 2020, 530–540
2020
-
[152]
Universal and data-adaptive algorithms for model selection in linear contextual bandits
Muthukumar, V .K.; Krishnamurthy, A. Universal and data-adaptive algorithms for model selection in linear contextual bandits. Proceedings of the 39th International Conference on Machine Learning 2022, 16197–16222
2022
-
[153]
Model Selection for Contextual Bandits and Reinforcement Learning
Pacchiano Camacho, A. Model Selection for Contextual Bandits and Reinforcement Learning. Ph.D. Thesis, UC Berkeley, Berkeley, CA, USA, 2021. Under review
2021
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.