Pith. sign in

REVIEW 7 minor 299 references

Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation

T0 review · 0 major / 7 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read This paper claims that cost-aware multi-objective bandits can achieve the same logarithmic regret as single-objective budgeted bandits, using a hypervolume-per-cost index and cost-aware Pareto gap elimination.

desk verdict A solid, honestly-written extension of budgeted bandits to multi-objective evaluation; the flagged Hoeffding concern is not a real flaw, and the zero-gap limitation is assumed away transparently. read the letter →

arxiv 2608.04333 v1 pith:QCPXTMVO submitted 2026-08-05 cs.LG

classification cs.LG MSC 62L0568W2790C29
keywords cost-awaremulti-armedbanditsmulti-objectivehypervolumeregretParetosetidentificationLLMconfigurationevaluationbudgetedempiricalgapelimination
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to establish that LLM configuration evaluation can be treated as a cost-aware multi-objective bandit problem, and that both online selection and fixed-budget Pareto identification admit algorithms with strong guarantees despite configuration-dependent evaluation costs. For online selection, it proposes CoHV-UCB, which pulls arms by an optimistic hypervolume-per-cost index, and proves hypervolume-efficiency regret of order $O\left(\sum_{i\neq i^*}\frac{\log B}{\Delta_i}\right)$, retaining the logarithmic budget dependence of single-objective budgeted bandits. For fixed-budget identification, it proposes CoPSI, a cost-aware gap-elimination algorithm whose misidentification probability decays exponentially with the budget at a rate governed by a cost-aware complexity $H_{\mu,c}$. The framework also connects to concrete LLM configuration choices, where accuracy and efficiency are competing objectives and token costs vary across configurations.

What carries the argument

The central machinery is the hypervolume-efficiency index for online selection and cost-aware empirical gap elimination for identification. Hypervolume efficiency $\nu_i = H(\mu_i)/\mu^c_i$ converts a vector objective and a scalar cost into a single per-cost utility; CoHV-UCB's optimism step constructs an upper confidence bound on the hypervolume and a lower confidence bound on cost to form an optimistic efficiency index, which is what forces the optimal arm to remain competitive and limits suboptimal pulls to $O(\log B/\Delta_i^2)$. For Pareto identification, CoPSI uses empirical pairwise margins $\hat m_r(i,j)$ and $\hat M_r(i,j)$ to define a classification gap $\hat\gamma_{i,r}$ for each active arm, and its cost-aware sampling target $n_r$ depends on the aggregate cost $C_r$ of the active set; this couples the statistical gap $\gamma_{(k)}$ to the budget via $H_{\mu,c} = \max_A C(A)/\gamma^2_{(|A|)}$ and makes the exponential error decay follow from Hoeffding concentration at adaptive targets.

What would settle it

Simulate CoPSI on a valid instance with known strictly positive gaps, for example three arms in two objectives with gaps 0.001 and 0.002 and costs differing by a factor of four, and measure the log misidentification probability over many repetitions at the budget values specified by the theorem; if the empirical error curve does not lie below $2K^2D\exp(-B/(256L_{K,\lambda}H_{\mu,c}))$ at those budgets, the claimed exponential rate fails. Alternatively, run CoPSI on an instance with two arms sharing exactly the same mean vector, where $\gamma_i = 0$; the paper's own assumption excludes this, and the error probability will remain bounded away from zero no matter the budget, confirming that the guarantee depends on the strictly-positive-gap condition.

Watch

Extended reading notes

Core claim

The central claim is that cost awareness and multi-objectivity can be combined without sacrificing classical bandit efficiency. CoHV-UCB defines hypervolume efficiency $\nu_i = H(\mu_i)/\mu^c_i$ as the value of arm $i$ per unit expected cost, builds an optimistic index $H(U_{i,t})/c_{i,t}$ with UCB-adjusted reward coordinates and a lower confidence bound on cost, and shows that the expected number of pulls of any suboptimal arm is $O(\log B / \Delta_i^2)$, yielding regret $O\left(\sum_{i\neq i^*} \frac{\log B}{\Delta_i}\right)$. CoPSI instead allocates samples with a phase-dependent target $n_r = \lfloor B/(L_{K,\lambda} C_r)\rfloor$ that depends on the aggregate cost $C_r$ of the active configuration set, and removes the active arm with the largest empirical classification gap; its error probability is bounded by $2K^2D\exp(-B/(256L_{K,\lambda}H_{\mu,c}))$, where $H_{\mu,c}$ is the worst-case ratio of active-set cost to squared classification gap. Together these theorems assert that configuration-dependent evaluation costs can be handled with the same logarithmic regret and exponential-error decay as cost-blind or cost-unaware settings, and the experiments on GSM8K and PIQA support that the proposed indices and elimination rules save budget in practice.

Load-bearing premise

The fixed-budget guarantee for CoPSI requires every configuration to have a strictly positive classification gap $\gamma_i > 0$, so that exact Pareto identification is statistically well posed; if two arms have identical or arbitrarily close mean reward vectors, the exponential error bound no longer applies, and the experiments resort to a curated 14-arm subset or an approximate F1 score.

Editorial extensions

If this is right

  • CoHV-UCB gives an online policy for LLM configuration selection that balances accuracy and efficiency, with a regret bound that grows only logarithmically in the evaluation budget, so efficiency regret diminishes quickly as more tokens are spent.
  • CoPSI gives a principled cost-aware way to decide which configurations to keep evaluating under a fixed budget, and its error probability decays exponentially with the budget at a rate set by a cost-aware complexity measure.
  • When all evaluation costs are identical, CoPSI's guarantee reduces to the standard fixed-budget Pareto set identification rate, so the new algorithm is a strict generalization of existing methods.
  • The standard cumulative hypervolume regret also obeys the same $O(\sum_{i\neq i^*} \log B/\Delta_i)$ bound, so the cost-aware index does not sacrifice the usual multi-objective quality measure.
  • The complexity $H_{\mu,c}$ makes the cost of identifying a Pareto front concrete: configurations that are cheap but hard to classify can receive more samples without exceeding the total budget.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The hypervolume-per-cost construction could transfer to other resource-constrained multi-objective choice problems, such as prompt routing or model selection in production, where cost is monetary or latency; the paper does not explore this transfer.
  • Because the fixed-budget theorem assumes known deterministic costs, an untested natural extension is to unknown or random evaluation costs in the identification setting; the paper's stochastic-cost robustness experiment gives preliminary evidence but no theorem.
  • The near-tied empirical means in the full datasets suggest that exact Pareto set identity is often a knife-edge notion in real LLM evaluation, and an approximate-recovery criterion such as Pareto F1 may be the more operational target; a theory of approximate cost-aware Pareto identification would be a direct follow-up.
  • The regret bound's dependence on the efficiency gap $\Delta_i$ implies that near-tied configurations dominate the guarantee, so a practical prefilter that removes nearly equivalent arms could improve allocation, though the paper does not propose such a prefilter.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 7 minor

Summary. The paper introduces a cost-aware multi-objective bandit formulation for LLM configuration evaluation, where each arm pull yields a vector-valued reward in [0,1]^D and incurs an arm-dependent cost. For online configuration selection, it proposes CoHV-UCB, which selects the arm maximizing an optimistic product-hypervolume divided by a lower confidence bound on cost, and proves an O(Σ_{i≠i*} C^2_{λ,H} log B / Δ_i) bound on hypervolume-efficiency regret (Theorem 1), together with a standard-hypervolume regret corollary (Corollary 1). For fixed-budget Pareto identification, it proposes CoPSI, a successive-elimination algorithm with cost-aware phase targets, and proves an error probability of order exp(-B/(256 L_{K,λ} H_{μ,c})) (Theorem 2). Experiments on GSM8K and PIQA configuration sets compare the methods against cost-insensitive and single-objective baselines; full-set experiments use F1 because of near-tied means, while exact recovery is demonstrated on a curated 14-arm subset.

Significance. If the proofs are correct, this is the first theoretical treatment of cost-aware multi-objective bandits, and the two bounds match the best-known single-objective budgeted rates in their budget and gap dependence. The proofs are standard but careful: Theorem 1 is a clean UCB analysis with a clipped cost LCB and a Lipschitz argument for the product hypervolume; Theorem 2 adapts EGE-SR with a budget-feasible sampling schedule and a time-uniform Hoeffding bound. I checked the time-uniform step flagged in the review: the application to the data-dependent target n_r is legitimate because n_r is deterministically lower-bounded by B/(2 L_{K,λ} H_{μ,c} γ^2_{k_r}), so the random-time deviation is a subset of the deterministic supremum event. The paper is unusually transparent about the main limitation of the Pareto-identification theorem: exact recovery requires strictly positive classification gaps γ_i>0, and the full LLM datasets contain near-tied means, so exact recovery is only demonstrated on a curated 14-arm subset while full sets are evaluated with F1.

minor comments (7)
  1. [§3.2 and Appendix D] The definition of Δ_i^+ contains a stray double plus sign in the expression 'M(j,i) + + (Δ_j^-)^+'; it should read M(j,i)^+ + (Δ_j^-)^+, and the same typo appears in the proof of Lemma 2.
  2. [§4.1 and Theorem 1] The relationship among T(B), the while-loop in Algorithm 1, and the possible boundary pull is not stated precisely: as written, Algorithm 1 can perform one evaluation whose cost pushes the cumulative cost above B, whereas R_HV^B sums only over t≤T(B). Please clarify whether n_i(B) counts the boundary pull and state the convention in the theorem.
  3. [Appendix A] The experimental confidence radii use profiling-calibrated scale factors (s_r=s_c=0.01) that are much smaller than the uncalibrated radius required by Theorem 1 (α≥2); please add a sentence noting that the experiments use calibrated radii and that the theoretical guarantee applies to the uncalibrated choice.
  4. [Figures] Several figure captions and axis labels appear as Unicode/encoding artifacts in the manuscript text, with sequences such as '/uni00000013/...' replacing readable labels; these need to be regenerated with a proper text encoding.
  5. [§5.1] The main text should state explicitly that the reward vector has D=2 coordinates (accuracy and normalized inverse latency); this detail currently appears only in the appendix.
  6. [Abstract and §1] The claim of matching single-objective budgeted bandits 'in budget and gap dependence' should be qualified as matching up to the problem-dependent constant C_{λ,H}, since the dependence on λ and D enters the stated bound.
  7. [Conclusion] The conclusion repeats the future-work item about small or unreliable gaps; given that the full datasets contain near-tied means, this limitation could also be stated in the abstract or introduction for a reader who only sees the headline guarantees.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: both main theorems are proven from stated assumptions with self-contained concentration and elimination arguments.

full rationale

The paper's central claims, Theorem 1 and Theorem 2, are derived rather than assumed. The proof of Theorem 1 (Appendix B) builds a uniform Hoeffding concentration event, shows optimism of the optimal arm, upper-bounds the index of each suboptimal arm by ν_i + C_{λ,H} β_{i,t}, and converts this into a pull-count bound; the regret bound then follows from the definition of hypervolume-efficiency regret. Corollary 1 is a separate regret decomposition using Lemma 1's benchmark upper bound, not a restatement of Theorem 1. The proof of Theorem 2 (Appendix D) contains a deterministic correctness lemma for the EGE-style elimination rule (Lemma 2) and a cost-aware sampling calculation showing budget feasibility and the phase-wise lower bound n_r γ²_(k_r) ≥ B/(2 L_{K,λ} H_{μ,c}); the exponential error bound follows from Ville's inequality for time-uniform Hoeffding supermartingales. No parameter is fitted to the data whose prediction is then reported: the regret and error bounds are expressed in terms of instance-dependent gaps and costs, and the experimental confidence radii (s_r, s_c, η) are selected on a disjoint 200-instance profiling split. The only self-citations (Qiu et al. 2024; Hong et al. 2026) appear in related-work discussion and are not load-bearing for either theorem. The paper also transparently states the zero-gap limitation and handles it experimentally with a curated 14-arm subset and F1 on full sets. Consequently, no step in the derivation reduces by definition or by self-citation to its own inputs.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central claims rest on standard stochastic bandit assumptions plus a strict positive-gap condition and a specific product scalarization. The only fitted quantities are experimental confidence radii chosen on a holdout profiling split; the theory itself introduces no fitted constants beyond standard exploration constants.

free parameters (3)
  • s_r = 0.01
    Reward confidence radius scale in CoHV-UCB experiments, selected by grid search on a 200-instance profiling split; affects reported regret numbers.
  • s_c = 0.01
    Cost confidence radius scale in CoHV-UCB experiments, selected by grid search on the same profiling split.
  • eta = 0.05
    Initial sample fraction in CoHV-UCB experiments, selected by grid search on the profiling split.
assumptions (5)
  • domain assumption Rewards and costs are i.i.d. with bounded supports [0,1] and [lambda,1]
    Assumed in Section 3; the entire concentration analysis relies on this.
  • domain assumption All Pareto classification gaps are strictly positive (gamma_i > 0)
    Assumed in Section 3.2; exact identification is not well-posed with zero gaps, and the full datasets contain near-ties.
  • domain assumption Costs are deterministic and known in the Pareto identification setting
    Section 3 states this cost model for CoPSI; the robustness experiment relaxes it but without a matching guarantee.
  • standard math Hoeffding and Ville inequalities for bounded random variables
    Used throughout Appendices B and D for concentration of empirical means.
  • ad hoc to paper Product hypervolume is a valid scalarization of multi-objective preference
    The paper defines H_i as the product of mean objectives and builds the whole online regret around it; this is a modeling choice, not forced by the multi-objective problem.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation." pith.science (2026). https://pith.science/paper/QCPXTMVO

@misc{pith2026260804333,
  author       = {Pith},
  title        = {Pith review of: Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QCPXTMVO}},
  note         = {Machine review of arXiv:2608.04333}
}
abstract

Large language model (LLM) configuration evaluation is challenging due to limited evaluation budgets, varying costs, and multiple competing objectives. In this paper, we formulate LLM configuration evaluation as a cost-aware multi-objective bandit problem, where each configuration evaluation incurs a configuration-dependent cost and yields a noisy vector-valued outcome. Under this framework, we study two fundamental problems: online configuration selection and Pareto configuration identification. For online configuration selection, we propose a hypervolume-based UCB algorithm that optimizes an optimistic hypervolume-per-cost index. We establish a budgeted regret bound of order $O\bigl(\sum_{i\ne i^\star}\frac{\log B}{\Delta_i}\bigr)$, where $B$ is the evaluation budget, $i^\star$ is the optimal configuration in terms of hypervolume efficiency, and $\Delta_i$ is the corresponding efficiency gap of configuration $i$. This bound retains the logarithmic budget dependence of classical single-objective budgeted bandits. For fixed-budget Pareto identification, we develop a cost-aware empirical gap elimination algorithm and prove that its error probability is of order $O\bigl(\exp(-\frac{B}{H_{\mu,c}})\bigr)$, where $H_{\mu,c}$ is a cost-aware Pareto identification complexity depending on configuration costs and Pareto classification gaps. This error probability decays exponentially with the evaluation budget and recovers the standard Pareto set identification guarantee when all configuration costs are identical. Experiments on LLM configuration evaluation tasks demonstrate that the proposed framework enables efficient online decision-making and accurate cost-aware Pareto identification under limited budgets.

Figures

Figures reproduced from arXiv: 2608.04333 by the authors.

Figure 1
Figure 1. Configuration landscapes on GSM8K and PIQA. [PITH_FULL_IMAGE:figures/full_fig_p011_1.png] view at source ↗
Figure 2
Figure 2. Main experimental results. Panels (a)–(b) show hypervolume-efficiency regret over 100 [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Standard hypervolume regret versus actual token budget. [PITH_FULL_IMAGE:figures/full_fig_p019_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Online decision trajectories at the fixed budget [PITH_FULL_IMAGE:figures/full_fig_p020_4.png]
Figure 5
Figure 5. Figure 5: Pareto F1 on all 108/132 configurations, where ties make exact identification uninformative. [PITH_FULL_IMAGE:figures/full_fig_p021_5.png]
Figure 6
Figure 6. Figure 6: Pareto identification error at ρ = 5000 under deterministic and stochastic token costs over 500 paired repetitions. Values above the bars are empirical error probabilities. its number of complete sampling rounds is determined directly by the realized costs. Under stoch…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

299 extracted references · 74 canonical work pages

  1. [1]

    Proceedings of the 29th International Joint Conference on Artificial Intelligence , pages =

    Nearly Optimal Regret for Stochastic Linear Bandits with Heavy-Tailed Payoffs , author =. Proceedings of the 29th International Joint Conference on Artificial Intelligence , pages =

  2. [2]

    arXiv preprint arXiv:2407.17466 , year=

    Traversing pareto optimal policies: Provably efficient multi-objective reinforcement learning , author=. arXiv preprint arXiv:2407.17466 , year=

  3. [3]

    Efficient Algorithms for Generalized Linear Bandits with Heavy-tailed Rewards , year =

    Xue, Bo and Wang, Yimu and Wan, Yuanyu and Yi, Jinfeng and Zhang, Lijun , booktitle =. Efficient Algorithms for Generalized Linear Bandits with Heavy-tailed Rewards , year =

  4. [4]

    Multiobjective Lipschitz Bandits under Lexicographic Ordering , journal=

    Xue, Bo and Cheng, Ji and Liu, Fei and Wang, Yimu and Zhang, Qingfu , year=. Multiobjective Lipschitz Bandits under Lexicographic Ordering , journal=

  5. [5]

    Proceedings of the 34th International Joint Conference on Artificial Intelligence , pages =

    Problem-dependent Regret for Lexicographic Multi-Armed Bandits with Adversarial Corruptions , author =. Proceedings of the 34th International Joint Conference on Artificial Intelligence , pages =

  6. [6]

    Multiple Trade-offs: An Improved Approach for Lexicographic Linear Bandits , booktitle=

    Xue, Bo and Lin, Xi and Zhang, Xiaoyuan and Zhang, Qingfu , year=. Multiple Trade-offs: An Improved Approach for Lexicographic Linear Bandits , booktitle=

  7. [7]

    Proceedings of the 42nd International Conference on Machine Learning , year=

    Multi-objective Linear Reinforcement Learning with Lexicographic Rewards , author=. Proceedings of the 42nd International Conference on Machine Learning , year=

  8. [8]

    Hierarchize Pareto Dominance in Multi-Objective Stochastic Linear Bandits , booktitle=

    Cheng, Ji and Xue, Bo and Yi, Jiaxiang and Zhang, Qingfu , year=. Hierarchize Pareto Dominance in Multi-Objective Stochastic Linear Bandits , booktitle=

Show all 299 references
  1. [9]

    Proceedings of the 34th International Joint Conference on Artificial Intelligence , pages =

    Multi-Objective Neural Bandits with Random Scalarization , author =. Proceedings of the 34th International Joint Conference on Artificial Intelligence , pages =

  2. [10]

    Using Confidence Bounds for Exploitation-Exploration Trade-offs , journal =

    Auer, Peter , year =. Using Confidence Bounds for Exploitation-Exploration Trade-offs , journal =

  3. [11]

    Bandits With Heavy Tail , year=

    Bubeck, Sébastien and Cesa-Bianchi, Nicolò and Lugosi, Gábor , journal=. Bandits With Heavy Tail , year=

  4. [12]

    Weighted sums of certain dependent random variables

    Azuma, Kazuoki. Weighted sums of certain dependent random variables. Tohoku Mathematical Journal. 1967

  5. [13]

    Annales de l'I.H.P

    Catoni, Olivier , title =. Annales de l'I.H.P. Probabilit\'es et statistiques , volume=. 2012 , pages =

  6. [14]

    Some aspects of the sequential design of experiments

    Robbins, Herbert. Some aspects of the sequential design of experiments. Bulletin of the American Mathematical Society. 1952

  7. [15]

    and Long, Philip M

    Abe, Naoki and Biermann, Alan W. and Long, Philip M. , title=. Algorithmica , year=

  8. [16]

    The heavy tail of the human brain

    James A Roberts and Tjeerd W Boonstra and Michael Breakspear. The heavy tail of the human brain. Current Opinion in Neurobiology. 2015

  9. [17]

    , title =

    Li, Lihong and Chu, Wei and Langford, John and Schapire, Robert E. , title =. Proceedings of the 19th International Conference on World Wide Web , year =

  10. [18]

    Proceedings of the 33rd International Conference on International Conference on Machine Learning , pages =

    Medina, Andres Munoz and Yang, Scott , title =. Proceedings of the 33rd International Conference on International Conference on Machine Learning , pages =

  11. [19]

    Proceedings of the 14th International Conference on Artificial Intelligence and Statistics , pages =

    Contextual Bandits with Linear Payoff Functions , author =. Proceedings of the 14th International Conference on Artificial Intelligence and Statistics , pages =

  12. [20]

    Proceedings of the 31st International Conference on Machine Learning , pages =

    Heavy-tailed regression with a generalized median-of-means , author =. Proceedings of the 31st International Conference on Machine Learning , pages =

  13. [21]

    Proceedings of the 31st Conference On Learning Theory , pages =

    Efficient Contextual Bandits in Non-stationary Worlds , author =. Proceedings of the 31st Conference On Learning Theory , pages =

  14. [22]

    Proceedings of the 31st International Conference on Machine Learning , pages =

    Taming the Monster: A Fast and Simple Algorithm for Contextual Bandits , author =. Proceedings of the 31st International Conference on Machine Learning , pages =

  15. [23]

    _1 -regression with Heavy-tailed Distributions , booktitle =

    Lijun Zhang and Zhi. _1 -regression with Heavy-tailed Distributions , booktitle =

  16. [24]

    Online Stochastic Linear Optimization under One-bit Feedback , booktitle =

    Lijun Zhang and Tianbao Yang and Rong Jin and Yichi Xiao and Zhi. Online Stochastic Linear Optimization under One-bit Feedback , booktitle =

  17. [25]

    , title =

    Shao, Han and Yu, Xiaotian and King, Irwin and Lyu, Michael R. , title =. Advances in Neural Information Processing Systems 31 , pages =

  18. [26]

    Advances in Neural Information Processing Systems 24 , pages =

    Improved Algorithms for Linear Stochastic Bandits , author =. Advances in Neural Information Processing Systems 24 , pages =

  19. [27]

    Proceedings of the 15th International Conference on Artificial Intelligence and Statistics , pages =

    Online-to-Confidence-Set Conversions and Application to Sparse Stochastic Bandits , author =. Proceedings of the 15th International Conference on Artificial Intelligence and Statistics , pages =

  20. [28]

    Advances in Neural Information Processing Systems 24 , pages =

    Linear Submodular Bandits and their Application to Diversified Retrieval , author =. Advances in Neural Information Processing Systems 24 , pages =

  21. [29]

    Sergey Foss and Dmitry Korshunov and Stan Zachary , TITLE =

  22. [30]

    Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems , JOURNAL =

    S\'. Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems , JOURNAL =. 2012 , volume =

  23. [31]

    Asymptotically efficient adaptive allocation rules

    Tze Leung Lai and Herbert Robbins. Asymptotically efficient adaptive allocation rules. Advances in Applied Mathematics. 1985

  24. [32]

    Proceedings of the 23rd Annual Conference on Learning Theory , pages =

    Audibert, Jean-Yves and Bubeck, Sébastien , title =. Proceedings of the 23rd Annual Conference on Learning Theory , pages =

  25. [33]

    Best Arm Identification: A Unified Approach to Fixed Budget and Fixed Confidence , year =

    Gabillon, Victor and Ghavamzadeh, Mohammad and Lazaric, Alessandro , booktitle =. Best Arm Identification: A Unified Approach to Fixed Budget and Fixed Confidence , year =

  26. [34]

    Bandit problems: sequential allocation of experiments , author=

  27. [35]

    and Kalai, Adam Tauman and McMahan, H

    Flaxman, Abraham D. and Kalai, Adam Tauman and McMahan, H. Brendan , title =. Proceedings of the 16th Annual ACM-SIAM Symposium on Discrete Algorithms , year =

  28. [36]

    Proceedings of the 40th Annual ACM Symposium on Theory of Computing , year =

    Kleinberg, Robert and Slivkins, Aleksandrs and Upfal, Eli , title =. Proceedings of the 40th Annual ACM Symposium on Theory of Computing , year =

  29. [37]

    Advances in Applied Probability , YEAR =

    Rajeev Agrawal , TITLE =. Advances in Applied Probability , YEAR =

  30. [38]

    Finite-time Analysis of the Multiarmed Bandit Problem , journal =

    Auer, Peter and Cesa-Bianchi, Nicol\`. Finite-time Analysis of the Multiarmed Bandit Problem , journal =

  31. [39]

    IEEE Journal of Selected Topics in Signal Processing , pages=

    Deterministic sequencing of exploration and exploitation for multi-armed bandit problems , author=. IEEE Journal of Selected Topics in Signal Processing , pages=. 2013 , volume=

  32. [40]

    Hayes and Sham M

    Varsha Dani and Thomas P. Hayes and Sham M. Kakade , TITLE =. Proceedings of the 21st Annual Conference on Learning , YEAR =

  33. [41]

    The Nonstochastic Multiarmed Bandit Problem , journal =

    Auer, Peter and Cesa-Bianchi, Nicol\`. The Nonstochastic Multiarmed Bandit Problem , journal =. 2002 , pages =

  34. [42]

    and Thomas, Joy A

    Cover, Thomas M. and Thomas, Joy A. , title =. 2006 , publisher =

  35. [43]

    IEEE Transactions on Information Theory , volume=

    PAC-Bayesian inequalities for martingales , author=. IEEE Transactions on Information Theory , volume=

  36. [44]

    Robust linear least squares regression

    Audibert, Jean-Yves and Catoni, Olivier. Robust linear least squares regression. The Annals of Statistics. 2011

  37. [45]

    Empirical risk minimization for heavy-tailed losses

    Brownlees, Christian and Joly, Emilien and Lugosi, G \'a bor. Empirical risk minimization for heavy-tailed losses. The Annals of Statistics. 2015

  38. [46]

    Journal of Machine Learning Research , year =

    Daniel Hsu and Sivan Sabato , title =. Journal of Machine Learning Research , year =

  39. [47]

    Proceedings of the 36th International Conference on Machine Learning , pages =

    Optimal Algorithms for Lipschitz Bandits with Heavy-tailed Rewards , author =. Proceedings of the 36th International Conference on Machine Learning , pages =

  40. [48]

    and Van Loan,, Charles F

    Golub,, Gene H. and Van Loan,, Charles F. , title =. 1996 , publisher =

  41. [49]

    and Hsu, Daniel and Kakade, Sham M

    Agarwal, Alekh and Foster, Dean P. and Hsu, Daniel and Kakade, Sham M. and Rakhlin, Alexander , title =. SIAM Journal on Optimization , volume =

  42. [50]

    2014 , journal =

    From Bandits to Monte-Carlo Tree Search: The Optimistic Principle Applied to Optimization and Planning , pages =. 2014 , journal =

  43. [51]

    Lyu and Irwin King , title =

    Xiaotian Yu and Han Shao and Michael R. Lyu and Irwin King , title =. 2018 , journal =

  44. [52]

    Lipschitz Bandits Without the Lipschitz Constant , booktitle =

    Bubeck, S. Lipschitz Bandits Without the Lipschitz Constant , booktitle =. 2011 , pages =

  45. [53]

    SIAM Journal on Control and Optimization , volume =

    Agrawal, Rajeev , title =. SIAM Journal on Control and Optimization , volume =

  46. [54]

    Bandit Convex Optimization:

    Sébastien Bubeck and Ofer Dekel and Tomer Koren and Yuval Peres , booktitle =. Bandit Convex Optimization:

  47. [55]

    Bandit Algorithms , publisher=

    Lattimore, Tor and Szepesvári, Csaba , year=. Bandit Algorithms , publisher=

  48. [56]

    Advances in Neural Information Processing Systems 23 , pages =

    Parametric Bandits: The Generalized Linear Case , author =. Advances in Neural Information Processing Systems 23 , pages =

  49. [57]

    Advances in Neural Information Processing Systems 30 , pages =

    Jun, Kwang-Sung and Bhargava, Aniruddha and Nowak, Robert and Willett, Rebecca , title =. Advances in Neural Information Processing Systems 30 , pages =

  50. [58]

    McCullagh, John A

    P. McCullagh, John A. Nelder , title =

  51. [59]

    Hubert and Shi, Elaine and Song, Dawn , title =

    Chan, T.-H. Hubert and Shi, Elaine and Song, Dawn , title =. Proceedings of the 37th International Colloquium Conference on Automata, Languages and Programming: Part II , year =

  52. [60]

    Proceedings of the 28th International Joint Conference on Artificial Intelligence , pages =

    Multi-Objective Generalized Linear Bandits , author =. Proceedings of the 28th International Joint Conference on Artificial Intelligence , pages =

  53. [61]

    Logarithmic Regret Algorithms for Online Convex Optimization , journal =

    Hazan, Elad and Agarwal, Amit and Kale, Satyen , year =. Logarithmic Regret Algorithms for Online Convex Optimization , journal =

  54. [62]

    Proceedings of the 34th International Conference on Machine Learning , pages =

    Provably Optimal Algorithms for Generalized Linear Contextual Bandits , author =. Proceedings of the 34th International Conference on Machine Learning , pages =

  55. [63]

    Breaking the Moments Condition Barrier: No-Regret Algorithm for Bandits with Super Heavy-Tailed Payoffs , year =

    Zhong, Han and Huang, Jiayi and Yang, Lin and Wang, Liwei , booktitle =. Breaking the Moments Condition Barrier: No-Regret Algorithm for Bandits with Super Heavy-Tailed Payoffs , year =

  56. [64]

    The Annals of Statistics , number =

    G. The Annals of Statistics , number =

  57. [65]

    Bayesian Optimization under Heavy-tailed Payoffs , year =

    Ray Chowdhury, Sayak and Gopalan, Aditya , booktitle =. Bayesian Optimization under Heavy-tailed Payoffs , year =

  58. [66]

    IEEE transactions on pattern analysis and machine intelligence , pages=

    The apolloscape open dataset for autonomous driving and its application , author=. IEEE transactions on pattern analysis and machine intelligence , pages=

  59. [67]

    Proceedings of the 29th International Joint Conference on Artificial Intelligence , year =

    Lijun Zhang , title =. Proceedings of the 29th International Joint Conference on Artificial Intelligence , year =

  60. [68]

    and Nowe, Ann , booktitle=

    Drugan, Madalina M. and Nowe, Ann , booktitle=. Designing multi-objective multi-armed bandits algorithms: A study , year=

  61. [69]

    Proceedings of the 19th International Conference on Artificial Intelligence and Statistics , pages =

    Pareto Front Identification from Stochastic Bandit Feedback , author =. Proceedings of the 19th International Conference on Artificial Intelligence and Statistics , pages =

  62. [70]

    2005 , publisher =

    Ehrgott, Matthias , title =. 2005 , publisher =

  63. [71]

    Multi-objective multi-armed bandit with lexicographically ordered and satisficing objectives , journal =

    Alihan H. Multi-objective multi-armed bandit with lexicographically ordered and satisficing objectives , journal =

  64. [72]

    Proceedings of the 21st International Conference on Artificial Intelligence and Statistics , pages=

    Multi-objective contextual bandit problem with similarity information , author=. Proceedings of the 21st International Conference on Artificial Intelligence and Statistics , pages=

  65. [73]

    Multi-objective

    Van Moffaert, Kristof and Van Vaerenbergh, Kevin and Vrancx, Peter and Nowe, Ann , booktitle=. Multi-objective. 2014 , pages=

  66. [74]

    Proceedings of the 39th International Conference on Machine Learning , pages =

    Stochastic Contextual Dueling Bandits under Linear Stochastic Transitivity Models , author =. Proceedings of the 39th International Conference on Machine Learning , pages =

  67. [75]

    2016 , booktitle =

    Joseph, Matthew and Kearns, Michael and Morgenstern, Jamie and Roth, Aaron , title =. 2016 , booktitle =

  68. [76]

    2012 , booktitle =

    Rodriguez, Mario and Posse, Christian and Zhang, Ethan , title =. 2012 , booktitle =

  69. [77]

    2004 , publisher=

    Convex optimization , author=. 2004 , publisher=

  70. [78]

    Breaking the

    Ghosh, Avishek and Sankararaman, Abishek , booktitle =. Breaking the

  71. [79]

    Proceedings of the 39th International Conference on Machine Learning , pages =

    Adaptive Best-of-Both-Worlds Algorithm for Heavy-Tailed Multi-Armed Bandits , author =. Proceedings of the 39th International Conference on Machine Learning , pages =

  72. [80]

    Proceedings of the 39th International Conference on Machine Learning , pages =

    A Simple Unified Framework for High Dimensional Bandit Problems , author =. Proceedings of the 39th International Conference on Machine Learning , pages =

  73. [81]

    Proceedings of the 39th International Conference on Machine Learning , pages =

    Linear Bandit Algorithms with Sublinear Time Complexity , author =. Proceedings of the 39th International Conference on Machine Learning , pages =

  74. [82]

    Proceedings of the 39th International Conference on Machine Learning , pages =

    Contextual Bandits with Smooth Regret: Efficient Learning in Continuous Action Spaces , author=. Proceedings of the 39th International Conference on Machine Learning , pages =

  75. [83]

    A Best-of-Both-Worlds Algorithm for Bandits with Delayed Feedback , year =

    Masoudian, Saeed and Zimmert, Julian and Seldin, Yevgeny , booktitle =. A Best-of-Both-Worlds Algorithm for Bandits with Delayed Feedback , year =

  76. [84]

    Lipschitz Bandits with Batched Feedback , year =

    Feng, Yasong and Huang, Zengfeng and Wang, Tianyu , booktitle =. Lipschitz Bandits with Batched Feedback , year =

  77. [85]

    Finite-Time Regret of Thompson Sampling Algorithms for Exponential Family Multi-Armed Bandits , year =

    Jin, Tianyuan and Xu, Pan and Xiao, Xiaokui and Anandkumar, Anima , booktitle =. Finite-Time Regret of Thompson Sampling Algorithms for Exponential Family Multi-Armed Bandits , year =

  78. [86]

    Communication Efficient Federated Learning for Generalized Linear Bandits , year =

    Li, Chuanhao and Wang, Hongning , booktitle =. Communication Efficient Federated Learning for Generalized Linear Bandits , year =

  79. [87]

    Outlier-Robust Sparse Mean Estimation for Heavy-Tailed Distributions , year =

    Diakonikolas, Ilias and Kane, Daniel and Lee, Jasper and Pensia, Ankit , booktitle =. Outlier-Robust Sparse Mean Estimation for Heavy-Tailed Distributions , year =

  80. [88]

    Clipped Stochastic Methods for Variational Inequalities with Heavy-Tailed Noise , year =

    Gorbunov, Eduard and Danilova, Marina and Dobre, David and Dvurechenskii, Pavel and Gasnikov, Alexander and Gidel, Gauthier , booktitle =. Clipped Stochastic Methods for Variational Inequalities with Heavy-Tailed Noise , year =

  81. [89]

    Proceedings of the 39th International Conference on Machine Learning , pages =

    Improved Rates for Differentially Private Stochastic Convex Optimization with Heavy-Tailed Data , author =. Proceedings of the 39th International Conference on Machine Learning , pages =

  82. [90]

    Proceedings of the 39th International Conference on Machine Learning , pages =

    High Probability Guarantees for Nonconvex Stochastic Gradient Descent with Heavy Tails , author =. Proceedings of the 39th International Conference on Machine Learning , pages =

  83. [91]

    Multi-armed Bandit Models for the Optimal Design of Clinical Trials: Benefits and Challenges , volume =

    Sof. Multi-armed Bandit Models for the Optimal Design of Clinical Trials: Benefits and Challenges , volume =. Statistical Science , number =

  84. [92]

    MOEA/D: A Multiobjective Evolutionary Algorithm Based on Decomposition , year=

    Zhang, Qingfu and Li, Hui , journal=. MOEA/D: A Multiobjective Evolutionary Algorithm Based on Decomposition , year=

  85. [93]

    Proceedings of the 38th International Conference on Machine Learning , pages =

    Near-Optimal Representation Learning for Linear Bandits and Linear RL , author =. Proceedings of the 38th International Conference on Machine Learning , pages =

  86. [94]

    Proceedings of the 38th International Conference on Machine Learning , pages =

    Robust Pure Exploration in Linear Bandits with Limited Budget , author =. Proceedings of the 38th International Conference on Machine Learning , pages =

  87. [95]

    2012 , volume =

    Foundations and Trends in Machine Learning , title =. 2012 , volume =

  88. [96]

    and Barto, Andrew G

    Sutton, Richard S. and Barto, Andrew G. , edition =. Reinforcement Learning: An Introduction , year =

  89. [97]

    Proceedings of the 24th International Conference on Artificial Intelligence and Statistics , pages =

    An Efficient Algorithm For Generalized Linear Bandit: Online Stochastic Gradient Descent and Thompson Sampling , author =. Proceedings of the 24th International Conference on Artificial Intelligence and Statistics , pages =

  90. [98]

    PAC-Bayesian Inequalities for Martingales , journal =

    Yevgeny Seldin and Fran. PAC-Bayesian Inequalities for Martingales , journal =

  91. [99]

    Learning in Generalized Linear Contextual Bandits with Stochastic Delays , year =

    Zhou, Zhengyuan and Xu, Renyuan and Blanchet, Jose , booktitle =. Learning in Generalized Linear Contextual Bandits with Stochastic Delays , year =

  92. [100]

    Advances in Neural Information Processing Systems 34 , pages =

    Generalized Linear Bandits with Local Differential Privacy , author=. Advances in Neural Information Processing Systems 34 , pages =

  93. [101]

    John Myles White , TITLE =

  94. [102]

    Physics in Medicine & Biology , year=

    Lexicographic ordering: intuitive multicriteria optimization for IMRT , author=. Physics in Medicine & Biology , year=

  95. [103]

    Integrated Assessment and Decision Support - Proceedings of the 1st Biennial Meeting of the International Environmental Modelling and Software Society , year=

    Lexicographic Optimisation for Water Resources Planning: the Case of Lake Verbano, Italy , author=. Integrated Assessment and Decision Support - Proceedings of the 1st Biennial Meeting of the International Environmental Modelling and Software Society , year=

  96. [104]

    Proceedings of the 27th Conference on Learning Theory , pages =

    Lipschitz Bandits: Regret Lower Bound and Optimal Algorithms , author =. Proceedings of the 27th Conference on Learning Theory , pages =

  97. [105]

    Advances in Neural Information Processing Systems 17 , year=

    Nearly Tight Bounds for the Continuum-armed Bandit Problem , author=. Advances in Neural Information Processing Systems 17 , year=

  98. [106]

    X-Armed Bandits , journal =

    S. X-Armed Bandits , journal =. 2011 , volume =

  99. [107]

    Improved Rates for the Stochastic Continuum-Armed Bandit Problem

    Auer, Peter and Ortner, Ronald and Szepesv \'a ri, Csaba. Improved Rates for the Stochastic Continuum-Armed Bandit Problem. Proceedings of the 20th Annual Conference on Learning Theory. 2007

  100. [108]

    Online Optimization in X-Armed Bandits , year =

    Bubeck, S\'. Online Optimization in X-Armed Bandits , year =. Advances in Neural Information Processing Systems 21 , pages =

  101. [109]

    2020 , booktitle =

    Wang, Tianyu and Ye, Weicheng and Geng, Dawei and Rudin, Cynthia , title =. 2020 , booktitle =

  102. [110]

    Yahyaa, Saba and M

    Q. Yahyaa, Saba and M. Drugan, Madalina and Manderick, Bernard , title =. 2014 , booktitle =

  103. [111]

    Multi-objective Contextual Multi-armed Bandit With a Dominant Objective , year=

    Tekin, Cem and Turgay, Eralp , journal=. Multi-objective Contextual Multi-armed Bandit With a Dominant Objective , year=

  104. [112]

    2018 , booktitle =

    Ma, Xiao and Zhao, Liqin and Huang, Guan and Wang, Zhi and Hu, Zelin and Zhu, Xiaoqiang and Gai, Kun , title =. 2018 , booktitle =

  105. [113]

    Proceedings of 34th Conference on Learning Theory , pages =

    Adaptive Discretization for Adversarial Lipschitz Bandits , author =. Proceedings of 34th Conference on Learning Theory , pages =

  106. [114]

    Proceedings of the 40th International Conference on International Conference on Machine Learning , pages=

    Pareto Regret Analyses in Multi-objective Multi-armed Bandit , author=. Proceedings of the 40th International Conference on International Conference on Machine Learning , pages=

  107. [115]

    An Empirical Evaluation of Thompson Sampling , year =

    Chapelle, Olivier and Li, Lihong , booktitle =. An Empirical Evaluation of Thompson Sampling , year =

  108. [116]

    Proceedings of the Workshop on On-line Trading of Exploration and Exploitation 2 , pages=

    An unbiased offline evaluation of contextual bandit algorithms with generalized linear models , author=. Proceedings of the Workshop on On-line Trading of Exploration and Exploitation 2 , pages=

  109. [117]

    Proceedings of the 31st International Joint Conference on Artificial Intelligence , pages =

    Lexicographic Multi-Objective Reinforcement Learning , author =. Proceedings of the 31st International Joint Conference on Artificial Intelligence , pages =

  110. [118]

    2015 , booktitle =

    Wray, Kyle Hollins and Zilberstein, Shlomo , title =. 2015 , booktitle =

  111. [119]

    Fair and Efficient Allocations under Lexicographic Preferences , booktitle=

    Hosseini, Hadi and Sikdar, Sujoy and Vaish, Rohit and Xia, Lirong , year=. Fair and Efficient Allocations under Lexicographic Preferences , booktitle=

  112. [120]

    Multi-Objective MDPs with Conditional Lexicographic Reward Preferences , booktitle=

    Wray, Kyle and Zilberstein, Shlomo and Mouaddib, Abdel-Illah , pages =. Multi-Objective MDPs with Conditional Lexicographic Reward Preferences , booktitle=

  113. [121]

    Stochastic Contextual Bandits with Long Horizon Rewards , booktitle=

    Qin, Yuzhen and Li, Yingcong and Pasqualetti, Fabio and Fazel, Maryam and Oymak, Samet , year=. Stochastic Contextual Bandits with Long Horizon Rewards , booktitle=

  114. [122]

    Proceedings of the 34st Conference On Learning Theory , pages =

    Chara Podimata and Alex Slivkins , title =. Proceedings of the 34st Conference On Learning Theory , pages =

  115. [123]

    Advances in Neural Information Processing Systems 32 , pages =

    Nirandika Wanigasekara and Christina Lee Yu , title =. Advances in Neural Information Processing Systems 32 , pages =

  116. [124]

    Multi-Objective contextual bandits with a dominant objective , year=

    Tekin, Cem and Turgay, Eralp , booktitle=. Multi-Objective contextual bandits with a dominant objective , year=

  117. [125]

    Customer Acquisition via Display Advertising Using Multi-Armed Bandit Experiments , number =

    Schwartz, Eric and Bradlow, Eric and Fader, Peter , year =. Customer Acquisition via Display Advertising Using Multi-Armed Bandit Experiments , number =

  118. [126]

    Resource Allocation for Multi-source Multi-relay Wireless Networks

    Khansa, Ali Al and Visoz, Raphael and Hayel, Yezekael and Lasaulce, Samson. Resource Allocation for Multi-source Multi-relay Wireless Networks. Ubiquitous Networking. 2021

  119. [127]

    Multi-objective optimization for supply chain management problem: A literature review , volume =

    Trisna, Trisna and Marimin, Marimin and Arkeman, Yandra and Sunarti, Titi , year =. Multi-objective optimization for supply chain management problem: A literature review , volume =

  120. [128]

    ACM Transactions on Information Systems , pages=

    Wang, Yifan and Ma, Weizhi and Zhang, Min and Liu, Yiqun and Ma, Shaoping , title =. ACM Transactions on Information Systems , pages=. 2023 , volume =

  121. [129]

    2022 , author =

    Multiobjective optimization under uncertainty: A multiobjective robust (relative) regret approach , journal =. 2022 , author =

  122. [130]

    2018 , booktitle =

    Lykouris, Thodoris and Mirrokni, Vahab and Paes Leme, Renato , title =. 2018 , booktitle =

  123. [131]

    2009 , journal =

    Click Fraud , author =. 2009 , journal =

  124. [132]

    Proceedings of the 32nd Conference on Learning Theory , pages =

    Better Algorithms for Stochastic Bandits with Adversarial Corruptions , author =. Proceedings of the 32nd Conference on Learning Theory , pages =

  125. [133]

    Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics , pages =

    Corruption-Tolerant Gaussian Process Bandit Optimization , author =. Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics , pages =

  126. [134]

    Nearly Optimal Algorithms for Linear Contextual Bandits with Adversarial Corruptions , year =

    He, Jiafan and Zhou, Dongruo and Zhang, Tong and Gu, Quanquan , booktitle =. Nearly Optimal Algorithms for Linear Contextual Bandits with Adversarial Corruptions , year =

  127. [135]

    Advances in Neural Information Processing Systems 36 , year=

    Robust Lipschitz Bandits to Adversarial Corruptions , author=. Advances in Neural Information Processing Systems 36 , year=

  128. [136]

    Stochastic Graphical Bandits with Adversarial Corruptions , booktitle=

    Lu, Shiyin and Wang, Guanghui and Zhang, Lijun , year=. Stochastic Graphical Bandits with Adversarial Corruptions , booktitle=

  129. [137]

    Proceedings of the 23nd International Conference on Artificial Intelligence and Statistics , pages =

    An Optimal Algorithm for Stochastic and Adversarial Bandits , author =. Proceedings of the 23nd International Conference on Artificial Intelligence and Statistics , pages =

  130. [138]

    Proceedings of the 25th International Conference on Artificial Intelligence and Statistics , pages =

    Robust Stochastic Linear Contextual Bandits Under Adversarial Attacks , author =. Proceedings of the 25th International Conference on Artificial Intelligence and Statistics , pages =

  131. [139]

    and Hu, Timothy Y

    Guan, Ziwei and Ji, Kaiyi and Bucci Jr., Donald J. and Hu, Timothy Y. and Palombo, Joseph and Liston, Michael and Liang, Yingbin , year=. Robust Stochastic Bandit Algorithms under Probabilistic Unbounded Adversarial Attack , booktitle=

  132. [140]

    Journal of Machine Learning Research , year =

    Jason Altschuler and Victor-Emmanuel Brunel and Alan Malek , title =. Journal of Machine Learning Research , year =

  133. [141]

    Mean-based Best Arm Identification in Stochastic Bandits under Reward Contamination , year =

    Mukherjee, Arpan and Tajer, Ali and Chen, Pin-Yu and Das, Payel , booktitle =. Mean-based Best Arm Identification in Stochastic Bandits under Reward Contamination , year =

  134. [142]

    2023 , eprint=

    Adversarial Attacks on Combinatorial Multi-Armed Bandits , author=. 2023 , eprint=

  135. [143]

    37th Conference on Neural Information Processing Systems , year=

    Corruption-Robust Offline Reinforcement Learning with General Function Approximation , author=. 37th Conference on Neural Information Processing Systems , year=

  136. [144]

    Proceedings of the 25th International Conference on Artificial Intelligence and Statistics , pages =

    Corruption-robust Offline Reinforcement Learning , author =. Proceedings of the 25th International Conference on Artificial Intelligence and Statistics , pages =

  137. [145]

    Proceedings of the 38th International Conference on Machine Learning , pages =

    On Reinforcement Learning with Adversarial Corruption and Its Application to Block MDP , author =. Proceedings of the 38th International Conference on Machine Learning , pages =

  138. [146]

    Proceedings of the 26th International Conference on Artificial Intelligence and Statistics , pages =

    Vector Optimization with Stochastic Bandit Feedback , author =. Proceedings of the 26th International Conference on Artificial Intelligence and Statistics , pages =

  139. [147]

    Advances in Neural Information Processing Systems 36 , year=

    Adaptive Algorithms for Relaxed Pareto Set Identification , author=. Advances in Neural Information Processing Systems 36 , year=

  140. [148]

    No-regret Algorithms for Multi-task

    Chowdhury, Sayak Ray and Gopalan, Aditya , booktitle =. No-regret Algorithms for Multi-task

  141. [149]

    Proceedings of the 36th International Conference on Machine Learning , pages =

    Data Poisoning Attacks on Stochastic Bandits , author =. Proceedings of the 36th International Conference on Machine Learning , pages =

  142. [150]

    The End of Optimism?

    Lattimore, Tor and Szepesvari, Csaba , booktitle =. The End of Optimism?

  143. [151]

    Proceedings of 34th Conference on Learning Theory , pages =

    Fine-Grained Gap-Dependent Bounds for Tabular MDPs via Adaptive Multi-Step Bootstrap , author =. Proceedings of 34th Conference on Learning Theory , pages =

  144. [152]

    Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics , pages =

    Adaptive Exploration in Linear Contextual Bandit , author =. Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics , pages =

  145. [153]

    Exploration in Structured Reinforcement Learning , year =

    Ok, Jungseul and Proutiere, Alexandre and Tranos, Damianos , booktitle =. Exploration in Structured Reinforcement Learning , year =

  146. [154]

    Non-Asymptotic Gap-Dependent Regret Bounds for Tabular MDPs , year =

    Simchowitz, Max and Jamieson, Kevin G , booktitle =. Non-Asymptotic Gap-Dependent Regret Bounds for Tabular MDPs , year =

  147. [155]

    2016 , booktitle =

    Chen, Wei and Hu, Wei and Li, Fu and Li, Jian and Liu, Yu and Lu, Pinyan , title =. 2016 , booktitle =

  148. [156]

    Interactive Submodular Bandit , year =

    Chen, Lin and Krause, Andreas and Karbasi, Amin , booktitle =. Interactive Submodular Bandit , year =

  149. [157]

    Proceedings of the 36th International Conference on Machine Learning , pages =

    Tighter Problem-Dependent Regret Bounds in Reinforcement Learning without Domain Knowledge using Value Function Bounds , author =. Proceedings of the 36th International Conference on Machine Learning , pages =

  150. [158]

    Exploration–exploitation Tradeoff Using Variance Estimates in Multi-armed Bandits , journal =

  151. [159]

    Improved Regret Analysis for Variance-Adaptive Linear Bandits and Horizon-Free Linear Mixture MDPs , year =

    Kim, Yeoneung and Yang, Insoon and Jun, Kwang-Sung , booktitle =. Improved Regret Analysis for Variance-Adaptive Linear Bandits and Horizon-Free Linear Mixture MDPs , year =

  152. [160]

    Improved Variance-Aware Confidence Sets for Linear Bandits and Linear Mixture MDP , year =

    Zhang, Zihan and Yang, Jiaqi and Ji, Xiangyang and Du, Simon S , booktitle =. Improved Variance-Aware Confidence Sets for Linear Bandits and Linear Mixture MDP , year =

  153. [161]

    Variance-Aware Off-Policy Evaluation with Linear Function Approximation , year =

    Min, Yifei and Wang, Tianhao and Zhou, Dongruo and Gu, Quanquan , booktitle =. Variance-Aware Off-Policy Evaluation with Linear Function Approximation , year =

  154. [162]

    Llorens , title =

    Xin-Qiang Cai and Pushi Zhang and Li Zhao and Bian Jiang and Masashi Sugiyama and Ashley J. Llorens , title =. Advances in Neural Information Processing Systems 36 , year =

  155. [163]

    1974 , author =

    A two-armed bandit theory of market pricing , journal =. 1974 , author =

  156. [164]

    2023 , author =

    A multi-objective home healthcare delivery model and its solution using a branch-and-price algorithm and a two-stage meta-heuristic algorithm , journal =. 2023 , author =

  157. [165]

    International Conference on Agents and Artificial Intelligence , year=

    Thompson Sampling in the Adaptive Linear Scalarized Multi Objective Multi Armed Bandit , author=. International Conference on Agents and Artificial Intelligence , year=

  158. [166]

    Yahyaa, Saba and M

    Q. Yahyaa, Saba and M. Drugan, Madalina and Manderick, Bernard , title =. 2014 , pages =

  159. [167]

    and Zintgraf, Luisa M

    Roijers, Diederik M. and Zintgraf, Luisa M. and Nowe, Ann , title =. 2017 , booktitle =

  160. [168]

    and Zintgraf, Luisa M

    Roijers, Diederik M. and Zintgraf, Luisa M. and Libin, Pieter and Reymond, Mathieu and Bargiacchi, Eugenio and Now\'. Interactive Multi-Objective Reinforcement Learning in Multi-Armed Bandits with Gaussian Process Utility Models , year =

  161. [169]

    Lexicographic Actor-Critic Deep Reinforcement Learning for Urban Autonomous Driving , year=

    Zhang, Hengrui and Lin, Youfang and Han, Sheng and Lv, Kai , journal=. Lexicographic Actor-Critic Deep Reinforcement Learning for Urban Autonomous Driving , year=

  162. [170]

    Nonstochastic Multi-Armed Bandits with Graph-structured Feedback

    Noga Alon and Nicolo Cesa-Bianchi and Claudio Gentile and Shie Mannor and Yishay Mansour and Ohad Shamir. Nonstochastic Multi-Armed Bandits with Graph-structured Feedback. SIAM Journal on Computing. 2017

  163. [171]

    From Bandits to Experts: On the Value of Side-Observations , year =

    Mannor, Shie and Shamir, Ohad , booktitle =. From Bandits to Experts: On the Value of Side-Observations , year =

  164. [172]

    , edition = 2, publisher =

    West, Douglas B. , edition = 2, publisher =. Introduction to Graph Theory , year =

  165. [173]

    Leveraging Side Observations in Stochastic Bandits , year =

    Caron, St\'. Leveraging Side Observations in Stochastic Bandits , year =. Proceedings of the 28th Conference on Uncertainty in Artificial Intelligence , pages =

  166. [174]

    , title =

    Buccapatnam, Swapna and Eryilmaz, Atilla and Shroff, Ness B. , title =. 2014 , booktitle =

  167. [175]

    Shroff , title =

    Swapna Buccapatnam and Fang Liu and Atilla Eryilmaz and Ness B. Shroff , title =. Journal of Machine Learning Research , year =

  168. [176]

    Proceedings of the 33rd International Conference on Machine Learning , pages =

    Online Learning with Feedback Graphs Without the Graphs , author =. Proceedings of the 33rd International Conference on Machine Learning , pages =

  169. [177]

    Thompson Sampling for Stochastic Bandits with Graph Feedback , booktitle=

    Tossou, Aristide and Dimitrakakis, Christos and Dubhashi, Devdatt , pages =. Thompson Sampling for Stochastic Bandits with Graph Feedback , booktitle=

  170. [178]

    Information Directed Sampling for Stochastic Bandits With Graph Feedback , booktitle=

    Liu, Fang and Buccapatnam, Swapna and Shroff, Ness , pages =. Information Directed Sampling for Stochastic Bandits With Graph Feedback , booktitle=

  171. [179]

    Shroff , title =

    Fang Liu and Zizhan Zheng and Ness B. Shroff , title =. Proceedings of the 34th Conference on Uncertainty in Artificial Intelligence , pages =

  172. [180]

    Proceedings of the 35th Uncertainty in Artificial Intelligence Conference , pages =

    Problem-dependent Regret Bounds for Online Learning with Feedback Graphs , author =. Proceedings of the 35th Uncertainty in Artificial Intelligence Conference , pages =

  173. [181]

    Proceedings of the 31st International Conference on Algorithmic Learning Theory , pages =

    Feedback graph regret bounds for Thompson Sampling and UCB , author =. Proceedings of the 31st International Conference on Algorithmic Learning Theory , pages =

  174. [182]

    Proceedings of the 39th Conference on Uncertainty in Artificial Intelligence , pages =

    Stochastic Graphical Bandits with Heavy-Tailed Rewards , author =. Proceedings of the 39th Conference on Uncertainty in Artificial Intelligence , pages =

  175. [183]

    Journal of Machine Learning Research , year =

    Eyal Even-Dar and Shie Mannor and Yishay Mansour , title =. Journal of Machine Learning Research , year =

  176. [184]

    Thompson , journal =

    William R. Thompson , journal =. On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples , volume =

  177. [185]

    Efficient learning by implicit exploration in bandit problems with side observations , year =

    Koc\'. Efficient learning by implicit exploration in bandit problems with side observations , year =. Advances in Neural Information Processing Systems 27 , pages =

  178. [186]

    Proceedings of the 28th Conference on Learning Theory , pages =

    Online Learning with Feedback Graphs: Beyond Bandits , author =. Proceedings of the 28th Conference on Learning Theory , pages =

  179. [187]

    Proceedings of 33rd Conference on Learning Theory , pages =

    A Closer Look at Small-loss Bounds for Bandits with Graph Feedback , author =. Proceedings of 33rd Conference on Learning Theory , pages =

  180. [188]

    and Mohri, Mehryar , title =

    Arora, Raman and Marinov, Teodor V. and Mohri, Mehryar , title =. 2019 , booktitle =

  181. [189]

    Proceedings of the 31st Conference On Learning Theory , pages =

    Small-loss bounds for online learning with partial information , author =. Proceedings of the 31st Conference On Learning Theory , pages =

  182. [190]

    Proceedings of the 35th AAAI Conference on Artificial Intelligence , year =

    Shiyin Lu and Yao Hu and Lijun Zhang , title =. Proceedings of the 35th AAAI Conference on Artificial Intelligence , year =

  183. [191]

    2000 , journal =

    HERD BEHAVIOR AND AGGREGATE FLUCTUATIONS IN FINANCIAL MARKETS , author =. 2000 , journal =

  184. [192]

    Nearly Optimal Best-of-Both-Worlds Algorithms for Online Learning with Feedback Graphs , year =

    Ito, Shinji and Tsuchiya, Taira and Honda, Junya , booktitle =. Nearly Optimal Best-of-Both-Worlds Algorithms for Online Learning with Feedback Graphs , year =

  185. [193]

    Towards Best-of-All-Worlds Online Learning with Feedback Graphs , year =

    Erez, Liad and Koren, Tomer , booktitle =. Towards Best-of-All-Worlds Online Learning with Feedback Graphs , year =

  186. [194]

    1999 , Address =

    Nonlinear Multiobjective Optimization , Author =. 1999 , Address =

  187. [195]

    2021 , author =

    On the estimation of pareto front and dimensional similarity in many-objective evolutionary algorithm , journal =. 2021 , author =

  188. [196]

    2018 , author =

    Approximating the irregularly shaped Pareto front of multi-objective reservoir flood control operation problem , journal =. 2018 , author =

  189. [197]

    An Evolutionary Many-Objective Optimization Algorithm Using Reference-Point-Based Nondominated Sorting Approach, Part I: Solving Problems With Box Constraints , year=

    Deb, Kalyanmoy and Jain, Himanshu , journal=. An Evolutionary Many-Objective Optimization Algorithm Using Reference-Point-Based Nondominated Sorting Approach, Part I: Solving Problems With Box Constraints , year=

  190. [198]

    An Evolutionary Many-Objective Optimization Algorithm Based on Dominance and Decomposition , year=

    Li, Ke and Deb, Kalyanmoy and Zhang, Qingfu and Kwong, Sam , journal=. An Evolutionary Many-Objective Optimization Algorithm Based on Dominance and Decomposition , year=

  191. [199]

    and Santana-Quintero, Luis V

    Hernández-Díaz, Alfredo G. and Santana-Quintero, Luis V. and Coello Coello, Carlos A. and Molina, Julián , title = ". Evolutionary Computation , volume =

  192. [200]

    Multi-objective Bandits: Optimizing the Generalized

    R. Multi-objective Bandits: Optimizing the Generalized. Proceedings of the 34th International Conference on Machine Learning , pages =

  193. [201]

    Annals of Operations Research , year=2022, volume=

    Maciej Nowak and Tadeusz Trzaskalik , title=. Annals of Operations Research , year=2022, volume=

  194. [202]

    Keeney , journal =

    Ralph L. Keeney , journal =. Common Mistakes in Making Value Trade-Offs , volume =

  195. [203]

    A. D. Athanassopoulos and V. V. Podinovski , journal =. Dominance and Potential Optimality in Multiple Criteria Decision Analysis with Imprecise Information , volume =

  196. [204]

    Podinovski , number =

    Victor V. Podinovski , number =. A DSS for multiple criteria decision analysis with imprecisely specified trade-offs , journal =

  197. [205]

    Ruiz and Francisco Ruiz and Kaisa Miettinen and Laura Delgado-Antequera and Vesa Ojalehto , title=

    Ana B. Ruiz and Francisco Ruiz and Kaisa Miettinen and Laura Delgado-Antequera and Vesa Ojalehto , title=. Journal of Global Optimization , year=2019, volume=

  198. [206]

    1999 , author =

    Searching for psychologically stable solutions of multiple criteria decision problems , journal =. 1999 , author =

  199. [207]

    2000 , author =

    Using trade-off information in decision-making algorithms , journal =. 2000 , author =

  200. [208]

    TopRank: A practical algorithm for online stochastic ranking , year =

    Lattimore, Tor and Kveton, Branislav and Li, Shuai and Szepesvari, Csaba , booktitle =. TopRank: A practical algorithm for online stochastic ranking , year =

  201. [209]

    IEEE/ACM Transactions on Networking , pages =

    Gai, Yi and Krishnamachari, Bhaskar and Jain, Rahul , title =. IEEE/ACM Transactions on Networking , pages =. 2012 , volume =

  202. [210]

    Proceedings of the 30th International Conference on Machine Learning , pages =

    Combinatorial Multi-Armed Bandit: General Framework and Applications , author =. Proceedings of the 30th International Conference on Machine Learning , pages =

  203. [211]

    Journal of Machine Learning Research , year =

    Wei Chen and Yajun Wang and Yang Yuan and Qinshi Wang , title =. Journal of Machine Learning Research , year =

  204. [212]

    Proceedings of the 35th International Conference on Machine Learning , pages =

    Thompson Sampling for Combinatorial Semi-Bandits , author =. Proceedings of the 35th International Conference on Machine Learning , pages =

  205. [213]

    Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence , pages =

    Combinatorial semi-bandit in the non-stationary environment , author =. Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence , pages =

  206. [214]

    Hybrid Regret Bounds for Combinatorial Semi-Bandits and Adversarial Linear Bandits , year =

    Ito, Shinji , booktitle =. Hybrid Regret Bounds for Combinatorial Semi-Bandits and Adversarial Linear Bandits , year =

  207. [215]

    Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =

    Further Adaptive Best-of-Both-Worlds Algorithm for Combinatorial Semi-Bandits , author =. Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =

  208. [216]

    Proceedings of the 18th International Conference on Artificial Intelligence and Statistics , pages =

    Tight Regret Bounds for Stochastic Combinatorial Semi-Bandits , author =. Proceedings of the 18th International Conference on Artificial Intelligence and Statistics , pages =

  209. [217]

    Proceedings of the 40th International Conference on Machine Learning , pages =

    Probably Anytime-Safe Stochastic Combinatorial Semi-Bandits , author =. Proceedings of the 40th International Conference on Machine Learning , pages =

  210. [218]

    Neural Computation , volume =

    Kuroki, Yuko and Xu, Liyuan and Miyauchi, Atsushi and Honda, Junya and Sugiyama, Masashi , title =. Neural Computation , volume =

  211. [219]

    Combinatorial Pure Exploration with Full-Bandit or Partial Linear Feedback , journal=

    Du, Yihan and Kuroki, Yuko and Chen, Wei , year=. Combinatorial Pure Exploration with Full-Bandit or Partial Linear Feedback , journal=

  212. [220]

    Proceedings of the 38th Conference on Uncertainty in Artificial Intelligence , pages =

    An explore-then-commit algorithm for submodular maximization under full-bandit feedback , author =. Proceedings of the 38th Conference on Uncertainty in Artificial Intelligence , pages =

  213. [221]

    Discover Artificial Intelligence , volume=

    Multi-objective optimization for autonomous driving strategy based on Deep Q Network , author=. Discover Artificial Intelligence , volume=

  214. [222]

    Proceedings of the 28th International Joint Conference on Artificial Intelligence , pages =

    Learning Multi-Objective Rewards and User Utility Function in Contextual Bandits for Personalized Ranking , author =. Proceedings of the 28th International Joint Conference on Artificial Intelligence , pages =

  215. [223]

    Proceedings of The 27th International Conference on Artificial Intelligence and Statistics , pages =

    Sequential learning of the. Proceedings of The 27th International Conference on Artificial Intelligence and Statistics , pages =

  216. [224]

    Near-Optimal Regret Bounds for Contextual Combinatorial Semi-Bandits with Linear Payoff Functions , booktitle =

    Kei Takemura and Shinji Ito and Daisuke Hatano and Hanna Sumita and Takuro Fukunaga and Naonori Kakimura and Ken. Near-Optimal Regret Bounds for Contextual Combinatorial Semi-Bandits with Linear Payoff Functions , booktitle =

  217. [225]

    Combinatorial Bandits with Linear Constraints: Beyond Knapsacks and Fairness , year =

    Liu, Qingsong and Xu, Weihang and Wang, Siwei and Fang, Zhixuan , booktitle =. Combinatorial Bandits with Linear Constraints: Beyond Knapsacks and Fairness , year =

  218. [226]

    2022 , pages =

    Hu, Xinyan and Ngo, Dung Daniel and Slivkins, Aleksandrs and Wu, Zhiwei Steven , title =. 2022 , pages =

  219. [227]

    2016 , eprint=

    Influence Maximization with Bandits , author=. 2016 , eprint=

  220. [228]

    Proceedings of the 40th International Conference on Machine Learning , pages =

    A Framework for Adapting Offline Algorithms to Solve Combinatorial Multi-Armed Bandit Problems with Bandit Feedback , author =. Proceedings of the 40th International Conference on Machine Learning , pages =

  221. [229]

    Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =

    Randomized Greedy Learning for Non-monotone Stochastic Submodular Maximization Under Full-bandit Feedback , author =. Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =

  222. [230]

    2024 , booktitle =

    Fourati, Fares and Alouini, Mohamed-Slim and Aggarwal, Vaneet , title =. 2024 , booktitle =

  223. [231]

    Combinatorial Stochastic-Greedy Bandit , journal=

    Fourati, Fares and Quinn, Christopher John and Alouini, Mohamed-Slim and Aggarwal, Vaneet , year=. Combinatorial Stochastic-Greedy Bandit , journal=

  224. [232]

    When Combinatorial Thompson Sampling meets Approximation Regret , year =

    Perrault, Pierre , booktitle =. When Combinatorial Thompson Sampling meets Approximation Regret , year =

  225. [233]

    A Unified Confidence Sequence for Generalized Linear Models, with Applications to Bandits , year =

    Lee, Jungyhun and Yun, Se-Young and Jun, Kwang-Sung , booktitle =. A Unified Confidence Sequence for Generalized Linear Models, with Applications to Bandits , year =

  226. [234]

    Neural Combinatorial Clustered Bandits for Recommendation Systems , journal=

    Atalar, Baran and Joe-Wong, Carlee , year=. Neural Combinatorial Clustered Bandits for Recommendation Systems , journal=

  227. [235]

    Nemhauser, G. L. and Wolsey, L. A. and Fisher, M. L. , title =. Mathematical Programming , pages =. 1978 , volume =

  228. [236]

    Journal of Machine Learning Research , year =

    Jianqing Fan and Bai Jiang and Qiang Sun , title =. Journal of Machine Learning Research , year =

  229. [237]

    2013 , publisher =

    Bach, Francis , title =. 2013 , publisher =

  230. [238]

    Proceedings of the 31st International Conference on Algorithmic Learning Theory , pages =

    Top- k Combinatorial Bandits with Full-Bandit Feedback , author =. Proceedings of the 31st International Conference on Algorithmic Learning Theory , pages =

  231. [239]

    Sadegh and Proutiere, Alexandre and Lelarge, Marc , title =

    Combes, Richard and Talebi, M. Sadegh and Proutiere, Alexandre and Lelarge, Marc , title =. 2015 , booktitle =

  232. [240]

    Combinatorial bandits , year =

    Cesa-Bianchi, Nicol\`. Combinatorial bandits , year =. Journal of Computer and System Sciences , pages =

  233. [241]

    IEEE Transactions on Information Theory , pages =

    Fang, Guanhua and Li, Ping and Samorodnitsky, Gennady , title =. IEEE Transactions on Information Theory , pages =. 2025 , volume =

  234. [242]

    IEEE Transactions on Information Theory , pages =

    Khaleghi, Azadeh , title =. IEEE Transactions on Information Theory , pages =. 2025 , volume =

  235. [243]

    Karthik, P. N. and Reddy, Kota Srinivas and Tan, Vincent Y. F. , title =. IEEE Transactions on Information Theory , pages =. 2023 , volume =

  236. [244]

    Multi-Armed Bandits With Correlated Arms , year =

    Gupta, Samarth and Chaudhari, Shreyas and Joshi, Gauri and Ya. Multi-Armed Bandits With Correlated Arms , year =. IEEE Transactions on Information Theory , pages =

  237. [245]

    The 13th International Conference on Learning Representations , year=

    On Speeding Up Language Model Evaluation , author=. The 13th International Conference on Learning Representations , year=

  238. [246]

    2024 , eprint=

    Sample-Efficient Alignment for LLMs , author=. 2024 , eprint=

  239. [247]

    Bandit-Based Prompt Design Strategy Selection Improves Prompt Optimizers

    Ashizawa, Rin and Hirose, Yoichi and Yoshinari, Nozomu and Uchida, Kento and Shirakawa, Shinichi. Bandit-Based Prompt Design Strategy Selection Improves Prompt Optimizers. Findings of the Association for Computational Linguistics: ACL 2025. 2025

  240. [248]

    2025 , eprint=

    Online Multi-LLM Selection via Contextual Bandits under Unstructured Context Evolution , author=. 2025 , eprint=

  241. [249]

    Efficient Prompt Optimization Through the Lens of Best Arm Identification , year =

    Shi, Chengshuai and Yang, Kun and Chen, Zihan and Li, Jundong and Yang, Jing and Shen, Cong , booktitle =. Efficient Prompt Optimization Through the Lens of Best Arm Identification , year =

  242. [250]

    ACM Computing Surveys , numpages =

    Liu, Pengfei and Yuan, Weizhe and Fu, Jinlan and Jiang, Zhengbao and Hayashi, Hiroaki and Neubig, Graham , title =. ACM Computing Surveys , numpages =. 2023 , volume =

  243. [251]

    Advances in Neural Information Processing Systems 35 , pages =

    Wei, Jason and Wang, Xuezhi and Schuurmans, Dale and Bosma, Maarten and ichter, brian and Xia, Fei and Chi, Ed and Le, Quoc V and Zhou, Denny , title=. Advances in Neural Information Processing Systems 35 , pages =

  244. [252]

    A Thorough Examination of Decoding Methods in the Era of LLM s

    Shi, Chufan and Yang, Haoran and Cai, Deng and Zhang, Zhisong and Wang, Yifan and Yang, Yujiu and Lam, Wai. A Thorough Examination of Decoding Methods in the Era of LLM s. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024

  245. [253]

    2024 , eprint=

    GPT-4 Technical Report , author=. 2024 , eprint=

  246. [254]

    2021 , booktitle =

    Reynolds, Laria and McDonell, Kyle , title =. 2021 , booktitle =

  247. [255]

    and Yang, Qiang and Xie, Xing , title =

    Chang, Yupeng and Wang, Xu and Wang, Jindong and Wu, Yuan and Yang, Linyi and Zhu, Kaijie and Chen, Hao and Yi, Xiaoyuan and Wang, Cunxiang and Wang, Yidong and Ye, Wei and Zhang, Yue and Chang, Yi and Yu, Philip S. and Yang, Qiang and Xie, Xing , title =. 2024 , volume =

  248. [256]

    From Generation to Judgment: Opportunities and Challenges of LLM -as-a-judge

    Li, Dawei and Jiang, Bohan and Huang, Liangjie and Beigi, Alimohammad and Zhao, Chengshuai and Tan, Zhen and Bhattacharjee, Amrita and Jiang, Yuxuan and Chen, Canyu and Wu, Tianhao and Shu, Kai and Cheng, Lu and Liu, Huan. From Generation to Judgment: Opportunities and Challen...

  249. [257]

    Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena , year =

    Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric and Zhang, Hao and Gonzalez, Joseph E and Stoica, Ion , booktitle =. Judging LLM-as-a-Judge with MT-Bench and C...

  250. [258]

    Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition

    Feng, Kehua and Ding, Keyan and Hongzhi, Tan and Ma, Kede and Wang, Zhihua and Guo, Shuangquan and Yuzhou, Cheng and Sun, Ge and Zheng, Guozhou and Zhang, Qiang and Chen, Huajun. Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition. Pr...

  251. [259]

    2023 , eprint=

    Evaluating Large Language Models: A Comprehensive Survey , author=. 2023 , eprint=

  252. [260]

    Data Intelligence , volume =

    Li, Linhan and Zhang, Huaping and Li, Chunjin and You, Haowen and Cui, Wenyao , title =. Data Intelligence , volume =

  253. [261]

    Proceedings of The 27th Conference on Learning Theory , pages =

    lil' UCB : An Optimal Exploration Algorithm for Multi-Armed Bandits , author =. Proceedings of The 27th Conference on Learning Theory , pages =

  254. [262]

    2021 , eprint=

    Training Verifiers to Solve Math Word Problems , author=. 2021 , eprint=

  255. [263]

    Proceedings of the 34th AAAI Conference on Artificial Intelligence , year =

    Yonatan Bisk and Rowan Zellers and Ronan Le Bras and Jianfeng Gao and Yejin Choi , title =. Proceedings of the 34th AAAI Conference on Artificial Intelligence , year =

  256. [264]

    Proceedings of The 33rd International Conference on Machine Learning , pages =

    Anytime optimal algorithms in stochastic multi-armed bandits , author =. Proceedings of The 33rd International Conference on Machine Learning , pages =

  257. [265]

    Proceedings of the 36th International Conference on Machine Learning , pages =

    Bilinear Bandits with Low-rank Structure , author =. Proceedings of the 36th International Conference on Machine Learning , pages =

  258. [266]

    Proceedings of the 38th International Conference on Machine Learning , pages =

    Improved Regret Bounds of Bilinear Bandits using Action Space Analysis , author =. Proceedings of the 38th International Conference on Machine Learning , pages =

  259. [267]

    Proceedings of The 24th International Conference on Artificial Intelligence and Statistics , pages =

    Low-Rank Generalized Linear Bandit Problems , author =. Proceedings of The 24th International Conference on Artificial Intelligence and Statistics , pages =

  260. [268]

    Efficient Frameworks for Generalized Low-Rank Matrix Bandit Problems , year =

    Kang, Yue and Hsieh, Cho-Jui and Lee, Thomas Chun Man , booktitle =. Efficient Frameworks for Generalized Low-Rank Matrix Bandit Problems , year =

  261. [269]

    2025 , archivePrefix=

    Generalized Low-Rank Matrix Contextual Bandits with Graph Information , author=. 2025 , archivePrefix=

  262. [270]

    2026 , archivePrefix=

    Low-Rank Contextual Reinforcement Learning from Heterogeneous Human Feedback , author=. 2026 , archivePrefix=

  263. [271]

    Proceedings of the Fortieth Conference on Uncertainty in Artificial Intelligence , pages =

    Low-rank Matrix Bandits with Heavy-tailed Rewards , author =. Proceedings of the Fortieth Conference on Uncertainty in Artificial Intelligence , pages =

  264. [272]

    Proceedings of the 41st International Conference on Machine Learning , pages =

    Efficient Low-Rank Matrix Estimation, Experimental Design, and Arm-Set-Dependent Low-Rank Bandits , author =. Proceedings of the 41st International Conference on Machine Learning , pages =

  265. [273]

    Multi-task Representation Learning for Pure Exploration in Bilinear Bandits , year =

    Mukherjee, Subhojyoti and Xie, Qiaomin and Hanna, Josiah and Nowak, Robert , booktitle =. Multi-task Representation Learning for Pure Exploration in Bilinear Bandits , year =

  266. [274]

    Advances in Neural Information Processing Systems 38 , year=

    Thompson Sampling for Multi-Objective Linear Contextual Bandit , author=. Advances in Neural Information Processing Systems 38 , year=

  267. [275]

    Optimal Scalarizations for Sublinear Hypervolume Regret , year =

    Zhang, Qiuyi (Richard) , booktitle =. Optimal Scalarizations for Sublinear Hypervolume Regret , year =

  268. [276]

    Prabhu , title =

    Alperen Tercan, Vinayak S. Prabhu , title =. Proceedings of the 27th European Conference on Artificial Intelligence , pages =

  269. [277]

    2013 , booktitle =

    Ding, Wenkui and Qiny, Tao and Zhang, Xu-Dong and Liu, Tie-Yan , title =. 2013 , booktitle =

  270. [278]

    Transactions on Machine Learning Research , issn=

    Holistic Evaluation of Language Models , author=. Transactions on Machine Learning Research , issn=

  271. [279]

    2024 , eprint=

    A Survey on Efficient Inference for Large Language Models , author=. 2024 , eprint=

  272. [280]

    2015 , booktitle =

    Xia, Yingce and Li, Haifang and Qin, Tao and Yu, Nenghai and Liu, Tie-Yan , title =. 2015 , booktitle =

  273. [281]

    2016 , booktitle =

    Xia, Yingce and Qin, Tao and Ma, Weidong and Yu, Nenghai and Liu, Tie-Yan , title =. 2016 , booktitle =

  274. [282]

    2017 , booktitle =

    Li, Haifang and Xia, Yingce , title =. 2017 , booktitle =

  275. [283]

    Machine Learning , pages =

    Luedtke, Alex and Kaufmann, Emilie and Chambaz, Antoine , title =. Machine Learning , pages =. 2019 , volume =

  276. [284]

    Journal of the ACM , pages =

    Badanidiyuru, Ashwinkumar and Kleinberg, Robert and Slivkins, Aleksandrs , title =. Journal of the ACM , pages =. 2018 , volume =

  277. [285]

    Proceedings of The 27th Conference on Learning Theory , pages =

    Resourceful Contextual Bandits , author =. Proceedings of The 27th Conference on Learning Theory , pages =

  278. [286]

    Proceedings of the Twenty-First International Conference on Artificial Intelligence and Statistics , pages =

    Combinatorial Semi-Bandits with Knapsacks , author =. Proceedings of the Twenty-First International Conference on Artificial Intelligence and Statistics , pages =

  279. [287]

    Journal of the ACM , pages =

    Immorlica, Nicole and Sankararaman, Karthik and Schapire, Robert and Slivkins, Aleksandrs , title =. Journal of the ACM , pages =. 2022 , volume =

  280. [288]

    2022 , booktitle =

    Liu, Shang and Jiang, Jiashuo and Li, Xiaocheng , title =. 2022 , booktitle =

  281. [289]

    Proceedings of the 37th International Conference on Machine Learning , year =

    Random Hypervolume Scalarizations for Provable Multi-Objective Black Box Optimization , author =. Proceedings of the 37th International Conference on Machine Learning , year =

  282. [290]

    Kone, Cyrille and Kaufmann, Emilie and Richert, Laura , booktitle =. Bandit

  283. [291]

    Use Your

    Lin, Xiaoqiang and Wu, Zhaoxuan and Dai, Zhongxiang and Hu, Wenyang and Shu, Yao and Ng, See-Kiong and Jaillet, Patrick and Low, Bryan Kian Hsiang , booktitle =. Use Your

  284. [292]

    2024 , booktitle =

    Wu, Zhaoxuan and Lin, Xiaoqiang and Dai, Zhongxiang and Hu, Wenyang and Shu, Yao and Ng, See-Kiong and Jaillet, Patrick and Low, Bryan Kian Hsiang , title =. 2024 , booktitle =

  285. [293]

    Zhi Hong and Qian Zhang and Jiahang Sun and Zhiwei Shang and Mingze Kong and Xiangyi Wang and Yao Shu and Zhongxiang Dai , booktitle=

  286. [294]

    2025 , eprint=

    FedPOB: Sample-Efficient Federated Prompt Optimization via Bandits , author=. 2025 , eprint=

  287. [295]

    and Dai, Zhongxiang

    Sun, Jiahang and Wang, Zhiyong and Yang, Runhan and Xiao, Chenjun and Lui, John C.s. and Dai, Zhongxiang. Large Language Model-Enhanced Multi-Armed Bandits. Proceedings of the 64th Annual Meeting of the A ssociation for C omputational L inguistics (Volume 1: Long Papers). 2026

  288. [296]

    Automatic Prompt Optimization with ``Gradient Descent'' and Beam Search

    Pryzant, Reid and Iter, Dan and Li, Jerry and Lee, Yin and Zhu, Chenguang and Zeng, Michael. Automatic Prompt Optimization with ``Gradient Descent'' and Beam Search. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023

  289. [297]

    Chen, Lingjiao and Zaharia, Matei and Zou, James , journal =

  290. [298]

    and Kadous, M

    Ong, Isaac and Almahairi, Amjad and Wu, Vincent and Chiang, Wei-Lin and Wu, Tianhao and Gonzalez, Joseph E. and Kadous, M. Waleed and Stoica, Ion , booktitle =

  291. [299]

    2024 , booktitle =

    Efficient Multi-Prompt Evaluation of Large Language Models , author =. 2024 , booktitle =

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.