MNL-VQL is the first algorithm with regret bounds for combinatorial reinforcement learning with multinomial-logit preference feedback, and it is nearly minimax-optimal in linear MDPs.
Instance-wise minimax-optimal algorithms for logistic bandits
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Combinatorial Reinforcement Learning with Preference Feedback
MNL-VQL is the first algorithm with regret bounds for combinatorial reinforcement learning with multinomial-logit preference feedback, and it is nearly minimax-optimal in linear MDPs.