Under sub-exponential reward delays, Delayed NeuralUCB achieves regret O(d~√T logT + d~^{3/2}D+ log^{3/2}T), with D+ depending on the expected delay.
Contextual bandits for adapting treatment in a mouse model of de novo carcinogenesis,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Neural Contextual Bandits Under Delayed Feedback Constraints
Under sub-exponential reward delays, Delayed NeuralUCB achieves regret O(d~√T logT + d~^{3/2}D+ log^{3/2}T), with D+ depending on the expected delay.