A Pareto-optimal trade-off between regret and treatment-effect estimation is derived and achieved for bandits with network interference, by compressing the action space through exposure mapping.
Thompson sampling for contextual bandits with linear payoffs
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Online Experimental Design With Estimation-Regret Trade-off Under Network Interference
A Pareto-optimal trade-off between regret and treatment-effect estimation is derived and achieved for bandits with network interference, by compressing the action space through exposure mapping.