A single Q-network trained with envelope (convex-hull) updates over preferences can output near-optimal policies for any linear combination of objectives and infer hidden preferences from few samples.
On min-norm and min-max methods of multi-objective optimization
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Generalized Algorithm for Multi-Objective Reinforcement Learning and Policy Adaptation
A single Q-network trained with envelope (convex-hull) updates over preferences can output near-optimal policies for any linear combination of objectives and infer hidden preferences from few samples.