A reward model trained on sparse heart-emoji reactions predicted online Love Reaction rates (r=0.95 across ten models) and, when added to multi-objective RL, lifted Love Reactions by up to 28% in live A/B tests, with reward-hacking tradeoffs.
Deep learning recommendation model for personalization and recommendation systems
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Reinforcement Learning from User Feedback
A reward model trained on sparse heart-emoji reactions predicted online Love Reaction rates (r=0.95 across ten models) and, when added to multi-objective RL, lifted Love Reactions by up to 28% in live A/B tests, with reward-hacking tradeoffs.