Pith. sign in

Deep learning recommendation model for personalization and recommendation systems

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.AI 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Reinforcement Learning from User Feedback

cs.AI · 2025-05-20 · conditional · novelty 5.0

A reward model trained on sparse heart-emoji reactions predicted online Love Reaction rates (r=0.95 across ten models) and, when added to multi-objective RL, lifted Love Reactions by up to 28% in live A/B tests, with reward-hacking tradeoffs.

citing papers explorer

Showing 1 of 1 citing paper.

  • Reinforcement Learning from User Feedback cs.AI · 2025-05-20 · conditional · none · ref 7

    A reward model trained on sparse heart-emoji reactions predicted online Love Reaction rates (r=0.95 across ten models) and, when added to multi-objective RL, lifted Love Reactions by up to 28% in live A/B tests, with reward-hacking tradeoffs.