REVIEW 2 cited by
STEM: Unleashing the Power of Embeddings for Multi-task Recommendation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Multi-task learning (MTL) has gained significant popularity in recommender systems as it enables simultaneous optimization of multiple objectives. A key challenge in MTL is negative transfer, but existing studies explored negative transfer on all samples, overlooking the inherent complexities within them. We split the samples according to the relative amount of positive feedback among tasks. Surprisingly, negative transfer still occurs in existing MTL methods on samples that receive comparable feedback across tasks. Existing work commonly employs a shared-embedding paradigm, limiting the ability of modeling diverse user preferences on different tasks. In this paper, we introduce a novel Shared and Task-specific EMbeddings (STEM) paradigm that aims to incorporate both shared and task-specific embeddings to effectively capture task-specific user preferences. Under this paradigm, we propose a simple model STEM-Net, which is equipped with an All Forward Task-specific Backward gating network to facilitate the learning of task-specific embeddings and direct knowledge transfer across tasks. Remarkably, STEM-Net demonstrates exceptional performance on comparable samples, achieving positive transfer. Comprehensive evaluation on three public MTL recommendation datasets demonstrates that STEM-Net outperforms state-of-the-art models by a substantial margin. Our code is released at https://github.com/LiangcaiSu/STEM.
Forward citations
Cited by 2 Pith papers
-
Towards Unifying Feature Interaction Models for Click-Through Rate Prediction
Most explicit feature-interaction CTR models can be expressed as combinations of an interaction function, a layer pooling strategy, and a layer aggregator; the derived PFL model is competitive with state-of-the-art methods.
-
No More Tuning: Prioritized Multi-Task Learning with Lagrangian Differential Multiplier Methods
NMT optimizes lower-priority tasks under a Lagrangian penalty that keeps the primary task loss near its pre-trained optimum, with no manual balancing weights in the loss combination.
Discussion (0). Continue with ORCID to comment.