TM20K uses a one-time full-token teacher plus knowledge distillation to let cheaper student models merge tokens and serve 20K-length e-commerce sequences at near-5K cost, gaining +1.036% ADSS online.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.IR 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Teacher Retains Full Tokens, Student Merges Efficiently: TM20K for E-Commerce Sequence Modeling in Ad Recommendation
TM20K uses a one-time full-token teacher plus knowledge distillation to let cheaper student models merge tokens and serve 20K-length e-commerce sequences at near-5K cost, gaining +1.036% ADSS online.