REVIEW 8 cited by
Hiformer: Heterogeneous Feature Interactions Learning with Transformers for Recommender Systems
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Learning feature interaction is the critical backbone to building recommender systems. In web-scale applications, learning feature interaction is extremely challenging due to the sparse and large input feature space; meanwhile, manually crafting effective feature interactions is infeasible because of the exponential solution space. We propose to leverage a Transformer-based architecture with attention layers to automatically capture feature interactions. Transformer architectures have witnessed great success in many domains, such as natural language processing and computer vision. However, there has not been much adoption of Transformer architecture for feature interaction modeling in industry. We aim at closing the gap. We identify two key challenges for applying the vanilla Transformer architecture to web-scale recommender systems: (1) Transformer architecture fails to capture the heterogeneous feature interactions in the self-attention layer; (2) The serving latency of Transformer architecture might be too high to be deployed in web-scale recommender systems. We first propose a heterogeneous self-attention layer, which is a simple yet effective modification to the self-attention layer in Transformer, to take into account the heterogeneity of feature interactions. We then introduce \textsc{Hiformer} (\textbf{H}eterogeneous \textbf{I}nteraction Trans\textbf{former}) to further improve the model expressiveness. With low-rank approximation and model pruning, \hiformer enjoys fast inference for online deployment. Extensive offline experiment results corroborates the effectiveness and efficiency of the \textsc{Hiformer} model. We have successfully deployed the \textsc{Hiformer} model to a real world large scale App ranking model at Google Play, with significant improvement in key engagement metrics (up to +2.66\%).
Forward citations
Cited by 8 Pith papers
-
TransX: Scaling Transformer-based Recommendation via Behavioral and Serving Stream Crossings
A dual-stream encoder-decoder recommender with nearline-cached behavior encoding reports higher CTR/CVR than DLRM baselines at roughly one-fifth of the online compute in LinkedIn deployment.
-
SpecFormer: Mitigating Embedding and Attention Collapse via Spectral-Aware Transformer for Recommendation
SpecFormer is a spectral-aware Transformer that flattens the singular-value spectrum of embeddings to prevent embedding/attention collapse, outperforming baselines on CTR benchmarks and scaling with layer depth.
-
UniRank: Benchmarking Ranking Models for Unified Sequential Modeling and Feature Interaction
UniRank is an open benchmark that standardizes chronological autoregressive supervision, multi-task evaluation, and capacity controls for 15 unified ranking models on five large datasets.
-
From Scaling to Structured Expressivity: Rethinking Transformers for CTR Prediction
FAT specializes attention by semantic field and reports +0.51% AUC over baselines on Taobao data, but its power-law scaling law is an empirical fit, not a derived prediction.
-
RankMixer: Scaling Up Ranking Models in Industrial Recommenders
RankMixer scales an industrial ranking model to 1B dense parameters with 10x MFU improvement and unchanged latency, gaining 1.08% in app duration in Douyin A/B tests.
-
WHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture
Layer-wise attention fusion between Wukong-style feature crosses and HSTU-style behavior history improves recommendation quality over each backbone alone at matched FLOPs.
-
TMallGS: Scaling Unified Feature and Sequence Modeling for Generative E-commerce Search
TmallGS, a decoupled Transformer ranking architecture with per-field projections, gating, FiLM fusion, and progressive training, reports consistent offline and online gains on Tmall Search.
-
LO-FAR: A Cost-Aware Local Filter for Sparse Feature Ranking in Industrial Ad Recommendation
LO-FAR ranks sparse ID-list features by stand-alone held-out predictive signal and reports downstream NE gains competitive with shuffle importance and BSN at 100–400 retained features in about two CPU-hours.
Discussion (0). Continue with ORCID to comment.