REVIEW 16 cited by
DHEN: A Deep and Hierarchical Ensemble Network for Large-Scale Click-Through Rate Prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Learning feature interactions is important to the model performance of online advertising services. As a result, extensive efforts have been devoted to designing effective architectures to learn feature interactions. However, we observe that the practical performance of those designs can vary from dataset to dataset, even when the order of interactions claimed to be captured is the same. That indicates different designs may have different advantages and the interactions captured by them have non-overlapping information. Motivated by this observation, we propose DHEN - a deep and hierarchical ensemble architecture that can leverage strengths of heterogeneous interaction modules and learn a hierarchy of the interactions under different orders. To overcome the challenge brought by DHEN's deeper and multi-layer structure in training, we propose a novel co-designed training system that can further improve the training efficiency of DHEN. Experiments of DHEN on large-scale dataset from CTR prediction tasks attained 0.27\% improvement on the Normalized Entropy (NE) of prediction and 1.2x better training throughput than state-of-the-art baseline, demonstrating their effectiveness in practice.
Forward citations
Cited by 16 Pith papers
-
Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators
A Triton compiler backend, TorchInductor adaptations, and small language extensions let Meta's MTIA-2i run Triton kernels competitively with expert-tuned C++ in production.
-
ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation
ROCS restructures recommendation models so user-side computation is shared across all candidate items, yielding up to 3x serving throughput at equal or better prediction quality.
-
HCCL: Collective Communication for Meta Training and Inference Accelerators
HCCL offloads collective communication to MTIA 300's message engines, achieving up to 940 GB/s intra-rack bandwidth and sub-6µs latency for inference.
-
Probabilistic Residual Learning for Online Recommendations
PRL adds a cluster-aware, causality-adjusted residual correction layer to any base recommender, improving cold-start cross-domain recommendation accuracy in experiments.
-
UniRank: Benchmarking Ranking Models for Unified Sequential Modeling and Feature Interaction
UniRank is an open benchmark that standardizes chronological autoregressive supervision, multi-task evaluation, and capacity controls for 15 unified ranking models on five large datasets.
-
Bumblebee: Interleaved Mixed-Layer Building Blocks for Large-Scale Recommendation Systems
Interleaving sequence modeling with feature interaction in repeated blocks improves recommendation accuracy by about 0.2-1.4% NE over sequential baselines at matched parameter counts.
-
MixFormer: Co-Scaling Up Dense and Sequence in Industrial Recommenders
MixFormer unifies dense feature interaction and user-sequence modeling in a single Transformer-style backbone with a user-item decoupling speedup, reporting accuracy and efficiency gains over stacked and parallel reco...
-
KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta
An agentic kernel-coding system combining tree search with hardware-knowledge retrieval generated optimized Triton kernels for NVIDIA, AMD, and Meta's MTIA accelerators: 100% correctness on 480 operator-platform confi...
-
Request-Only Optimization for Recommendation Systems
A request-level training data format eliminates duplicate user features, increasing storage efficiency and training throughput while enabling larger recommendation architectures.
-
RankMixer: Scaling Up Ranking Models in Industrial Recommenders
RankMixer scales an industrial ranking model to 1B dense parameters with 10x MFU improvement and unchanged latency, gaining 1.08% in app duration in Douyin A/B tests.
-
Action is All You Need: Dual-Flow Generative Ranking Network for Recommendation
DFGR, a dual-flow generative ranking network, halves training cost and quarters inference cost versus HSTU-based generative ranking while reporting better offline AUC/G-AUC on public and industrial datasets.
-
Optimus: A Generic Operator-Level PyTorch Model Transformation Framework
Optimus rewrites atomic operator patterns in PT2 graphs via greedy search, delivering large QPS, memory, and compile-time gains on production recommendation models.
-
Large Foundation Model for Ads Recommendation
Tencent's LFM4Ads transfers user, item, and user-item cross representations from a pre-trained foundation model into downstream ad models via feature, module, and model-level mechanisms, reporting a 2.45% platform-wid...
-
Synergizing Implicit and Explicit User Interests: A Multi-Embedding Retrieval Framework at Pinterest
A production recommender framework combines a differentiable clustering module for implicit interests and conditional retrieval for explicit followed topics, deployed at Pinterest home feed.
-
Privacy Preserving Conversion Modeling in Data Clean Room
Batch-level aggregated gradients, LoRA adapters, and de-biased label differential privacy let advertisers and platforms train conversion models in a clean room with modest AUC loss and much lower communication cost.
-
Decoupled Entity Representation Learning for Pinterest Ads Ranking
Pre-computed user and Pin embeddings from multi-tower models improve Pinterest ad ranking by small but statistically significant margins.
Discussion (0). Continue with ORCID to comment.