REVIEW 15 cited by
Wukong: Towards a Scaling Law for Large-Scale Recommendation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Scaling laws play an instrumental role in the sustainable improvement in model quality. Unfortunately, recommendation models to date do not exhibit such laws similar to those observed in the domain of large language models, due to the inefficiencies of their upscaling mechanisms. This limitation poses significant challenges in adapting these models to increasingly more complex real-world datasets. In this paper, we propose an effective network architecture based purely on stacked factorization machines, and a synergistic upscaling strategy, collectively dubbed Wukong, to establish a scaling law in the domain of recommendation. Wukong's unique design makes it possible to capture diverse, any-order of interactions simply through taller and wider layers. We conducted extensive evaluations on six public datasets, and our results demonstrate that Wukong consistently outperforms state-of-the-art models quality-wise. Further, we assessed Wukong's scalability on an internal, large-scale dataset. The results show that Wukong retains its superiority in quality over state-of-the-art models, while holding the scaling law across two orders of magnitude in model complexity, extending beyond 100 GFLOP/example, where prior arts fall short.
Forward citations
Cited by 15 Pith papers
-
Sample Is Feature: Beyond Item-Level, Toward Sample-Level Tokens for Unified Large Recommender Models
SIF replaces item-ID history tokens with lossily compressed full-sample tokens and reports consistent CTR/CVR gains in offline and live recommender tests.
-
Scaling Transformers for Discriminative Recommendation via Generative Pretraining
Generative pretraining plus sparse-embedding freezing makes large Transformer ranking models scale consistently, following a power law from 13K to 0.3B dense parameters.
-
WhisperRec: Latent Reasoning for Efficient Foundation Recommendation Models
WhisperRec distills multi-view chain-of-thought rationale into three latent tokens, beating explicit-reasoning recommenders at about ten times the inference throughput.
-
Mosaic: A Fleet of User Embedding Specialists for Recommendation at Meta
Mosaic shows that a fleet of four heterogeneous user-embedding specialists, trained with redundancy-reduction and composite-label losses, improves downstream recommendation quality at Meta.
-
Hi-SAM: A Hierarchical Structure-Aware Multi-modal Framework for Large-Scale Recommendation
Hi-SAM improves semantic-ID multimodal recommendation by disentangling shared versus modality-specific item codes and by letting transformers access history only through compressed anchor tokens.
-
KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta
An agentic kernel-coding system combining tree search with hardware-knowledge retrieval generated optimized Triton kernels for NVIDIA, AMD, and Meta's MTIA accelerators: 100% correctness on 480 operator-platform confi...
-
From Scaling to Structured Expressivity: Rethinking Transformers for CTR Prediction
FAT specializes attention by semantic field and reports +0.51% AUC over baselines on Taobao data, but its power-law scaling law is an empirical fit, not a derived prediction.
-
Yambda-5B -- A Large-Scale Multi-modal Dataset for Ranking And Retrieval
A new open 4.79B-interaction music dataset from Yandex Music with an is_organic flag, audio embeddings, and a Global Temporal Split benchmark protocol.
-
MTGR: Industrial-Scale Generative Recommendation Framework in Meituan
MTGR augments an HSTU-style generative ranking model with DLRM cross features and user-level aggregation, and reports a successful industrial deployment at Meituan with offline and online gains.
-
Unifying Generative Recall and Multi-Objective Ranking in a Single Decoder-Only Sequence
A single decoder-only sequence with dual-query prefix-causal attention and ranking-side LoRA unifies generative SID recall and multi-objective ranking, with offline and online gains at Kuaishou.
-
CMSL: Constructive Multi-Sequence Learning for Recommendation Systems
CMSL uses a learnable module to disentangle user history into multiple pure sequences modeled with linear attention to improve recommendation performance over single-sequence approaches.
-
TMallGS: Scaling Unified Feature and Sequence Modeling for Generative E-commerce Search
TmallGS, a decoupled Transformer ranking architecture with per-field projections, gating, FiLM fusion, and progressive training, reports consistent offline and online gains on Tmall Search.
-
DLF: Enhancing Explicit-Implicit Interaction via Dynamic Low-Order-Aware Fusion for CTR Prediction
DLF is a CTR prediction architecture that combines low-rank, high-rank, and implicit interaction blocks with layer-wise attention fusion, reporting state-of-the-art results on Criteo, Avazu, Movielens, and Frappe.
-
Privacy Preserving Conversion Modeling in Data Clean Room
Batch-level aggregated gradients, LoRA adapters, and de-biased label differential privacy let advertisers and platforms train conversion models in a clean room with modest AUC loss and much lower communication cost.
-
Climber: Toward Efficient Scaling Laws for Large Recommendation Models
Climber reports that splitting user sequences by behavior type, adding adaptive temperature, and co-designed batching enable more efficient Transformer scaling in recommender systems.
Discussion (0). Continue with ORCID to comment.