Pith. sign in

REVIEW 15 cited by

Wukong: Towards a Scaling Law for Large-Scale Recommendation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.02545 v4 pith:5ETU7MH3 submitted 2024-03-04 cs.LG cs.AI

classification cs.LGcs.AI
keywords wukongmodelsscalingrecommendationdatasetsdomainlarge-scalelaws
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Scaling laws play an instrumental role in the sustainable improvement in model quality. Unfortunately, recommendation models to date do not exhibit such laws similar to those observed in the domain of large language models, due to the inefficiencies of their upscaling mechanisms. This limitation poses significant challenges in adapting these models to increasingly more complex real-world datasets. In this paper, we propose an effective network architecture based purely on stacked factorization machines, and a synergistic upscaling strategy, collectively dubbed Wukong, to establish a scaling law in the domain of recommendation. Wukong's unique design makes it possible to capture diverse, any-order of interactions simply through taller and wider layers. We conducted extensive evaluations on six public datasets, and our results demonstrate that Wukong consistently outperforms state-of-the-art models quality-wise. Further, we assessed Wukong's scalability on an internal, large-scale dataset. The results show that Wukong retains its superiority in quality over state-of-the-art models, while holding the scaling law across two orders of magnitude in model complexity, extending beyond 100 GFLOP/example, where prior arts fall short.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sample Is Feature: Beyond Item-Level, Toward Sample-Level Tokens for Unified Large Recommender Models

    cs.IR 2026-04 unverdicted novelty 7.0 of 10

    SIF replaces item-ID history tokens with lossily compressed full-sample tokens and reports consistent CTR/CVR gains in offline and live recommender tests.

  2. Scaling Transformers for Discriminative Recommendation via Generative Pretraining

    cs.IR 2025-06 conditional novelty 7.0 of 10

    Generative pretraining plus sparse-embedding freezing makes large Transformer ranking models scale consistently, following a power law from 13K to 0.3B dense parameters.

  3. WhisperRec: Latent Reasoning for Efficient Foundation Recommendation Models

    cs.IR 2026-07 conditional novelty 6.0 of 10

    WhisperRec distills multi-view chain-of-thought rationale into three latent tokens, beating explicit-reasoning recommenders at about ten times the inference throughput.

  4. Mosaic: A Fleet of User Embedding Specialists for Recommendation at Meta

    cs.IR 2026-07 conditional novelty 6.0 of 10

    Mosaic shows that a fleet of four heterogeneous user-embedding specialists, trained with redundancy-reduction and composite-label losses, improves downstream recommendation quality at Meta.

  5. Hi-SAM: A Hierarchical Structure-Aware Multi-modal Framework for Large-Scale Recommendation

    cs.AI 2026-02 conditional novelty 6.0 of 10

    Hi-SAM improves semantic-ID multimodal recommendation by disentangling shared versus modality-specific item codes and by letting transformers access history only through compressed anchor tokens.

  6. KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta

    cs.LG 2025-12 conditional novelty 6.0 of 10

    An agentic kernel-coding system combining tree search with hardware-knowledge retrieval generated optimized Triton kernels for NVIDIA, AMD, and Meta's MTIA accelerators: 100% correctness on 480 operator-platform confi...

  7. From Scaling to Structured Expressivity: Rethinking Transformers for CTR Prediction

    cs.IR 2025-11 reject novelty 6.0 of 10

    FAT specializes attention by semantic field and reports +0.51% AUC over baselines on Taobao data, but its power-law scaling law is an empirical fit, not a derived prediction.

  8. Yambda-5B -- A Large-Scale Multi-modal Dataset for Ranking And Retrieval

    cs.IR 2025-05 conditional novelty 6.0 of 10

    A new open 4.79B-interaction music dataset from Yandex Music with an is_organic flag, audio embeddings, and a Global Temporal Split benchmark protocol.

  9. MTGR: Industrial-Scale Generative Recommendation Framework in Meituan

    cs.IR 2025-05 conditional novelty 6.0 of 10

    MTGR augments an HSTU-style generative ranking model with DLRM cross features and user-level aggregation, and reports a successful industrial deployment at Meituan with offline and online gains.

  10. Unifying Generative Recall and Multi-Objective Ranking in a Single Decoder-Only Sequence

    cs.IR 2026-07 conditional novelty 5.5 of 10

    A single decoder-only sequence with dual-query prefix-causal attention and ranking-side LoRA unifies generative SID recall and multi-objective ranking, with offline and online gains at Kuaishou.

  11. CMSL: Constructive Multi-Sequence Learning for Recommendation Systems

    cs.IR 2026-06 unverdicted novelty 5.5 of 10

    CMSL uses a learnable module to disentangle user history into multiple pure sequences modeled with linear attention to improve recommendation performance over single-sequence approaches.

  12. TMallGS: Scaling Unified Feature and Sequence Modeling for Generative E-commerce Search

    cs.IR 2026-07 conditional novelty 5.0 of 10

    TmallGS, a decoupled Transformer ranking architecture with per-field projections, gating, FiLM fusion, and progressive training, reports consistent offline and online gains on Tmall Search.

  13. DLF: Enhancing Explicit-Implicit Interaction via Dynamic Low-Order-Aware Fusion for CTR Prediction

    cs.IR 2025-05 conditional novelty 5.0 of 10

    DLF is a CTR prediction architecture that combines low-rank, high-rank, and implicit interaction blocks with layer-wise attention fusion, reporting state-of-the-art results on Criteo, Avazu, Movielens, and Frappe.

  14. Privacy Preserving Conversion Modeling in Data Clean Room

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Batch-level aggregated gradients, LoRA adapters, and de-biased label differential privacy let advertisers and platforms train conversion models in a clean room with modest AUC loss and much lower communication cost.

  15. Climber: Toward Efficient Scaling Laws for Large Recommendation Models

    cs.IR 2025-02 conditional novelty 4.0 of 10

    Climber reports that splitting user sequences by behavior type, adding adaptive temperature, and co-designed batching enable more efficient Transformer scaling in recommender systems.

Pith tools