REVIEW 4 major objections 5 minor 1 cited by
Generating Long Semantic IDs in Parallel for Recommendation
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read By predicting every token of a long semantic ID in parallel, a recommender beats autoregressive generative baselines by 12.6% on NDCG@10 while keeping inference cost independent of catalog size.
desk verdict RPG is genuinely new—parallel generation of long OPQ semantic IDs with a multi-token-prediction loss plus graph-constrained decoding—and the experiments mostly hold up, but the abstract oversells m=64 and the paper never quantifies the cost of its factorized scoring. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the factorized logit sum: for a candidate ID $(c_1,\ldots,c_m)$, the recommendation score is $\sum_{j=1}^m \log p^{(j)}_{c_j}$, where $p^{(j)}$ is the cached softmax vector over the $j$-th codebook after one sequence encoder pass. This identity turns ranking into a nearest-neighbor-style search over a product space and motivates the decoding graph, whose nodes are valid IDs and whose edges connect IDs with high summed token-embedding similarity. Iterative graph propagation with a beam of size $b$ and $q$ steps explores only $O(bqk)$ candidate IDs, making inference complexity independent of the item count.
What would settle it
Run a held-out test in which two code positions are deliberately made informative only through their co-occurrence, such as a brand subvector and a category subvector that jointly identify a niche product class, and compare the factorized scorer against a scorer that uses the full token combination. If the joint scorer improves NDCG@10, or if brute-force scoring of every valid ID reorders RPG's graph-decoded top-$K$, the central independence and parallel-decoding claim is refuted.
Extended reading notes
Core claim
RPG's core claim is that next-item prediction can be done directly in the space of long, unordered semantic IDs. Each item is encoded by optimized product quantization into $m$ tokens drawn from $m$ separate codebooks; conditional on the user sequence, the model assumes the tokens are independent, so the probability of a candidate ID factors as $P(c_1,\dots,c_m|s)=\prod_{j=1}^m P^{(j)}(c_j|s)$. The training loss is the negative sum of per-token log probabilities, and at inference the score of each candidate ID is the sum of cached per-token logits. Because decoding is parallel, ID length can grow to 64 without adding autoregressive steps, and the model avoids invalid token combinations by propagating a small beam through a graph whose edges connect IDs that differ in few tokens. The paper's empirical claim is that this design outranks all compared baselines on 11 of 12 metric-dataset combinations, with an average 12.6% NDCG@10 improvement over the best generative baseline, while keeping inference cost independent of catalog size.
Load-bearing premise
The load-bearing premise is that, once the user's history is fixed, the separate parts of the next item's code do not influence each other, so the model can rank by adding up independent per-part scores; if combinations of code parts carry extra meaning about what a user wants, the ranking objective is misspecified.
Editorial extensions
If this is right
- A single sequence encoder forward pass supplies all token logits, so the number of decoding steps does not grow with semantic ID length; long IDs are therefore no longer a latency burden.
- The decoding graph can be precomputed once after training, so serving-time memory and latency stay flat as the catalog grows, unlike retrieval-based recommenders that must scan item embeddings.
- The multi-token prediction objective lets the model learn from sub-item semantic structure directly, which is what the paper credits for its gains on cold-start items with few interactions.
- Because candidate scoring is parallel, the same architecture can consume richer item encoders and longer IDs without changing the decoder, which the paper shows translates into better NDCG@10 on larger datasets.
Reading between the lines
- A direct test of the independence assumption would be to compare RPG's factorized score with a jointly trained scorer that sees full token combinations; if the joint scorer wins on held-out ranking, some of RPG's signal is being left on the table.
- The same parallel-scoring scheme should transfer to other generative retrieval settings where items are represented by product-quantized codes, such as document or product search, provided a valid-ID graph can be built.
- Because the graph is static once built, RPG would need an incremental graph update to stay correct in a dynamic catalog where items are added or re-embedded; the paper does not address that serving scenario.
- The reported 12.6% gain bundles the effect of the longer ID with the new objective and decoder; a length-controlled comparison that varies only the ID length while holding everything else fixed would separate the two contributions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes RPG, a semantic ID-based sequential recommendation framework. Items are tokenized into long, unordered semantic IDs using optimized product quantization (up to 64 tokens), and a Transformer decoder is trained with a multi-token prediction (MTP) objective that factorizes the probability of the next item's ID tokens as a product of per-codebook marginals. At inference, RPG computes per-token logits in parallel and decodes a top-K list using graph-constrained beam propagation over a graph that connects semantically similar IDs. Experiments on four Amazon categories compare RPG with item-ID and semantic-ID baselines, report ablations, analyze scaling with ID length, and measure inference memory and time as the item pool grows.
Significance. If the central claims hold, RPG is a useful contribution to generative semantic-ID recommendation: it replaces autoregressive token-by-token beam search with a single encoder pass and parallel token scoring, enabling much longer IDs (16-64 tokens) than the 4-token IDs used by TIGER. The paper includes informative ablations (random vs. OPQ IDs, RQ IDs, projection-head variants, with/without graph constraints), an efficiency comparison with dummy items, and a cold-start analysis, and the code is released. The main caveats are that the conditional-independence factorization in Eq. (1) is asserted rather than validated, that the abstract overstates the role of m=64, and that the graph decoder's approximation quality relative to exhaustive scoring is never quantified.
major comments (4)
- [Sec. 2.2.1, Eq. (1)] The factorization P(c_{t,1},...,c_{t,m}|s) = \prod_j P^{(j)}(c_{t,j}|s) is the load-bearing modeling assumption for the parallel design, but the paper does not validate it. Figure 2 only shows that the factorized score from Eq. (3) changes monotonically with the number of differing digits; since that score is a sum of per-token marginal logits, the figure is nearly a restatement of the score definition and does not establish that the product approximates the true conditional distribution over token combinations. If user preferences induce correlations across OPQ subvector tokens (e.g., a user who prefers running shoes should jointly elevate both the 'sport' and 'shoe-type' tokens), the MTP model is optimizing a misspecified objective and rankings from Eq. (3) can be biased. Please provide a direct validation, for example by comparing Eq. (3) scores with an autoregressive or jointly trained model on the same OPQ IDs, or by measuring inter-token dependencies in the learned conditional distribution.
- [Abstract and Sec. 3.5.1 (Table 6, Fig. 4)] The abstract states that scaling semantic ID length to 64 enables RPG to outperform generative baselines by an average of 12.6% on NDCG@10, but Table 6 reports optimal semantic ID lengths of 16, 32, 16, and 64 for Sports, Beauty, Toys, and CDs, respectively, and Figure 4 shows performance saturating or declining for the smaller datasets. The reported 12.6% average gain is therefore obtained under per-dataset optimal lengths, not under m=64 uniformly. Please rephrase the abstract and Section 3.2 to describe the per-dataset optimal lengths and the average gain under those settings, and soften the claim that m=64 is the source of the improvement for all datasets.
- [Sec. 2.3.3 and Table 6] The paper claims that graph-constrained decoding has time and memory complexity independent of catalog size, but Table 6 shows that RPG visits 10.90%, 24.79%, 20.97%, and 23.28% of the item pool on the four datasets under the best hyperparameters. Because the decoder is approximate, the paper should report the NDCG@10 that would be obtained by exhaustive scoring of all candidate items using Eq. (3). At the catalog sizes used here exhaustive scoring is feasible and would give an upper bound on the model's ranking quality; without it, the reader cannot tell how much of RPG's effectiveness comes from the learned scores versus from the graph-propagation heuristic, and the efficiency/effectiveness trade-off cannot be properly assessed.
- [Sec. 3.3 and Table 8] The inference-efficiency comparison with TIGER is informative, but the headline memory/time savings depend on treating graph construction as outside the online inference budget. Table 8 shows that RPG has total storage O(N(m+k)) and that inference fetches only O(bqkm) storage, but the one-time cost of building the decoding graph from pairwise token-embedding similarities is neither reported nor discussed in the complexity analysis. Please state explicitly whether graph construction is included in the experimental memory/time measurements, and quantify its one-time cost, so that the 'independent of the number of items' claim is scoped to online inference rather than the full system.
minor comments (5)
- [Sec. 3.1] The evaluation metrics paragraph contains a typo: 'NDCD@K' should be 'NDCG@K'.
- [Sec. 2.2.1 and 2.2.2] Equation (1) defines L as a loss, but Section 2.2.2 says 'we compute L in Equation (1) as the logits.' Using the same symbol for the loss and the negative log-likelihood score is confusing; please introduce a separate notation for the inference score.
- [Table 4] The row labeled 'sentence-t5-base 16 ~ 64' is ambiguous: it is unclear whether this denotes the best length in a range, a single tuned length, or aggregation over several lengths. Please specify the exact setting.
- [Figure 2] The y-axis label '|logits|' should be expanded to 'absolute difference in factorized sequence scores' and the plot should state which dataset, model checkpoint, and evaluation split are used.
- [Sec. 2.4.2] The statement that RPG reduces sequence encoder forward passes 'from approximately O(bm) to just O(1)' should clarify that the O(1) refers to encoder forward passes only, not to the total decoding operations, which include graph propagation steps of cost O(bqkm).
Circularity Check
No circular derivation: RPG's performance claims rest on external benchmarks and ablations, though one motivating observation is partly a consequence of its additive scoring rule.
full rationale
The paper's central claim (RPG outperforms generative baselines by 12.6% NDCG@10 and is more efficient) is established by external comparisons on four Amazon datasets and by ablations, not by fitting the target metric into the loss. The MTP objective in Eq. (1) is a stated modeling choice, not a hidden restatement of the evaluation. The conditional-independence factorization is explicitly introduced as an assumption (Section 2.2.1) and is not derived from the evaluation target; the paper even reports that RPG underperforms TIGER at 4-digit IDs and explains why its advantage comes from longer IDs, which is an empirical finding rather than a circular construction. The graph-constrained decoding is an approximate search heuristic whose benefit is validated by ablation (Table 3), including a no-graph variant and a beam-search variant; it is not claimed as the source of the target result. The only construction-adjacent element is the observation in Figure 2 that semantic IDs differing in few digits have similar logits, which follows partly from the additive score in Eq. (3); however, the paper uses this only to motivate building a search graph, and the actual recommendation quality is measured against held-out ground truth, so this does not close a derivation loop. Self-citations such as VQ-Rec for OPQ tokenization are methodological and are backed by external OPQ references and by ablations against RQ-based IDs, so they are not load-bearing circular support. Overall, no step reduces a prediction to its own input by construction.
Assumptions & free parameters
free parameters (5)
- semantic ID length m =
Sports 16, Beauty 32, Toys 16, CDs 64
- softmax temperature tau =
0.03 for all datasets
- beam size b =
10 for all datasets
- edges per node k =
Sports 100, Beauty 100, Toys 50, CDs 500
- iteration steps q =
Sports 2, Beauty 3, Toys 5, CDs 3
assumptions (4)
- domain assumption Tokens in the target semantic ID are conditionally independent given the user sequence: P(c_t,1,...,c_t,m | s) = product_j P^(j)(c_t,j | s).
- domain assumption OPQ produces semantic IDs that preserve enough item semantics and are effectively collision-free for long IDs.
- domain assumption The decoding graph, with edge weights equal to sums of token embedding dot products, connects items whose model logits are close, so beam propagation visits the top items after bqk node visits.
- domain assumption The pretrained text encoder provides a semantic representation suitable for product-quantized item indexing.
Cite this review
Pith. "Pith review of Generating Long Semantic IDs in Parallel for Recommendation." pith.science (2026). https://pith.science/paper/CHYNKSL6
@misc{pith2026250605781,
author = {Pith},
title = {Pith review of: Generating Long Semantic IDs in Parallel for Recommendation},
year = {2026},
howpublished = {\url{https://pith.science/paper/CHYNKSL6}},
note = {Machine review of arXiv:2506.05781}
}
read the original abstract
Semantic ID-based recommendation models tokenize each item into a small number of discrete tokens that preserve specific semantics, leading to better performance, scalability, and memory efficiency. While recent models adopt a generative approach, they often suffer from inefficient inference due to the reliance on resource-intensive beam search and multiple forward passes through the neural sequence model. As a result, the length of semantic IDs is typically restricted (e.g. to just 4 tokens), limiting their expressiveness. To address these challenges, we propose RPG, a lightweight framework for semantic ID-based recommendation. The key idea is to produce unordered, long semantic IDs, allowing the model to predict all tokens in parallel. We train the model to predict each token independently using a multi-token prediction loss, directly integrating semantics into the learning objective. During inference, we construct a graph connecting similar semantic IDs and guide decoding to avoid generating invalid IDs. Experiments show that scaling up semantic ID length to 64 enables RPG to outperform generative baselines by an average of 12.6% on the NDCG@10, while also improving inference efficiency. Code is available at: https://github.com/facebookresearch/RPG_KDD2025.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
GR-LLMs: Recent Advances in Generative Recommendation Based on Large Language Models
A survey of LLM-based generative recommendation systems, covering application settings, training pipelines, industrial deployment challenges, and future directions.
Reference graph
Works this paper leans on
-
[1]
Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, and Tri Dao. 2024. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.arXiv preprint arXiv:2401.10774(2024)
arXiv 2024
-
[2]
Jianxin Chang, Chen Gao, Yu Zheng, Yiqun Hui, Yanan Niu, Yang Song, Depeng Jin, and Yong Li. 2021. Sequential Recommendation with Graph Neural Networks. InSIGIR
work page 2021
-
[3]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. 2020. A Simple Framework for Contrastive Learning of Visual Representations. In ICML
work page 2020
-
[4]
Jeffrey Dean and Sanjay Ghemawat. 2008. MapReduce: simplified data processing on large clusters.Commun. ACM51, 1 (2008), 107–113
2008
-
[5]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL
work page 2019
-
[6]
Yijie Ding, Yupeng Hou, Jiacheng Li, and Julian McAuley. 2024. Inductive Generative Recommendation via Retrieval-based Speculation.arXiv preprint arXiv:2410.02939(2024)
arXiv 2024
-
[7]
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. 2024. The faiss library.arXiv preprint arXiv:2401.08281(2024)
arXiv 2024
-
[8]
Tiezheng Ge, Kaiming He, Qifa Ke, and Jian Sun. 2014. Optimized Product Quantization.IEEE Trans. Pattern Anal. Mach. Intell.36, 4 (2014), 744–755
2014
Show all 73 references
-
[9]
Shijie Geng, Shuchang Liu, Zuohui Fu, Yingqiang Ge, and Yongfeng Zhang. 2022. Recommendation as Language Processing (RLP): A Unified Pretrain, Personalized Prompt & Predict Paradigm (P5). InRecSys
2022
-
[10]
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, and Gabriel Synnaeve. 2024. Better & faster large language models via multi-token prediction.arXiv preprint arXiv:2404.19737(2024)
2024 arXiv
-
[11]
Alex Graves. 2012. Sequence transduction with recurrent neural networks.arXiv preprint arXiv:1211.3711(2012)
2012 arXiv
-
[12]
Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk
-
[13]
Chris Hokamp and Qun Liu. 2017. Lexically Constrained Decoding for Sequence Generation Using Grid Beam Search. InACL. 1535–1546
2017
-
[14]
Yupeng Hou, Zhankui He, Julian McAuley, and Wayne Xin Zhao. 2023. Learning vector-quantized item representation for transferable sequential recommenders. InWWW. 1162–1171
2023
-
[15]
Yupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li, Bolin Ding, and Ji-Rong Wen. 2022. Towards universal sequence representation learning for recommender systems. InKDD. 585–593
2022
-
[16]
Chi, Julian McAuley, and Derek Zhiyuan Cheng
Yupeng Hou, Jianmo Ni, Zhankui He, Noveen Sachdeva, Wang-Cheng Kang, Ed H. Chi, Julian McAuley, and Derek Zhiyuan Cheng. 2025. ActionPiece: Contextually Tokenizing Action Sequences for Generative Recommendation. InICML
2025
-
[17]
Yupeng Hou, An Zhang, Leheng Sheng, Zhengyi Yang, Xiang Wang, Tat-Seng Chua, and Julian McAuley. 2025. Generative Recommendation Models: Progress and Directions. InCompanion Proceedings of the ACM on Web Conference 2025. 13–16
2025
-
[18]
Wenyue Hua, Shuyuan Xu, Yingqiang Ge, and Yongfeng Zhang. 2023. How to Index Item IDs for Recommendation Foundation Models. InSIGIR-AP
2023
-
[19]
Hervé Jégou, Matthijs Douze, and Cordelia Schmid. 2011. Product Quantization for Nearest Neighbor Search.IEEE Trans. Pattern Anal. Mach. Intell.33, 1 (2011), 117–128
2011
-
[20]
Bowen Jin, Hansi Zeng, Guoyin Wang, Xiusi Chen, Tianxin Wei, Ruirui Li, Zhengyang Wang, Zheng Li, Yang Li, Hanqing Lu, Suhang Wang, Jiawei Han, and Xianfeng Tang. 2024. Language Models As Semantic Indexers. InICML
2024
-
[21]
Wang-Cheng Kang and Julian J. McAuley. 2018. Self-Attentive Sequential Rec- ommendation. InICDM
2018
-
[22]
Guanghan Li, Xun Zhang, Yufei Zhang, Yifan Yin, Guojun Yin, and Wei Lin. 2025. Semantic Convergence: Harmonizing Recommender Systems via Two-Stage Alignment and Behavioral Semantic Tokenization. InAAAI
2025
-
[23]
Jing Li, Pengjie Ren, Zhumin Chen, Zhaochun Ren, Tao Lian, and Jun Ma. 2017. Neural Attentive Session-based Recommendation. InCIKM
2017
-
[24]
Jiacheng Li, Ming Wang, Jin Li, Jinmiao Fu, Xin Shen, Jingbo Shang, and Julian McAuley. 2023. Text is all you need: Learning language representations for sequential recommendation. InKDD
2023
-
[25]
Xiaoxi Li, Jiajie Jin, Yujia Zhou, Yuyao Zhang, Peitian Zhang, Yutao Zhu, and Zhicheng Dou. 2024. From Matching to Generation: A Survey on Generative Information Retrieval.arXiv preprint arXiv:2404.14851(2024)
2024 arXiv
-
[26]
Yongqi Li, Xinyu Lin, Wenjie Wang, Fuli Feng, Liang Pang, Wenjie Li, Liqiang Nie, Xiangnan He, and Tat-Seng Chua. 2024. A Survey of Generative Search and Recommendation in the Era of Large Language Models.arXiv preprint arXiv:2404.16924(2024)
2024 arXiv
-
[27]
Xinyu Lin, Haihan Shi, Wenjie Wang, Fuli Feng, Qifan Wang, See-Kiong Ng, and Tat-Seng Chua. 2025. Order-agnostic Identifier for Large Language Model-based Generative Recommendation. InSIGIR
2025
-
[28]
Xinyu Lin, Chaoqun Yang, Wenjie Wang, Yongqi Li, Cunxiao Du, Fuli Feng, See- Kiong Ng, and Tat-Seng Chua. 2024. Efficient Inference for Large Language Model- based Generative Recommendation.arXiv preprint arXiv:2410.05165(2024)
2024 arXiv
-
[29]
Enze Liu, Bowen Zheng, Cheng Ling, Lantao Hu, Han Li, and Wayne Xin Zhao
-
[30]
Han Liu, Yinwei Wei, Xuemeng Song, Weili Guan, Yuan-Fang Li, and Liqiang Nie. 2024. MMGRec: Multimodal Generative Recommendation with Transformer Model.arXiv preprint arXiv:2404.16555(2024)
2024 arXiv
-
[31]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach.arXiv:1907.11692(2019)
2019 arXiv
-
[32]
Zihan Liu, Yupeng Hou, and Julian McAuley. 2024. Multi-Behavior Generative Recommendation. InCIKM
2024
-
[33]
Chen Ma, Peng Kang, and Xue Liu. 2019. Hierarchical Gating Networks for Sequential Recommendation. InKDD
2019
-
[34]
McAuley, Christopher Targett, Qinfeng Shi, and Anton van den Hengel
Julian J. McAuley, Christopher Targett, Qinfeng Shi, and Anton van den Hengel
-
[35]
Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith Hall, Daniel Cer, and Yinfei Yang. 2022. Sentence-T5: Scalable Sentence Encoders from Pre- trained Text-to-Text Models. InFindings of ACL
2022
-
[36]
Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, et al . 2022. Large Dual Encoders Are Generalizable Retrievers. InEMNLP. 9844–9855
2022
-
[37]
Aleksandr V Petrov and Craig Macdonald. 2023. Generative sequential recom- mendation with gptrec.arXiv preprint arXiv:2306.11114(2023)
2023 arXiv
-
[38]
Aleksandr V Petrov and Craig Macdonald. 2024. RecJPQ: training large-catalogue sequential recommenders. InProceedings of the 17th ACM International Conference on Web Search and Data Mining. 538–547
2024
-
[39]
Tran, Jonah Samost, Maciej Kula, Ed H
Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Q. Tran, Jonah Samost, Maciej Kula, Ed H. Chi, and Maheswaran Sathiamoorthy. 2023. Recommender Systems with Generative Retrieval. InNeurIPS
2023
-
[40]
Steffen Rendle, Christoph Freudenthaler, and Lars Schmidt-Thieme. 2010. Factor- izing personalized Markov chains for next-basket recommendation. InWWW
2010
-
[41]
Zihua Si, Zhongxiang Sun, Jiale Chen, Guozhang Chen, Xiaoxue Zang, Kai Zheng, Yang Song, Xiao Zhang, Jun Xu, and Kun Gai. 2024. Generative retrieval with semantic tree-structured item identifiers via contrastive learning. InSIGIR-AP
2024
-
[42]
Chi, and Xinyang Yi
Anima Singh, Trung Vu, Nikhil Mehta, Raghunandan Keshavan, Maheswaran Sathiamoorthy, Yilin Zheng, Lichan Hong, Lukasz Heldt, Li Wei, Devansh Tandon, Ed H. Chi, and Xinyang Yi. 2024. Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations. InRecSys
2024
-
[43]
Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang
-
[44]
Juntao Tan, Shuyuan Xu, Wenyue Hua, Yingqiang Ge, Zelong Li, and Yongfeng Zhang. 2024. Idgenrec: Llm-recsys alignment with textual id learning. InSIGIR
2024
-
[45]
Jiaxi Tang and Ke Wang. 2018. Personalized Top-N Sequential Recommendation via Convolutional Sequence Embedding. InWSDM
2018
-
[46]
Cohen, and Donald Metzler
Yi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Prakash Gupta, Tal Schuster, William W. Cohen, and Donald Metzler. 2022. Transformer memory as a differentiable search index. In NeurIPS
2022
-
[47]
Wenjie Wang, Honghui Bao, Xinyu Lin, Jizhi Zhang, Yongqi Li, Fuli Feng, See- Kiong Ng, and Tat-Seng Chua. 2024. Learnable Tokenizer for LLM-based Genera- tive Recommendation. InCIKM
2024
-
[48]
Yidan Wang, Zhaochun Ren, Weiwei Sun, Jiyuan Yang, Zhixiang Liang, Xin Chen, Ruobing Xie, Su Yan, Xu Zhang, Pengjie Ren, Zhumin Chen, and Xin Xin
-
[49]
Ye Wang, Jiahao Xun, Minjie Hong, Jieming Zhu, Tao Jin, Wang Lin, Haoyuan Li, Linjun Li, Yan Xia, Zhou Zhao, and Zhenhua Dong. 2024. EAGER: Two- Stream Generative Recommender with Behavior-Semantic Collaboration. In KDD. 3245–3254
2024
-
[50]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020
-
[51]
Shu Wu, Yuyuan Tang, Yanqiao Zhu, Liang Wang, Xing Xie, and Tieniu Tan. 2019. Session-Based Recommendation with Graph Neural Networks. InAAAI
2019
-
[52]
Liu Yang, Fabian Paischer, Kaveh Hassani, Jiacheng Li, Shuai Shao, Zhang Gabriel Li, Yun He, Xue Feng, Nima Noorshams, Sem Park, Bo Long, Robert D Nowak, Xiaoli Gao, and Hamid Eghbalzadeh. 2024. Unifying Generative and Dense Retrieval for Sequential Recommendation.arXiv prepri...
2024 arXiv
-
[53]
Enhanced Generative Recommendation via Content and Collaboration Integration.arXiv preprint arXiv:2403.18480(2024)
2024 arXiv
-
[54]
Zhenrui Yue, Yueqi Wang, Zhankui He, Huimin Zeng, Julian McAuley, and Dong Wang. 2024. Linear recurrent units for sequential recommendation. InWSDM. 930–938
2024
-
[55]
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. 2022. SoundStream: An End-to-End Neural Audio Codec.IEEE ACM Trans. Audio Speech Lang. Process.30 (2022), 495–507
2022
-
[56]
Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhao- jie Gong, Fangda Gu, Michael He, Yinghai Lu, and Yu Shi. 2024. Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations. InICML
2024
-
[57]
Taolin Zhang, Junwei Pan, Jinpeng Wang, Yaohua Zha, Tao Dai, Bin Chen, Ruisheng Luo, Xiaoxiang Deng, Yuan Wang, Ming Yue, et al . 2024. To- wards Scalable Semantic Representation for Recommendation.arXiv preprint arXiv:2410.09560(2024)
2024 arXiv
-
[58]
Zheng Yuan, Fajie Yuan, Yu Song, Youhua Li, Junchen Fu, Fei Yang, Yunzhu Pan, and Yongxin Ni. 2023. Where to go next for recommender systems? id-vs. KDD ’25, August 3–7, 2025, Toronto, ON, Canada Yupeng Hou et al. Table 5: Notations and explanations. Notation Explaination 𝑖1 T...
2023
-
[59]
Wayne Xin Zhao, Zihan Lin, Zhichao Feng, Pengfei Wang, and Ji-Rong Wen. 2022. A revisiting study of appropriate offline evaluation for top-N recommendation algorithms.ACM Transactions on Information Systems41, 2 (2022), 1–41
2022
-
[60]
Wayne Xin Zhao, Shanlei Mu, Yupeng Hou, Zihan Lin, Yushuo Chen, Xingyu Pan, Kaiyuan Li, Yujie Lu, Hui Wang, Changxin Tian, Yingqian Min, Zhichao Feng, Xinyan Fan, Xu Chen, Pengfei Wang, Wendi Ji, Yaliang Li, Xiaoling Wang, and Ji-Rong Wen. 2021. RecBole: Towards a Unified, Com...
2021
-
[61]
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong W...
2023 arXiv
-
[62]
Bowen Zheng, Yupeng Hou, Hongyu Lu, Yu Chen, Wayne Xin Zhao, and Ji- Rong Wen. 2024. Adapting Large Language Models by Integrating Collaborative Semantics for Recommendation. InICDE
2024
-
[63]
Sheng, Jiajie Xu, De- qing Wang, Guanfeng Liu, and Xiaofang Zhou
Tingting Zhang, Pengpeng Zhao, Yanchi Liu, Victor S. Sheng, Jiajie Xu, De- qing Wang, Guanfeng Liu, and Xiaofang Zhou. 2019. Feature-level Deeper Self-Attention Network for Sequential Recommendation. InIJCAI
2019
-
[64]
Bowen Zheng, Hongyu Lu, Yu Chen, Wayne Xin Zhao, and Ji-Rong Wen. 2025. Universal Item Tokenization for Transferable Generative Recommendation.arXiv preprint arXiv:2504.04405(2025)
2025 arXiv
-
[65]
Bowen Zheng, Junjie Zhang, Hongyu Lu, Yu Chen, Ming Chen, Wayne Xin Zhao, and Ji-Rong Wen. 2024. Enhancing Graph Contrastive Learning with Reliable and Informative Augmentation for Recommendation.arXiv preprint arXiv:2409.05633 (2024)
2024 arXiv
-
[66]
Kun Zhou, Hui Wang, Wayne Xin Zhao, Yutao Zhu, Sirui Wang, Fuzheng Zhang, Zhongyuan Wang, and Ji-Rong Wen. 2020. S3-Rec: Self-Supervised Learning for Sequential Recommendation with Mutual Information Maximization. InCIKM
2020
-
[67]
Jieming Zhu, Mengqun Jin, Qijiong Liu, Zexuan Qiu, Zhenhua Dong, and Xiu Li
-
[68]
Bowen Zheng, Enze Liu, Zhongfu Chen, Zhongrui Ma, Yue Wang, Wayne Xin Zhao, and Ji-Rong Wen. 2025. Pre-training Generative Recommender with Multi- Identifier Item Tokenization.arXiv preprint arXiv:2504.04400(2025)
2025 arXiv
-
[73]
learning rate
CoST: Contrastive Quantization based Semantic Tokenization for Genera- tive Recommendation. InRecSys. Appendices A Notations We summarize the notations used throughout the paper in Table 5. B Additional Implementation Details All experiments were conducted on a single NVIDIA R...
2025
-
[2015]
Image-Based Recommendations on Styles and Substitutes. InSIGIR
-
[2016]
Session-based Recommendations with Recurrent Neural Networks. In ICLR
-
[2019]
BERT4Rec: Sequential Recommendation with Bidirectional Encoder Repre- sentations from Transformer. InCIKM
-
[2024]
arXiv preprint arXiv:2409.05546(2024)
End-to-End Learnable Item Tokenization for Generative Recommendation. arXiv preprint arXiv:2409.05546(2024)
2024 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.