REVIEW 3 major objections 4 minor 1 cited by
Empowering Large Language Model for Sequential Recommendation via Multimodal Embeddings and Semantic IDs
T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read MME-SID claims that feeding an LLM both multimodal embeddings and semantic-ID code embeddings initialized from a trained multimodal RQ-VAE fixes embedding collapse and catastrophic forgetting in LLM-based sequential recommendation, with con
desk verdict A solid empirical systems paper with a genuinely new combination of multimodal embeddings and semantic IDs for LLM-based sequential recommendation, but the theoretical and diagnostic narrative is over-claimed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is MM-RQ-VAE, a per-modality residual-quantized variational autoencoder trained with a characteristic-kernel maximum mean discrepancy reconstruction loss and a contrastive loss that aligns collaborative codes with textual and visual codes. Its trained code embeddings become the initial embeddings of the semantic-ID tokens in the LLM, which is fine-tuned with LoRA; item representations enter as a concatenation of projected raw embeddings and summed code embeddings (Eq. 16), and the final score is a frequency-aware fusion over modalities (Eq. 18).
What would settle it
Replace the MMD reconstruction loss in MM-RQ-VAE with a loss that directly enforces the pairwise distance ordering (maximizing Kendall's tau) while keeping all other components fixed; if recommendation performance does not improve alongside tau, the distance-preservation mechanism is not the causal driver. Alternatively, sweep the unspecified Gaussian kernel bandwidth; if performance collapses across a range of bandwidths, the reported MMD advantage is not well-defined.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the two pathologies can be addressed simultaneously: treating semantic IDs as symbols to be re-learned from scratch is the wrong move, and treating pre-trained collaborative embeddings as the only vector input is also the wrong move. The right design is to give the LLM both kinds of signal—multimodal embeddings and quantized-semantic-ID embeddings—and to preserve the geometry learned during quantization by initializing the downstream ID embeddings with the trained code embeddings, while using an MMD-based reconstruction loss that keeps the quantized embeddings' pairwise distance structure close to the originals. With this recipe, MME-SI
Load-bearing premise
The explanation of why the method works rests on Kendall's tau between pairwise distance orderings being a faithful measure of how much information an embedding retains; if that proxy is not meaningful, the forgetting-and-preservation narrative loses support even if the ranking results stand.
Editorial extensions
If this is right
- Semantic-ID methods for generative retrieval should keep and reuse trained code embeddings instead of discarding them, since re-initialization is what the paper identifies as the cause of catastrophic forgetting.
- MMD as a reconstruction loss for vector quantization appears to preserve more pairwise distance information than MSE, so any quantization-based tokenizer could adopt it.
- Combining low-dimensional collaborative embeddings with high-dimensional multimodal and quantized embeddings expands the effective rank of the LLM input, reducing dimensional collapse without enlarging the collaborative embedding table.
- The frequency-aware fusion implies that treating all target items uniformly is suboptimal; head and tail items can benefit from different modality weighting.
- Fine-tuning only about 0.19% of parameters via LoRA suffices to reach the reported accuracy, which matters for industrial-scale item sets.
Reading between the lines
- The paper's Kendall-tau measure of forgetting is a proxy, not a guarantee: preserving the ordering of Euclidean distances between behavioral and target items may not be the actual driver of recommendation quality, and a direct ordinal-preserving loss could test this.
- The code-embedding-initialization trick may transfer to other tokenization-heavy tasks, such as generative retrieval over documents or images, where a quantization stage is followed by downstream fine-tuning.
- The Gaussian kernel bandwidth for the MMD loss is not specified in the paper, so the reconstruction signal is underdetermined; performance may be sensitive to it, and a sweep would clarify whether the MMD advantage is robust or bandwidth-dependent.
- If the claims hold, LLM-based recommenders could drop the autoregressive, code-by-code retrieval used by TIGER-style models, scoring all target items in a single forward pass and cutting inference cost dramatically.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MME-SID, a two-stage framework for LLM-based sequential recommendation. In the encoding stage, collaborative, textual, and visual item embeddings are obtained (from SASRec and LLM2CLIP) and quantized by a multimodal RQ-VAE (MM-RQ-VAE) that uses maximum mean discrepancy as the reconstruction loss and contrastive learning for cross-modal alignment. In the fine-tuning stage, Llama3-8B-instruct is tuned with LoRA; the item representation concatenates projected original embeddings with semantic-ID embeddings initialized from the trained code embeddings, and a frequency-aware fusion module weights the target-item scores. Experiments on Amazon Beauty, Toys & Games, and Sports & Outdoors show consistent improvements over baselines, with reported statistical significance. The paper interprets these gains as evidence that MME-SID alleviates embedding collapse and catastrophic forgetting.
Significance. If the mechanism claims are established, this is a valuable contribution to LLM4SR: it identifies two practically important failure modes and proposes a concrete integration of multimodal embeddings and semantic IDs. The empirical results are strong and consistent across three datasets, and the authors commit to releasing code and data, which supports reproducibility. The main significance is currently conditional because the paper's explanatory narrative—that MMD preserves per-item distance information and that Kendall's tau measures information retention—is not rigorously supported. With those points addressed, the paper could be an influential systems contribution.
major comments (3)
- [§4.2.2, Eq. (12)] The reconstruction loss L_Recon = Σ_b Σ_j MMD²_K(SG(s^b_j), \s^b_j) is a batch-level distributional discrepancy. It is invariant under permutation of the reconstructed items within a batch, so it does not constrain the correspondence between a specific original item embedding and its reconstruction. The paper's mechanism for mitigating forgetting is that the quantized embedding preserves the partial order of distances between behavioral and target items (Sec. 3.2, Sec. 5.4), which is a per-item, identity-dependent property. The reported tau gain (0.4436 vs 0.3714) is therefore not explained by this loss. Moreover, with MMD as the only reconstruction signal, the decoder is not trained to reconstruct each item's embedding. Please add a per-item alignment analysis or modify the loss, and revise the causal claim accordingly.
- [§3.2 and §5.5] Kendall's tau is used as a measure of "information preserved", and tau=0.0550 is translated into "94.50% of the previously learned information is forgotten". Tau is a rank correlation, not an information measure; it can be negative and has no interpretation as a percentage of information content. Furthermore, for a user with two behavioral items the distance variable has only two elements, so the per-user tau is trivially 1, -1, or 0. The aggregate tau may be sensitive to the sampling of users/pairs. Since RQ3 and RQ4 conclusions rest on tau comparisons, please justify this proxy, report its distribution/uncertainty, or soften the information-retention claims.
- [§5.1.4 / Appendix B] The Gaussian kernel in the MMD loss is written as k(e,e')=exp(-||e-e'||²/2σ²), but the value (or selection procedure) of σ is never reported. MMD with a Gaussian kernel is highly sensitive to bandwidth; without σ the reconstruction objective in Eq. (12) is underdetermined and the MMD-vs-MSE comparison in Sec. 5.4 cannot be reproduced or assessed. Please report σ, and preferably include a sensitivity analysis over σ.
minor comments (4)
- [§3.1, Eq. (7)-(8)] The statement rank(A+B) < rank(A)+rank(B) is not always strict; the correct inequality is ≤. Replacing < with ≤ still yields the intended bound rank(WE_c+b) ≤ rank(E_c)+1, so the collapse conclusion survives, but the formula as written is false.
- [Fig. 2 caption] The red dashed arrow is described twice; one description should refer to the alignment arrow and the other to the reconstruction path.
- [Tab. 3] Although three runs are averaged and significance is reported, no standard deviations are given. Reporting them would help assess stability.
- [Contributions] The claim of being "the first work" to identify these issues is difficult to verify and may be unnecessarily strong; consider softening the wording.
Circularity Check
No significant circularity: the central claims are supported by held-out experiments and the proposed losses are not by-construction equal to the metrics used to evaluate them.
full rationale
The paper is an empirical systems paper. Its main claim—that MME-SID improves sequential recommendation—is evaluated on held-out test sets against multiple baselines on three public datasets. The supporting analyses for embedding collapse and catastrophic forgetting are post-hoc measurements, not quantities directly optimized by the training objective. The MMD reconstruction loss (Eq. 12) is a distributional discrepancy, not Kendall's tau or a per-item distance-order loss; therefore the observed higher tau for MMD-trained embeddings is an empirical result, not a tautology. The rank argument in Eq. 8 is a direct linear-algebra consequence, not an imported uniqueness theorem. The paper cites prior work by the same authors for the concepts of embedding collapse, but the paper also provides its own derivation, so those citations are not load-bearing. The frequency-aware fusion and code-embedding initialization are architectural choices whose effects are tested via ablations on held-out data. No fitted parameter is renamed as a prediction, and no self-citation chain forces the conclusion. The reported performance improvements are therefore independent of the paper's explanatory narrative, even if that narrative's proxy metrics are debatable.
Assumptions & free parameters
free parameters (6)
- Gaussian kernel bandwidth sigma for MMD =
not specified in the paper
- Alignment loss weight beta =
1e-3
- RQ-VAE loss weight gamma =
1
- Contrastive temperature epsilon =
not specified
- Codebook size S =
256 (300 for Sports & Outdoors)
- Number of codebook levels L =
4
assumptions (5)
- standard math rank(W E_c + b) <= rank(E_c) + 1 for any linear projection with bias
- standard math A characteristic kernel preserves all statistics of a distribution in MMD
- ad hoc to paper Kendall's tau between behavioral-target distance variables measures information retention
- domain assumption LLM2CLIP maps textual and visual embeddings into the same representational space
- domain assumption Pre-trained SASRec collaborative embeddings capture useful collaborative signal
invented entities (1)
-
Multimodal semantic ID code embeddings (from MM-RQ-VAE)
Cite this review
Pith. "Pith review of Empowering Large Language Model for Sequential Recommendation via Multimodal Embeddings and Semantic IDs." pith.science (2026). https://pith.science/paper/GO7QXKRY
@misc{pith2026250902017,
author = {Pith},
title = {Pith review of: Empowering Large Language Model for Sequential Recommendation via Multimodal Embeddings and Semantic IDs},
year = {2026},
howpublished = {\url{https://pith.science/paper/GO7QXKRY}},
note = {Machine review of arXiv:2509.02017}
}
read the original abstract
Sequential recommendation (SR) aims to capture users' dynamic interests and sequential patterns based on their historical interactions. Recently, the powerful capabilities of large language models (LLMs) have driven their adoption in SR. However, we identify two critical challenges in existing LLM-based SR methods: 1) embedding collapse when incorporating pre-trained collaborative embeddings and 2) catastrophic forgetting of quantized embeddings when utilizing semantic IDs. These issues dampen the model scalability and lead to suboptimal recommendation performance. Therefore, based on LLMs like Llama3-8B-instruct, we introduce a novel SR framework named MME-SID, which integrates multimodal embeddings and quantized embeddings to mitigate embedding collapse. Additionally, we propose a Multimodal Residual Quantized Variational Autoencoder (MM-RQ-VAE) with maximum mean discrepancy as the reconstruction loss and contrastive learning for alignment, which effectively preserve intra-modal distance information and capture inter-modal correlations, respectively. To further alleviate catastrophic forgetting, we initialize the model with the trained multimodal code embeddings. Finally, we fine-tune the LLM efficiently using LoRA in a multimodal frequency-aware fusion manner. Extensive experiments on three public datasets validate the superior performance of MME-SID thanks to its capability to mitigate embedding collapse and catastrophic forgetting. The implementation code and datasets are publicly available for reproduction: https://github.com/Applied-Machine-Learning-Lab/MME-SID.
Figures
Forward citations
Cited by 1 Pith paper
-
The Best of the Two Worlds: Harmonizing Semantic and Hash IDs for Sequential Recommendation
A dual-branch recommender that merges hash-ID and semantic-ID representations outperforms baselines while improving tail-item accuracy without losing head-item accuracy.
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
-
[2]
Keqin Bao, Jizhi Zhang, Yang Zhang, Wenjie Wang, Fuli Feng, and Xiangnan He. 2023. Tallrec: An effective and efficient tuning framework to align large language model with recommendation. InProceedings of the 17th ACM Conference on Recommender Systems. 1007–1014
2023
-
[3]
Jingtong Gao, Xiangyu Zhao, Muyang Li, Minghao Zhao, Runze Wu, Ruocheng Guo, Yiding Liu, and Dawei Yin. 2024. Smlp4rec: An efficient all-mlp architecture for sequential recommendations. ACM Transactions on Information Systems 42, 3 (2024), 1–23
work page 2024
-
[4]
Xingzhuo Guo, Junwei Pan, Ximei Wang, Baixu Chen, Jie Jiang, and Mingsheng Long. 2024. On the Embedding Collapse when Scaling up Recommendation Models. In Proceedings of the 41st International Conference on Machine Learning
work page 2024
-
[5]
Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recom- mendation. In 2018 IEEE international conference on data mining (ICDM) . IEEE, 197–206
2018
-
[6]
Maurice G Kendall. 1938. A new measure of rank correlation. Biometrika 30, 1-2 (1938), 81–93
work page 1938
-
[7]
Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han. 2022. Autoregressive image generation using residual quantization. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 11523–11532
2022
-
[8]
Chengxi Li, Yejing Wang, Qidong Liu, Xiangyu Zhao, Wanyu Wang, Yiqi Wang, Lixin Zou, Wenqi Fan, and Qing Li. 2023. STRec: Sparse transformer for sequential recommendations. In Proceedings of the 17th ACM conference on recommender systems. 101–111
work page 2023
Show all 73 references
-
[9]
Xiangyang Li, Bo Chen, Lu Hou, and Ruiming Tang. 2023. CTRL: Connect Collaborative and Language Model for CTR Prediction. ACM Transactions on Recommender Systems (2023)
2023
-
[10]
Xinhang Li, Chong Chen, Xiangyu Zhao, Yong Zhang, and Chunxiao Xing. 2023. E4srec: An elegant effective efficient extensible solution of large language models for sequential recommendation. arXiv preprint arXiv:2312.02443 (2023)
2023 arXiv
-
[11]
Xiaopeng Li, Fan Yan, Xiangyu Zhao, Yichao Wang, Bo Chen, Huifeng Guo, and Ruiming Tang. 2023. Hamur: Hyper adapter for multi-domain recommendation. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management. 1268–1277
2023
-
[12]
Youhua Li, Hanwen Du, Yongxin Ni, Yuanqi He, Junchen Fu, Xiangyan Liu, and Qi Guo. 2024. An Empirical Study of Training ID-Agnostic Multi-modal Sequential Recommenders. arXiv preprint arXiv:2403.17372 (2024)
2024 arXiv
-
[13]
Jiayi Liao, Sihang Li, Zhengyi Yang, Jiancan Wu, Yancheng Yuan, Xiang Wang, and Xiangnan He. 2024. Llara: Large language-recommendation assistant. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 1785–1795
2024
-
[14]
Weilin Lin, Xiangyu Zhao, Yejing Wang, Yuanshao Zhu, and Wanyu Wang. 2023. Autodenoise: Automatic data instance denoising for recommendations. In Pro- ceedings of the ACM Web Conference 2023 . 1003–1011
2023
-
[15]
Zhutian Lin, Junwei Pan, Haibin Yu, Xi Xiao, Ximei Wang, Zhixiang Feng, Shifeng Wen, Shudong Huang, Lei Xiao, and Jie Jiang. 2024. Disentangled Representation with Cross Experts Covariance Loss for Multi-Domain Recommendation. arXiv preprint arXiv:2405.12706 (2024)
2024 arXiv
-
[16]
Langming Liu, Liu Cai, Chi Zhang, Xiangyu Zhao, Jingtong Gao, Wanyu Wang, Yifu Lv, Wenqi Fan, Yiqi Wang, Ming He, et al. 2023. Linrec: Linear attention mechanism for long-term sequential recommender systems. In Proceedings of the 46th International ACM SIGIR Conference on Rese...
2023
-
[17]
Qidong Liu, Jiaxi Hu, Yutian Xiao, Xiangyu Zhao, Jingtong Gao, Wanyu Wang, Qing Li, and Jiliang Tang. 2024. Multimodal recommender systems: A survey. Comput. Surveys 57, 2 (2024), 1–17
2024
-
[18]
Qidong Liu, Xian Wu, Wanyu Wang, Yejing Wang, Yuanshao Zhu, Xiangyu Zhao, Feng Tian, and Yefeng Zheng. 2025. Llmemb: Large language model can be a good embedding generator for sequential recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 39...
2025
-
[19]
Qidong Liu, Xian Wu, Yejing Wang, Zijian Zhang, Feng Tian, Yefeng Zheng, and Xiangyu Zhao. 2024. Llm-esr: Large language models enhancement for long- tailed sequential recommendation. Advances in Neural Information Processing Systems 37 (2024), 26701–26727
2024
-
[20]
Qidong Liu, Xian Wu, Xiangyu Zhao, Yejing Wang, Zijian Zhang, Feng Tian, and Yefeng Zheng. 2024. Large language models enhanced sequential recommenda- tion for long-tail user and item. arXiv e-prints (2024), arXiv–2405
2024
-
[21]
Qidong Liu, Xian Wu, Xiangyu Zhao, Yuanshao Zhu, Zijian Zhang, Feng Tian, and Yefeng Zheng. 2024. Large language model distilling medication recommendation model. arXiv preprint arXiv:2402.02803 (2024)
2024 arXiv
-
[22]
Qidong Liu, Xiangyu Zhao, Yejing Wang, Zijian Zhang, Howard Zhong, Chong Chen, Xiang Li, Wei Huang, and Feng Tian. 2025. Bridge the Domains: Large Lan- guage Models Enhanced Cross-domain Sequential Recommendation. In Proceed- ings of the 48th International ACM SIGIR Conference...
2025
-
[23]
Shuchang Liu, Qingpeng Cai, Bowen Sun, Yuhao Wang, Ji Jiang, Dong Zheng, Peng Jiang, Kun Gai, Xiangyu Zhao, and Yongfeng Zhang. 2023. Exploration and regularization of the latent action space in recommendation. In Proceedings of the ACM Web Conference 2023. 833–844
2023
-
[24]
Yifan Liu, Kangning Zhang, Xiangyuan Ren, Yanhua Huang, Jiarui Jin, Yingjie Qin, Ruilong Su, Ruiwen Xu, Yong Yu, and Weinan Zhang. 2024. AlignRec: Aligning and Training in Multimodal Recommendations. In Proceedings of the 33rd ACM International Conference on Information and Kn...
2024
-
[25]
Ziwei Liu, Qidong Liu, Yejing Wang, Wanyu Wang, Pengyue Jia, Maolin Wang, Zitao Liu, Yi Chang, and Xiangyu Zhao. 2024. Bidirectional gated mamba for sequential recommendation. arXiv e-prints (2024), arXiv–2408
2024
-
[26]
Ziwei Liu, Qidong Liu, Yejing Wang, Wanyu Wang, Pengyue Jia, Maolin Wang, Zitao Liu, Yi Chang, and Xiangyu Zhao. 2025. SIGMA: Selective Gated Mamba for Sequential Recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 12264–12272
2025
-
[27]
Ziru Liu, Shuchang Liu, Zijian Zhang, Qingpeng Cai, Xiangyu Zhao, Kesen Zhao, Lantao Hu, Peng Jiang, and Kun Gai. 2024. Sequential recommendation for optimizing both immediate feedback and long-term retention. In Proceedings of the 47th International ACM SIGIR Conference on Re...
2024
-
[28]
Mingsheng Long, Yue Cao, Jianmin Wang, and Michael Jordan. 2015. Learning transferable features with deep adaptation networks. In International conference on machine learning. PMLR, 97–105
2015
-
[29]
Ilya Loshchilov, Frank Hutter, et al. 2017. Fixing weight decay regularization in adam. arXiv preprint arXiv:1711.05101 (2017)
2017 arXiv
-
[30]
Xinchen Luo, Jiangxia Cao, Tianyu Sun, Jinkai Yu, Rui Huang, Wei Yuan, Hezheng Lin, Yichen Zheng, Shiyao Wang, Qigen Hu, et al . 2024. QARM: Quantita- tive Alignment Multi-Modal Recommendation at Kuaishou. arXiv preprint arXiv:2411.11739 (2024)
2024 arXiv
-
[31]
Julian McAuley, Christopher Targett, Qinfeng Shi, and Anton Van Den Hengel
-
[32]
Junwei Pan, Wei Xue, Ximei Wang, Haibin Yu, Xun Liu, Shijie Quan, Xueming Qiu, Dapeng Liu, Lei Xiao, and Jie Jiang. 2024. Ads recommendation in a collapsed and entangled world. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 5566–5577
2024
-
[33]
Yoon-Joo Park and Alexander Tuzhilin. 2008. The long tail of recommender systems and how to leverage it. In Proceedings of the 2008 ACM conference on Recommender systems. 11–18
2008
-
[34]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[35]
Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Tran, Jonah Samost, et al
-
[36]
Dino Sejdinovic, Bharath Sriperumbudur, Arthur Gretton, and Kenji Fukumizu
-
[37]
Anirvan M Sengupta and Partha P Mitra. 1999. Distributions of singular values for some random matrices. Physical Review E 60, 3 (1999), 3389
1999
-
[38]
Jianhong Shen. 2001. On the singular values of Gaussian random matrices.Linear Algebra Appl. 326, 1-3 (2001), 1–14
2001
-
[39]
Liangcai Su, Junwei Pan, Ximei Wang, Xi Xiao, Shijie Quan, Xihua Chen, and Jie Jiang. 2024. STEM: Unleashing the Power of Embeddings for Multi-task Recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 9002–9010
2024
-
[40]
Weiwei Sun, Lingyong Yan, Zheng Chen, Shuaiqiang Wang, Haichao Zhu, Pengjie Ren, Zhumin Chen, Dawei Yin, Maarten Rijke, and Zhaochun Ren. 2024. Learning to tokenize for generative retrieval. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[41]
Hanbing Wang, Xiaorui Liu, Wenqi Fan, Xiangyu Zhao, Venkataramana Kini, Devendra Yadav, Fei Wang, Zhen Wen, Jiliang Tang, and Hui Liu. 2024. Rethinking large language model architectures for sequential recommendations. arXiv preprint arXiv:2402.09543 (2024)
2024 arXiv
-
[42]
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al
-
[43]
Wenjie Wang, Honghui Bao, Xinyu Lin, Jizhi Zhang, Yongqi Li, Fuli Feng, See- Kiong Ng, and Tat-Seng Chua. 2024. Learnable item tokenization for generative recommendation. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management . 2400–2409
2024
-
[44]
Yuhao Wang. 2024. Multi-Granularity Modeling in Recommendation: from the Multi-Scenario Perspective. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management . 5491–5494
2024
-
[45]
Yejing Wang, Zhaocheng Du, Xiangyu Zhao, Bo Chen, Huifeng Guo, Ruiming Tang, and Zhenhua Dong. 2023. Single-shot feature selection for multi-task recommendations. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval...
2023
-
[46]
Yuhao Wang, Ha Tsz Lam, Yi Wong, Ziru Liu, Xiangyu Zhao, Yichao Wang, Bo Chen, Huifeng Guo, and Ruiming Tang. 2023. Multi-task deep recommender systems: A survey. arXiv preprint arXiv:2302.03525 (2023)
2023 arXiv
-
[47]
Yuhao Wang, Ziru Liu, Yichao Wang, Xiangyu Zhao, Bo Chen, Huifeng Guo, and Ruiming Tang. 2024. Diff-MSR: A diffusion model enhanced paradigm for cold-start multi-scenario recommendation. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining . 779–787
2024
-
[48]
Yuhao Wang, Junwei Pan, Pengyue Jia, Wanyu Wang, Maolin Wang, Zhixiang Feng, Xiaotian Li, Jie Jiang, and Xiangyu Zhao. 2025. Pre-train, Align, and Disentangle: Empowering Sequential Recommendation with Large Language Models. In Proceedings of the 48th International ACM SIGIR C...
2025
-
[49]
Yidan Wang, Zhaochun Ren, Weiwei Sun, Jiyuan Yang, Zhixiang Liang, Xin Chen, Ruobing Xie, Su Yan, Xu Zhang, Pengjie Ren, et al. 2024. Content-Based Collaborative Generation for Recommender Systems. In Proceedings of the 33rd ACM International Conference on Information and Know...
2024
-
[50]
Yuhao Wang, Yichao Wang, Zichuan Fu, Xiangyang Li, Wanyu Wang, Yuyang Ye, Xiangyu Zhao, Huifeng Guo, and Ruiming Tang. 2024. LLM4MSR: An LLM- Enhanced Paradigm for Multi-Scenario Recommendation. In Proceedings of the 33rd ACM International Conference on Information and Knowled...
2024
-
[51]
Yuhao Wang, Xiangyu Zhao, Bo Chen, Qidong Liu, Huifeng Guo, Huanshuo Liu, Yichao Wang, Rui Zhang, and Ruiming Tang. 2023. PLATE: A prompt-enhanced paradigm for multi-scenario recommendations. In Proceedings of the 46th In- ternational ACM SIGIR Conference on Research and Devel...
2023
-
[52]
Aoqi Wu, Yifan Yang, Xufang Luo, Yuqing Yang, Chunyu Wang, Liang Hu, Xiyang Dai, Dongdong Chen, Chong Luo, Lili Qiu, et al . 2024. LLM2CLIP: Powerful Language Model Unlock Richer Visual Representation. InNeurIPS 2024 Workshop: Self-Supervised Learning-Theory and Practice
2024
-
[53]
Zheng Yuan, Fajie Yuan, Yu Song, Youhua Li, Junchen Fu, Fei Yang, Yunzhu Pan, and Yongxin Ni. 2023. Where to go next for recommender systems? id- vs. modality-based recommender models revisited. In Proceedings of the 46th International ACM SIGIR Conference on Research and Deve...
2023
-
[54]
Chi Zhang, Yantong Du, Xiangyu Zhao, Qilong Han, Rui Chen, and Li Li. 2022. Hierarchical item inconsistency signal learning for sequence denoising in se- quential recommendation. In Proceedings of the 31st ACM international conference on information & knowledge management . 2508–2518
2022
-
[55]
Chi Zhang, Qilong Han, Rui Chen, Xiangyu Zhao, Peng Tang, and Hongtao Song
-
[56]
Kangning Zhang, Jiarui Jin, Yingjie Qin, Ruilong Su, Jianghao Lin, Yong Yu, and Weinan Zhang. 2024. Learning ID-free Item Representation with Token Crossing for Multimodal Recommendation. arXiv preprint arXiv:2410.19276 (2024)
2024 arXiv
-
[57]
Sheng Zhang, Maolin Wang, Wanyu Wang, Jingtong Gao, Xiangyu Zhao, Yu Yang, Xuetao Wei, Zitao Liu, and Tong Xu. 2025. Glint-ru: Gated lightweight intelligent recurrent units for sequential recommender systems. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discov...
2025
-
[58]
Taolin Zhang, Junwei Pan, Jinpeng Wang, Yaohua Zha, Tao Dai, Bin Chen, Ruisheng Luo, Xiaoxiang Deng, Yuan Wang, Ming Yue, et al . 2024. To- wards Scalable Semantic Representation for Recommendation. arXiv preprint arXiv:2410.09560 (2024)
2024 arXiv
-
[59]
Yang Zhang, Fuli Feng, Jizhi Zhang, Keqin Bao, Qifan Wang, and Xiangnan He
-
[60]
Zijian Zhang, Shuchang Liu, Jiaao Yu, Qingpeng Cai, Xiangyu Zhao, Chunxu Zhang, Ziru Liu, Qidong Liu, Hongwei Zhao, Lantao Hu, et al . 2024. M3oe: Multi-domain multi-task mixture-of experts recommendation framework. In Proceedings of the 47th International ACM SIGIR Conference...
2024
-
[61]
Kesen Zhao, Shuchang Liu, Qingpeng Cai, Xiangyu Zhao, Ziru Liu, Dong Zheng, Peng Jiang, and Kun Gai. 2023. KuaiSim: A comprehensive simulator for recom- mender systems. Advances in Neural Information Processing Systems 36 (2023), 44880–44897
2023
-
[62]
Kesen Zhao, Xiangyu Zhao, Zijian Zhang, and Muyang Li. 2022. Mae4rec: Storage- saving transformer for sequential recommendations. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management . 2681– 2690
2022
-
[63]
Xiangyu Zhao, Long Xia, Liang Zhang, Zhuoye Ding, Dawei Yin, and Jiliang Tang. 2018. Deep reinforcement learning for page-wise recommendations. In Proceedings of the 12th ACM conference on recommender systems . 95–103
2018
-
[64]
Xiangyu Zhao, Long Xia, Lixin Zou, Hui Liu, Dawei Yin, and Jiliang Tang. 2020. Whole-chain recommendations. In Proceedings of the 29th ACM international conference on information & knowledge management . 1883–1891
2020
-
[65]
Xiangyu Zhao, Liang Zhang, Zhuoye Ding, Long Xia, Jiliang Tang, and Dawei Yin
-
[66]
Bowen Zheng, Yupeng Hou, Hongyu Lu, Yu Chen, Wayne Xin Zhao, Ming Chen, and Ji-Rong Wen. 2024. Adapting large language models by integrating collaborative semantics for recommendation. In 2024 IEEE 40th International Conference on Data Engineering (ICDE) . IEEE, 1435–1448
2024
-
[2013]
The annals of statistics (2013), 2263–2291
Equivalence of distance-based and RKHS-based statistics in hypothesis testing. The annals of statistics (2013), 2263–2291
2013
-
[2015]
In Proceedings of the 38th international ACM SIGIR conference on research and development in information retrieval
Image-based recommendations on styles and substitutes. In Proceedings of the 38th international ACM SIGIR conference on research and development in information retrieval. 43–52
-
[2018]
In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining
Recommendations with negative feedback via pairwise deep reinforcement learning. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining . 1040–1048
-
[2022]
arXiv preprint arXiv:2208.10442 (2022)
Image as a foreign language: Beit pretraining for all vision and vision- language tasks. arXiv preprint arXiv:2208.10442 (2022). Empowering Large Language Model for Sequential Recommendation via Multimodal Embeddings and Semantic IDs CIKM ’25, November 10–14, 2025, Seoul, Repu...
2022 arXiv
-
[2023]
Advances in Neural Information Processing Systems 36 (2023), 10299–10315
Recommender systems with generative retrieval. Advances in Neural Information Processing Systems 36 (2023), 10299–10315
2023
-
[2024]
In 2024 IEEE 40th International Conference on Data Engineering (ICDE)
Ssdrec: Self-augmented sequence denoising for sequential recommendation. In 2024 IEEE 40th International Conference on Data Engineering (ICDE) . IEEE, 803– 815
2024
-
[2025]
IEEE Transactions on Knowledge and Data Engineering (2025)
Collm: Integrating collaborative embeddings into large language models for recommendation. IEEE Transactions on Knowledge and Data Engineering (2025)
2025
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.