REVIEW 3 major objections 6 minor 63 references
IMFuse: Instance-Aware Multi-Layer Fusion for LLM-Enhanced Sequential Recommendation
T0 review · 3 major / 6 minor · reviewed 2026-07-30 · grok-4.5
Pith's one-line read Final-layer LLM item embeddings are not enough: adaptively fusing all layers, item by item, lifts sequential recommendation by about 6.7%.
desk verdict Solid incremental systems paper: multi-layer LLM fusion with item-routed experts gives consistent mid-single-digit gains and low overhead; worth a referee, not a paradigm shift. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
IMFuse: a global learnable score matrix over layers (softmaxed per embedding dimension), modulated per item by a soft mixture of shared layer-preference expert templates whose routing comes from the final-layer embedding, then used to weight-sum transformed multi-layer embeddings into one semantic vector.
What would settle it
On the same four datasets and backbones, replace the final-layer router with a fixed or randomly routed mixture (or drop instance modulation entirely) and check whether the reported gains over strong final-layer enhancers and over simple multi-layer averages disappear; or show that shallower layers never improve next-item ranking once dimensional collapse is controlled.
Extended reading notes
Core claim
Relying solely on the final LLM layer for item semantics is suboptimal for sequential recommendation because that layer suffers stronger dimensional collapse, intermediate layers carry non-redundant signals, and items differ in which depths help them. An instance-aware fusion that combines global dimension-wise layer weights with light expert modulation of those weights produces better item representations and improves standard enhancement pipelines by roughly 6.7% on average.
Load-bearing premise
That an item’s final-layer embedding is a good enough summary to decide how much each earlier layer should matter for that item.
Editorial extensions
If this is right
- LLM-enhanced sequential recommenders should store or cache intermediate item hidden states, not only the last layer.
- Existing final-layer adapters (alignment, SVD init, spectral transforms) can be upgraded by swapping in multi-layer fused semantics without redesigning the recommender.
- Layer usefulness is item-dependent, so uniform depth selection or plain averaging leaves accuracy on the table.
- The same global-plus-instance pattern can transfer across LLM encoders (e.g., LLaMA-3 and Qwen3) with modest retuning.
Reading between the lines
- If final-layer routing is only a convenient proxy, user- or session-conditioned routing could further personalize which semantic depths matter.
- The spectral-collapse diagnosis suggests multi-layer fusion may matter most for long-tail or fine-grained catalog items whose distinctions live in subordinate singular directions.
- Caching all layers is the practical bottleneck; learned sparse layer subsets per catalog could keep the gains while cutting memory.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that LLM-enhanced sequential recommenders underuse intermediate-layer item embeddings by defaulting to the final layer, which exhibits stronger dimensional (spectral) collapse and is not uniformly optimal across items. Motivated by layer-wise singular-value concentration, inter-layer cosine similarity, and cluster-wise trajectory differences (Figs. 1–2; §3.1), it proposes IMFuse: a global dimension-wise layer-preference matrix W (Eqs. 8–11) combined with instance-aware expert modulation that routes from the final-layer embedding to modulate those preferences per item. The fused semantic embedding plugs into existing ID–semantic pipelines without changing the SR backbone. Experiments on four Amazon datasets, SASRec and HSTU, four enhancement methods (RLMRec, LLM-ESR, LLMInit, SpecTran), multi-layer fusion baselines, ablations, efficiency, and a Qwen3-8B transfer study report consistent gains (claimed average relative improvement 6.72%) with small parameter/runtime overhead.
Significance. If the empirical gains hold under standard reproducibility checks, this is a useful, model-agnostic systems contribution for LLM-enhanced sequential recommendation: it reframes “which LLM layer to use” as a recommendation-supervised, dimension- and item-adaptive fusion problem rather than a fixed final-layer choice. Strengths include clear motivation analyses tied to the method, broad coverage (4 datasets × 2 backbones × multiple enhancers), ablations that move in the expected direction (Table 4: last/mean/w/o IM/rand/perm), efficiency numbers (Table 5), cross-LLM transfer (Table 6), and qualitative evidence of non-monotonic global weights and heterogeneous expert routing (Fig. 4). The overhead is modest, which matters for practical adoption. The work is incremental relative to multi-layer embedding fusion outside recsys, but the recommendation-specific design (global dim-wise preferences + item modulation under next-item loss) is a legitimate and timely delta.
major comments (3)
- [§4.2, Tables 2–3; Abstract] Tables 2–3 and the Abstract’s 6.72% average relative improvement are the central empirical claim, yet the manuscript does not report multi-seed means/stds, confidence intervals, or significance tests. Relative lifts on sparse Amazon splits (especially Office) can be sensitive to initialization and early stopping. Please add at least 3–5 random seeds (or equivalent paired tests) for the main SASRec/HSTU + SpecTran/RLMRec settings and state whether the average improvement remains stable; without this, the headline percentage is harder to trust at journal standard.
- [§3.3, Eq. (9); Appendix A.1, Table 7] Eq. (9) routes instance-aware modulation exclusively from the final-layer embedding E^L_i, while §3.1 and Fig. 1 argue that deeper layers are more collapsed and that useful depth varies by item. Appendix Table 7 also shows standalone NDCG strongly favoring the final layer. Table 4’s w/o IM ablation supports modulation in aggregate, but does not test whether the router input is adequate (e.g., shallow/mid summary, concatenated layer stats, or ID embedding as router features). A small ablation on router input would either validate the weakest design assumption or show the gains are mostly from global W rather than true instance-depth matching.
- [§3.1; §4.6, Figure 4; Table 4] Motivation treats spectral collapse and inter-layer dissimilarity as evidence of complementary recommendation-useful signal (§3.1, Eqs. 4–7). Mean fusion is sometimes competitive and sometimes weaker (Table 4; Appendix Table 8), which is consistent with non-uniform layer value but does not directly show that subordinate singular components improve ranking rather than adding noise that the supervised fusion happens to reweight. Briefly linking learned A_i / G mass to layers with less collapse (or to cluster trajectories in Fig. 2) would tighten the causal story between the analyses and the claimed mechanism.
minor comments (6)
- [Title page; References] Venue/metadata placeholders remain throughout (Conference acronym ’XX, Woodstock, NY, 2018 copyright, https://doi.org/XXXXXXX.XXXXXXX). Several cited venues are dated 2025–2026 (e.g., SpecTran SIGIR’26, CASE-MLP EACL’26, VA-HS ICML’26); ensure bibliographic accuracy before camera-ready.
- [§2.2–§3.3] Notation switches among E_l, eE_l, E_{f,i}, E_sem and uses both L+1 layers (including 0) and “32” as final in figures; define whether layer 0 is embeddings or first block once in §2.2.1.
- [Figure 2] Figure 1–2 captions and axis labels are readable but the dual-panel similarity heatmaps in Fig. 2(b) are not explained in the main text beyond a short clause; add one sentence on what structural difference the two groups illustrate.
- [§4.2, Table 3] Table 3 reports multi-layer baselines only on Clothing and Toy with two enhancers; a sentence on why Beauty/Office were omitted (space vs. cost) would help.
- [§4.1.4; §4.3] Hyperparameters η ∈ {5,10}, β ∈ {0.1,0.2,0.3}, M=3 are stated (§4.1.4) but sensitivity is only partly covered via w/ Rand. and w/ Perm.; a one-row sensitivity on M would strengthen the efficiency narrative.
- [§1; §3.1] Minor prose issues: “inevitably leaves” (§1), “dimensional collapse” vs “spectral collapse” used interchangeably—pick one primary term after first definition.
Circularity Check
No significant circularity: empirical multi-layer fusion trained and evaluated on held-out splits; gains are not forced by definition or self-citation.
full rationale
IMFuse is a systems/empirical paper. Its load-bearing claim is that recommendation-supervised fusion of multi-layer LLM item states (global dimension-wise preferences W plus instance-aware expert modulation via a router on E^L_i) improves next-item ranking over final-layer baselines and generic multi-layer aggregators. Layer weights and experts are optimized under InfoNCE on training sequences and measured by HR/NDCG on chronological leave-one-out test splits across four datasets, two backbones, four enhancement pipelines, and a second LLM encoder. That is ordinary supervised learning, not a derivation that reduces the reported lift to a fitted identity. Motivating analyses (singular-value collapse, inter-layer cosine similarity, item-group trajectories) are observational diagnostics, not uniqueness theorems. Depth-biased initialization of W and final-layer routing are design choices; ablations (w/ Last, w/ Mean, w/o IM, w/ Rand., w/ Perm.) and learned non-monotonic weights / heterogeneous routing (Figure 4) test them rather than smuggle the result in by construction. Overlapping-author citations (e.g., SpecTran) supply baselines and related methods, not a load-bearing uniqueness or ansatz that forces the 6.72% claim. No self-definitional loop, fitted-input-as-prediction, or renaming of a known identity appears in the equation chain (Eqs. 8–11).
Assumptions & free parameters
free parameters (5)
- depth-bias strength η =
chosen in {5, 10}
- modulation strength β =
chosen in {0.1, 0.2, 0.3}
- number of experts M =
3
- recommendation embedding dim d and sequence/model hyperparameters =
d=128; history=10; batch=256
- global layer-score matrix W and template bank B =
trained end-to-end (not closed form)
assumptions (5)
- domain assumption Frozen multi-layer LLM encodings of item text are valid, fixed semantic features for collaborative sequential ranking.
- domain assumption Next-item prediction under chronological leave-one-out on Amazon reviews with InfoNCE (64 negatives) is an adequate test of recommendation quality.
- ad hoc to paper Final-layer item embedding is a sufficient router input for choosing layer mixtures for that item.
- ad hoc to paper Dimensional collapse / inter-layer dissimilarity imply complementary recommendation-useful signal rather than noise or task-irrelevant features.
- standard math Standard linear algebra SVD/cosine and K-means++ on embeddings are appropriate probes of representation geometry.
invented entities (2)
-
IMFuse global dimension-wise layer preference matrix G/W
-
Instance-aware expert modulation (template bank B + MLP router)
Cite this review
Pith. "Pith review of IMFuse: Instance-Aware Multi-Layer Fusion for LLM-Enhanced Sequential Recommendation." pith.science (2026). https://pith.science/paper/HBRWWQCC
@misc{pith2026260727002,
author = {Pith},
title = {Pith review of: IMFuse: Instance-Aware Multi-Layer Fusion for LLM-Enhanced Sequential Recommendation},
year = {2026},
howpublished = {\url{https://pith.science/paper/HBRWWQCC}},
note = {Machine review of arXiv:2607.27002}
}
read the original abstract
Recent advancements in Large Language Models (LLMs) have significantly enhanced sequential recommendation by encoding rich item textual information into semantic representations. However, existing methods typically rely on the final-layer hidden states of LLMs, overlooking potentially useful semantic signals encoded in other layers. Through empirical analysis, we reveal the limitations of this practice: final-layer representations often suffer from dimensional collapse, whereas intermediate layers preserve complementary, coarse-to-fine semantic knowledge. Furthermore, we observe that different items exhibit heterogeneous layer-wise representation evolution, making a uniform layer selection sub-optimal. To bridge this gap, we propose IMFuse, an instance-aware multi-layer fusion strategy designed for LLM-enhanced recommendation. Instead of relying on a single layer, IMFuse adaptively aggregates multi-layer semantic information by learning global dimension-wise layer preferences to capture general semantic contributions. To address item-level heterogeneity, IMFuse introduces an instance-aware expert modulation mechanism that dynamically adjusts these global preferences, generating personalized, item-specific semantic representations. Extensive experiments across four real-world datasets demonstrate the effectiveness of IMFuse. It consistently outperforms state-of-the-art baselines with an average relative improvement of 6.72%, while introducing limited parameter and computational overhead.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
David Arthur and Sergei Vassilvitskii. 2007. K-means++: The Advantages of Careful Seeding. InProceedings of the Eighteenth Annual ACM-SIAM Symposium on Discrete Algorithms. 1027–1035
2007
-
[2]
Keqin Bao, Jizhi Zhang, Yang Zhang, Wenjie Wang, Fuli Feng, and Xiangnan He. 2023. TALLRec: An Effective and Efficient Tuning Framework to Align Large Language Model with Recommendation. InProceedings of the 17th ACM Conference on Recommender Systems. 1007–1014. doi:10.1145/3604915.3608857
arXiv 2023
-
[4]
Junyi Chen, Lu Chi, Bingyue Peng, and Zehuan Yuan. 2024. HLLM: Enhancing Sequential Recommendations via Hierarchical Large Language Models for Item and User Modeling. arXiv:2409.12740 doi:10.48550/arXiv.2409.12740
-
[5]
Sishuo Chen, Xiaohan Bi, Rundong Gao, and Xu Sun. 2022. Holistic Sentence Em- beddings for Better Out-of-Distribution Detection. InFindings of the Association for Computational Linguistics: EMNLP. 6676–6686. doi:10.18653/v1/2022.findings- emnlp.497
-
[6]
Yu Cui, Feng Liu, Pengbo Wang, Bohao Wang, Heng Tang, Yi Wan, Jun Wang, and Jiawei Chen. 2024. Distillation Matters: Empowering Sequential Recommenders to Match the Performance of Large Language Models. InProceedings of the 18th ACM Conference on Recommender Systems. 507–517. doi:10.1145/3640457.3688118
arXiv 2024
-
[7]
Yu Cui, Feng Liu, Zhaoxiang Wang, Changwang Zhang, Jun Wang, Can Wang, and Jiawei Chen. 2026. SpecTran: Spectral-Aware Transformer-Based Adapter for LLM-Enhanced Sequential Recommendation. InProceedings of the 49th In- ternational ACM SIGIR Conference on Research and Development in Information Retrieval. 234–245. doi:10.1145/3805712.3809701
arXiv 2026
-
[8]
Kounianhua Du, Jizheng Chen, Jianghao Lin, Yunjia Xi, Hangyu Wang, Xinyi Dai, Bo Chen, Ruiming Tang, and Weinan Zhang. 2024. DisCo: Towards Har- monious Disentanglement and Collaboration between Tabular and Semantic Space for Recommendation. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 666–676. doi:10.1145/363752...
arXiv 2024
-
[9]
Shijie Geng, Shuchang Liu, Zuohui Fu, Yingqiang Ge, and Yongfeng Zhang. 2022. Recommendation as Language Processing (RLP): A Unified Pretrain, Personalized Prompt & Predict Paradigm (P5). InProceedings of the 16th ACM Conference on Recommender Systems. 299–315. doi:10.1145/3523227.3546767
arXiv 2022
Show all 63 references
- [10]
- [11]
-
[12]
Yingzhi He, Xiaohao Liu, An Zhang, Yunshan Ma, and Tat-Seng Chua. 2025. LLM2Rec: Large Language Models Are Powerful Embedding Models for Sequen- tial Recommendation. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2. 896–907. doi:10.114...
2025
-
[13]
Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk
-
[14]
Yupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li, Bolin Ding, and Ji-Rong Wen. 2022. Towards Universal Sequence Representation Learning for Recom- mender Systems. InProceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 585–593. doi:10.1145/353...
2022
-
[15]
Yupeng Hou, Junjie Zhang, Zihan Lin, Hongyu Lu, Ruobing Xie, Julian McAuley, and Wayne Xin Zhao. 2024. Large Language Models Are Zero-Shot Rankers for Recommender Systems. InAdvances in Information Retrieval. 364–381. doi:10. 1007/978-3-031-56060-6_24
2024
-
[16]
Guoqing Hu, An Zhang, Shuo Liu, Zhibo Cai, Xun Yang, and Xiang Wang. 2025. AlphaFuse: Learn ID Embeddings for Sequential Recommendation in Null Space of Language Embeddings. InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information...
2025
-
[17]
Kalervo Järvelin and Jaana Kekäläinen. 2002. Cumulated Gain-Based Evaluation of IR Techniques.ACM Transactions on Information Systems20, 4 (2002), 422–446. doi:10.1145/582415.582418
2002
-
[18]
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019. What Does BERT Learn about the Structure of Language?. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 3651–3657. doi:10.18653/v1/P19-1356
2019 doi
-
[19]
Jian Jia, Yipei Wang, Yan Li, Honggang Chen, Xuehan Bai, Zhaocheng Liu, Jian Liang, Quan Chen, Han Li, Peng Jiang, and Kun Gai. 2025. LEARN: Knowledge Adaptation from Large Language Model to Recommendation for Practical Indus- trial Application. InProceedings of the AAAI Confe...
2025 doi
-
[20]
Li Jing, Pascal Vincent, Yann LeCun, and Yuandong Tian. 2022. Understanding Dimensional Collapse in Contrastive Self-Supervised Learning. InInternational Conference on Learning Representations
2022
-
[21]
Tianjie Ju, Weiwei Sun, Wei Du, Xinwei Yuan, Zhaochun Ren, and Gongshen Liu. 2024. How Large Language Models Encode Context Knowledge? A Layer- Wise Probing Study. InProceedings of LREC-COLING. 8235–8246. doi:10.63317/ 2IB4REXE35CA
2024
-
[22]
Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recom- mendation. In2018 IEEE international conference on data mining (ICDM). IEEE, Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Zheng et al. 197–206
2018
-
[23]
Sein Kim, Hongseok Kang, Seungyoon Choi, Donghyun Kim, Minchul Yang, and Chanyoung Park. 2024. Large Language Models Meet Collaborative Filtering: An Efficient All-Round LLM-Based Recommender System. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Da...
2024
-
[24]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Opti- mization. InInternational Conference on Learning Representations
2015
-
[27]
Xinyu Lin, Wenjie Wang, Yongqi Li, Fuli Feng, See-Kiong Ng, and Tat-Seng Chua. 2024. Bridging Items and Language: A Transition Paradigm for Large Language Model-Based Recommendation. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 1816–1...
2024
-
[28]
Liu, Matt Gardner, Yonatan Belinkov, Matthew E
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019. Linguistic Knowledge and Transferability of Contextual Represen- tations. InProceedings of NAACL-HLT. 1073–1094. doi:10.18653/v1/N19-1112
2019 doi
-
[30]
Qidong Liu, Xian Wu, Yejing Wang, Zijian Zhang, Feng Tian, Yefeng Zheng, and Xiangyu Zhao. 2024. LLM-ESR: Large Language Models Enhancement for Long-Tailed Sequential Recommendation. InAdvances in Neural Information Processing Systems, Vol. 37. 26701–26727. doi:10.52202/079017-0839
2024 doi
-
[31]
Julian McAuley, Christopher Targett, Qinfeng Shi, and Anton van den Hengel
-
[32]
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Distributed Representations of Words and Phrases and their Compositionality. In Proceedings of the 26th International Conference on Neural Information Processing Systems. 3111–3119
2013
-
[33]
Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, ...
2019
-
[34]
Lutz Prechelt. 1998. Early Stopping—But When? InNeural Networks: Tricks of the Trade. Springer, 55–69. doi:10.1007/3-540-49430-8_3
1998 doi
- [35]
-
[36]
Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. 3982–3992. doi:10.18653/v1/D19-1410
2019 doi
-
[37]
Xubin Ren, Wei Wei, Lianghao Xia, Lixin Su, Suqi Cheng, Junfeng Wang, Dawei Yin, and Chao Huang. 2024. Representation Learning with Large Language Models for Recommendation. InProceedings of the ACM Web Conference. 3464–
2024
-
[38]
Yankun Ren, Zhongde Chen, Xinxing Yang, Longfei Li, Cong Jiang, Lei Cheng, Bo Zhang, Linjian Mo, and Jun Zhou. 2024. Enhancing Sequential Recommenders with Augmented Knowledge from Aligned Large Language Models. InProceedings of the 47th International ACM SIGIR Conference on R...
2024
-
[39]
Leheng Sheng, An Zhang, Yi Zhang, Yuxin Chen, Xiang Wang, and Tat-Seng Chua
-
[40]
Wentao Shi, Xiangnan He, Yang Zhang, Chongming Gao, Xinyue Li, Jizhi Zhang, Qifan Wang, and Fuli Feng. 2024. Large Language Models Are Learnable Planners for Long-Term Recommendation. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in I...
2024
-
[41]
Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang. 2019. BERT4Rec: Sequential Recommendation with Bidirectional En- coder Representations from Transformer. InProceedings of the 28th ACM In- ternational Conference on Information and Knowledge Managemen...
2019
-
[42]
Zhongxiang Sun, Zihua Si, Xiaoxue Zang, Kai Zheng, Yang Song, Xiao Zhang, and Jun Xu. 2024. Large Language Models Enhanced Collaborative Filtering. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management. 2178–2188. doi:10.1145/3627673.3679558
2024
-
[43]
Juntao Tan, Shuyuan Xu, Wenyue Hua, Yingqiang Ge, Zelong Li, and Yongfeng Zhang. 2024. IDGenRec: LLM-RecSys Alignment with Textual ID Learning. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 355–364. doi:10.11...
2024
-
[44]
Jiaxi Tang and Ke Wang. 2018. Personalized Top-N Sequential Recommendation via Convolutional Sequence Embedding. InProceedings of the Eleventh ACM International Conference on Web Search and Data Mining. 565–573. doi:10.1145/ 3159652.3159656
2018
- [45]
-
[46]
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019. BERT Rediscovers the Classical NLP Pipeline. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 4593–4601. doi:10.18653/v1/P19-1452
2019 doi
- [47]
-
[48]
Bin Wang and C.-C. Jay Kuo. 2020. SBERT-WK: A Sentence Embedding Method by Dissecting BERT-Based Word Models.IEEE/ACM Transactions on Audio, Speech, and Language Processing28 (2020), 2146–2157. doi:10.1109/TASLP.2020.3008390
2020
-
[49]
Bohao Wang, Feng Liu, Jiawei Chen, Xingyu Lou, Changwang Zhang, Jun Wang, Yuegang Sun, Yan Feng, Chun Chen, and Can Wang. 2025. MSL: Not All Tokens Are What You Need for Tuning LLM as a Recommender. InProceedings of the 48th International ACM SIGIR Conference on Research and D...
2025
-
[50]
Bohao Wang, Feng Liu, Changwang Zhang, Jiawei Chen, Yudi Wu, Sheng Zhou, Xingyu Lou, Jun Wang, Yan Feng, Chun Chen, and Can Wang. 2025. LLM4DSR: Leveraging Large Language Model for Denoising Sequential Recommendation. ACM Transactions on Information Systems44, 1 (2025), 1–32. ...
2025 doi
-
[51]
Yuling Wang, Changxin Tian, Binbin Hu, Yanhua Yu, Ziqi Liu, Zhiqiang Zhang, Jun Zhou, Liang Pang, and Xiao Wang. 2024. Can Small Language Models Be Good Reasoners for Sequential Recommendation?. InProceedings of the ACM Web Conference 2024. 3876–3887. doi:10.1145/3589334.3645671
2024
-
[52]
Wei Wei, Xubin Ren, Jiabin Tang, Qinyong Wang, Lixin Su, Suqi Cheng, Junfeng Wang, Dawei Yin, and Chao Huang. 2024. LLMRec: Large Language Models with Graph Augmentation for Recommendation. InProceedings of the 17th ACM International Conference on Web Search and Data Mining. 8...
2024
-
[53]
Yunjia Xi, Weiwen Liu, Jianghao Lin, Xiaoling Cai, Hong Zhu, Jieming Zhu, Bo Chen, Ruiming Tang, Weinan Zhang, and Yong Yu. 2024. Towards Open-World Recommendation with Knowledge Augmentation from Large Language Models. InProceedings of the 18th ACM Conference on Recommender S...
2024 doi
-
[54]
Xu Xie, Fei Sun, Zhaoyang Liu, Shiwen Wu, Jinyang Gao, Jiandong Zhang, Bolin Ding, and Bin Cui. 2022. Contrastive Learning for Sequential Recommendation. In2022 IEEE 38th International Conference on Data Engineering. 1259–1273. doi:10. 1109/ICDE53745.2022.00099
2022
-
[55]
Wujiang Xu, Qitian Wu, Zujie Liang, Jiaojiao Han, Xuying Ning, Yunxiao Shi, Wenfang Lin, and Yongfeng Zhang. 2025. SLMRec: Distilling Large Language Models into Small for Sequential Recommendation. InInternational Conference on Learning Representations
2025
-
[56]
Zhengyi Yang, Xiangnan He, Jizhi Zhang, Jiancan Wu, Xin Xin, Jiawei Chen, and Xiang Wang. 2023. A Generic Learning Framework for Sequential Recom- mendation with Distribution Shifts. InProceedings of the 46th International ACM SIGIR Conference on Research and Development in In...
2023
-
[57]
Zheng Yuan, Fajie Yuan, Yu Song, Youhua Li, Junchen Fu, Fei Yang, Yunzhu Pan, and Yongxin Ni. 2023. Where to Go Next for Recommender Systems? ID- vs. Modality-Based Recommender Models Revisited. InProceedings of the 46th International ACM SIGIR Conference on Research and Devel...
2023
-
[58]
Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhaojie Gong, Fangda Gu, Jiayuan He, Yinghai Lu, and Yu Shi. 2024. Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations. InProceedings of the 41st In...
2024
-
[59]
Gaifan Zhang, Yi Zhou, and Danushka Bollegala. 2026. CASE: Condition-Aware Sentence Embeddings for Conditional Semantic Textual Similarity Measurement. InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long...
2026
-
[61]
Weizhi Zhang, Liangwei Yang, Wooseong Yang, Henry Peng Zou, Yuqing Liu, Ke Xu, Sourav Medya, and Philip S. Yu. 2025. LLMInit: A Free Lunch from Large Language Models for Selective Initialization of Recommendation. InProceedings of the 2025 Conference on Empirical Methods in Na...
2025 doi
-
[62]
Yeqin Zhang, Yunfei Wang, Jiaxuan Chen, Ke Qin, Yizheng Zhao, and Cam- Tu Nguyen. 2026. LLM-Based Embeddings: Attention Values Encode Sentence Semantics Better Than Hidden States. InProceedings of the 43rd International Conference on Machine Learning. To appear
2026
-
[63]
Zihuai Zhao, Wenqi Fan, Jiatong Li, Yunqing Liu, Xiaowei Mei, Yiqi Wang, Zhen Wen, Fei Wang, Xiangyu Zhao, Jiliang Tang, and Qing Li. 2024. Recommender Systems in the Era of Large Language Models (LLMs).IEEE Transactions on Knowledge and Data Engineering36, 11 (2024), 6889–690...
2024
-
[64]
Bowen Zheng, Yupeng Hou, Hongyu Lu, Yu Chen, Wayne Xin Zhao, Ming Chen, and Ji-Rong Wen. 2024. Adapting Large Language Models by Integrating Collaborative Semantics for Recommendation. In2024 IEEE 40th International Conference on Data Engineering. 1435–1448. doi:10.1109/ICDE60...
2024
-
[2015]
InProceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval
Image-Based Recommendations on Styles and Substitutes. InProceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval. 43–52. doi:10.1145/2766462.2767755
-
[2016]
In International Conference on Learning Representations
Session-Based Recommendations with Recurrent Neural Networks. In International Conference on Learning Representations
-
[2025]
InInternational Conference on Learning Representations
Language Representations Can Be What Recommenders Need: Findings and Potentials. InInternational Conference on Learning Representations
-
[3475]
doi:10.1145/3589334.3645458
Reviewed July 30, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.