REVIEW 4 major objections 5 minor 55 references
Large Language Model driven Policy Exploration for Recommender Systems
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Using an LLM's item preferences to pre-train a recommender policy improves both cold-start quality and long-term returns once the policy is adapted online.
desk verdict A sensible warm-start recipe for RL recommenders, but the paper doesn't isolate whether the LLM label source is what makes it work. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is LLM-based preference distillation used as an offline reward and action generator. Given a user state encoded by SASRec, the actor samples k candidate actions, the LLM is prompted to pick one or 'None', the picked action receives reward 1, and this reward feeds the actor-critic update. The second key mechanism is the adaptive online blend: A-iALP_ap acts according to (1-alpha)pi_theta + alpha pi_beta, with the pretrained policy pi_theta frozen and the learnable policy pi_beta eventually taking over; this preserves pretrained behavior early while allowing the agent to escape any bad preferences the LLM encoded.
What would settle it
Measure agreement between the LLM's top-1 item choice and the actual next item in a held-out slice of the LFM or Industry interaction logs; if agreement is at or below random-item chance, then the reward signal in Equation (3) carries no user-preference information and the claimed pre-training advantage should disappear.
Extended reading notes
Core claim
The paper's central discovery is that preference signals distilled from an LLM are a viable offline pre-training signal for RL-based recommendation, and that the resulting policy can be transferred online without the usual cold-start penalty. Concretely, the authors construct prompts that ask the LLM to choose among ten candidate items given the user's history; a chosen item gets reward 1, 'None' gets 0, and these state-action-reward triples train an A2C actor-critic whose state encoder is SASRec. The resulting iALP policy produces markedly better initial sequences than randomly initialized DQN, PG, or A2C. To move online, A-iALP_ft fine-tunes the same network on simulated user rewards, while A-iALP_ap freezes the pretrained policy and learns a new policy alongside it, mixing the two with a weight alpha that shifts from 0 to 1 as training progresses. Across LFM, Industry, and Coat, A-iALP_ap reports the highest cumulative returns, and both adaptive variants converge faster and more stably than training from scratch.
Load-bearing premise
The load-bearing premise is that an LLM prompted with item attributes and a user's history picks the items the user would actually pick, so its 1/0 choices can serve as a trustworthy pre-training reward.
Editorial extensions
If this is right
- LLM-pretrained policies can be dropped into an online recommender to cut the initial poor-recommendation phase, addressing a main reason online RL recommenders are rarely deployed.
- Fine-tuning the LLM-pretrained policy on real feedback already improves over offline training alone, so online adaptation is complementary to LLM distillation.
- The adaptive blend scheme avoids catastrophic forgetting of pretrained behavior when the pretrained policy's preferences mismatch the real environment.
- A-iALP outperforms directly querying the LLM online, since the learned policy continues improving while the LLM's choices stay static and expensive.
- Under all tested exploration strategies, the pretrained start gives faster convergence and higher stable returns than A2C.
Reading between the lines
- If LLM preferences are good enough, this recipe could extend to other interactive domains with textual item attributes but no obvious reward model, such as dialogue or tutoring systems.
- The paper fine-tunes the LLM only on response format, not on user preferences; a testable extension is to fine-tune on real feedback when available, which might fix the failure mode visible in the Coat environment where iALP alone underperforms from-scratch A2C.
- Because the simulated environments share data with the reward models used for evaluation, a harder test is to evaluate against held-out user behavior or a different platform's reward model; the paper does not show how the gains transfer.
- The alpha schedule of A-iALP_ap is scenario-dependent; making alpha adaptive to the online policy's measured performance, rather than a fixed time schedule, is a natural next step.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes iALP, a policy for recommender systems pretrained offline using rewards and actions distilled from an LLM prompted with user interaction histories and item attributes. It then introduces two online adaptation strategies: A-iALP_ft, which fine-tunes the pretrained policy with simulated online feedback, and A-iALP_ap, which blends a frozen pretrained policy with a learnable online policy using a time-varying weight alpha. Experiments on three simulated environments (LFM, Industry, Coat) compare against DQN, PG, A2C, and frozen iALP in terms of Return, Length, and Average Reward, with claims of substantial improvements in initial and long-term performance.
Significance. The core idea of using LLM judgments as an offline pretraining signal to improve cold-start and long-term recommendation is practically relevant and the adaptive alpha-blending mechanism is a sensible way to transition from LLM-derived behavior to environment-driven behavior. The paper gives a clear formulation of the pretraining and online phases, and the three-environment evaluation is broader than many prior works. However, the empirical support for the central causal claim is incomplete: the LLM preference signal is never validated against real user behavior or the reward models, there is no non-LLM pretraining control, and the reported margins lack error bars or significance tests. If these gaps are addressed, the method could be a useful contribution; as it stands, the evidence is suggestive rather than conclusive.
major comments (4)
- [§4.2.1, §5.1, Table 5] The central claim that LLM-distilled preferences cause the observed improvements is not supported because there is no ablation isolating the label source. The pretraining phase in Eqs. (2)-(3) and (6)-(10) uses LLM choices as both action labels and rewards, and Section 5.4 states that the LoRA fine-tuning on 1000 samples only learns the response format, not user preferences. The paper never checks whether LLM judgments agree with held-out user interactions or with the reward models used for evaluation. The authors' own Table 5 makes the attribution doubtful: frozen iALP is worse than every online baseline in all three environments (e.g., LFM 11.2 vs A2C 28.1; Industry 25.3 vs 46.3; Coat 31.2 vs 81.7). The gains of A-iALP_ap could therefore come from the pretrained state encoder, the pretrained initialization of the online policy, or the alpha-blending schedule, rather than from the semantic content of LLM preferences. A control with a non-LLM pretraining signal (e.g., random labels, popularity-based labels, or reward-model labels) is needed to attribute the improvement to the LLM.
- [§5.2, Table 3] The evaluation is entirely against simulated reward models trained on the same public datasets used to build the recommendation scenarios. No calibration or agreement analysis is reported between these reward models and actual user behavior, and no validation of the LLM's preference judgments against the reward model is provided. Consequently, the reported Returns are internal consistency checks within a simulator, not measures of real-world user satisfaction. In addition, Table 5 reports single numbers without variance across seeds, confidence intervals, or significance tests; margins such as 33.1 vs 28.1 on LFM or 51.8 vs 46.3 on Industry could be within run-to-run noise. At minimum, multiple random seeds with standard deviations and a significance test are needed for the headline claims.
- [§4.2.3, Eq. (14), Algorithm 2] The RQ1 comparison at epoch 0 is not a meaningful competition: iALP is a fully pretrained policy, while DQN, PG, and A2C are randomly initialized at epoch 0. The statement that iALP 'significantly outperforms' these baselines is true by construction and is not accompanied by any statistical test. A more informative comparison would report the number of online steps required for each baseline to reach the initial return of iALP, or would compare all methods from the same initialization schedule. The current framing overstates the contribution of the LLM signal.
- [§4.2.3, Eq. (14), Algorithm 2] The adaptive scheme A-iALP_ap depends on a weight alpha that is stated to increase to 1 as training proceeds, but the exact schedule is never specified. Section 4.2.3 says only that the initial value depends on the scenario, and Algorithm 2 does not list alpha as a tunable parameter or provide its values for LFM, Industry, or Coat. Since alpha controls the relative contribution of the frozen pretrained policy versus the learnable policy, the reported results are not reproducible without this information, and the sensitivity of the method to the alpha schedule is unknown.
minor comments (5)
- [§1, §2.1] The policy parameter is denoted as psi in Eq. (1) and as theta in Section 4.1.2; the notation should be made consistent throughout.
- [§4.2.1, §5.1] There are several typos and inconsistent abbreviations: 'A-iALT_ap' appears in the introduction, 'iAPT' appears in the contributions, and 'A-iALP_ap' is sometimes written as 'A-iALP_ap'. These should be corrected for clarity.
- [§5.4, §6.3] The description of the reward models is incomplete: for Coat, the text says 'Matrix Factorization (DeepFM)', which conflates two distinct model families, and for LFM/Industry the reward model is said to follow 'a sequential recommender method' from [24] without specifying the architecture, training objective, or hyperparameters.
- [§6.4] The LLMOnline baseline used for RQ3 is not fully specified: it is unclear whether it uses the same Mistral 7B model, the same LoRA format tuning, the same prompt template, and the same candidate sampling procedure as iALP. Without this information, the comparison in Figures 8 and 9 is difficult to interpret.
- [§6.4] The exploration-strategy experiments are reported only for LFM and Industry (Figures 10 and 11); no results are given for Coat, despite the paper claiming 'three simulated environments'. The absence of Coat in RQ1, RQ2, and RQ4 should be acknowledged or addressed.
Circularity Check
No significant circularity: the empirical return comparisons are independent of the method's definitions.
full rationale
The paper's derivation chain is empirical rather than definitional. The pre-training phase (Section 4.1) uses Eq. (6), Eq. (8), and Eq. (10) to fit the actor and critic to LLM-selected actions and LLM-defined rewards, so iALP's epoch-0 behavior is indeed a re-expression of the LLM's choices. However, this is not a hidden circularity: the paper does not present that epoch-0 comparison as a derived prediction, and its central claim concerns A-iALP after online adaptation, which is evaluated in Tables 4-5 against external baselines (DQN, PG, A2C) using reward models trained from public interaction data. The adaptive losses (Eqs. 12-15) and Algorithm 2 are standard offline-to-online fine-tuning; the combining weight alpha in Eq. (14) is a scenario-dependent schedule, not a fitted parameter that is later renamed as a result. The one self-citation, reference [39], supports the motivational premise that LLMs can capture user objectives; it is a published, externally checkable paper and is not the load-bearing evidence for the empirical comparisons. The absence of a non-LLM pretraining control is a real experimental-design weakness and limits causal attribution to LLM preferences, but it does not make any equation or reported number equal to its input by construction. No specific circular step of the kinds enumerated was found, so the score is 0.
Assumptions & free parameters
free parameters (5)
- alpha schedule =
unspecified
- candidate action count k =
10
- LLM format-tuning sample count =
1000
- pretraining epochs =
100
- online training steps =
50k
assumptions (4)
- domain assumption LLM can predict user preferences from item attribute text
- domain assumption Simulated reward models faithfully represent user feedback
- domain assumption SASRec state encoder is a valid representation for RL recommender states
- standard math Actor-critic with TD Q-loss is an appropriate learning rule
Cite this review
Pith. "Pith review of Large Language Model driven Policy Exploration for Recommender Systems." pith.science (2026). https://pith.science/paper/SJMMFR5R
@misc{pith2026250113816,
author = {Pith},
title = {Pith review of: Large Language Model driven Policy Exploration for Recommender Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/SJMMFR5R}},
note = {Machine review of arXiv:2501.13816}
}
abstract
Recent advancements in Recommender Systems (RS) have incorporated Reinforcement Learning (RL), framing the recommendation as a Markov Decision Process (MDP). However, offline RL policies trained on static user data are vulnerable to distribution shift when deployed in dynamic online environments. Additionally, excessive focus on exploiting short-term relevant items can hinder exploration, leading to suboptimal recommendations and negatively impacting long-term user gains. Online RL-based RS also face challenges in production deployment, due to the risks of exposing users to untrained or unstable policies. Large Language Models (LLMs) offer a promising solution to mimic user objectives and preferences for pre-training policies offline to enhance the initial recommendations in online settings. Effectively managing distribution shift and balancing exploration are crucial for improving RL-based RS, especially when leveraging LLM-based pre-training. To address these challenges, we propose an Interaction-Augmented Learned Policy (iALP) that utilizes user preferences distilled from an LLM. Our approach involves prompting the LLM with user states to extract item preferences, learning rewards based on feedback, and updating the RL policy using an actor-critic framework. Furthermore, to deploy iALP in an online scenario, we introduce an adaptive variant, A-iALP, that implements a simple fine-tuning strategy (A-iALP$_{ft}$), and an adaptive approach (A-iALP$_{ap}$) designed to mitigate issues with compromised policies and limited exploration. Experiments across three simulated environments demonstrate that A-iALP introduces substantial performance improvements
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Keqin Bao, Jizhi Zhang, Yang Zhang, Wenjie Wang, Fuli Feng, and Xiangnan He
-
[2]
Jin Chen, Zheng Liu, Xu Huang, Chenwang Wu, Qi Liu, Gangwei Jiang, Yuanhao Pu, Yuxuan Lei, Xiaolong Chen, Xingmei Wang, et al. 2023. When large language models meet personalization: Perspectives of challenges and opportunities. arXiv preprint arXiv:2307.16376 (2023)
arXiv 2023
-
[3]
Yingpeng Du, Di Luo, Rui Yan, Hongzhi Liu, Yang Song, Hengshu Zhu, and Jie Zhang. 2023. Enhancing job recommendation through llm-based generative adversarial networks. arXiv preprint arXiv:2307.10747 (2023)
arXiv 2023
-
[4]
Wenqi Fan, Zihuai Zhao, Jiatong Li, Yunqing Liu, Xiaowei Mei, Yiqi Wang, Jiliang Tang, and Qing Li. 2023. Recommender systems in the era of large language models (llms). arXiv preprint arXiv:2307.02046 (2023)
arXiv 2023
-
[5]
Amir massoud Farahmand, Rémi Munos, and Csaba Szepesvári. 2010. Error propagation for Approximate Policy and Value Iteration. In Proceedings of the 23rd International Conference on Neural Information Processing Systems - Volume 1 (Vancouver, British Columbia, Canada) (NIPS’10). Curran Associates Inc., Red Hook, NY, USA, 568–576
work page 2010
-
[6]
Junchen Fu, Fajie Yuan, Yu Song, Zheng Yuan, Mingyue Cheng, Shenghui Cheng, Jiaqi Zhang, Jie Wang, and Yunzhu Pan. 2024. Exploring adapter-based transfer learning for recommender systems: Empirical studies and practical insights. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining. 208–217
2024
-
[7]
Shijie Geng, Shuchang Liu, Zuohui Fu, Yingqiang Ge, and Yongfeng Zhang. 2022. Recommendation as language processing (rlp): A unified pretrain, personalized prompt & predict paradigm (p5). In Proceedings of the 16th ACM Conference on Recommender Systems. 299–315
2022
-
[8]
Dalin Guo, Sofia Ira Ktena, Pranay Kumar Myana, Ferenc Huszar, Wenzhe Shi, Alykhan Tejani, Michael Kneier, and Sourav Das. 2020. Deep bayesian bandits: Exploring in online personalized recommendations. In Proceedings of the 14th ACM Conference on Recommender Systems . 456–461
work page 2020
Show all 55 references
-
[9]
Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017. DeepFM: a factorization-machine based neural network for CTR prediction.arXiv preprint arXiv:1703.04247 (2017)
2017 arXiv
-
[10]
Zhankui He, Zhouhang Xie, Rahul Jha, Harald Steck, Dawen Liang, Yesu Feng, Bodhisattwa Prasad Majumder, Nathan Kallus, and Julian McAuley. 2023. Large language models as zero-shot conversational recommenders. InProceedings of the 32nd ACM international conference on informatio...
2023
-
[11]
Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk
-
[12]
Yupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li, Bolin Ding, and Ji-Rong Wen. 2022. Towards universal sequence representation learning for recommender systems. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Dis- covery and Data Mining . 585–593
2022
-
[13]
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685 (2021)
2021 arXiv
-
[14]
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al . 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023)
2023 arXiv
-
[15]
Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recom- mendation. In 2018 IEEE international conference on data mining (ICDM) . IEEE, 197–206
2018
-
[16]
Vijay Konda and John Tsitsiklis. 1999. Actor-critic algorithms.Advances in neural information processing systems 12 (1999)
1999
-
[17]
Volodymyr Kuleshov and Doina Precup. 2014. Algorithms for multi-armed bandit problems. arXiv preprint arXiv:1402.6028 (2014)
2014 arXiv
-
[18]
Aviral Kumar, Justin Fu, George Tucker, and Sergey Levine. 2019. Stabilizing off-policy Q-learning via bootstrapping error reduction . Curran Associates Inc., Red Hook, NY, USA
2019
-
[19]
Minae Kwon, Sang Michael Xie, Kalesha Bullard, and Dorsa Sadigh. 2023. Re- ward Design with Language Models. In The Eleventh International Conference on Learning Representations. https://openreview.net/forum?id=10uNUgI5Kl
2023
-
[20]
Seunghyun Lee, Younggyo Seo, Kimin Lee, Pieter Abbeel, and Jinwoo Shin. 2021. Offline-to-Online Reinforcement Learning via Balanced Replay and Pessimistic Q-Ensemble. In Annual Conference on Robot Learning
2021
-
[21]
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971 (2015)
2015 arXiv
-
[22]
Jianghao Lin, Xinyi Dai, Yunjia Xi, Weiwen Liu, Bo Chen, Xiangyang Li, Chenxu Zhu, Huifeng Guo, Yong Yu, Ruiming Tang, et al . 2023. How Can Recom- mender Systems Benefit from Large Language Models: A Survey. arXiv preprint arXiv:2306.05817 (2023)
2023 arXiv
-
[23]
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. Comput. Surveys 55, 9 (2023), 1–35
2023
-
[24]
Shuchang Liu, Qingpeng Cai, Bowen Sun, Yuhao Wang, Ji Jiang, Dong Zheng, Peng Jiang, Kun Gai, Xiangyu Zhao, and Yongfeng Zhang. 2023. Exploration and regularization of the latent action space in recommendation. In Proceedings of the ACM Web Conference 2023. 833–844
2023
-
[25]
Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heintz, and Dan Roth. 2023. Recent advances in natural language processing via large pre-trained language models: A survey. Comput. Surveys 56, 2 (2023), 1–40
2023
-
[26]
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602 (2013)
2013 arXiv
-
[27]
Jianmo Ni, Jiacheng Li, and Julian McAuley. 2019. Justifying recommendations using distantly-labeled reviews and fine-grained aspects. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natura...
2019
-
[28]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems 35...
2022
-
[29]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research 21, 1 (2020), 5485–5551
2020
-
[30]
Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan H Keshavan, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Q Tran, Jonah Samost, et al. 2023. Recommender Systems with Generative Retrieval.arXiv preprint arXiv:2305.05065 (2023)
2023 arXiv
-
[31]
Zhaochun Ren, Na Huang, Yidan Wang, Pengjie Ren, Jun Ma, Jiahuan Lei, Xinlei Shi, Hengliang Luo, Joemon Jose, and Xin Xin. 2023. Contrastive State Augmen- tations for Reinforcement Learning-Based Recommender Systems. In Proceedings of the 46th International ACM SIGIR Conferenc...
2023
-
[32]
Shideh Rezaeifar, Robert Dadashi, Nino Vieillard, Léonard Hussenot, Olivier Bachem, Olivier Pietquin, and Matthieu Geist. 2022. Offline Reinforcement Learning as Anti-exploration. In AAAI Conference on Artificial Intelligence
2022
-
[33]
Markus Schedl. 2016. The lfm-1b dataset for music retrieval and recommendation. In Proceedings of the 2016 ACM on international conference on multimedia retrieval . 103–110
2016
-
[34]
Tobias Schnabel, Adith Swaminathan, Ashudeep Singh, Navin Chandak, and Thorsten Joachims. 2016. Recommendations as treatments: Debiasing learning and evaluation. In international conference on machine learning . PMLR, 1670– 1679
2016
-
[35]
Younggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee
-
[36]
Wentao Shi, Xiangnan He, Yang Zhang, Chongming Gao, Xinyue Li, Jizhi Zhang, Qifan Wang, and Fuli Feng. 2024. Large Language Models are Learnable Planners for Long-Term Recommendation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in ...
2024
-
[37]
Jiaxi Tang and Ke Wang. 2018. Personalized top-n sequential recommenda- tion via convolutional sequence embedding. In Proceedings of the eleventh ACM international conference on web search and data mining . 565–573
2018
-
[38]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)
2023 arXiv
-
[39]
Jie Wang, Alexandros Karatzoglou, Ioannis Arapakis, and Joemon M Jose. 2024. Reinforcement Learning-based Recommender Systems with Large Language Models for State Reward and Action Modeling. In Proceedings of the 47th In- ternational ACM SIGIR Conference on Research and Develo...
2024
-
[40]
Jie Wang, Alexandros Karatzoglou, Ioannis Arapakis, Xin Xin, Xuri Ge, and Joemon M Jose. 2024. Sparks of Surprise: Multi-objective Recommendations with Hierarchical Decision Transformers for Diversity, Novelty, and Serendipity. In Proceedings of the 33rd ACM International Conf...
2024
-
[41]
Jie Wang, Fajie Yuan, Mingyue Cheng, Joemon M Jose, Chenyun Yu, Beibei Kong, Xiangnan He, Zhijin Wang, Bo Hu, and Zang Li. 2022. Transrec: Learning transferable recommendation from mixture-of-modality feedback. arXiv preprint arXiv:2206.06190 (2022)
2022
-
[42]
Pengfei Wang, Yu Fan, Long Xia, Wayne Xin Zhao, ShaoZhang Niu, and Jimmy Huang. 2020. KERL: A knowledge-guided reinforcement learning model for WSDM ’25, March 10–14, 2025, Hannover, Germany Jie Wang, Alexandros Karatzoglou, Ioannis Arapakis, and Joemon M. Jose sequential reco...
2020
-
[43]
Ronald J Williams. 1992. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning 8 (1992), 229–256
1992
-
[44]
Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, and Joemon M Jose. 2020. Self-supervised reinforcement learning for recommender systems. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval. 931–940
2020
-
[45]
Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, and Joemon M Jose. 2022. Supervised advantage actor-critic for recommender systems. In Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining . 1186– 1196
2022
-
[46]
Xin Xin, Tiago Pimentel, Alexandros Karatzoglou, Pengjie Ren, Konstantina Christakopoulou, and Zhaochun Ren. 2022. Rethinking Reinforcement Learning for Recommendation: A Prompt Perspective. In Proceedings of the 45th Inter- national ACM SIGIR Conference on Research and Develo...
2022
-
[47]
Yuanqing Yu, Chongming Gao, Jiawei Chen, Heng Tang, Yuefeng Sun, Qian Chen, Weizhi Ma, and Min Zhang. 2024. EasyRL4Rec: A User-Friendly Code Library for Reinforcement Learning Based Recommender Systems. arXiv preprint arXiv:2402.15164 (2024)
2024 arXiv
-
[48]
Fajie Yuan, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M Jose, and Xi- angnan He. 2019. A simple convolutional generative network for next item recommendation. In Proceedings of the twelfth ACM international conference on web search and data mining . 582–590
2019
-
[49]
Gangyi Zhang. 2023. User-Centric Conversational Recommendation: Adapting the Need of User with Large Language Models. In Proceedings of the 17th ACM Conference on Recommender Systems . 1349–1354
2023
-
[50]
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223 (2023)
2023 arXiv
-
[51]
Xiangyu Zhao, Liang Zhang, Zhuoye Ding, Long Xia, Jiliang Tang, and Dawei Yin
-
[2015]
arXiv preprint arXiv:1511.06939 (2015)
Session-based recommendations with recurrent neural networks. arXiv preprint arXiv:1511.06939 (2015)
2015 arXiv
-
[2018]
In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining
Recommendations with negative feedback via pairwise deep reinforcement learning. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining . 1040–1048
-
[2021]
In International Conference on Machine Learning
State entropy maximization with random encoders for efficient exploration. In International Conference on Machine Learning . PMLR, 9443–9454
-
[2023]
arXiv preprint arXiv:2305.00447 (2023)
Tallrec: An effective and efficient tuning framework to align large language model with recommendation. arXiv preprint arXiv:2305.00447 (2023)
2023 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.