REVIEW 4 major objections 4 minor 56 references
The paper claims that separating 'what the user wants' from 'how to act' — and validating the latter by outcome feedback — improves generative recommendation while keeping online serving LLM-free.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 01:21 UTC pith:TARXR7PX
load-bearing objection The 'outcome-grounded' policy discovery rests on a reward target that is ambiguous between a validation-target leak and a recency signal; the rest of the framework is solid and worth reviewing. the 4 major comments →
From Understanding to Action: Feedback-Grounded Policy Discovery for Generative Recommendation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Central claim: recommendation decisions should be grounded in outcome feedback, not linguistic plausibility. For each candidate policy the paper computes an advantage — the semantic similarity between the policy-conditioned predicted item and the last item in the user's training sequence, minus the intent-only score. That advantage selects among K candidates and drives group-wise iterative critique; only policies with positive verified advantage become 'policy knowledge.' The paper further claims this knowledge complements intent (intent localizes the broad semantic region, policy sharpens deeper SID levels), and that both can be compressed into two latent tokens via relational distillation
What carries the argument
The central machinery is the advantage-based policy evaluation in Eq. (5): the incremental utility of a policy is the gain in cosine similarity between the executor's prediction and M+_u (the last item in the user's training sequence) over an intent-only baseline. This dense semantic reward replaces sparse rank metrics, which give zero feedback when the target is outside the top-K. Around it the paper builds a dual-agent loop — a policy agent proposes K hypotheses, a feedback agent critiques the group based on execution outcomes, and evolution continues while the best advantage improves — and then a latent intent-policy chain (Intent Token, Policy Token) with dual-space relational distillati
Load-bearing premise
The whole pipeline depends on treating the semantic similarity between the predicted next item and the last item in the user's training sequence as a valid dense reward for recommendation utility; if that similarity misjudges utility — or if, under the paper's data split, that 'last training item' is actually the validation target — then the discovered policies are not truly grounded in outcomes.
What would settle it
Compare two runs of the pipeline: one where the reward target is the last training interaction as defined in the paper, and one where the reward target is a genuinely held-out validation item excluded from all prompts and training signals. If the second run's policies give the same or better downstream Recall@10, the reward is not the source of the gain; if the first run's performance collapses without the validation-target reward, the protocol is leaking.
If this is right
- Policy knowledge validated by advantage over an intent-only baseline yields consistent gains across datasets and backbones, so the Understanding–Action Gap is empirically real, not just conceptual.
- Intent and policy have a division of labor: intent localizes the semantic region at coarse SID levels, while policy sharpens fine-grained SID-level discrimination.
- The gains come from outcome-grounded selection and iterative refinement: removing evolution or picking policies randomly degrades performance, as does replacing the policy with the target item's description.
- Relational distillation of intent and policy structures outperforms direct embedding matching, and both tokens together recover more than either alone.
- The approach stays deployable: two latent tokens add about 9 ms per sample versus the base model, while direct LLM inference is orders of magnitude slower.
Where Pith is reading between the lines
- If the cosine-similarity reward is a valid proxy for utility, the same advantage-over-baseline loop could be ported to other generative decision problems — retrieval, autocomplete, or even non-recommendation generation — wherever a cheap intent-only baseline exists.
- The paper's reward target is the 'last item in the user's training sequence'; under the stated leave-one-out split (last interaction = test, second-to-last = validation), that item may coincide with the validation target. If so, the 'validated policy' claim would rest on a protocol leak, and the reported gains would need to be re-measured with a truly held-out reward target.
- The group-wise feedback mechanism is essentially prompt-level policy improvement: it uses execution outcomes to refine natural-language policies rather than updating weights. That suggests a cheap alternative to RL fine-tuning that could be tested on tasks where LLM weights are frozen.
- The explicit 'rejection boundary' component could be repurposed as a guardrail for content safety or diversity in production recommender systems, since it is a separate, audit-able decision rule.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a framework for generative recommendation that separates intent understanding from recommendation policy. An LLM-based intent agent induces a textual task-oriented intent; a policy agent generates candidate decision policies; a shared executor predicts a next-item profile with and without each policy; and a feedback agent critiques the candidate set. Policies are scored by their cosine-similarity advantage over the intent-only baseline (Eq. 5), refined for up to R rounds, and only policies with positive advantage are used as supervision. The intent and validated policy texts are then transferred into two latent tokens (Intent Token and Policy Token) of a Semantic-ID generator via dual-space first- and higher-order relational distillation, allowing LLM-free serving. Offline experiments on three Amazon categories with TIGER and LETTER backbones report consistent improvements, and a seven-day online A/B test in Kuaishou's local-services advertising system reports Revenue +4.506% and ADVV +4.621% at the 95% confidence level.
Significance. If the result holds, the paper makes a useful conceptual contribution: distinguishing intent knowledge from policy knowledge and grounding the latter in incremental recommendation utility rather than linguistic plausibility. The offline ablations include a sensible Target-as-Policy control that partially addresses the concern that gains come from direct target-text injection, and the deployment design is practical: LLMs are used offline as teachers, and the student remains lightweight for online serving. The clean separation of the two latent tokens and the relational (rather than coordinate-wise) distillation is a plausible mechanism. However, the central claim of 'outcome-validated policy discovery' depends on the definition of the reward target in Eq. (5), which is ambiguous and internally inconsistent with the stated data split. The statistical reporting also needs strengthening, since most offline gains are small point estimates without variance or significance measures.
major comments (4)
- [§3.3.2, Eq. (5); Appendix A.1; Appendix C.2] The identity of M+_u is contradictory and load-bearing. Eq. (5) defines M+_u as 'the textual profile of the last item in the user's training sequence.' Under the leave-one-out protocol in A.1, the training sequence contains all interactions before the second-to-last, so the last training item is v_{T-2}, while the validation target is v_{T-1} and the test target is v_T. §3.3.3 asserts that 'Validation and test targets are never used for policy generation, evaluation, or refinement,' and C.2 states that the ground-truth next item is 'used only by the external reward environment' while also saying the Feedback Agent does not observe 'the last item in the user's training sequence.' If M+_u is actually v_{T-1}, then policy selection is computed against the very label the policy later helps predict, making the 'outcome-validated policy' claim circular. If M+_u is v_{T-2}, the reward is simila
- [Tables 1-6, 8-13; Appendix A.2] The manuscript reports only point estimates, despite A.2 stating that 'All experiments are repeated five times, and the average results are reported.' Several decisive comparisons are numerically small: e.g., in Table 1 on Beauty, LETTER+Ours vs. LETTER differ by roughly 0.004-0.006 in Recall@10 and 0.004-0.005 in NDCG@10; in Table 4, the difference between Full Model and w/o Policy Evolution is 0.0011 in Recall@10 on Beauty. Without standard deviations, confidence intervals, or paired significance tests, the reader cannot tell whether these differences are stable across runs. Please report variance/CI for the main tables and run significance tests for the full-model versus key-ablation comparisons.
- [§4.4.2 and Table 4] The online A/B test is reported as a single pair of lift numbers (Revenue +4.506%, ADVV +4.621%) with 'statistically significant at the 95% confidence level,' but no confidence intervals, number of observations per metric, variance estimates, or test procedure are given. Moreover, the online baseline is described only as 'the production baseline,' so it is unclear whether the online treatment corresponds to TIGER+Ours, LETTER+Ours, or a different internal model. Given that the offline gains over LETTER are much smaller than the online lifts, the online claim would be stronger with more detail. At minimum, state the CI width and the test used, and clarify how the deployed model maps to the offline architecture. Also, in Table 4, 'Target-as-Policy' uses the target-item description in the Policy Token; please specify whether the target item is the validation item v_{T-1} or the test item v_
- [§3.3.3 and Figure 3] Figure 3 plots 'Reward' and 'Recall@10' as a function of evolution round. The reward is computed on the policy-discovery split, but it is not specified on which split Recall@10 is measured (training, validation, or test). If Recall@10 is evaluated on the same split used for policy selection, the trend in Figure 3 is not evidence of generalization. Please state the split for each curve, or replot the figure with a validation/test evaluation. This matters because the paper's main claim is that the reward is aligned with recommendation quality, not merely with the training objective.
minor comments (4)
- [Throughout] Notation is occasionally inconsistent: e.g., Eq. (12) uses a nonstandard symbol for stop-gradient; Table 3's 'R1:ℓ @10' is not defined before the table; and the figure captions refer to 'IntentLLM' without explanation. Please proofread and unify notation.
- [Appendix C.2] The sentence 'It does not observe the last item in the user's training sequence' appears to conflict with the earlier statement that the Feedback Agent does not see the ground-truth next item. If this is a typo, please correct; if intentional, explain the exact identity of 'last item in the user's training sequence' in the context of the case study.
- [§4.1.2 / Appendix A.3] The Direct LLM baseline is described only briefly. It would be useful to state the inference-time cost and how the embedding matching was performed, since a reviewer would want to compare the offline quality and the serving latency of that baseline against the proposed method.
- [§4.2 / Table 1] The main table lacks error bars, and it is not clear whether the same five seeds were used for all baselines and for the proposed method. Please state the seeds and whether paired comparisons were performed.
Circularity Check
No demonstrated circularity; reward-target ambiguity flagged as correctness risk, not a reduction-by-construction.
full rationale
Eq. (3) derives intent from history; Eqs. (4)-(9) discover policies via a similarity reward; Eqs. (10)-(14) distill into latent tokens and train with next-item loss. No equation defines its output in terms of itself. The reward target M+_u is defined in Eq. (5) as 'the textual profile of the last item in the user's training sequence,' while case study C.2 refers to the 'ground-truth next item' being used by the external reward environment. This ambiguity could indicate a protocol leak if M+ is the validation/test target, but the paper explicitly states in Sec. 3.3.3 that 'Validation and test targets are never used for policy generation, evaluation, or refinement,' and the Target-as-Policy ablation (Table 4) shows that replacing the discovered policy with the target-item description hurts performance, which would not be expected if the policy were merely a repackaged target. No reduction-by-construction can be exhibited from the paper's equations. The few self-citations (refs. [3], [29], [46]) are background and not load-bearing. Score 2 reflects the unresolved reward-target ambiguity, not a demonstrated circular step.
Axiom & Free-Parameter Ledger
free parameters (7)
- λ_I (intent distillation weight) =
selected from {0.5, 1.0}
- λ_P (policy distillation weight) =
selected from {0.5, 1.0}
- γ (higher-order distillation coefficient) =
0.1
- K (candidate policies per round) =
5
- R (max refinement rounds) =
2
- K1, K2 (neighborhood sizes) =
5
- τ_ν (softmax temperature) =
not stated
axioms (5)
- domain assumption Cosine similarity between embeddings of the predicted next-item profile and the 'last item in the user's training sequence' is a dense proxy for recommendation utility (Eq. 5).
- domain assumption The LLM executor's profile prediction is comparable across intent-only and policy-conditioned branches.
- domain assumption The last item in the user's training sequence can serve as the outcome without using validation/test targets.
- domain assumption Relational KL distillation transfers semantics across heterogeneous teacher and student spaces.
- domain assumption SID code prefixes correspond to coarse-to-fine semantic levels.
invented entities (2)
-
Intent Token (q_I)
no independent evidence
-
Policy Token (q_P)
no independent evidence
read the original abstract
Semantic-ID-based generative recommenders enable efficient next-item generation, but their item-level supervision mainly captures behavioral co-occurrence and local transitions. Large language models (LLMs) can complement these models by reasoning over heterogeneous interaction histories to understand the user's current demand. However, LLMs are not inherently trained with recommendation-specific outcome feedback, and linguistically plausible reasoning therefore does not necessarily lead to effective recommendation decisions. We term this mismatch the Understanding-Action Gap. Accordingly, we distinguish intent knowledge, which captures the user's current demand, from policy knowledge, which specifies the recommendation direction and rejection boundary under that demand. To bridge this gap, we propose a feedback-driven agent framework that first induces task-oriented intent and then discovers recommendation policies according to their incremental utility over an intent-only baseline. Candidate policies are evaluated and refined using outcome-derived feedback rather than linguistic plausibility. We further transfer the resulting intent and policy knowledge into two latent tokens of a lightweight Semantic-ID generator through dual-space relational distillation, enabling LLM-free online inference. Experiments on public benchmarks show consistent improvements over baselines, while large-scale online A/B tests achieve gains of 4.506% in Revenue and 4.621% in ADVV.
Figures
Reference graph
Works this paper leans on
-
[1]
Seunghwan Bang and Hwanjun Song. 2025. Llm-based user profile management for recommender system.arXiv preprint arXiv:2502.14541(2025)
Pith/arXiv arXiv 2025
-
[2]
Junyi Chen, Lu Chi, Bingyue Peng, and Zehuan Yuan. 2024. Hllm: Enhancing sequential recommendations via hierarchical large language models for item and user modeling.arXiv preprint arXiv:2409.12740(2024)
Pith/arXiv arXiv 2024
-
[3]
Jiaju Chen, Chongming Gao, Shuai Yuan, Shuchang Liu, Qingpeng Cai, and Peng Jiang. 2025. DLCRec: A Novel Approach for Managing Diversity in LLM- Based Recommender Systems. InProceedings of the Eighteenth ACM International Conference on Web Search and Data Mining(Hannover, Germany)(WSDM ’25). Association for Computing Machinery, New York, NY, USA, 857–865....
arXiv 2025
-
[4]
Yu Cui, Feng Liu, Pengbo Wang, Bohao Wang, Heng Tang, Yi Wan, Jun Wang, and Jiawei Chen. 2024. Distillation matters: empowering sequential recommenders to match the performance of large language models. InProceedings of the 18th ACM Conference on Recommender Systems. 507–517
2024
-
[5]
Jiaxin Deng, Shiyao Wang, Kuo Cai, Lejian Ren, Qigen Hu, Weifeng Ding, Qiang Luo, and Guorui Zhou. 2025. Onerec: Unifying retrieve and rank with generative recommender and iterative preference alignment.arXiv preprint arXiv:2502.18965 (2025)
Pith/arXiv arXiv 2025
-
[6]
Shijie Geng, Shuchang Liu, Zuohui Fu, Yingqiang Ge, and Yongfeng Zhang. 2022. Recommendation as language processing (rlp): A unified pretrain, personalized prompt & predict paradigm (p5). InProceedings of the 16th ACM conference on recommender systems. 299–315
2022
-
[7]
Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk
-
[8]
Yupeng Hou, Jiacheng Li, Ashley Shin, Jinsung Jeon, Abhishek Santhanam, Wei Shao, Kaveh Hassani, Ning Yao, and Julian McAuley. 2025. Generating long semantic ids in parallel for recommendation. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 956–966
2025
-
[9]
Yupeng Hou, Jianmo Ni, Zhankui He, Noveen Sachdeva, Wang-Cheng Kang, Ed H Chi, Julian McAuley, and Derek Zhiyuan Cheng. 2025. Actionpiece: Contextually tokenizing action sequences for generative recommendation.arXiv preprint arXiv:2502.13581(2025)
Pith/arXiv arXiv 2025
-
[10]
Kalervo Järvelin and Jaana Kekäläinen. 2002. Cumulated gain-based evaluation of IR techniques.ACM Transactions on Information Systems (TOIS)20, 4 (2002), 422–446
2002
-
[11]
Jian Jia, Yipei Wang, Yan Li, Honggang Chen, Xuehan Bai, Zhaocheng Liu, Jian Liang, Quan Chen, Han Li, Peng Jiang, et al. 2025. LEARN: knowledge adaptation from large language model to recommendation for practical industrial application. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 11861– 11869
2025
-
[12]
Clark Mingxuan Ju, Liam Collins, Leonardo Neves, Bhuvesh Kumar, Louis Yufeng Wang, Tong Zhao, and Neil Shah. 2025. Generative Recommendation with Seman- tic IDs: A Practitioner’s Handbook. InProceedings of the 34th ACM International Conference on Information and Knowledge Management
2025
-
[13]
Wang-Cheng Kang and Julian McAuley. 2018. Self-Attentive Sequential Recom- mendation. In2018 IEEE International Conference on Data Mining. 197–206
2018
-
[14]
Jianghao Lin, Xinyi Dai, Yunjia Xi, Weiwen Liu, Bo Chen, Hao Zhang, Yong Liu, Chuhan Wu, Xiangyang Li, Chenxu Zhu, et al . 2025. How can recommender systems benefit from large language models: A survey.ACM Transactions on Information Systems43, 2 (2025), 1–47
2025
-
[15]
Enze Liu, Bowen Zheng, Cheng Ling, Lantao Hu, Han Li, and Wayne Xin Zhao
-
[16]
Enze Liu, Bowen Zheng, Xiaolei Wang, Wayne Xin Zhao, Jinpeng Wang, Sheng Chen, and Ji-Rong Wen. 2025. Lares: Latent reasoning for sequential recommen- dation.arXiv preprint arXiv:2505.16865(2025)
Pith/arXiv arXiv 2025
-
[17]
Han Liu, Yinwei Wei, Xuemeng Song, Weili Guan, Yuan-Fang Li, and Liqiang Nie. 2024. Mmgrec: Multimodal generative recommendation with transformer model.arXiv preprint arXiv:2404.16555(2024)
arXiv 2024
-
[18]
Zhanyu Liu, Shiyao Wang, Xingmei Wang, Rongzhou Zhang, Jiaxin Deng, Honghui Bao, Jinghao Zhang, Wuchao Li, Pengfei Zheng, Xiangyu Wu, et al
-
[19]
Chen Ma, Peng Kang, and Xue Liu. 2019. Hierarchical Gating Networks for Se- quential Recommendation. InProceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 825–833
2019
-
[20]
Julian McAuley, Christopher Targett, Qinfeng Shi, and Anton van den Hengel
-
[21]
Onerec-think: In-text reasoning for generative recommendation.arXiv preprint arXiv:2510.11639(2025)
arXiv 2025
-
[22]
Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Tran, Jonah Samost, et al
-
[23]
Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang
-
[24]
Jiakai Tang, Sunhao Dai, Teng Shi, Jun Xu, Xu Chen, Wen Chen, Jian Wu, and Yuning Jiang. 2026. Think before recommend: Unleashing the latent reasoning power for sequential recommendation.IEEE Transactions on Knowledge and Data Engineering(2026)
2026
-
[25]
Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho. 2019. Relational Knowledge Distillation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 3967–3976
2019
-
[26]
Zhen Tao, Riwei Lai, Chenyun Yu, Weixin Chen, Li Chen, Beibei Kong, Lei Cheng, Chengxiang Zhuo, Zang Li, and Qingqiang Sun. 2026. SAGER: Self-Evolving User Policy Skills for Recommendation Agent.arXiv preprint arXiv:2604.14972 (2026)
Pith/arXiv arXiv 2026
-
[27]
Frederick Tung and Greg Mori. 2019. Similarity-preserving knowledge distillation. InProceedings of the IEEE/CVF international conference on computer vision. 1365– 1374
2019
-
[28]
Lu Wang, Di Zhang, Fangkai Yang, Pu Zhao, Jianfeng Liu, Yuefeng Zhan, Hao Sun, Qingwei Lin, Weiwei Deng, Dongmei Zhang, et al. 2025. Lettingo: Explore user profile generation for recommendation system. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 2985–2995
2025
-
[29]
Minmao Wang, Xingchen Liu, Shijie Yi, Likang Wu, Hongke Zhao, Fei Pan, Qingpeng Cai, and Peng Jiang. 2026. Hierarchical Semantic RL: Tackling the Problem of Dynamic Action Space for RL-based Recommendations. InProceedings of the ACM Web Conference 2026(United Arab Emirates)(WWW ’26). Association for Computing Machinery, New York, NY, USA, 7857–7868. doi:1...
doi:10.1145/3774904 2026
-
[30]
Wenjie Wang, Honghui Bao, Xinyu Lin, Jizhi Zhang, Yongqi Li, Fuli Feng, See- Kiong Ng, and Tat-Seng Chua. 2024. Learnable Item Tokenization for Generative Recommendation. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management
2024
-
[31]
Jiaxi Tang and Ke Wang. 2018. Personalized Top-N Sequential Recommendation via Convolutional Sequence Embedding. InProceedings of the Eleventh ACM International Conference on Web Search and Data Mining. 565–573
2018
-
[32]
Likang Wu, Zhi Li, Hongke Zhao, Zhenya Huang, Yongqiang Han, Junji Jiang, and Enhong Chen. 2024. Supporting Your Idea Reasonably: A Knowledge-Aware Topic Reasoning Strategy for Citation Recommendation.IEEE Trans. on Knowl. and Data Eng.36, 8 (Aug. 2024), 4275–4289. doi:10.1109/TKDE.2024.3365508
arXiv 2024
-
[33]
Likang Wu, Zhaopeng Qiu, Zhi Zheng, Hengshu Zhu, and Enhong Chen. 2024. Exploring large language model for graph data understanding in online job recommendations. InProceedings of the Thirty-Eighth AAAI Conference on Artifi- cial Intelligence and Thirty-Sixth Conference on Innovative Applications of Arti- ficial Intelligence and Fourteenth Symposium on Ed...
-
[34]
Likang Wu, Zhi Zheng, Zhaopeng Qiu, Hao Wang, Hongchao Gu, Tingjia Shen, Chuan Qin, Chen Zhu, Hengshu Zhu, Qi Liu, Hui Xiong, and Enhong Chen. 2024. A survey on large language models for recommendation.World Wide Web27, 5 (Aug. 2024), 31 pages. doi:10.1007/s11280-024-01291-2
-
[35]
Wujiang Xu, Yunxiao Shi, Zujie Liang, Xuying Ning, Kai Mei, Kun Wang, Xi Zhu, Min Xu, and Yongfeng Zhang. 2025. iagent: Llm agent as a shield between user and recommender systems. InFindings of the Association for Computational Linguistics: ACL 2025. 18056–18084
2025
-
[36]
Xiaochuan Xu, Zeqiu Xu, Peiyang Yu, and Jiani Wang. 2025. Enhancing user intent for recommendation systems via large language models. InInternational Conference on Artificial Intelligence and Machine Learning Research (CAIMLR 2024), Vol. 13635. SPIE, 46–54
2025
-
[37]
Zhipeng Wei, Kuo Cai, Junda She, Jie Chen, Minghao Chen, Yang Zeng, Qiang Luo, Wencong Zeng, Ruiming Tang, Kun Gai, et al. 2026. Oneloc: Geo-aware generative recommender systems for local life service. InProceedings of the Nineteenth ACM International Conference on Web Search and Data Mining. 735–744
2026
-
[38]
Liu Yang, Fabian Paischer, Kaveh Hassani, Jiacheng Li, Shuai Shao, Zhang Gabriel Li, Yun He, Xue Feng, Nima Noorshams, Sem Park, et al . 2024. Unifying gen- erative and dense retrieval for sequential recommendation.arXiv preprint arXiv:2411.18814(2024)
Pith/arXiv arXiv 2024
-
[39]
Chao Yi, Dian Chen, Gaoyang Guo, Jiakai Tang, Jian Wu, Jing Yu, Mao Zhang, Wen Chen, Wenjun Yang, Yujie Luo, et al. 2025. RecGPT-V2 Technical Report. arXiv preprint arXiv:2512.14503(2025)
arXiv 2025
-
[40]
Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhao- jie Gong, Fangda Gu, Michael He, et al. 2024. Actions speak louder than words: Trillion-parameter sequential transducers for generative recommendations.arXiv preprint arXiv:2402.17152(2024). Conference’17, July 2017, Washington, DC, USA Chen et al
Pith/arXiv arXiv 2024
-
[41]
Jianyang Zhai, Zi-Feng Mai, Chang-Dong Wang, Feidiao Yang, Xiawu Zheng, Hui Li, and Yonghong Tian. 2025. Multimodal quantitative language for generative recommendation.arXiv preprint arXiv:2504.05314(2025)
Pith/arXiv arXiv 2025
-
[42]
Fuwei Zhang, Xiaoyu Liu, Dongbo Xi, Jishen Yin, Huan Chen, Peng Yan, Fuzhen Zhuang, and Zhao Zhang. 2026. Multi-aspect cross-modal quantization for generative recommendation. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 16271–16279
2026
-
[43]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. 2025. Qwen3 technical report.arXiv preprint arXiv:2505.09388(2025)
Pith/arXiv arXiv 2025
-
[44]
Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, Fei Huang, and Jingren Zhou
-
[45]
Yang Zhang, Wenxin Xu, Xiaoyan Zhao, Wenjie Wang, Fuli Feng, Xiangnan He, and Tat-Seng Chua. 2025. Reinforced latent reasoning for llm-based recommen- dation.arXiv preprint arXiv:2505.19092(2025)
arXiv 2025
-
[46]
Zijian Zhang, Shuchang Liu, Ziru Liu, Rui Zhong, Qingpeng Cai, Xiangyu Zhao, Chunxu Zhang, Qidong Liu, and Peng Jiang. 2025. LLM-powered user simulator for recommender system. InProceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence and Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence and Fifteenth Sympos...
-
[47]
Keyu Zhao, Fengli Xu, and Yong Li. 2025. Reason-to-recommend: Using interaction-of-thought reasoning to enhance llm recommendation.arXiv preprint arXiv:2506.05069(2025)
Pith/arXiv arXiv 2025
-
[48]
Guorui Zhou, Na Mou, Ying Fan, Qi Pi, Weijie Bian, Chang Zhou, Xiaoqiang Zhu, and Kun Gai. 2019. Deep interest evolution network for click-through rate prediction. InProceedings of the AAAI conference on artificial intelligence, Vol. 33. 5941–5948
2019
-
[49]
Junjie Zhang, Beichen Zhang, Wenqi Sun, Hongyu Lu, Wayne Xin Zhao, Yu Chen, and Ji-Rong Wen. 2025. Slow thinking for sequential recommendation.arXiv preprint arXiv:2504.09627(2025)
Pith/arXiv arXiv 2025
-
[51]
Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models.arXiv preprint arXiv:2506.05176(2025)
Pith/arXiv arXiv 2025
-
[56]
Kun Zhou, Hui Wang, Wayne Xin Zhao, Yutao Zhu, Sirui Wang, Fuzheng Zhang, Zhongyuan Wang, and Ji-Rong Wen. 2020. S3-Rec: Self-Supervised Learning for Sequential Recommendation with Mutual Information Maximization. InPro- ceedings of the 29th ACM International Conference on Information and Knowledge Management. 1893–1902. doi:10.1145/3340531.3411954 A Addi...
arXiv 2020
-
[2015]
InProceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval
Image-Based Recommendations on Styles and Substitutes. InProceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval. 43–52
-
[2016]
In International Conference on Learning Representations
Session-Based Recommendations with Recurrent Neural Networks. In International Conference on Learning Representations
-
[2019]
InProceedings of the 28th ACM International Conference on Information and Knowledge Management
BERT4Rec: Sequential Recommendation with Bidirectional Encoder Rep- resentations from Transformer. InProceedings of the 28th ACM International Conference on Information and Knowledge Management. 1441–1450
-
[2023]
Recommender systems with generative retrieval.Advances in Neural Information Processing Systems36 (2023), 10299–10315
2023
-
[2025]
Generative Recommender with End-to-End Learnable Item Tokenization. InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval. 729–739. doi:10.1145/3726302.3729989
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.