Pith. sign in

REVIEW 4 major objections 6 minor 105 references

Full-Stack Optimized Large Language Models for Lifelong Sequential Behavior Comprehension in Recommendation

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read ReLLaX claims LLMs fail to comprehend long user behavior sequences, and fixes it with a three-level framework: semantic retrieval, soft prompt augmentation, and a fully interactive LoRA variant.

desk verdict Solid, niche-important paper with correct containment algebra and a plausible new LoRA design, but the claim that the full interaction matrix drives the gains is not isolated by the ablations. read the letter →

arxiv 2501.13344 v1 pith:JPZSDK4I submitted 2025-01-23 cs.IR cs.AI

classification cs.IRcs.AI
keywords lifelongsequentialbehaviorincomprehensionLLM4RecCTRpredictionsemanticuserretrievalsoftpromptaugmentationcomponentfully-interactiveLoRAlow-rankadaptationcollaborativeknowledgeinjection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that large language models lose the ability to extract useful signal from user behavior sequences once those sequences grow long, even when the token count stays well below the context window. To fix this, the authors propose ReLLaX, a framework that intervenes at three levels: data, prompt, and parameters. At the data level, semantic user behavior retrieval (SUBR) replaces the most recent behaviors with the most semantically relevant ones toward the target item. At the prompt level, soft prompt augmentation (SPA) injects collaborative knowledge from a conventional recommender into the LLM's token embeddings. At the parameter level, component fully-interactive LoRA (CFLoRA) lets all atom components of the LoRA matrices interact, going beyond the constrained interaction patterns of prior LoRA-based LLM4Rec methods. If the claims hold, ReLLaX offers both a practical recipe for long-history CTR prediction and a unified theoretical view of how LoRA is used across the LLM4Rec literature.

What carries the argument

The central object is the LoRA interaction matrix $W \in \mathbb{R}^{r \times r}$ in the composite form $\Delta\Theta = B W A$. In the decomposed view, $\Delta\Theta = \sum_{i,j} w_{ij} B_i A_j$, each pair of rank-one atom components interacts with weight $w_{ij}$; vanilla LoRA fixes $W = I$, personalized LoRA methods fix $W$ as block-diagonal with per-block scalars, and CFLoRA leaves $W$ fully general. $W$ is generated per data sample by a two-layer MLP that projects the conventional recommender's final representation $h$ (from Equation 21), which is how the method injects collaborative, lifelong-sequence information into the otherwise frozen LLM parameters. This object carries both the expressiveness argument and the theoretical unification of prior LLM4Rec LoRA uses.

What would settle it

Run a matched comparison on a sparse dataset such as BookCrossing: train ReLLaX once with $W$ generated by the conventional recommender and once with $W$ fixed to the identity matrix (vanilla LoRA), keeping all other settings identical. If the AUC gap does not persist across random seeds and especially on cold-user subsets, the claim that full atom-component interaction drives the gain collapses. A second check is to shuffle the conventional recommender's embeddings before generating $W$; if ReLLaX's performance does not drop, the $W$ generator is not actually using collaborative signal.

Watch

Extended reading notes

Core claim

The paper identifies a specific failure mode, lifelong sequential behavior incomprehension, in which LLM-based pointwise CTR scorers peak at short histories (around 15 items for MovieLens-1M) and then decline as the sequence lengthens, despite the context staying within the model's window. The proposed fix, ReLLaX, combines three mechanisms: SUBR homogenizes the behavior sequence by retrieving the most semantically relevant items for the target; SPA appends soft prompt tokens derived from a pretrained conventional recommender's item embeddings, aligning the language space with collaborative signal; and CFLoRA extends LoRA by inserting a per-sample interaction matrix $W$ between the down- and up-projections, so the update becomes $\Delta\Theta = B W A$. The paper further argues theoretically that existing LoRA-based LLM4Rec methods are degraded versions of CFLoRA: vanilla LoRA uses $W = I$, and personalized variants like iLoRA use a block-diagonal $W$, whereas CFLoRA leaves $W$ fully general and learns it from the conventional recommender's final representation. Experiments on BookCrossing, MovieLens-1M, and MovieLens-25M show ReLLaX achieving the best AUC among all baselines, including ReLLa, iLoRA, and CoLLM, and the ablation study attributes the gain to all three components working together.

Load-bearing premise

The whole CFLoRA argument rests on the assumption that the conventional recommender's final representation $h$, trained with plain binary cross-entropy, is informative enough to generate a per-sample interaction matrix $W$ that generalizes to users and items the recommender has never seen.

Editorial extensions

If this is right

  • If ReLLaX is correct, LLMs can be made to keep improving as user history grows, matching conventional models like SIM instead of peaking and degrading.
  • The CFLoRA formulation gives a common language for LoRA-based LLM4Rec: every existing method becomes a specific constraint on the interaction matrix $W$, which makes comparing methods a matter of comparing matrix structures.
  • SPA shows that injecting collaborative embeddings as soft tokens can help an LLM reason about item relationships, pointing to a general design where conventional recommenders supply side information to language models.
  • The framework still operates under few-shot instruction tuning, so the reported gains are achievable with relatively little training data while beating full-shot ID-based baselines.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's unification claim suggests a practical benchmark recipe: report every LoRA variant by its constraint on $W$ (identity, block-diagonal, or full). This would make future LLM4Rec comparisons more structural than the paper itself demonstrates.
  • The heterogeneity metric used to justify SUBR (number of unique movie genres) is only a proxy; a direct measure of token-level attention or perplexity over retrieved versus recent sequences would make the comprehension benefit more inspectable.
  • Because the LLM baselines are trained with only a few thousand samples while the conventional recommenders use the full training set, part of ReLLaX's margin may reflect sample efficiency rather than comprehension per se; a full-shot LLM comparison would test that interpretation.
  • If CFLoRA's gains come from full interaction, a natural extension is to learn $W$ end-to-end without a conventional recommender, which would test whether the collaborative signal is necessary or merely a convenient generator.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper formulates the 'lifelong sequential behavior incomprehension' problem for LLM-based CTR prediction, in which LLMs fail to exploit long textual user behavior sequences even within their context window. It proposes ReLLaX, a three-level framework: SUBR (semantic retrieval of user behaviors), SPA (soft prompts derived from a conventional recommendation model), and CFLoRA (LoRA with a per-sample full interaction matrix W between down/up projection components). The authors show algebraically that existing LoRA-based LLM4Rec methods correspond to constrained versions of CFLoRA (W=I or block-diagonal W), and report experiments on BookCrossing, MovieLens-1M, and MovieLens-25M where ReLLaX outperforms ID-based and LM-based baselines.

Significance. If the empirical claims hold, the paper makes two contributions: a reusable diagnosis of why long textual sequences hurt LLM recommenders, and a unified view of LoRA variants in LLM4Rec. The theoretical containment result (Eqs. 13-20) is clean and correct: identity, block-diagonal, and full W form a clear hierarchy, and the 'degraded versions' framing is a useful organizing perspective. The framework is accompanied by code and evaluated on three public datasets. However, the practical significance of the full-interaction W is not yet established: the key ablation replaces CFLoRA with vanilla LoRA (W=I) rather than with a per-sample diagonal W, so the empirical advantage of the full matrix over intermediate constraints is untested. The lack of error bars and the imported ReLLa baseline further weaken the state-of-the-art claim as currently presented.

major comments (4)
  1. [3.4.2 / Table 4] The central claim that the full interaction matrix W is the source of CFLoRA's gains is not directly tested. The ablation 'ReLLaX (w/o CFLoRA)' falls back to vanilla LoRA, i.e. W=I; it does not replace W with a per-sample diagonal (or block-diagonal) matrix, which is exactly the iLoRA/RecLoRA constraint. Thus the comparison cannot separate the effect of full W from the effect of per-sample adaptation. Please add an ablation that keeps SUBR and SPA fixed and compares full W against diagonal W (and, if feasible, block-diagonal W) generated by the same projector. If a diagonal W matches full W, the 'degraded versions' claim is mathematically clean but has no demonstrated performance consequence.
  2. [4.2 / Table 3 / Appendix B.2] The ReLLa baseline results are imported verbatim from the prior conference paper, while ReLLaX and the other baselines are executed in the present setup. Because ReLLa is one of the two strongest LM baselines in Table 3 (e.g., 0.8033 vs 0.8091 AUC on MovieLens-1M), a difference in hardware, library versions, prompt formatting, or data preprocessing can bias the comparison in either direction. Please re-run ReLLa under the same code and environment as ReLLaX and report both sets of numbers.
  3. [4.2 / Table 3] No significance-test procedure is described. The table reports '*' for p<0.001, but does not state which test was used, whether the comparison is paired, or how many seeds/runs are aggregated. Given the small margins at stake (e.g., 0.0005 AUC between some Table 4 variants), please report means and standard deviations over at least 3 random seeds for the main table and ablations, and specify the test used for the asterisks.
  4. [3.4.3 / Eq. (21) / Table 2] On BookCrossing the CRM is trained on only 17,714 samples but must produce per-sample interaction matrices W for a dataset with 278,858 users. The paper does not analyze how CFLoRA behaves for cold users or when the CRM's final representation h is unreliable. Please report performance broken down by user history length or include a cold-start experiment; at minimum, discuss this limitation explicitly.
minor comments (6)
  1. [2.1] The sentence 'Each user u has a has a chronological interaction sequence' contains a duplicated phrase, and 'the objective of is to predict' is missing a word.
  2. [3.4.1 / Eq. (16)] The notation alpha_{j//N} is not defined; clarify how the block index maps to the r1-dimension blocks after merging N LoRA sets.
  3. [Figure 1] Figure 1 uses two different y-axes (left for SIM, right for the LLMs) without explicit axis labels in the caption; the reader cannot immediately tell the scales.
  4. [4.4.1] The text says ReLLa (few-shot) achieves continuous improvement as K increases; consider stating explicitly the K ranges for which this holds, since the curves in Figure 6 are not trivially monotonic in all panels.
  5. [Appendix B.2] The statement that ReLLa results are taken directly from the original paper should be moved to or repeated in the main text as a limitation of the comparison.
  6. [Throughout] There are several typos ('comparision', 'sequnce', 'Ragnaro', 'but also but also'); a careful proofread is needed.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity; the 'degraded versions' theorem is a definitional containment, and the SOTA claim rests on independent experiments.

full rationale

The paper's central empirical claim—that ReLLaX, combining SUBR, SPA and CFLoRA, achieves state-of-the-art CTR prediction—is supported by direct experiments (Tables 3 and 4) against external baselines and ablations, and none of the reported metrics reduce to a fitted parameter or to a self-citation. The 'theoretical demonstration' that existing LoRA-based methods are degraded versions of CFLoRA (Section 3.4.2, Eqs. 18-19) is a definitional containment: CFLoRA is defined as ΔΘ = BWA with an unconstrained W, so setting W = I (vanilla LoRA) or W = block-diagonal (iLoRA/RecLoRA) yields the existing forms by construction. This is a valid algebraic reformulation rather than a circular derivation, and the paper does not use this containment alone to infer performance; the performance case rests on the independent experiments. The only notable self-citation is the reuse of ReLLa baseline results from the authors' own conference paper (Appendix B), but this is a transparent baseline choice and is corroborated by re-run ReLLa variants in Table 4, so it is not load-bearing. The skeptic's concern about the missing ablation that swaps full W for a per-sample diagonal W while holding SUBR and SPA fixed is a legitimate confound in attributing the gain to the full interaction matrix, but it is an experimental-design issue, not a circularity. Overall, the derivation chain is self-contained; no prediction reduces to its input by construction.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central empirical claim rests on a pretrained conventional recommender (CRM) whose representation h drives both the soft prompt projectors (SPA) and the per-sample W in CFLoRA (Eq. 21); if h is uninformative, both modules lose value. SUBR also assumes LLM semantic cosine similarity approximates recommendation relevance. The algebraic unification is self-contained matrix algebra and adds no hidden parameters.

free parameters (4)
  • PCA dimension d_q = 512
    Section 3.2: dimension set by hand for semantic item vectors; central to SUBR's cosine retrieval.
  • Textual sequence length K per dataset = 60 (BookCrossing), 30 (MovieLens-1M), 30 (MovieLens-25M)
    Section 4.1.4: chosen as the largest length fitting within about 700 tokens; performance results in Table 3 depend on these lengths.
  • LoRA hyperparameters = rank 8, alpha 16, dropout 0.05
    Section 4.1.4: values taken from prior work; CFLoRA adds a per-sample r by r matrix W, so the effective trainable parameters depend on these choices.
  • Training data subsample sizes = 2048/65536/65536 for Table 3; 1024/8192/8192 for Table 4
    All LLM baselines and ReLLaX are trained on few-shot subsets, which affects absolute performance and the "data efficiency" claim.
assumptions (4)
  • domain assumption Averaged last-layer LLM hidden states plus PCA encode item semantics usable for cosine similarity
    Section 3.2: semantic item encoding; assumed to be a valid relevance signal for SUBR.
  • domain assumption The pretrained CRM's representation h carries enough collaborative signal to generate both soft prompts (SPA) and the interaction matrix W (CFLoRA)
    Sections 3.3 and 3.4: if h is noisy, both modules degrade.
  • domain assumption CRM embeddings and h generalize to unseen users and items, including cold-start cases
    The method assumes that the pretrained CRM trained on the training set remains informative for all test samples; sparse datasets like BookCrossing stress this.
  • standard math Standard matrix algebra: LoRA update is BA, and CFLoRA update is BWA with W from a projector
    Section 3.4.1-3.4.2: algebraic identity used for the unification argument.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Full-Stack Optimized Large Language Models for Lifelong Sequential Behavior Comprehension in Recommendation." pith.science (2026). https://pith.science/paper/JPZSDK4I

@misc{pith2026250113344,
  author       = {Pith},
  title        = {Pith review of: Full-Stack Optimized Large Language Models for Lifelong Sequential Behavior Comprehension in Recommendation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JPZSDK4I}},
  note         = {Machine review of arXiv:2501.13344}
}
read the original abstract

In this paper, we address the lifelong sequential behavior incomprehension problem in large language models (LLMs) for recommendation, where LLMs struggle to extract useful information from long user behavior sequences, even within their context limits. To tackle this, we propose ReLLaX (Retrieval-enhanced Large Language models Plus), a framework offering optimization across data, prompt, and parameter levels. At the data level, we introduce Semantic User Behavior Retrieval (SUBR) to reduce sequence heterogeneity, making it easier for LLMs to extract key information. For prompt-level enhancement, we employ Soft Prompt Augmentation (SPA) to inject collaborative knowledge, aligning item representations with recommendation tasks and improving LLMs's exploration of item relationships. Finally, at the parameter level, we propose Component Fully-interactive LoRA (CFLoRA), which enhances LoRA's expressiveness by enabling interactions between its components, allowing better capture of sequential information. Moreover, we present new perspectives to compare current LoRA-based LLM4Rec methods, i.e. from both a composite and a decomposed view. We theoretically demonstrate that the ways they employ LoRA for recommendation are degraded versions of our CFLoRA, with different constraints on atom component interactions. Extensive experiments on three public datasets demonstrate ReLLaX's superiority over existing baselines and its ability to mitigate lifelong sequential behavior incomprehension effectively.

Figures

Figures reproduced from arXiv: 2501.13344 by the authors.

Figure 1
Figure 1. The illustration of lifelong sequential behavior incomprehension problem for LLMs. We report the AUC [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Comprehensive comparison of the ways different LLM4Rec methods employ LoRA for finetuning. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. The overall framework of our proposed ReLLaX, which proposes a full-stack optimization for LLMs to [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Illustration of descriptive text for an item (movie). [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Examples of prompt templates for the MovieLens-1M dataset. [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 2
Figure 2. Figure 2: 3.4.1 Decomposed View. First, we examine the LoRA matrices from a decomposed view, where we focus on the vectors composing the matrices. As stated in Section 2.4, the vanilla LoRA approximates the update of weights by the multiplication of the up-projection matrix 𝐵 ∈ …
Figure 6
Figure 6. Figure 6: The AUC performance of different models w.r.t. different length of user behavior sequence [PITH_FULL_IMAGE:figures/full_fig_p017_6.png]
Figure 7
Figure 7. Figure 7: The performance of ReLLaX w.r.t. different [PITH_FULL_IMAGE:figures/full_fig_p018_7.png]
Figure 8
Figure 8. Figure 8: Case study of ReLLaX on the MovieLens-25M dataset. We visualize the attention scores for the items [PITH_FULL_IMAGE:figures/full_fig_p019_8.png]
Figure 9
Figure 9. Figure 9: Examples of vanilla prompt templates for three datasets [PITH_FULL_IMAGE:figures/full_fig_p026_9.png]
Figure 10
Figure 10. Figure 10: Examples of retrieval-enhanced prompt templates for three datasets [PITH_FULL_IMAGE:figures/full_fig_p027_10.png]
Figure 11
Figure 11. Figure 11: Examples of comprehensive prompt templates for three datasets [PITH_FULL_IMAGE:figures/full_fig_p028_11.png]
Figure 12
Figure 12. Figure 12: Examples of hard prompt templates of item descriptions for three datasets. The textual description is [PITH_FULL_IMAGE:figures/full_fig_p028_12.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

105 extracted references · 31 canonical work pages

  1. [1]

    Keqin Bao, Ming Yan, Yang Zhang, Jizhi Zhang, Wenjie Wang, Fuli Feng, and Xiangnan He. 2024. Real-Time Person- alization for LLM-based Recommendation with Customized In-Context Learning. arXiv preprint arXiv:2410.23136 (2024)

  2. [2]

    Keqin Bao, Jizhi Zhang, Yang Zhang, Wenjie Wang, Fuli Feng, and Xiangnan He. 2023. Tallrec: An effective and efficient tuning framework to align large language model with recommendation. arXiv preprint arXiv:2305.00447 (2023)

  3. [3]

    Vadim Borisov, Kathrin Seßler, Tobias Leemann, Martin Pawelczyk, and Gjergji Kasneci. 2022. Language models are realistic tabular data generators. arXiv preprint arXiv:2210.06280 (2022)

  4. [4]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901

  5. [5]

    Aldo Gael Carranza, Rezsa Farahani, Natalia Ponomareva, Alex Kurakin, Matthew Jagielski, and Milad Nasr. 2023. Privacy-Preserving Recommender Systems with Synthetic Query Generation using Differentially Private Large J. ACM, Vol. 1, No. 1, Article . Publication date: January 2025. 22 Rong Shan et al. Language Models. arXiv preprint arXiv:2305.05973 (2023)

  6. [6]

    Yin-Wen Chang, Cho-Jui Hsieh, Kai-Wei Chang, Michael Ringgaard, and Chih-Jen Lin. 2010. Training and testing low-degree polynomial data mappings via linear SVM. Journal of Machine Learning Research 11, 4 (2010)

  7. [7]

    Bo Chen, Yichao Wang, Zhirong Liu, Ruiming Tang, Wei Guo, Hongkun Zheng, Weiwei Yao, Muyu Zhang, and Xiuqiang He. 2021. Enhancing explicit and implicit feature interactions via information sharing for parallel deep ctr models. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management . 3757–3766

  8. [8]

    Zheng Chen. 2023. PALR: Personalization Aware LLMs for Recommendation. arXiv preprint arXiv:2305.07622 (2023)

Show all 105 references
  1. [9]

    Zhenyi Lu Chenghao Fan and Jie Tian. 2023. Chinese-Vicuna: A Chinese Instruction-following LLaMA-based Model. https://github.com/Facico/Chinese-Vicuna

  2. [10]

    Gonzalez, Ion Stoica, and Eric P

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. https://lmsys.org/blog/2023-...

  3. [11]

    Konstantina Christakopoulou, Alberto Lalama, Cj Adams, Iris Qu, Yifat Amir, Samer Chucri, Pierce Vollucci, Fabio Soldo, Dina Bseiso, Sarah Scodel, et al . 2023. Large Language Models for User Interest Journeys. arXiv preprint arXiv:2305.15498 (2023)

  4. [12]

    Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. 2014. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555 (2014)

  5. [13]

    Zeyu Cui, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022. M6-rec: Generative pretrained language models are open-ended recommender systems. arXiv preprint arXiv:2205.08084 (2022)

  6. [14]

    Sunhao Dai, Ninglu Shao, Haiyuan Zhao, Weijie Yu, Zihua Si, Chen Xu, Zhongxiang Sun, Xiao Zhang, and Jun Xu

  7. [15]

    Xinyi Dai, Jianghao Lin, Weinan Zhang, Shuai Li, Weiwen Liu, Ruiming Tang, Xiuqiang He, Jianye Hao, Jun Wang, and Yong Yu. 2021. An adversarial imitation click model for information retrieval. InProceedings of the Web Conference

  8. [16]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)

  9. [17]

    Junchen Fu, Fajie Yuan, Yu Song, Zheng Yuan, Mingyue Cheng, Shenghui Cheng, Jiaqi Zhang, Jie Wang, and Yunzhu Pan. 2023. Exploring Adapter-based Transfer Learning for Recommender Systems: Empirical Studies and Practical Insights. arXiv preprint arXiv:2305.15036 (2023)

  10. [18]

    Lingyue Fu, Jianghao Lin, Weiwen Liu, Ruiming Tang, Weinan Zhang, Rui Zhang, and Yong Yu. 2023. An F-shape Click Model for Information Retrieval on Multi-block Mobile Pages. In Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining . 1057–1065

  11. [20]

    Shijie Geng, Shuchang Liu, Zuohui Fu, Yingqiang Ge, and Yongfeng Zhang. 2022. Recommendation as language processing (rlp): A unified pretrain, personalized prompt & predict paradigm (p5). In Proceedings of the 16th ACM Conference on Recommender Systems . 299–315

  12. [21]

    Shijie Geng, Juntao Tan, Shuchang Liu, Zuohui Fu, and Yongfeng Zhang. 2023. VIP5: Towards Multimodal Foundation Models for Recommendation. arXiv preprint arXiv:2305.14302 (2023)

  13. [22]

    Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017. DeepFM: a factorization-machine based neural network for CTR prediction. In Proceedings of the 26th International Joint Conference on Artificial Intelligence . 1725–1731

  14. [23]

    Balázs Hidasi and Alexandros Karatzoglou. 2018. Recurrent neural networks with top-k gains for session-based recommendations. CIKM (2018)

  15. [24]

    Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk. 2016. Session-based recommendations with recurrent neural networks. In ICLR

  16. [25]

    Yupeng Hou, Zhankui He, Julian McAuley, and Wayne Xin Zhao. 2023. Learning vector-quantized item representation for transferable sequential recommenders. In Proceedings of the ACM Web Conference 2023 . 1162–1171

  17. [26]

    Yupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li, Bolin Ding, and Ji-Rong Wen. 2022. Towards universal sequence representation learning for recommender systems. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 585–593

  18. [27]

    Yupeng Hou, Junjie Zhang, Zihan Lin, Hongyu Lu, Ruobing Xie, Julian McAuley, and Wayne Xin Zhao. 2023. Large language models are zero-shot rankers for recommender systems. arXiv preprint arXiv:2305.08845 (2023)

  19. [28]

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

  20. [29]

    Wenyue Hua, Yingqiang Ge, Shuyuan Xu, Jianchao Ji, and Yongfeng Zhang. 2023. UP5: Unbiased Foundation Model for Fairness-aware Recommendation. arXiv preprint arXiv:2305.12090 (2023)

  21. [30]

    Wenyue Hua, Shuyuan Xu, Yingqiang Ge, and Yongfeng Zhang. 2023. How to Index Item IDs for Recommendation Foundation Models. arXiv preprint arXiv:2305.06569 (2023)

  22. [31]

    Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023)

  23. [32]

    Wang-Cheng Kang and Julian McAuley. 2018. Self-Attentive Sequential Recommendation. ICDM (2018)

  24. [33]

    Wang-Cheng Kang, Jianmo Ni, Nikhil Mehta, Maheswaran Sathiamoorthy, Lichan Hong, Ed Chi, and Derek Zhiyuan Cheng. 2023. Do LLMs Understand User Preferences? Evaluating LLMs On User Rating Prediction. arXiv preprint arXiv:2305.06474 (2023)

  25. [34]

    Xiaoyu Kong, Jiancan Wu, An Zhang, Leheng Sheng, Hui Lin, Xiang Wang, and Xiangnan He. 2024. Customizing Language Models with Instance-wise LoRA for Sequential Recommendation. arXiv preprint arXiv:2408.10159 (2024)

  26. [35]

    Chen Li, Yixiao Ge, Jiayong Mao, Dian Li, and Ying Shan. 2023. TagGPT: Large Language Models are Zero-shot Multimodal Taggers. arXiv preprint arXiv:2304.03022 (2023)

  27. [36]

    Jiacheng Li, Ming Wang, Jin Li, Jinmiao Fu, Xin Shen, Jingbo Shang, and Julian McAuley. 2023. Text Is All You Need: Learning Language Representations for Sequential Recommendation. arXiv preprint arXiv:2305.13731 (2023)

  28. [37]

    Ruyu Li, Wenhao Deng, Yu Cheng, Zheng Yuan, Jiaqi Zhang, and Fajie Yuan. 2023. Exploring the Upper Limits of Text- Based Collaborative Filtering Using Large Language Models: Discoveries and Insights. arXiv preprint arXiv:2305.11700 (2023)

  29. [38]

    Xinyi Li, Yongfeng Zhang, and Edward C Malthouse. 2023. PBNR: Prompt-based News Recommender System. arXiv preprint arXiv:2304.07862 (2023)

  30. [39]

    Zeyu Li, Wei Cheng, Yang Chen, Haifeng Chen, and Wei Wang. 2020. Interpretable click-through rate prediction through hierarchical attention. In Proceedings of the 13th International Conference on Web Search and Data Mining . 313–321

  31. [40]

    Jianxun Lian, Xiaohuan Zhou, Fuzheng Zhang, Zhongxia Chen, Xing Xie, and Guangzhong Sun. 2018. xdeepfm: Combining explicit and implicit feature interactions for recommender systems. In KDD. 1754–1763

  32. [41]

    Jiayi Liao, Sihang Li, Zhengyi Yang, Jiancan Wu, Yancheng Yuan, Xiang Wang, and Xiangnan He. 2024. Llara: Large language-recommendation assistant. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 1785–1795

  33. [42]

    Jianghao Lin, Bo Chen, Hangyu Wang, Yunjia Xi, Yanru Qu, Xinyi Dai, Kangning Zhang, Ruiming Tang, Yong Yu, and Weinan Zhang. 2024. ClickPrompt: CTR Models are Strong Prompt Generators for Adapting Language Models to CTR Prediction. In Proceedings of the ACM on Web Conference 2...

  34. [43]

    Jianghao Lin, Xinyi Dai, Yunjia Xi, Weiwen Liu, Bo Chen, Hao Zhang, Yong Liu, Chuhan Wu, Xiangyang Li, Chenxu Zhu, Huifeng Guo, Yong Yu, Ruiming Tang, and Weinan Zhang. 2024. How Can Recommender Systems Benefit from Large Language Models: A Survey. ACM Trans. Inf. Syst. (jul 2...

  35. [44]

    Jianghao Lin, Weiwen Liu, Xinyi Dai, Weinan Zhang, Shuai Li, Ruiming Tang, Xiuqiang He, Jianye Hao, and Yong Yu

  36. [45]

    Jianghao Lin, Yanru Qu, Wei Guo, Xinyi Dai, Ruiming Tang, Yong Yu, and Weinan Zhang. 2023. MAP: A Model- agnostic Pretraining Framework for Click-through Rate Prediction. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 1384–1395

  37. [46]

    Jianghao Lin, Rong Shan, Chenxu Zhu, Kounianhua Du, Bo Chen, Shigang Quan, Ruiming Tang, Yong Yu, and Weinan Zhang. 2024. Rella: Retrieval-enhanced large language models for lifelong sequential behavior comprehension in recommendation. In Proceedings of the ACM on Web Conferen...

  38. [47]

    In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval

    A Graph-Enhanced Click Model for Web Search. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval . 1259–1268

  39. [48]

    Guang Liu, Jie Yang, and Ledell Wu. 2022. PTab: Using the Pre-trained Language Model for Modeling Tabular Data. arXiv preprint arXiv:2209.08060 (2022)

  40. [49]

    Junling Liu, Chao Liu, Renjie Lv, Kang Zhou, and Yan Zhang. 2023. Is chatgpt a good recommender? a preliminary study. arXiv preprint arXiv:2304.10149 (2023)

  41. [50]

    Bin Liu, Ruiming Tang, Yingzhi Chen, Jinkai Yu, Huifeng Guo, and Yuzhou Zhang. 2019. Feature generation by convolutional neural network for click-through rate prediction. In WWW. 1119–1129

  42. [51]

    Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101 (2017)

  43. [52]

    Zhiming Mao, Huimin Wang, Yiming Du, and Kam-fai Wong. 2023. UniTRec: A Unified Text-to-Text Transformer and Joint Contrastive Learning Framework for Text-based Recommendation. arXiv preprint arXiv:2305.15756 (2023). J. ACM, Vol. 1, No. 1, Article . Publication date: January 2...

  44. [53]

    Qijiong Liu, Nuo Chen, Tetsuya Sakai, and Xiao-Ming Wu. 2023. A First Look at LLM-Powered Generative News Recommendation. arXiv preprint arXiv:2305.06566 (2023)

  45. [54]

    Aashiq Muhamed, Iman Keivanloo, Sujan Perera, James Mracek, Yi Xu, Qingjun Cui, Santosh Rajagopalan, Belinda Zeng, and Trishul Chilimbi. 2021. CTR-BERT: Cost-effective knowledge distillation for billion-parameter teacher models. In NeurIPS Efficient Natural Language and Speech...

  46. [55]

    Sheshera Mysore, Andrew McCallum, and Hamed Zamani. 2023. Large Language Model Augmented Narrative Driven Recommendations. arXiv preprint arXiv:2306.02250 (2023)

  47. [57]

    Rajvardhan Patil and Venkat Gudivada. 2024. A review of current trends, techniques, and challenges in large language models (llms). Applied Sciences 14, 5 (2024), 2074

  48. [58]

    Aleksandr V Petrov and Craig Macdonald. 2023. Generative Sequential Recommendation with GPTRec. arXiv preprint arXiv:2306.11114 (2023)

  49. [59]

    Weitong Ou, Bo Chen, Yingxuan Yang, Xinyi Dai, Weiwen Liu, Weinan Zhang, Ruiming Tang, and Yong Yu. 2023. Deep Landscape Forecasting in Multi-Slot Real-Time Bidding. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 4685–4695

  50. [60]

    Qi Pi, Guorui Zhou, Yujing Zhang, Zhe Wang, Lejian Ren, Ying Fan, Xiaoqiang Zhu, and Kun Gai. 2020. Search-based user interest modeling with lifelong sequential behavior data for click-through rate prediction. In Proceedings of the 29th ACM International Conference on Informat...

  51. [61]

    Jiarui Qin, Weinan Zhang, Rong Su, Zhirong Liu, Weiwen Liu, Ruiming Tang, Xiuqiang He, and Yong Yu. 2021. Retrieval & Interaction Machine for Tabular Data Prediction. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining . 1379–1389

  52. [62]

    Qi Pi, Weijie Bian, Guorui Zhou, Xiaoqiang Zhu, and Kun Gai. 2019. Practice on long sequential user behavior modeling for click-through rate prediction. In KDD. 2671–2679

  53. [63]

    Zhaopeng Qiu, Xian Wu, Jingyue Gao, and Wei Fan. 2021. U-BERT: Pre-training user representations for improved recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 4320–4327

  54. [64]

    Yanru Qu, Han Cai, Kan Ren, Weinan Zhang, Yong Yu, Ying Wen, and Jun Wang. 2016. Product-based neural networks for user response prediction. In 2016 IEEE 16th international conference on data mining (ICDM) . IEEE, 1149–1154

  55. [65]

    Libo Qin, Qiguang Chen, Xiachong Feng, Yang Wu, Yongheng Zhang, Yinghui Li, Min Li, Wanxiang Che, and Philip S Yu. 2024. Large language models meet nlp: A survey. arXiv preprint arXiv:2405.12819 (2024)

  56. [66]

    Kan Ren, Jiarui Qin, Yuchen Fang, Weinan Zhang, Lei Zheng, Weijie Bian, Guorui Zhou, Jian Xu, Yong Yu, Xiaoqiang Zhu, et al. 2019. Lifelong Sequential Modeling with Personalized Memorization for User Response Prediction. SIGIR

  57. [67]

    Ruiyang Ren, Yuhao Wang, Kun Zhou, Wayne Xin Zhao, Wenjie Wang, Jing Liu, Ji-Rong Wen, and Tat-Seng Chua

  58. [68]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research 21, 1 (2020), 5485–5551

  59. [69]

    Rong Shan, Jianghao Lin, Chenxu Zhu, Bo Chen, Menghui Zhu, Kangning Zhang, Jieming Zhu, Ruiming Tang, Yong Yu, and Weinan Zhang. 2024. An Automatic Graph Construction Framework based on Large Language Models for Recommendation. arXiv preprint arXiv:2412.18241 (2024)

  60. [70]

    Jonathon Shlens. 2014. A tutorial on principal component analysis. arXiv preprint arXiv:1404.1100 (2014)

  61. [71]

    Weiping Song, Chence Shi, Zhiping Xiao, Zhijian Duan, Yewen Xu, Ming Zhang, and Jian Tang. 2019. Autoint: Automatic feature interaction learning via self-attentive neural networks. InProceedings of the 28th ACM International Conference on Information and Knowledge Management ....

  62. [72]

    Steffen Rendle. 2010. Factorization machines. In ICDM

  63. [73]

    Jiaxi Tang and Ke Wang. 2018. Personalized top-n sequential recommendation via convolutional sequence embedding. In Proceedings of the eleventh ACM international conference on web search and data mining . 565–573

  64. [74]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)

  65. [75]

    Hangyu Wang, Jianghao Lin, Xiangyang Li, Bo Chen, Chenxu Zhu, Ruiming Tang, Weinan Zhang, and Yong Yu. 2024. FLIP: Fine-grained Alignment between ID-based Models and Pretrained Language Models for CTR Prediction. In Proceedings of the 18th ACM Conference on Recommender Systems...

  66. [76]

    Weiwei Sun, Lingyong Yan, Xinyu Ma, Pengjie Ren, Dawei Yin, and Zhaochun Ren. 2023. Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agent. arXiv preprint arXiv:2304.09542 (2023)

  67. [77]

    Lei Wang and Ee-Peng Lim. 2023. Zero-Shot Next-Item Recommendation using Large Pretrained Language Models. arXiv preprint arXiv:2304.03153 (2023)

  68. [78]

    Ruoxi Wang, Bin Fu, Gang Fu, and Mingliang Wang. 2017. Deep & cross network for ad click predictions. In Proceedings of the ADKDD’17 . 1–7

  69. [79]

    Ruoxi Wang, Rakesh Shivanna, Derek Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed Chi. 2021. Dcn v2: Improved deep & cross network and practical lessons for web-scale learning to rank systems. InProceedings of the Web Conference

  70. [80]

    Jie Wang, Fajie Yuan, Mingyue Cheng, Joemon M Jose, Chenyun Yu, Beibei Kong, Xiangnan He, Zhijin Wang, Bo Hu, and Zang Li. 2022. TransRec: Learning Transferable Recommendation from Mixture-of-Modality Feedback. arXiv preprint arXiv:2206.06190 (2022). J. ACM, Vol. 1, No. 1, Art...

  71. [81]

    Wenjie Wang, Fuli Feng, Xiangnan He, Liqiang Nie, and Tat-Seng Chua. 2021. Denoising implicit feedback for recommendation. In Proceedings of the 14th ACM international conference on web search and data mining . 373–381

  72. [82]

    Yancheng Wang, Ziyan Jiang, Zheng Chen, Fan Yang, Yingxue Zhou, Eunah Cho, Xing Fan, Xiaojiang Huang, Yanbin Lu, and Yingzhen Yang. 2023. Recmind: Large language model powered agent for recommendation. arXiv preprint arXiv:2308.14296 (2023)

  73. [83]

    Zongwei Wang, Min Gao, Wentao Li, Junliang Yu, Linxin Guo, and Hongzhi Yin. 2023. Efficient bi-level optimization for recommendation denoising. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2502–2511

  74. [84]

    Wenjie Wang, Honghui Bao, Xinyu Lin, Jizhi Zhang, Yongqi Li, Fuli Feng, See-Kiong Ng, and Tat-Seng Chua. 2024. Learnable item tokenization for generative recommendation. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management . 2400–2409

  75. [85]

    Yunjia Xi, Weiwen Liu, Jianghao Lin, Jieming Zhu, Bo Chen, Ruiming Tang, Weinan Zhang, Rui Zhang, and Yong Yu

  76. [86]

    Jun Xiao, Hao Ye, Xiangnan He, Hanwang Zhang, Fei Wu, and Tat-Seng Chua. 2017. Attentional factorization machines: learning the weight of feature interactions via attention networks. In Proceedings of the 26th International Joint Conference on Artificial Intelligence . 3119–3125

  77. [87]

    Xin Xin, Bo Chen, Xiangnan He, Dong Wang, Yue Ding, and Joemon M Jose. 2019. CFM: Convolutional Factorization Machines for Context-Aware Recommendation.. In IJCAI, Vol. 19. 3926–3932

  78. [88]

    Yunjia Xi, Jianghao Lin, Weiwen Liu, Xinyi Dai, Weinan Zhang, Rui Zhang, Ruiming Tang, and Yong Yu. 2023. A Bird’s-eye View of Reranking: from List Level to Page Level. In Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining . 1075–1083

  79. [89]

    Zheng Yuan, Fajie Yuan, Yu Song, Youhua Li, Junchen Fu, Fei Yang, Yunzhu Pan, and Yongxin Ni. 2023. Where to go next for recommender systems? id-vs. modality-based recommender models revisited. arXiv preprint arXiv:2303.13835 (2023)

  80. [90]

    arXiv preprint arXiv:2306.10933 (2023)

    Towards Open-World Recommendation with Knowledge Augmentation from Large Language Models. arXiv preprint arXiv:2306.10933 (2023)

  81. [91]

    Junjie Zhang, Ruobing Xie, Yupeng Hou, Wayne Xin Zhao, Leyu Lin, and Ji-Rong Wen. 2023. Recommendation as instruction following: A large language model empowered recommendation approach.arXiv preprint arXiv:2305.07001 (2023)

  82. [92]

    Kai Zhang, Fubang Zhao, Yangyang Kang, and Xiaozhong Liu. 2023. Memory-augmented llm personalization with short-and long-term memory coordination. arXiv preprint arXiv:2309.11696 (2023)

  83. [93]

    Yang Yu, Fangzhao Wu, Chuhan Wu, Jingwei Yi, and Qi Liu. 2021. Tiny-newsrec: Effective and efficient plm-based news recommendation. arXiv preprint arXiv:2112.00944 (2021)

  84. [94]

    Xinyang Zhang, Yury Malkov, Omar Florez, Serim Park, Brian McWilliams, Jiawei Han, and Ahmed El-Kishky. 2022. TwHIN-BERT: a socially-enriched pre-trained language model for multilingual Tweet representations. arXiv preprint arXiv:2209.07562 (2022)

  85. [95]

    Jizhi Zhang, Keqin Bao, Yang Zhang, Wenjie Wang, Fuli Feng, and Xiangnan He. 2023. Is chatgpt fair for recommen- dation? evaluating fairness in large language model recommendation. arXiv preprint arXiv:2305.07609 (2023)

  86. [96]

    Yang Zhang, Fuli Feng, Jizhi Zhang, Keqin Bao, Qifan Wang, and Xiangnan He. 2023. Collm: Integrating collaborative embeddings into large language models for recommendation. arXiv preprint arXiv:2310.19488 (2023)

  87. [97]

    Zizhuo Zhang and Bang Wang. 2023. Prompt learning for news recommendation. arXiv preprint arXiv:2304.05263 (2023)

  88. [98]

    Weinan Zhang, Jiarui Qin, Wei Guo, Ruiming Tang, and Xiuqiang He. 2021. Deep learning for click-through rate estimation. IJCAI (2021)

  89. [99]

    Guorui Zhou, Xiaoqiang Zhu, Chenru Song, Ying Fan, Han Zhu, Xiao Ma, Yanghui Yan, Junqi Jin, Han Li, and Kun Gai

  90. [100]

    Yuhui Zhang, Hao Ding, Zeren Shui, Yifei Ma, James Zou, Anoop Deoras, and Hao Wang. 2021. Language models as recommender systems: Evaluations and limitations. (2021)

  91. [103]

    Guorui Zhou, Na Mou, Ying Fan, Qi Pi, Weijie Bian, Chang Zhou, Xiaoqiang Zhu, and Kun Gai. 2019. Deep interest evolution network for click-through rate prediction. In AAAI, Vol. 33. 5941–5948

  92. [106]

    Jiachen Zhu, Jianghao Lin, Xinyi Dai, Bo Chen, Rong Shan, Jieming Zhu, Ruiming Tang, Yong Yu, and Weinan Zhang

  93. [107]

    Thor: Ragnaro

    Lifelong personalized low-rank adaptation of large language models for recommendation. arXiv preprint arXiv:2408.03533 (2024). A PROMPT ILLUSTRATION We demonstrate several examples to illustrate the hard prompt templates used for ReLLaX on all three datasets. Figure 12 demonst...

  94. [2018]

    In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining

    Deep interest network for click-through rate prediction. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining . 1059–1068. J. ACM, Vol. 1, No. 1, Article . Publication date: January 2025. 26 Rong Shan et al. Input: The user's loca...

  95. [2021]

    arXiv preprint arXiv:2106.09685 (2021)

    Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685 (2021). J. ACM, Vol. 1, No. 1, Article . Publication date: January 2025. Full-Stack Optimized Large Language Models for Lifelong Sequential Behavior Comprehension in Recommendation 23

  96. [2023]

    arXiv preprint arXiv:2305.02182 (2023)

    Uncovering ChatGPT’s Capabilities in Recommender Systems. arXiv preprint arXiv:2305.02182 (2023)

  97. [2024]

    arXiv preprint arXiv:2411.04602 (2024)

    Self-Calibrated Listwise Reranking with Large Language Models. arXiv preprint arXiv:2411.04602 (2024)

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.