Pith. sign in

REVIEW 5 major objections 4 minor 2 cited by

A Survey on Large Language Models in Multimodal Recommender Systems

T0 review · 5 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A new taxonomy maps how LLMs are changing multimodal recommendation.

desk verdict Useful LLM-centric taxonomy of multimodal recommender papers, but the 'comprehensive' claim is undercut by a missing selection protocol and a pile of editorial inconsistencies; worth refereeing with heavy revision. read the letter →

arxiv 2505.09777 v1 pith:LYYZQCDK submitted 2025-05-14 cs.IR cs.CL

classification cs.IRcs.CL
keywords largelanguagemodelsmultimodalrecommendersystemstaxonomypromptingparameter-efficientfine-tuningdataadaptationknowledgegraphsevaluationmetrics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey tries to organize the fast-growing intersection of large language models and multimodal recommender systems. Its central claim is that LLM-based systems are best understood through a new taxonomy organized around LLM-specific functions—prompting, training, and data-type adaptation—rather than the encoder- and loss-centric categories used by earlier surveys. It argues that LLMs change the role of the recommender from a fixed encoder–decoder into a dynamic, prompt-conditioned reasoner, and it backs that view by classifying recent methods, importing transferable techniques from adjacent recommendation domains, and compiling expanded lists of datasets and evaluation metrics. A sympathetic reader would care because the survey offers a structured map of design choices and points to where research is converging and where gaps remain.

What carries the argument

The central object is the taxonomy itself, a three-part classification scheme for LLM-based multimodal recommender systems. Its first axis groups LLM methods by how the model is controlled—prompting strategies from fixed hard templates to learnable soft prompts and multi-stage control logic, training strategies from LoRA-style parameter-efficient tuning to fully frozen zero-tuning and agent-based systems, and data-type adaptation that converts graphs, IDs, tables, images, and behavior logs into LLM-compatible text or token sequences. Its second axis regroups the field's long-standing problems—disentanglement, alignment, and fusion—and shows how LLM-era work addresses them through contrastive learning, adapters, projection layers, post-hoc refinement, and attention-based fusion. The taxonomy does the argument's work by making each integration pattern a separate design axis, so a reader can compare methods across prompting, training, and modality adaptation without conflating them.

What would settle it

One concrete check is to run an independent, keyword-based search of preprint and conference proceedings from 2023 to 2025 covering LLM-based multimodal recommendation and count how many papers cannot be placed in any of the survey's categories without stretching definitions. A second check is to take two methods the taxonomy places in the same cell, say two soft-prompt systems, and show that they differ more in behavior than two methods placed in different cells; that would indicate the classification does not track the design choices that matter.

Watch

Extended reading notes

Core claim

On its own terms, the paper's contribution is a classification framework: it divides LLM–MRS integration first into LLM methods—prompting (hard, soft, hybrid, control-logic), training strategies (parameter adaptation, zero-tuning, pretrained multimodal LLMs, agents), and data-type adaptation (knowledge-graph conversion, semantic IDs, tabular-to-text, image summarization, behavior-to-text, prompt-based fusion, adapted multimodal fusion)—and then into MRS-specific challenges revisited through LLMs: disentanglement, alignment, and fusion. The taxonomy is meant to capture what is new when LLMs enter the picture: flexible input handling through prompts, reasoning and in-context learning, and the ability to act as agents or orchestrators. It also claims that techniques from sequential, textual, and knowledge-aware recommendation are transferable to multimodal settings before they have been explicitly adapted there. The survey positions this as a departure from earlier encoder-centric surveys, and it supplies appendices of datasets and metrics to make the map usable.

Load-bearing premise

The survey assumes that the papers it chose, and the categories it assigns them to, fairly represent the whole field of LLM-based multimodal recommendation; a biased or incomplete corpus would make the taxonomy and its trend analysis unreliable.

Editorial extensions

If this is right

  • A researcher can use the taxonomy to locate a new method's design choices—prompt-only, LoRA-tuned, frozen-encoder, or agent-based—and compare it against the surveyed alternatives.
  • The survey's inclusion of transferable techniques from sequential, textual, and knowledge-aware recommendation means that methods not yet tried in multimodal settings are explicitly flagged as candidates for adaptation.
  • The dataset and metric appendices give a shared benchmark vocabulary, including NLP-derived and LLM-based evaluators alongside classic ranking metrics.
  • If the identified trends hold, future systems will increasingly combine knowledge graphs, soft prompting with adapters, and offline multimodal-LLM enrichment rather than full end-to-end multimodal LLM training.
  • A reader should expect the field's reported results to remain difficult to compare until common protocols for LLM-based evaluation and multimodal datasets are adopted.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy's design-axis framing suggests a combinatorial design space that the paper does not explicitly test: prompting strategy, training method, and data-adaptation format could be varied independently, and a systematic benchmark crossing those axes would give a sharper picture of what drives performance.
  • Because the survey selects papers without a stated search protocol, its trend claims (for example, that JSON-style and Python-class prompts are emerging, or that knowledge graphs act as a bridge) are editorial interpretations of a possibly biased corpus rather than measured field statistics.
  • The inclusion of works from adjacent domains implies a testable prediction: techniques such as graph-to-text conversion and code-like structural prompts will appear in explicitly multimodal recommenders within a short period.
  • The emphasis on LLM-based evaluation and human-study metrics suggests that the field's real bottleneck may shift from model accuracy to reliable evaluation protocols.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. The manuscript surveys recent work on large language models in multimodal recommender systems (MRS). It proposes a taxonomy built around three LLM-centric axes (prompting, training, data adaptation) and three MRS-specific axes (disentanglement, alignment, fusion), and it adds an evaluation-metrics appendix, a dataset appendix, and a list of future research directions. The survey's central claim is that it provides a comprehensive, LLM-focused review and a novel taxonomy of integration patterns, including transferable techniques from adjacent recommendation domains. The organizational idea is plausible and the reference collection is broad and current, but the manuscript as written does not yet make its coverage or category assignments checkable, and several internal statements and appendix resources are inconsistent.

Significance. If the taxonomy and coverage were properly validated, this survey would be a useful map for researchers entering LLM-based multimodal recommendation: it separates prompt-based, training-based, and data-adaptation strategies, distinguishes MRS-specific challenges such as disentanglement and alignment, and compiles a wider metric taxonomy than most prior surveys. The paper also deserves credit for including 2023-2025 work, for treating tabular and structured data as a first-class modality, and for discussing LLM-based evaluation and agent-based recommendation as forward-looking directions. There are no derivations or fitted quantities in the paper, so circularity is not a concern; the central risk is instead that the claimed 'comprehensive' coverage and 'novel taxonomy' rest on an undocumented corpus-selection rule and on category assignments that the manuscript itself sometimes contradicts. If the authors add a reproducible selection protocol and repair the internal inconsistencies, the survey's substantive contribution would stand.

major comments (5)
  1. [Sections 1.1-1.2] No literature search protocol or inclusion/exclusion criteria are given. The abstract's claim of a 'comprehensive review' is therefore not checkable: a reader cannot tell whether the selected corpus represents the field or the authors' prior selection. The authors should specify the venues, date range, search queries, screening rules, and an operational definition of 'transferable' used to admit methods from adjacent recommendation domains.
  2. [Sections 2.1.4, 2.3.1, 2.3.6] The taxonomy absorbs papers that the manuscript itself identifies as outside MRS or even outside recommender systems. GraphJudger is introduced as 'not an RS approach'; ADKGD is anomaly detection; TANS is graph classification; and P5 is declared 'not a multimodal model' in Section 2.2 but is later treated as a 'prompt-centric MRS' in Section 2.3.6. Without an explicit transferability criterion and a per-paper justification, these category assignments are not reliable and the bounds of the taxonomy cannot be validated.
  3. [Sections 2.3, 3.1.5, 3.2 and Figure 4] The internal consistency of the taxonomy needs repair. Section 2.3 announces eight data-adaptation categories but enumerates seven numbered categories plus a 'final mention'; Figure 4's caption states that 'Use-intent is not considered a category' while Section 3.1.5 contains a 'User Intent Disentanglement' subsection; and Section 3.2 says the alignment categories are 'not mutually exclusive' and intentionally overlap with Section 3.1. These choices may be defensible, but as written they prevent a reader from determining whether the classification is reproducible. The authors should state which categories are exclusive, what overlap is allowed, and provide a category-by-paper assignment table.
  4. [Appendix A.1, Tables 1-2] The dataset appendix is presented as a contribution, but Tables 1 and 2 are identical, and many entries contain placeholder or unusable links ('Link', 'old link', 'Link:scraping'). This makes the 'extensive dataset list' impossible to use as a resource. The tables should be deduplicated and filled with complete, working URLs and accurate modality/domain annotations.
  5. [Section 3.3.1] This subsection contains a placeholder citation, '(cites)', and refers to 'TripletFusion [81]' where the cited work is TMF ('Triple Modality Fusion'). Such artifacts break the reference chain and need to be corrected before the survey can be considered complete.
minor comments (4)
  1. [Section 2.1.1] The sentence beginning 'Similar in spirit to CoT, leverages prompt-based task reformulation...' lacks a grammatical subject and should be rewritten.
  2. [Appendix A.3, Notation] The notation table defines 'Large Language Model (VLM)', but VLM should be 'Vision Language Model'; several entries are also general concepts rather than abbreviations, and the table should be split or relabeled for clarity.
  3. [References] The reference list contains duplicate entries: MMGCN appears as both [125] and [126], and 'A Tale of Two Graphs' appears as both [157] and [158]. These duplicates should be consolidated.
  4. [Section 2.1.3] The description of hybrid prompting is broad enough to cover almost any method with a learned component; a short decision rule or a comparison table would make the boundary between hybrid prompting and adapter-based fusion easier for readers to apply.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the survey makes no fitted predictions or derived quantitative claims; its taxonomy is organizational, and the sole author self-citation is background, not load-bearing.

full rationale

This manuscript is a literature survey, not a derivation or empirical study. It proposes a taxonomy for organizing LLM-based multimodal recommender systems, reviews papers, and lists datasets and metrics. There is no fitted parameter, no prediction, and no equation whose output is equivalent to an input by construction. The 'novel taxonomy' is a classification scheme whose categories are asserted, not derived from the surveyed papers, so none of the circularity patterns apply. The only author-overlapping citation is [75], which appears in Section 3 as one of several background references for the observation that recent MRS models increasingly use GNN and Transformer architectures (e.g., 'recent MRS models increasingly adopt Graph Neural Networks and Transformer architectures to capture better complex user–item and modality interactions, as extensively studied in prior works [47, 67, 75, 94, 107, 135, 153]'). This self-citation is not the basis of any central claim, uniqueness theorem, or ansatz; it merely points to prior work on sequential recommendation as related background. The survey's coverage choices and category boundaries could be validated more rigorously, but that is a question of completeness or methodology, not circularity. No circular step can be exhibited by quoting an equation or a fitted input, and therefore the circularity score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

No quantitative derivations or fitted parameters appear; the paper's only postulates are the representativeness of its literature sample and the sufficiency of its taxonomy categories.

assumptions (2)
  • domain assumption The selected corpus is representative of LLM-based multimodal recommendation research.
    The survey's value rests on complete and accurate coverage, but no retrieval protocol or coverage metrics are provided (Sections 1.1-1.4).
  • domain assumption The taxonomy categories (prompting, training, data adaptation; disentanglement, alignment, fusion) are sufficient for organizing the literature.
    The taxonomy is asserted as novel in Section 1.3 without validation of exhaustiveness; overlaps between categories are acknowledged in Sections 3.1 and 3.2.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey on Large Language Models in Multimodal Recommender Systems." pith.science (2026). https://pith.science/paper/LYYZQCDK

@misc{pith2026250509777,
  author       = {Pith},
  title        = {Pith review of: A Survey on Large Language Models in Multimodal Recommender Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LYYZQCDK}},
  note         = {Machine review of arXiv:2505.09777}
}
read the original abstract

Multimodal recommender systems (MRS) integrate heterogeneous user and item data, such as text, images, and structured information, to enhance recommendation performance. The emergence of large language models (LLMs) introduces new opportunities for MRS by enabling semantic reasoning, in-context learning, and dynamic input handling. Compared to earlier pre-trained language models (PLMs), LLMs offer greater flexibility and generalisation capabilities but also introduce challenges related to scalability and model accessibility. This survey presents a comprehensive review of recent work at the intersection of LLMs and MRS, focusing on prompting strategies, fine-tuning methods, and data adaptation techniques. We propose a novel taxonomy to characterise integration patterns, identify transferable techniques from related recommendation domains, provide an overview of evaluation metrics and datasets, and point to possible future directions. We aim to clarify the emerging role of LLMs in multimodal recommendation and support future research in this rapidly evolving field.

Figures

Figures reproduced from arXiv: 2505.09777 by the authors.

Figure 1
Figure 1. Prompting strategies in LLM-based MRS. 2.1.1 Hard prompting. Hard prompting. Hard prompting ([9]) is the earliest and most widely adopted prompting paradigm. It involves crafting fixed in￾put templates—manually or programmatically, guiding the LLM’s behavior during inference without updating its internal weights. These prompts are human-readable and often designed through domain knowledge, intuition, or task-specifi… view at source ↗
Figure 2
Figure 2. Training strategies for adapting LLMs in Multimodal Recommender Systems. [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Classification of modality adaptation strategies for integrating structured and multimodal inputs into LLMs. [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Disentangle categories. Use-intent is not considered a category; it is just a note at the end of the section. [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]
Figure 5
Figure 5. Figure 5: Alignment strategies. recommender systems highlights how adapters can flexibly map se￾mantic information into structured predictive spaces. Additionally, ADKGD [128] explores adapter-like designs by enforcing disen￾tangled representations of knowledge graph triplets vi…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RecGOAT: Graph Optimal Adaptive Transport for LLM-Enhanced Multimodal Recommendation with Dual Semantic Alignment

    cs.IR 2026-01 reject novelty 6.0 of 10

    RecGOAT aligns LLM and vision item features with collaborative ID embeddings via instance-level contrastive learning and distribution-level optimal transport, reporting state-of-the-art results on three Amazon benchmarks.

  2. MMRM: A Multiplex Multimodal Representation Model for Product Ranking in E-commerce Search

    cs.IR 2026-07 conditional novelty 5.0 of 10

    A shared MLLM backbone with task-specific tokens learns four collaborative signals simultaneously and feeds multiplex embeddings into multitask search ranking, improving GAUC and online metrics at JD.

Reference graph

Works this paper leans on

161 extracted references · 14 canonical work pages · cited by 2 Pith papers

  1. [81]

    Luyi Ma, Xiaohan Li, Zezhong Fan, Jianpeng Xu, Jason Cho, Praveen Kanumala, Kaushiki Nag, Sushant Kumar, and Kannan Achan. 2024. Triple Modality Fusion: Aligning Visual, Textual, and Graph Data with Large Language Models for Multi-Behavior Recommendations. arXiv:2410.12228 [cs.IR] https://arxiv. org/abs/2410.12228

  2. [1]

    Abdullah Alchihabi and Yuhong Guo. 2021. Dual GNNs: Graph Neural Network Learning with Limited Supervision. arXiv:2106.15755 [cs.LG] https://arxiv.org/ abs/2106.15755

  3. [2]

    Vineeta Anand and Ashish Kumar Maurya. 2024. A Survey on Recommender Systems Using Graph Neural Network. ACM Trans. Inf. Syst. 43, 1, Article 9 (Nov. 2024), 49 pages. https://doi.org/10.1145/3694784

  4. [3]

    Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural Machine Translation by Jointly Learning to Align and Translate. CoRR abs/1409.0473 (2014). https://api.semanticscholar.org/CorpusID:11212020

  5. [4]

    Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, Jade Goldstein, Alon Lavie, Chin-Yew Lin, and Clare Voss (Eds.). Association for Computational...

  6. [5]

    Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bah- danau, Nicolas Chapados, and Siva Reddy. 2024. LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders. arXiv:2404.05961 [cs.CL] https://arxiv.org/abs/2404.05961

  7. [6]

    Weizhen Bian, Siyan Liu, Yubo Zhou, Dezhi Chen, Yijie Liao, Zhenzhen Fan, and Aobo Wang. 2024. IntellectSeeker: A Personalized Literature Man- agement System with the Probabilistic Model and Large Language Model. arXiv:2412.07213 [cs.IR] https://arxiv.org/abs/2412.07213

  8. [7]

    Ljubisa Bojic, Zorica Dodevska, Yashar Deldjoo, and Nenad Pantelic

Show all 161 references
  1. [8]

    Alexander Brinkmann, Roee Shraga, Reng Chiz Der, and Christian Bizer. 2023. Product Information Extraction using ChatGPT. arXiv:2306.14921 [cs.CL] https://arxiv.org/abs/2306.14921

  2. [9]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Pra- fulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey W...

  3. [10]

    Feiyu Chen, Junjie Wang, Yinwei Wei, Hai-Tao Zheng, and Jie Shao. 2022. Breaking Isolation: Multimodal Graph Fusion for Multimedia Recommendation by Edge-wise Modulation. 385–394. https://doi.org/10.1145/3503161.3548399

  4. [11]

    Hong Chen, Yudong Chen, Xin Wang, Ruobing Xie, Rui Wang, Feng Xia, and Wenwu Zhu. 2021. Curriculum Disentangled Recommendation with Noisy Multi-feedback. In Advances in Neural Information Processing Systems , M. Ran- zato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman...

  5. [12]

    Jiao Chen, Luyi Ma, Xiaohan Li, Nikhil Thakurdesai, Jianpeng Xu, Jason H. D. Cho, Kaushiki Nag, Evren Korpeoglu, Sushant Kumar, and Kannan Achan. 2023. Knowledge Graph Completion Models are Few-shot Learners: An Empirical Study of Relation Labeling in E-commerce with LLMs. arX...

  6. [13]

    Wen Chen, Pipei Huang, Jiaming Xu, Xin Guo, Cheng Guo, Fei Sun, Chao Li, Andreas Pfadler, Huan Zhao, and Binqiang Zhao. 2019. POG: Personalized Outfit Generation for Fashion Recommendation at Alibaba iFashion. In Proceedings of the 25th ACM SIGKDD International Conference on K...

  7. [14]

    Xi Chen, Yangsiyi Lu, Yuehai Wang, and Jianyi Yang. 2021. CMBF: Cross- Modal-Based Fusion Recommendation Algorithm. Sensors 21 (08 2021), 5275. https://doi.org/10.3390/s21165275

  8. [15]

    Yu Cheng, Yunzhu Pan, Jiaqi Zhang, Yongxin Ni, Aixin Sun, and Fajie Yuan. [n. d.]. An Image Dataset for Benchmarking Recommender Systems with Raw Pixels . 418–426. https://doi.org/10.1137/1.9781611978032.49 arXiv:https://epubs.siam.org/doi/pdf/10.1137/1.9781611978032.49

  9. [16]

    Chi, and Minmin Chen

    Konstantina Christakopoulou, Alberto Lalama, Cj Adams, Iris Qu, Yifat Amir, Samer Chucri, Pierce Vollucci, Fabio Soldo, Dina Bseiso, Sarah Scodel, Lucas Dixon, Ed H. Chi, and Minmin Chen. 2023. Large Language Models for User Interest Journeys. ArXiv abs/2305.15498 (2023). http...

  10. [17]

    Paul F Christiano, Jan Leike, Tom B Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep Reinforcement Learning from Human Preferences. arXiv preprint arXiv:1706.03741 (2017)

  11. [18]

    Yashar Deldjoo. 2024. Understanding Biases in ChatGPT-based Rec- ommender Systems: Provider Fairness, Temporal Stability, and Recency. arXiv:2401.10545 [cs.IR] https://arxiv.org/abs/2401.10545

  12. [19]

    Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. QLoRA: Efficient Finetuning of Quantized LLMs. ArXiv abs/2305.14314 (2023). https://api.semanticscholar.org/CorpusID:258841328

  13. [20]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human...

  14. [22]

    Yingpeng Du, Di Luo, Rui Yan, Xiaopei Wang, Hongzhi Liu, Hengshu Zhu, Yang Song, and Jie Zhang. 2024. Enhancing job recommendation through LLM-based generative adversarial networks. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth...

  15. [23]

    Miguel Escarda, Iñigo López-Riobóo Botana, Santiago Barro-Tojeiro, Lara Padrón-Cousillas, Sonia Gonzalez-Vázquez, Antonio Carreiro-Alonso, and Pablo Gómez-Area. 2024. LLMs on the Fly: Text-to-JSON for Custom API Calling

  16. [24]

    Chen Gao, Yu Zheng, Nian Li, Yinfeng Li, Yingrong Qin, Jinghua Piao, Yuhan Quan, Jianxin Chang, Depeng Jin, Xiangnan He, and Yong Li. 2023. A Survey of Graph Neural Networks for Recommender Systems: Challenges, Methods, and Directions. ACM Trans. Recomm. Syst. 1, 1, Article 3 ...

  17. [25]

    Jingtong Gao, Zhaocheng Du, Xiaopeng Li, Yichao Wang, Xiangyang Li, Huifeng Guo, Ruiming Tang, and Xiangyu Zhao. 2025. SampleLLM: Optimizing Tabular Data Synthesis in Recommendations. arXiv:2501.16125 [cs.IR] https://arxiv. org/abs/2501.16125

  18. [26]

    Xiaoyan Gao, Fuli Feng, Xiangnan He, Heyan Huang, Xinyu Guan, Chong Feng, Zhaoyan Ming, and Tat-Seng Chua. 2020. Hierarchical Attention Network for Visually-Aware Food Recommendation. IEEE Transactions on Multimedia 22, 6 (2020), 1647–1659. https://doi.org/10.1109/TMM.2019.2945180

  19. [27]

    Mouzhi Ge, Carla Delgado, and Dietmar Jannach. 2010. Beyond accuracy: Evaluating recommender systems by coverage and serendipity. RecSys’10 - Proceedings of the 4th ACM Conference on Recommender Systems, 257–260. https: //doi.org/10.1145/1864708.1864761

  20. [28]

    Shijie Geng, Shuchang Liu, Zuohui Fu, Yingqiang Ge, and Yongfeng Zhang

  21. [29]

    Shijie Geng, Juntao Tan, Shuchang Liu, Zuohui Fu, and Yongfeng Zhang. 2023. VIP5: Towards Multimodal Foundation Models for Recommendation. ArXiv abs/2305.14302 (2023). https://api.semanticscholar.org/CorpusID:258841635

  22. [30]

    Shivangi Gheewala, Shuxiang Xu, and Soonja Yeom. 2025. In-depth survey: deep learning in recommender systems—exploring prediction and ranking models, datasets, feature analysis, and emerging trends. Neural Computing and Applications N/A, N/A (March 2025), N/A. https://doi.org/...

  23. [31]

    Liu Guang, Jie Yang, and Ledell Wu. 2022. PTab: Using the Pre-trained Language Model for Modeling Tabular Data. (09 2022). https://doi.org/10.48550/arXiv. 2209.08060

  24. [32]

    Tengyue Han, Pengfei Wang, Shaozhang Niu, and Chenliang Li. 2022. Modality Matches Modality: Pretraining Modality-Disentangled Item Representations for Recommendation. In Proceedings of the ACM Web Conference 2022 (Virtual Event, Lyon, France) (WWW ’22). Association for Comput...

  25. [34]

    Jesse Harte, Wouter Zorgdrager, Panos Louridas, Asterios Katsifodimos, Dietmar Jannach, and Marios Fragkoulis. 2023. Leveraging Large Language Models for Sequential Recommendation. In Proceedings of the 17th ACM Conference on Recommender Systems (RecSys ’23) . ACM, 1096–1102. ...

  26. [35]

    Ruining He and Julian McAuley. 2016. VBPR: visual Bayesian Personalized Ranking from implicit feedback. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence (Phoenix, Arizona) (AAAI’16). AAAI Press, 144–150

  27. [36]

    Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi

  28. [37]

    Zheng, and Qi Liu

    Min Hou, Le Wu, Enhong Chen, Zhi Li, Vincent W. Zheng, and Qi Liu. 2019. Explainable fashion recommendation: a semantic attribute region guided ap- proach. In Proceedings of the 28th International Joint Conference on Artificial Intelligence (Macao, China) (IJCAI’19). AAAI Pres...

  29. [38]

    Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-Efficient Transfer Learning for NLP. ArXiv abs/1902.00751 (2019). https://api.semanticscholar.org/CorpusID:59599816

  30. [39]

    Enshuo Hsu and Kirk Roberts. 2025. LLM-IE: a python package for biomedical generative information extraction with large language models. JAMIA Open 8, 2 (03 2025), ooaf012. https://doi.org/10. 1093/jamiaopen/ooaf012 arXiv:https://academic.oup.com/jamiaopen/article- pdf/8/2/ooa...

  31. [40]

    Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen

    J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. 2021. LoRA: Low-Rank Adaptation of Large Language Models. ArXiv abs/2106.09685 (2021). https://api.semanticscholar.org/CorpusID: 235458009

  32. [41]

    Haoyu Huang, Chong Chen, Conghui He, Yang Li, Jiawei Jiang, and Wentao Zhang. 2025. Can LLMs be Good Graph Judger for Knowledge Graph Construc- tion? arXiv:2411.17388 [cs.CL] https://arxiv.org/abs/2411.17388

  33. [42]

    Dietmar Jannach, Lukas Lerche, Iman Kamehkhosh, and Michael Jugovac. 2015. What recommenders recommend: an analysis of recommendation biases and possible countermeasures. User Modeling and User-Adapted Interaction 25, 5 (Dec. 2015), 427–491. https://doi.org/10.1007/s11257-015-9165-3

  34. [43]

    Yanbiao Ji, Yue Ding, Dan Luo, Chang Liu, Jing Tong, Shaokai Wu, and Hongtao Lu. 2025. Generating Negative Samples for Multi-Modal Recommendation. arXiv:2501.15183 [cs.IR] https://arxiv.org/abs/2501.15183

  35. [44]

    Pengyue Jia, Zhaocheng Du, Yichao Wang, Xiangyu Zhao, Xiaopeng Li, Yuhao Wang, Qidong Liu, Huifeng Guo, and Ruiming Tang. 2024. AltFS: Agency-light Feature Selection with Large Language Models in Deep Recommender Systems. arXiv:2412.08516 [cs.IR] https://arxiv.org/abs/2412.08516

  36. [45]

    Hao Jiang, Wenjie Wang, Meng Liu, Liqiang Nie, Ling-Yu Duan, and Changsheng Xu. 2019. Market2Dish: A Health-aware Food Recommendation System. In Proceedings of the 27th ACM International Conference on Multimedia (Nice, France) (MM ’19). Association for Computing Machinery, New...

  37. [46]

    Yuezihan Jiang, Gaode Chen, Wenhan Zhang, Jingchi Wang, Yinjie Jiang, Qi Zhang, Jingjian Lin, Peng Jiang, and Kaigui Bian. 2024. Prompt Tuning for Item Cold-start Recommendation. arXiv:2412.18082 [cs.IR] https://arxiv.org/abs/ 2412.18082

  38. [47]

    Wang-Cheng Kang and Julian McAuley. 2018. Self-Attentive Sequential Rec- ommendation. In 2018 IEEE International Conference on Data Mining (ICDM) . A Survey on Large Language Models in Multimodal Recommender Systems 197–206. https://doi.org/10.1109/ICDM.2018.00035

  39. [48]

    Xiaoyu Kong, Jiancan Wu, An Zhang, Leheng Sheng, Hui Lin, Xiang Wang, and Xiangnan He. 2025. Customizing Language Models with Instance-wise LoRA for Sequential Recommendation. arXiv:2408.10159 [cs.IR] https://arxiv.org/ abs/2408.10159

  40. [49]

    Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The Power of Scale for Parameter-Efficient Prompt Tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih ...

  41. [50]

    Chen Li, Yixiao Ge, Jiayong Mao, Dian Li, and Ying Shan. 2023. TagGPT: Large Language Models are Zero-shot Multimodal Taggers. arXiv:2304.03022 [cs.IR] https://arxiv.org/abs/2304.03022

  42. [53]

    Lei Li, Yongfeng Zhang, and Li Chen. 2023. Personalized Prompt Learning for Explainable Recommendation. ACM Transactions on Information Systems 41 (03 2023), 1–26. https://doi.org/10.1145/3580488

  43. [54]

    Pang Li, Shahrul Noah, and Hafiz Sarim. 2024. A Survey on Deep Neural Networks in Collaborative Filtering Recommendation Systems. https://doi. org/10.48550/arXiv.2412.01378

  44. [55]

    Xiangyang Li, Bo Chen, Luyao Hou, and Ruiming Tang. 2023. CTRL: Con- nect Collaborative and Language Model for CTR Prediction. https://api. semanticscholar.org/CorpusID:259075336

  45. [56]

    Xinhang Li, Xiangyu Zhao, Jiaxing Xu, Yong Zhang, and Chunxiao Xing. 2023. IMF: Interactive Multimodal Fusion Model for Link Prediction. In Proceedings of the ACM Web Conference 2023 (Austin, TX, USA) (WWW ’23). Association for Computing Machinery, New York, NY, USA, 2572–2580...

  46. [57]

    Xiang Lisa Li and Percy Liang. 2021. Prefix-Tuning: Optimizing Continuous Prompts for Generation. In Proceedings of the 59th Annual Meeting of the Associ- ation for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: ...

  47. [58]

    Youhua Li, Hanwen Du, Yongxin Ni, Pengpeng Zhao, Qi Guo, Fajie Yuan, and Xiaofang Zhou. 2023. Multi-Modality is All You Need for Transferable Recom- mender Systems. 2024 IEEE 40th International Conference on Data Engineering (ICDE) (2023), 5008–5021. https://api.semanticschola...

  48. [59]

    Zixuan Li, Yutao Zeng, Yuxin Zuo, Weicheng Ren, Wenxuan Liu, Miao Su, Yucan Guo, Yantao Liu, Lixiang Lixiang, Zhilei Hu, Long Bai, Wei Li, Yidan Liu, Pan Yang, Xiaolong Jin, Jiafeng Guo, and Xueqi Cheng. 2024. KnowCoder: Coding Structured Knowledge into LLMs for Universal Info...

  49. [60]

    Jiayi Liao, Ruobing Xie, Sihang Li, Xiang Wang, Xingwu Sun, Zhanhui Kang, and Xiangnan He. 2025. PatchRec: Multi-Grained Patching for Efficient LLM- based Sequential Recommendation. arXiv:2501.15087 [cs.IR] https://arxiv.org/ abs/2501.15087

  50. [61]

    Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out. Association for Computational Linguistics, Barcelona, Spain, 74–81. https://aclanthology.org/W04-1013/

  51. [62]

    Yujie Lin, Pengjie Ren, Zhumin Chen, Zhaochun Ren, Jun Ma, and Maarten Rijke. 2019. Explainable Outfit Recommendation with Joint Outfit Matching and Comment Generation. IEEE Transactions on Knowledge and Data Engineering PP (03 2019), 1–1. https://doi.org/10.1109/TKDE.2019.2906190

  52. [63]

    Chang Liu, Xiaoguang Li, Guohao Cai, Zhenhua Dong, Hong Zhu, and Lifeng Shang. 2021. Non-invasive Self-attention for Side Information Fusion in Se- quential Recommendation. Proceedings of the AAAI Conference on Artificial Intelligence 35 (05 2021), 4249–4256. https://doi.org/1...

  53. [64]

    Fan Liu, Zhiyong Cheng, Huilin Chen, Anan Liu, Liqiang Nie, and Mohan Kankanhalli. 2022. Disentangled Multimodal Representation Learning for Rec- ommendation. (03 2022). https://doi.org/10.48550/arXiv.2203.05406

  54. [66]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruction Tuning. arXiv:2304.08485 [cs.CV] https://arxiv.org/abs/2304.08485

  55. [67]

    Langming Liu, Liu Cai, Chi Zhang, Xiangyu Zhao, Jingtong Gao, Wanyu Wang, Yifu Lv, Wenqi Fan, Yiqi Wang, Ming He, Zitao Liu, and Qing Li. 2023. LinRec: Linear Attention Mechanism for Long-term Sequential Recommender Systems. In Proceedings of the 46th International ACM SIGIR C...

  56. [68]

    Peng Liu, Lemei Zhang, and Jon Atle Gulla. 2023. Pre-train, Prompt, and Recommendation: A Comprehensive Survey of Language Modeling Paradigm Adaptations in Recommender Systems. Transactions of the Association for Computational Linguistics 11 (2023), 1553–1571. https://doi.org/...

  57. [69]

    Qidong Liu, Jiaxi Hu, Yutian Xiao, Xiangyu Zhao, Jingtong Gao, Wanyu Wang, Qing Li, and Jiliang Tang. 2024. Multimodal Recommender Systems: A Survey. ACM Comput. Surv. 57, 2, Article 26 (Oct. 2024), 17 pages. https://doi.org/10. 1145/3695461

  58. [70]

    Shang Liu, Zhenzhong Chen, Hongyi Liu, and Xinghai Hu. 2019. User-Video Co-Attention Network for Personalized Micro-video Recommendation. In The World Wide Web Conference (San Francisco, CA, USA) (WWW ’19). Association for Computing Machinery, New York, NY, USA, 3020–3026. htt...

  59. [71]

    Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2022. P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Sh...

  60. [72]

    Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Lam Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2022. P-Tuning v2: Prompt Tuning Can Be Comparable to Fine- tuning Universally Across Scales and Tasks. arXiv:2110.07602 [cs.CL] https: //arxiv.org/abs/2110.07602

  61. [73]

    Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2024. GPT understands, too. AI Open 5 (2024), 208–215. https: //doi.org/10.1016/j.aiopen.2023.08.012

  62. [74]

    Zhuang Liu, Yunpu Ma, Matthias Schubert, Yuanxin Ouyang, and Zhang Xiong

  63. [75]

    Alejo Lopez-Avila, Jinhua Du, Abbas Shimary, and Ze Li. 2024. Positional encod- ing is not the same as context: A study on positional encoding for Sequential recommendation. arXiv:2405.10436 [cs.IR] https://arxiv.org/abs/2405.10436

  64. [76]

    Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2016. Hierarchical Co-Attention for Visual Question Answering. (05 2016). https://doi.org/10. 48550/arXiv.1606.00061

  65. [77]

    Yucong Luo, Qitao Qin, Hao Zhang, Mingyue Cheng, Ruiran Yan, Kefan Wang, and Jie Ouyang. 2024. Molar: Multimodal LLMs with Collaborative Filtering Alignment for Enhanced Sequential Recommendation. arXiv:2412.18176 [cs.IR] https://arxiv.org/abs/2412.18176

  66. [78]

    Thang Luong, Hieu Pham, and Christopher D. Manning. 2015. Effective Ap- proaches to Attention-based Neural Machine Translation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing , Lluís Màrquez, Chris Callison-Burch, and Jian Su (Eds.). ...

  67. [79]

    InProceedings of the 2022 International Conference on Multimedia Retrieval (Newark, NJ, USA) (ICMR ’22)

    Multi-Modal Contrastive Pre-training for Recommendation. InProceedings of the 2022 International Conference on Multimedia Retrieval (Newark, NJ, USA) (ICMR ’22). Association for Computing Machinery, New York, NY, USA, 99–108. https://doi.org/10.1145/3512527.3531378

  68. [80]

    Jianxin Ma, Chang Zhou, Peng Cui, Hongxia Yang, and Wenwu Zhu. 2019. Learning disentangled representations for recommendation . Curran Associates Inc., Red Hook, NY, USA

  69. [82]

    Yunshan Ma, Yingzhi He, An Zhang, Xiang Wang, and Tat-Seng Chua. 2022. CrossCBR: Cross-view Contrastive Learning for Bundle Recommendation. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Washington DC, USA) (KDD ’22). Association for C...

  70. [83]

    Harshith Manjunath, Lucas Heublein, Tobias Feigl, and Felix Ott. 2025. Multimodal-to-Text Prompt Engineering in Large Language Models Alejo López-Ávila and Jinhua Du Using Feature Embeddings for GNSS Interference Characterization. arXiv:2501.05079 [cs.AI] https://arxiv.org/abs...

  71. [84]

    Hanjia Lyu, Song Jiang, Hanqing Zeng, Yinglong Xia, Qifan Wang, Si Zhang, Ren Chen, Chris Leung, Jiajie Tang, and Jiebo Luo. 2024. LLM-Rec: Personalized Recommendation via Prompting Large Language Models. In Findings of the Association for Computational Linguistics: NAACL 2024...

  72. [85]

    Daniel McKee, Justin Salamon, Josef Sivic, and Bryan Russell. 2023. Language- Guided Music Recommendation for Video via Prompt Analogies. https: //doi.org/10.48550/arXiv.2306.09327

  73. [86]

    Zongshen Mu, Yueting Zhuang, Jie Tan, Jun Xiao, and Siliang Tang. 2022. Learn- ing Hybrid Behavior Patterns for Multimedia Recommendation. In Proceedings of the 30th ACM International Conference on Multimedia (Lisboa, Portugal) (MM ’22). Association for Computing Machinery, Ne...

  74. [87]

    Fabian Paischer, Liu Yang, Linfeng Liu, Shuai Shao, Kaveh Hassani, Jiacheng Li, Ricky Chen, Zhang Gabriel Li, Xialo Gao, Wei Shao, Xue Feng, Nima Noorshams, Sem Park, Bo Long, and Hamid Eghbalzadeh. 2024. Preference Discerning with LLM-Enhanced Generative Retrieval. arXiv:2412...

  75. [88]

    Xingyu Pan, Yushuo Chen, Changxin Tian, Zihan Lin, Jinpeng Wang, He Hu, and Wayne Zhao. 2022. Multimodal Meta-Learning for Cold-Start Sequential Recommendation. 3421–3430. https://doi.org/10.1145/3511808.3557101

  76. [89]

    Javier Marin, Aritro Biswas, Ferda Ofli, Nicholas Hynes, Amaia Salvador, Yusuf Aytar, Ingmar Weber, and Antonio Torralba. 2019. Recipe1M+: A Dataset for Learning Cross-Modal Embeddings for Cooking Recipes and Food Images. arXiv:1810.06553 [cs.CV] https://arxiv.org/abs/1810.06553

  77. [90]

    Joshua Park and Yongfeng Zhang. 2025. AgentRec: Agent Recom- mendation Using Sentence Embeddings Aligned to Human Feedback. arXiv:2501.13333 [cs.LG] https://arxiv.org/abs/2501.13333

  78. [91]

    Tieyun Qian, Yile Liang, Qing Li, Xuan Ma, Ke Sun, and Zhiyong Peng. 2023. In- tent Disentanglement and Feature Self-Supervision for Novel Recommendation. IEEE Transactions on Knowledge and Data Engineering 35, 10 (2023), 9864–9877. https://doi.org/10.1109/TKDE.2022.3175536

  79. [92]

    Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners. https: //api.semanticscholar.org/CorpusID:160025533

  80. [93]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21, 1, Article 140 (Jan. 2020), 67 pages

  81. [94]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics , Pierre Isabelle, Eugene Charniak, and Dekang Lin (Eds...

  82. [95]

    Xubin Ren, Wei Wei, Lianghao Xia, Lixin Su, Suqi Cheng, Junfeng Wang, Dawei Yin, and Chao Huang. 2024. Representation Learning with Large Language Models for Recommendation. In Proceedings of the ACM Web Conference 2024 (Singapore, Singapore) (WWW ’24) . Association for Comput...

  83. [96]

    Ribeiro, Pedro H.P

    Leonardo F.R. Ribeiro, Pedro H.P. Saverese, and Daniel R. Figueiredo. 2017. struc2vec: Learning Node Representations from Structural Identity. In Proceed- ings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (Halifax, NS, Canada) (KDD ’17...

  84. [97]

    Teven Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander Rush, Stella Biderman, Albert Webson, Pawan Am- manamanchi, Thomas Wang, Benoît Sagot, Niklas Muenn...

  85. [98]

    Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning Robust Metrics for Text Generation. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Associ...

  86. [99]

    Ahmed Rashed, Shereen Elsayed, and Lars Schmidt-Thieme. 2022. Context and Attribute-Aware Sequential Recommendation via Cross-Attention. In Proceed- ings of the 16th ACM Conference on Recommender Systems (Seattle, WA, USA) (RecSys ’22). Association for Computing Machinery, New...

  87. [100]

    Younggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee. 2021. State Entropy Maximization with Random Encoders for Efficient Exploration. arXiv:2102.09430 [cs.LG] https://arxiv.org/abs/2102.09430

  88. [101]

    Xiang-Rong Sheng, Feifan Yang, Litong Gong, Biao Wang, Zhangming Chan, Yujing Zhang, Yueyao Cheng, Yong-Nan Zhu, Tiezheng Ge, Han Zhu, Yuning Jiang, Jian Xu, and Bo Zheng. 2024. Enhancing Taobao Display Advertising with Multimodal Representations: Challenges, Approaches and In...

  89. [102]

    Kaize Shi, Xueyao Sun, Dingxian Wang, Yinlin Fu, Guandong Xu, and Qing Li

  90. [103]

    Wentao Shi, Xiangnan He, Yang Zhang, Chongming Gao, Xinyue Li, Jizhi Zhang, Qifan Wang, and Fuli Feng. 2024. Large Language Models are Learnable Plan- ners for Long-Term Recommendation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development i...

  91. [104]

    Prithviraj Sen, Galileo Namata, Mustafa Bilgic, Lise Getoor, Brian Gallagher, and Tina Eliassi-Rad. 2008. Collective Classification in Network Data. AI Magazine 29, 3 (2008), 93–106. https://doi.org/10.1609/aimag.v29i3.2157 arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1609...

  92. [105]

    Xuemeng Song, Fuli Feng, Jinhuan Liu, Zekun Li, Liqiang Nie, and Jun Ma

  93. [106]

    Xuemeng Song, Xianjing Han, Yunkai Li, Jingyuan Chen, Xin-Shun Xu, and Liqiang Nie. 2019. GP-BPR: Personalized Compatibility Modeling for Clothing Matching. In Proceedings of the 27th ACM International Conference on Multimedia (Nice, France) (MM ’19). Association for Computing...

  94. [107]

    Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang

  95. [108]

    LLaMA-E: Empowering E-commerce Authoring with Object-Interleaved Instruction Following. In Proceedings of the 31st International Conference on Computational Linguistics, Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert (E...

  96. [109]

    Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023. Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language...

  97. [110]

    Damien Sileo, Wout Vossen, and Robbe Raymaekers. 2022. Zero-Shot Rec- ommendation as Language Modeling. In Advances in Information Retrieval: 44th European Conference on IR Research, ECIR 2022, Stavanger, Norway, April 10–14, 2022, Proceedings, Part II (Stavanger, Norway). Spr...

  98. [111]

    Zhulin Tao, Yinwei Wei, Xiang Wang, Xiangnan He, Xianglin Huang, and Tat-Seng Chua. 2020. MGAT: Multimodal Graph Attention Network for Rec- ommendation. Information Processing & Management 57, 5 (2020), 102277. https://doi.org/10.1016/j.ipm.2020.102277

  99. [112]

    Tsai, Adam Kraft, Long Jin, Chenwei Cai, Anahita Hosseini, Taibai Xu, Zemin Zhang, Lichan Hong, Ed H

    Alicia Y. Tsai, Adam Kraft, Long Jin, Chenwei Cai, Anahita Hosseini, Taibai Xu, Zemin Zhang, Lichan Hong, Ed H. Chi, and Xinyang Yi. 2024. Leveraging LLM Reasoning Enhances Personalized Recommender Systems. arXiv:2408.00802 [cs.IR] https://arxiv.org/abs/2408.00802

  100. [113]

    Jie Wang, Alexandros Karatzoglou, Ioannis Arapakis, and Joemon M. Jose. 2025. Large Language Model driven Policy Exploration for Recommender Systems. arXiv:2501.13816 [cs.IR] https://arxiv.org/abs/2501.13816

  101. [114]

    Jinpeng Wang, Ziyun Zeng, Yunxiao Wang, Yuting Wang, Xingyu Lu, Tianxiang Li, Jun Yuan, Rui Zhang, Hai-Tao Zheng, and Shu-Tao Xia. 2023. MISSRec: Pre- training and Transferring Multi-modal Interest-aware Sequence Representation for Recommendation. In Proceedings of the 31st AC...

  102. [115]

    Lingzhi Wang, Huang Hu, Lei Sha, Can Xu, Daxin Jiang, and Kam-Fai Wong

  103. [116]

    Rui Sun, Xuezhi Cao, Yan Zhao, Junchen Wan, Kun Zhou, Fuzheng Zhang, Zhongyuan Wang, and Kai Zheng. 2020. Multi-modal Knowledge Graphs for Recommender Systems. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management (Virtual Event, Ireland...

  104. [117]

    Shijie Wang, Wenqi Fan, Yue Feng, Xinyu Ma, Shuaiqiang Wang, and Dawei Yin. 2025. Knowledge Graph Retrieval-Augmented Generation for LLM-based Recommendation. arXiv:2501.02226 [cs.IR] https://arxiv.org/abs/2501.02226

  105. [118]

    Zhulin Tao, Xiaohao Liu, Yewei Xia, Xiang Wang, Lifang Yang, Huang Xi- anglin, and Tat-Seng Chua. 2022. Self-Supervised Learning for Multime- dia Recommendation. IEEE Transactions on Multimedia PP (01 2022), 1–10. https://doi.org/10.1109/TMM.2022.3187556

  106. [119]

    Xingyao Wang, Sha Li, and Heng Ji. 2023. Code4Struct: Code Generation for Few- Shot Event Structure Prediction. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okaz...

  107. [120]

    Yuhao Wang, Junwei Pan, Xiangyu Zhao, Pengyue Jia, Wanyu Wang, Yuan Wang, Yue Liu, Dapeng Liu, and Jie Jiang. 2024. Pre-train, Align, and Disentan- gle: Empowering Sequential Recommendation with Large Language Models. arXiv:2412.04107 [cs.IR] https://arxiv.org/abs/2412.04107

  108. [121]

    Zehong Wang, Sidney Liu, Zheyuan Zhang, Tianyi Ma, Chuxu Zhang, and Yanfang Ye. 2025. Can LLMs Convert Graphs to Text-Attributed Graphs? arXiv:2412.10136 [cs.CL] https://arxiv.org/abs/2412.10136

  109. [122]

    Tianxin Wei, Bowen Jin, Ruirui Li, Hansi Zeng, Zhengyang Wang, Jianhui Sun, Qingyu Yin, Hanqing Lu, Suhang Wang, Jingrui He, and Xianfeng Tang. 2024. Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and Beyond. ArXiv abs/2...

  110. [123]

    Wei Wei, Jiabin Tang, Yangqin Jiang, Lianghao Xia, and Chao Huang. 2024. PromptMM: Multi-Modal Knowledge Distillation for Recommendation with Prompt-Tuning. arXiv:2402.17188 [cs.IR] https://arxiv.org/abs/2402.17188

  111. [124]

    RecInDial: A Unified Framework for Conversational Recommendation with Pretrained Language Models. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the A Survey on Large Language Models in Multimodal Recommend...

  112. [125]

    Qifan Wang, Yinwei Wei, Jianhua Yin, Jianlong Wu, Xuemeng Song, and Liqiang Nie. 2021. DualGNN: Dual Graph Neural Network for Multimedia Recommen- dation. IEEE Transactions on Multimedia (2021)

  113. [126]

    Yinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He, Richang Hong, and Tat- Seng Chua. 2019. MMGCN: Multi-modal Graph Convolution Network for Personalized Recommendation of Micro-video. In Proceedings of the 27th ACM International Conference on Multimedia . 1437–1445

  114. [127]

    Siqi Wang, Chao Liang, Yunfan Gao, Yang Liu, Jing Li, and Haofen Wang. 2024. Decoding Urban Industrial Complexity: Enhancing Knowledge-Driven Insights via IndustryScopeGPT. In Proceedings of the 32nd ACM International Conference on Multimedia (Melbourne VIC, Australia) (MM ’24...

  115. [128]

    Jiayang Wu, Wensheng Gan, Jiahao Zhang, and Philip S. Yu. 2025. AD- KGD: Anomaly Detection in Knowledge Graphs with Dual-Channel Training. arXiv:2501.07078 [cs.AI] https://arxiv.org/abs/2501.07078

  116. [129]

    Yunjia Xi, Weiwen Liu, Jianghao Lin, Xiaoling Cai, Hong Zhu, Jieming Zhu, Bo Chen, Ruiming Tang, Weinan Zhang, and Yong Yu. 2024. Towards Open-World Recommendation with Knowledge Augmentation from Large Language Models. In Proceedings of the 18th ACM Conference on Recommender ...

  117. [130]

    Xiaochuan Xu, Zeqiu Xu, Peiyang Yu, and Jiani Wang. 2025. Enhanc- ing User Intent for Recommendation Systems via Large Language Models. arXiv:2501.10871 [cs.IR] https://arxiv.org/abs/2501.10871

  118. [131]

    An Yan, Zhankui He, Jiacheng Li, Tianyang Zhang, and Julian McAuley. 2023. Personalized Showcases: Generating Multi-Modal Explanations for Recom- mendations. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (Tai...

  119. [132]

    Li Yang, Qifan Wang, Zac Yu, Anand Kulkarni, Sumit Sanghai, Bin Shu, Jon Elsas, and Bhargav Kanagal. 2021. MAVE: A Product Dataset for Multi-source Attribute Value Extraction. arXiv:2112.08663 [cs.CL] https://arxiv.org/abs/2112.08663

  120. [133]

    Yinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He, and Tat-Seng Chua. 2020. Graph-Refined Convolutional Network for Multimedia Recommendation with Implicit Feedback. 3541–3549. https://doi.org/10.1145/3394171.3413556

  121. [134]

    Yinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He, Richang Hong, and Tat- Seng Chua. 2019. MMGCN: Multi-modal Graph Convolution Network for Personalized Recommendation of Micro-video. In Proceedings of the 27th ACM International Conference on Multimedia (Nice, France) (MM ’19)....

  122. [135]

    Yuhao Yang, Chao Huang, Lianghao Xia, Yuxuan Liang, Yanwei Yu, and Chen- liang Li. 2022. Multi-Behavior Hypergraph-Enhanced Transformer for Se- quential Recommendation. In Proceedings of the 28th ACM SIGKDD Confer- ence on Knowledge Discovery and Data Mining (Washington DC, US...

  123. [136]

    Chuhan Wu, Fangzhao Wu, Tao Qi, Chao Zhang, Yongfeng Huang, and Tong Xu. 2022. MM-Rec: Visiolinguistic Model Empowered Multimodal News Rec- ommendation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (Madrid, ...

  124. [137]

    Jing Yi and Zhenzhong Chen. 2021. Multi-Modal Variational Graph Auto- Encoder for Recommendation Systems. IEEE Transactions on Multimedia PP (09 2021), 1–1. https://doi.org/10.1109/TMM.2021.3111487

  125. [138]

    Zixuan Yi, Xi Wang, Iadh Ounis, and Craig Macdonald. 2022. Multi-modal Graph Contrastive Learning for Micro-video Recommendation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (Madrid, Spain) (SIGIR ’22). Ass...

  126. [139]

    Bin Yin, Junjie Xie, Yu Qin, Zixiang Ding, Zhichao Feng, Xiang Li, and Wei Lin. 2023. Heterogeneous Knowledge Fusion: A Novel Approach for Per- sonalized Recommendation via LLM. In Proceedings of the 17th ACM Con- ference on Recommender Systems (Singapore, Singapore) (RecSys ’...

  127. [140]

    Jun Yin, Zhengxin Zeng, Mingzheng Li, Hao Yan, Chaozhuo Li, Weihao Han, Jianjin Zhang, Ruochen Liu, Allen Sun, Denvy Deng, Feng Sun, Qi Zhang, Shirui Pan, and Senzhang Wang. 2024. Unleash LLMs Potential for Recom- mendation by Coordinating Twin-Tower Dynamic Semantic Token Gen...

  128. [141]

    Aron Yu and Kristen Grauman. 2014. Fine-Grained Visual Comparisons with Local Learning. In 2014 IEEE Conference on Computer Vision and Pattern Recog- nition. 192–199. https://doi.org/10.1109/CVPR.2014.32

  129. [142]

    Shenghao Yang, Chenyang Wang, Yankai Liu, Kangping Xu, Weizhi Ma, Yiqun Liu, Min Zhang, Haitao Zeng, Junlan Feng, and Chao Deng. 2023. Collaborative Word-based Pre-trained Item Representation for Transferable Recommendation. arXiv:2311.10501 [cs.IR] https://arxiv.org/abs/2311.10501

  130. [143]

    Yahe Yang and Chengyue Huang. 2025. Tree-based RAG-Agent Recommen- dation System: A Case Study in Medical Test Data. arXiv:2501.02727 [cs.IR] https://arxiv.org/abs/2501.02727

  131. [144]

    Zheng Yuan, Fajie Yuan, Yu Song, Youhua Li, Junchen Fu, Fei Yang, Yunzhu Pan, and Yongxin Ni. 2023. Where to Go Next for Recommender Systems? ID- vs. Modality-based Recommender Models Revisited. In Proceedings of the 46th International ACM SIGIR Conference on Research and Deve...

  132. [145]

    Yuyang Ye, Zhi Zheng, Yishan Shen, Tianshu Wang, Hengruo Zhang, Pei- jun Zhu, Runlong Yu, Kai Zhang, and Hui Xiong. 2024. Harnessing Multi- modal Large Language Models for Multimodal Sequential Recommendation. arXiv:2408.09698 [cs.IR] https://arxiv.org/abs/2408.09698

  133. [146]

    Jinghao Zhang, Yanqiao Zhu, Qiang Liu, Shu Wu, Shuhui Wang, and Liang Wang. 2021. Mining Latent Structures for Multimedia Recommendation. In Proceedings of the 29th ACM International Conference on Multimedia (Virtual Event, China) (MM ’21). Association for Computing Machinery,...

  134. [147]

    Jinghao Zhang, Yanqiao Zhu, Qiang Liu, Mengqi Zhang, Shu Wu, and Liang Wang. 2023. Latent Structure Mining With Contrastive Modality Fusion for Multimedia Recommendation. IEEE Transactions on Knowledge and Data Engi- neering 35, 9 (2023), 9154–9167. https://doi.org/10.1109/TKD...

  135. [148]

    Weinberger, and Yoav Artzi

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating Text Generation with BERT. arXiv:1904.09675 [cs.CL] https://arxiv.org/abs/1904.09675

  136. [149]

    Yin Zhang, Ziwei Zhu, Yun He, and James Caverlee. 2020. Content-Collaborative Disentanglement Representation Learning for Enhanced Recommendation. In Proceedings of the 14th ACM Conference on Recommender Systems (Virtual Event, Brazil) (RecSys ’20). Association for Computing M...

  137. [150]

    Zhi Zheng, Zhaopeng Qiu, Xiao Hu, Likang Wu, Hengshu Zhu, and Hui Xiong. 2023. Generative Job Recommendations with Large Language Model. arXiv:2307.02157 [cs.IR] https://arxiv.org/abs/2307.02157 Alejo López-Ávila and Jinhua Du

  138. [151]

    Aron Yu and Kristen Grauman. 2017. Semantic Jitter: Dense Supervision for Visual Comparisons via Synthetic Images. In 2017 IEEE International Conference on Computer Vision (ICCV) . 5571–5580. https://doi.org/10.1109/ICCV.2017.594

  139. [152]

    Yuanqing Yu, Chongming Gao, Jiawei Chen, Heng Tang, Yuefeng Sun, Qian Chen, Weizhi Ma, and Min Zhang. 2024. EasyRL4Rec: An Easy-to-use Library for Reinforcement Learning Based Recommender Systems. arXiv:2402.15164 [cs.IR] https://arxiv.org/abs/2402.15164

  140. [153]

    Peilin Zhou, Qichen Ye, Yueqi Xie, Jingqi Gao, Shoujin Wang, Jae Boum Kim, Chenyu You, and Sunghun Kim. 2023. Attention Calibration for Transformer- based Sequential Recommendation. InProceedings of the 32nd ACM International Conference on Information and Knowledge Management ...

  141. [154]

    Jizhi Zhang, Keqin Bao, Yang Zhang, Wenjie Wang, Fuli Feng, and Xiangnan He. 2023. Is ChatGPT Fair for Recommendation? Evaluating Fairness in Large Language Model Recommendation. 993–999. https://doi.org/10.1145/3604915. 3608860

  142. [155]

    Xiaoxia Zhou. 2023. MMRec: Simplifying Multimodal Recommendation. Pro- ceedings of the 5th ACM International Conference on Multimedia in Asia Work- shops (2023). https://api.semanticscholar.org/CorpusID:256627512

  143. [156]

    Xin Zhou and Chunyan Miao. 2024. Disentangled Graph Variational Auto- Encoder for Multimodal Recommendation With Interpretability. IEEE Transac- tions on Multimedia 26 (2024), 7543–7554. https://doi.org/10.1109/TMM.2024. 3369875

  144. [157]

    Xin Zhou and Zhiqi Shen. 2023. A Tale of Two Graphs: Freezing and Denoising Graph Structures for Multimodal Recommendation. In Proceedings of the 31st ACM International Conference on Multimedia (Ottawa ON, Canada) (MM ’23). Association for Computing Machinery, New York, NY, US...

  145. [158]

    Xin Zhou and Zhiqi Shen. 2023. A Tale of Two Graphs: Freezing and Denoising Graph Structures for Multimodal Recommendation. In Proceedings of the 31st ACM International Conference on Multimedia . 935–943. A Survey on Large Language Models in Multimodal Recommender Systems A AP...

  146. [160]

    Hongyu Zhou, Xin Zhou, Zhiwei Zeng, Lingzi Zhang, and Zhiqi Shen. 2023. A Comprehensive Survey on Multimodal Recommender Systems: Taxonomy, Evaluation, and Future Directions. ArXiv abs/2302.04473 (2023). https://api. semanticscholar.org/CorpusID:256697515

  147. [161]

    Hongyu Zhou, Xin Zhou, Lingzi Zhang, and Zhiqi Shen. 2023. Enhancing Dyadic Relations with Homogeneous Graphs for Multimodal Recommendation . https://doi.org/10.3233/FAIA230631

  148. [163]

    Tao Zhou, Zoltán Kuscsik, Jian-Guo Liu, Matúš Medo, Joseph Rushton Wake- ling, and Yi-Cheng Zhang. 2010. Solving the apparent diversity-accuracy dilemma of recommender systems. Proceedings of the National Academy of Sciences 107, 10 (2010), 4511–4515. https://doi.org/10.1073/p...

  149. [2017]

    In Proceedings of the 25th ACM International Conference on Multimedia (Mountain View, California, USA) (MM ’17)

    NeuroStylist: Neural Compatibility Modeling for Clothing Matching. In Proceedings of the 25th ACM International Conference on Multimedia (Mountain View, California, USA) (MM ’17). Association for Computing Machinery, New York, NY, USA, 753–761. https://doi.org/10.1145/3123266.3123314

  150. [2019]

    InProceedings of the 28th ACM International Conference on Information and Knowledge Management (Beijing, China) (CIKM ’19)

    BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer. InProceedings of the 28th ACM International Conference on Information and Knowledge Management (Beijing, China) (CIKM ’19). Association for Computing Machinery, New York, NY, USA, ...

  151. [2021]

    In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.)

    CLIPScore: A Reference-free Evaluation Metric for Image Captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.). Association for Computational Lingui...

  152. [2025]

    arXiv:2502.00055 [cs.SI] https://arxiv.org/abs/2502.00055

    Towards Recommender Systems LLMs Playground (RecSysLLMsP): Exploring Polarization and Engagement in Simulated Social Networks. arXiv:2502.00055 [cs.SI] https://arxiv.org/abs/2502.00055

  153. [3059]

    https://doi.org/10.18653/v1/2021.emnlp-main.243

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.