REVIEW 3 major objections 3 minor 6 cited by
The Future is Agentic: Definitions, Perspectives, and Open Challenges of Multi-Agent Recommender Systems
T0 review · 3 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Agentic recommenders are not uniformly better: the paper's controlled study shows single-shot ranking is Pareto-efficient on typical users, with decomposition and ensembling gaining 5.4–6.2 percent NDCG@3 only on high-diversity histories.
desk verdict A genuinely useful framework for agentic recommenders, with a pilot study whose headline empirical claim is real but not yet statistically established. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central formal objects are the recommender agent, defined as a tuple of model, input/output spaces, tools, and memory state, instantiated over a user, context, history, candidate set, and policy constraints, and the multi-agent system, defined as a triple of agents, a shared environment, and a communication protocol with a permission matrix and message schemas. Memory update and retrieval functions with retention, merge, and relevance-scoring operators carry user state across steps. The argumentative load is carried by the controlled study's contrast between a random cohort and a high-diversity cohort, where diversity is Eq. (2), the mean pairwise Jaccard distance between token sets of item titles in a user's history, and by seven workflows that vary only structure while holding autonomy fixed: single shot, profiler–ranker, planner–profiler–ranker, planner plus three rankers plus arbitrator, planner–profiler plus three rankers plus arbitrator, ranker–critic, and two-agent debate.
What would settle it
Re-run the seven workflows on the same Amazon-2023 cohorts but stratify users by a non-lexical complexity measure, such as the entropy of item categories or the number of latent intent clusters in the history; if the ensemble gains over the single-shot baseline vanish on the high-complexity cohort, the Jaccard proxy, and with it the paper's conditional conclusion, is an artifact of the title-token measure.
Extended reading notes
Core claim
Agentic recommendation is defined by the paper as a pipeline in which one or more stateful agents run an observe–plan–act–verify loop over users, item catalogs, candidate sets, and recommendation objectives, and a multi-agent recommender is formalized as a triple of agents, a shared environment, and a communication protocol. The empirical discovery is that this machinery is useful only conditionally. On a representative sample of Amazon-2023 users, every multi-agent workflow—profiler/ranker, planner/profiler/ranker, ensemble with arbitrator, ranker–critic, and two-agent debate—either ties or loses to the single-shot LLM call on NDCG@3 and NDCG@10 while costing three to six times more. On a high-diversity cohort stratified by the mean pairwise Jaccard distance of item-title tokens, the ensemble pipelines reach an NDCG@3 of 0.5162–0.5202 against 0.4898 for the baseline, while the ranker–critic pipeline remains below baseline, which the paper reads as ungrounded self-criticism amplifying error. The conclusion is that decomposition and ensembling pay off when the input is heterogeneous enough that a single pass under-resolves it, and iterative adversarial refinement does not.
Load-bearing premise
The empirical conclusion rests on equating 'high-diversity history' with the mean pairwise Jaccard distance between item-title token sets (Eq. 2); if that lexical proxy does not track the input complexity that actually makes decomposition and ensembling useful, the conditional design principle does not follow from the data.
Editorial extensions
If this is right
- Agentic orchestration should be treated as a routing decision: systems can estimate input diversity or complexity and choose whether to invoke decomposition or ensembling agents case by case.
- Evaluation of agentic recommenders should report a quality–latency–cost Pareto frontier rather than top-K accuracy alone, since added agents buy quality only on a subset of inputs.
- Closed-loop critique by an LLM without an external oracle is a likely source of error propagation, so verification should be grounded in tools, databases, or policy checkers.
- The profiler's history compression behaves like a working-memory mechanism: it helps exactly when the history is rich enough to make compression informative.
- The four task families and five challenge families give future work a shared vocabulary for attributing gains to the recommendation layer rather than to generic orchestration.
Reading between the lines
- The Jaccard-diversity proxy is testable: replacing it with a semantic or intent-based complexity measure, such as the number of latent interests inferred from embeddings, should preserve the cohort contrast if the underlying mechanism is input ambiguity rather than lexical heterogeneity.
- A natural production extension, not tested in the paper, is an adaptive router that measures history diversity and only then decides between the single-shot call and ensemble pipelines, using a cost budget.
- The closed-loop critique failure suggests that self-refinement benefits in other agentic settings may not transfer to ranking unless the critic has access to external evidence; this is an inference, not a paper claim.
- The same conditional logic likely applies to other recommendation tasks with variable input complexity, such as bundle construction or session-based recommendation, where the paper's task families sketch but do not test the transfer.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a formal framework for agentic recommender systems, defining recommender agents, multi-agent systems, memory update/retrieval functions, and a layered separation of recommendation, interaction, and orchestration concerns. It develops four representative task families and five challenge families tied to measurable RS signals, and reports a controlled next-item ranking study on Amazon-2023 comparing a single-shot LLM baseline with seven multi-agent workflows. The empirical study claims that multi-agent systems are not uniformly superior: on typical users the single-shot baseline is Pareto-efficient, while decomposition and ensemble pipelines help on high-diversity histories, supporting a conditional design principle for routing agentic complexity.
Significance. The paper's formal vocabulary and evaluation-oriented agenda are a useful contribution to the emerging area of agentic recommender systems, and the explicit separation of recommendation-layer outcomes from generic orchestration is a helpful corrective. The controlled study is well designed in terms of shared histories, candidate sets, prompts, and metrics, and the release of code and prompts supports reproducibility. However, the central empirical claim currently rests on point estimates without uncertainty quantification, and the per-category results show a reversal in one category; the framework part is solid, but the empirical evidence for the conditional design principle needs strengthening before the paper's headline claim is fully supported.
major comments (3)
- [Section 6.4, Tables 7–8] The central empirical claim that decomposition and ensembling help on high-diversity histories rests on point estimates without confidence intervals or significance tests. With one held-out relevant item per user, per-user NDCG@3 is a high-variance discrete variable; a conservative standard error for the pooled mean of 400 users is roughly 0.02–0.025, making the headline +6.2% NDCG@3 gain (PPEns vs. SA) approximately 1.2–1.5 standard errors. The paper should report paired bootstrap confidence intervals or a significance test before drawing the conditional design principle; as it stands, the observed differences could be sampling noise.
- [Appendix A, Table 10 (Toys and Games)] The high-diversity advantage is not uniform across categories: on Toys and Games, PEns NDCG@10 is below SA (0.6628 vs. 0.6772), and the pooled positive effect is driven mainly by Amazon Fashion and Appliances. Together with the absence of uncertainty quantification, this reversal means the abstract's claim that 'decomposition and ensemble agents become useful mainly for high-diversity histories' is stronger than the data support. The paper should either present per-category uncertainty and qualify the claim, or demonstrate that the pooled effect is statistically significant.
- [Section 6.1, Eq. (2)] The operationalization of 'high-diversity history' as the mean pairwise Jaccard distance between item-title token sets is an untested proxy for the input complexity that makes decomposition and ensembling useful. The entire cohort contrast rests on this measure; without a validation against alternative complexity notions (e.g., category entropy, number of distinct intents, or human annotation), the conditional design principle may not generalize beyond this lexical proxy.
minor comments (3)
- [Abstract] The abstract's statement that the results 'support a conditional design principle' should be qualified to indicate that this support comes from a pilot study with a single LLM and one task family, rather than a general empirical law.
- [Section 6.2] Consider reporting the raw sequential latency as well as the parallel critical-path latency, since the sequential figure conveys the inherent cost of composition even if the benchmark harness is single-threaded.
- [Section 7] The phrase 'pareto efficient' should be capitalized as 'Pareto-efficient' for consistency with the abstract and standard usage.
Circularity Check
No significant circularity: the conditional design principle rests on an external controlled comparison whose outcome metric is independent of the diversity definition, and self-citations are non-load-bearing.
full rationale
The paper's load-bearing empirical claim, that decomposition and ensemble agents help mainly on high-diversity histories, is established by a controlled comparison of seven fixed workflows on Amazon-2023 data under shared histories, candidate sets, prompts, and metrics. The cohort split is defined by the lexical diversity proxy Div(u) of Eq. 2, while the target outcome is next-item NDCG/MRR computed against a held-out positive that is not used in defining Div(u); the predictor and the outcome are independent, so the finding is contingent and falsifiable rather than true by construction. No parameter is fitted to the reported differences, and the cost tables are direct accounting, with the paper explicitly observing that the single-shot pipeline is the cost floor by construction rather than presenting this as an empirical discovery. Self-citations to the authors' prior work appear only as examples of tool-based or retrieval-augmented agentic designs in related-work positioning and do not carry the framework, the challenge agenda, or the experiment. Appendix A even reports that on Toys and Games the pure ensemble falls below the single-shot baseline at NDCG@10, a hedge that would be impossible if the conclusion were definitionally forced. The skeptic's concern about missing confidence intervals concerns statistical robustness, not circularity; the derivation chain contains no fitted input renamed as prediction and no imported uniqueness theorem.
Assumptions & free parameters
free parameters (3)
- high-diversity cohort size per category =
100 users
- minimum history length for eligibility =
7 items
- history cap for diversity computation =
100 titles
assumptions (5)
- domain assumption Lexical Jaccard diversity of item titles is a valid proxy for the input complexity that determines whether agentic decomposition helps (Eq. 2, Section 6.1).
- domain assumption A single held-out next purchase among ten candidates, with nine uniform-distractor items from the same category, is a meaningful evaluation of recommendation quality (Section 6.1).
- domain assumption gpt-5-mini-2025-08-07 is representative of LLM-based pipelines for comparing single-shot versus multi-agent workflows (Section 6.1).
- domain assumption The implemented workflows faithfully represent the role families: decomposition, ensembling, and iterative refinement (Table 5, Section 6.1).
- domain assumption Parallel critical-path latency is the correct basis for comparing workflow latency (Section 6.2).
Cite this review
Pith. "Pith review of The Future is Agentic: Definitions, Perspectives, and Open Challenges of Multi-Agent Recommender Systems." pith.science (2026). https://pith.science/paper/QDQY5LSN
@misc{pith2026250702097,
author = {Pith},
title = {Pith review of: The Future is Agentic: Definitions, Perspectives, and Open Challenges of Multi-Agent Recommender Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/QDQY5LSN}},
note = {Machine review of arXiv:2507.02097}
}
read the original abstract
Large language models (LLMs) are evolving from passive text generators into agentic systems that can plan, maintain state, invoke tools, and coordinate with other agents. This perspective paper examines what this shift means for recommender systems (RS). We define agentic recommender systems as pipelines in which one or more stateful agents observe, plan, call tools, and verify, rather than score in a single shot, while operating over users, item catalogs, candidate sets, and recommendation objectives. Their value is strongest when this machinery measurably improves recommendation-layer outcomes such as relevance, constraint satisfaction, bundle coherence, grounding, explanation faithfulness, or user effort, rather than merely because a pipeline contains an LLM or several modules. We introduce a recommender-specific formalism that models an agent by its state (user, context, history, candidate set), a reasoning core, tools, a hierarchical memory, and explicit policy constraints, and casts a multi-agent recommender as a triple of agents, a shared environment, and a communication protocol. Within this framework we develop four representative task families and an agenda tying five recurring challenge families to measurable RS signals. Finally, we run a controlled study comparing single-shot and multi-agent pipelines under shared histories, candidate sets, prompts, and metrics. A pilot next-item ranking study on Amazon-2023 shows multi-agent systems are not uniformly superior: on representative samples the single-shot baseline is Pareto-efficient, whereas decomposition and ensemble agents help mainly on high-diversity histories. This supports a conditional design principle: agentic complexity should be routed to cases where its marginal quality gain justifies the added latency, cost, and governance risk. Code: https://github.com/RezaYM/agenticrecsys.git
Figures
Figures from the paper (3 more)
Forward citations
Cited by 6 Pith papers
-
Attacking and Defending Multi-Agent Collaborative Filtering Systems Through Connectivity
In agent-based collaborative filtering, attack spread and privacy leakage grow with interaction connectivity, but the effect is asymmetric between user and item agents and differs between early and steady-state phases.
-
Autonomous Information Seeking: A Roadmap for Agentic Recommender Systems
Agentic recommender systems are organized by agent role (assisted, as-recommender, as-simulator) crossed with autonomy levels L2–L5, yielding a roadmap of architectures, evaluation limits, and open challenges.
-
RecoWorld: Building Simulated Environments for Agentic Recommender Systems
A design proposal, not a tested system: a dual-view simulation loop in which an LLM-simulated user issues reflective instructions when about to disengage, and an instruction-following recommender adapts to maximize si...
-
RAG-VisualRec: An Open Resource for Vision- and Text-Enhanced Retrieval-Augmented Generation in Recommendation
RAG-VisualRec is an open, auditable multimodal benchmark and pipeline for movie recommendation that fuses LLM-generated text with trailer embeddings and reports accuracy and beyond-accuracy metrics.
-
Agentic AI in 6G Software Businesses: A Layered Maturity Model
A preliminary thematic review distills 29 motivators and 27 demotivators for agentic AI adoption in 6G software businesses into five themes each and outlines a future maturity model.
-
OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems
A review-plus-demo claiming agentic capabilities emerge from system integration, backed by a 15-task benchmark whose C1→C3 performance gap is largely built into the test design.
Reference graph
Works this paper leans on
-
[1]
Swapnaja Achintalwar, Ioana Baldini, Djallel Bouneffouf, Joan Byamugisha, Maria Chang, Pierre Dognin, Eitan Farchi, Ndivhuwo Makondo, Aleksandra Mojsilović, Manish Nagireddy, et al. 2024. Alignment studio: Aligning large language models to particular contextual regulations.IEEE Internet Computing(2024)
2024
-
[2]
Abdul Basit Ahanger, Syed Wajid Aalam, Muzafar Rasool Bhat, and Assif Assad. 2022. Popularity bias in recommender systems-a review. InInternational Conference on Emerging Technologies in Computer Engineering. Springer, 431–444
2022
-
[3]
Chenxin An, Jun Zhang, Ming Zhong, Lei Li, Shansan Gong, Yao Luo, Jingjing Xu, and Lingpeng Kong. 2024. Why Does the Effective Context Length of LLMs Fall Short?arXiv preprint arXiv:2410.18745(2024)
arXiv 2024
-
[4]
Petr Anokhin, Nikita Semenov, Artyom Sorokin, Dmitry Evseev, Andrey Kravchenko, Mikhail Burtsev, and Evgeny Burnaev. 2024. Arigraph: Learning knowledge graph world models with episodic memory for llm agents.arXiv preprint arXiv:2407.04363(2024)
arXiv 2024
-
[5]
Ashmi Banerjee, Fitri Nur Aisyah, Adithi Satish, Wolfgang Wörndl, and Yashar Deldjoo. 2025. Collab-REC: An LLM-based Agentic Framework for Balancing Recommendations in Tourism.CoRRabs/2508.15030 (2025). doi:10. 48550/ARXIV.2508.15030 arXiv:2508.15030
-
[6]
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2022. Improving language models by retrieving from trillions of tokens. InInternational conference on machine learning. PMLR, 2206–2240
2022
-
[7]
Qiqi Cai, Jian Cao, Guandong Xu, and Nengjun Zhu. 2024. Distributed Recommendation Systems: Survey and Research Directions.ACM Transactions on Information Systems43, 1 (2024), 1–38
2024
-
[8]
Jiao Chen, Kehui Yao, Reza Yousefi Maragheh, Kai Zhao, Jianpeng Xu, Jason Cho, Evren Korpeoglu, Sushant Kumar, and Kannan Achan. 2025. CARTS: Collaborative Agents for Recommendation Textual Summarization.arXiv preprint arXiv:2506.17765(2025)
work page Pith review arXiv 2025
Show all 97 references
-
[9]
Luyu Chen, Quanyu Dai, Zeyu Zhang, Xueyang Feng, Mingyu Zhang, Pengcheng Tang, Xu Chen, Yue Zhu, and Zhenhua Dong. 2025. Recusersim: A realistic and diverse user simulator for evaluating conversational recommender systems. InCompanion Proceedings of the ACM on Web Conference 2...
2025
-
[10]
Pin-Yu Chen, Han Shen, Payel Das, and Tianyi Chen. 2025. Fundamental Safety-Capability Trade-offs in Fine-tuning Large Language Models.arXiv preprint arXiv:2503.20807(2025)
2025 arXiv
-
[11]
Konstantina Christakopoulou, Filip Radlinski, and Katja Hofmann. 2016. Towards conversational recommender systems. InProceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. 815–824
2016
-
[12]
Damien de Mijolla, Wen Yang, Philippa Duckett, Christopher Frye, and Mark Worrall. 2024. Language hooks: a modular framework for augmenting LLM reasoning that decouples tool usage from the model and its prompt.arXiv preprint arXiv:2412.05967(2024)
2024 arXiv
-
[13]
Yashar Deldjoo, Zhankui He, Julian McAuley, Anton Korikov, Scott Sanner, Arnau Ramisa, René Vidal, Maheswaran Sathiamoorthy, Atoosa Kasirzadeh, and Silvia Milano. 2024. A Review of Modern Recommender Systems using Generative Models (Gen-RecSys). InProceedings of the 30th ACM S...
2024
-
[14]
Yashar Deldjoo, Zhankui He, Julian McAuley, Anton Korikov, Scott Sanner, Arnau Ramisa, Rene Vidal, Maheswaran Sathiamoorthy, Atoosa Kasrizadeh, Silvia Milano, et al. 2024. Recommendation with Generative Models.arXiv preprint arXiv:2409.15173(2024)
2024 arXiv
-
[15]
Tenenbaum, and Igor Mordatch
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. 2024. Improving Factuality and Reasoning in Language Models through Multiagent Debate. InForty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 (Proce...
2024
-
[16]
Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim, Suhong Moon, Hiroki Furuta, Gopala Anumanchipalli, Kurt Keutzer, and Amir Gholami. 2025. Plan-and-act: Improving planning of agents for long-horizon tasks.arXiv preprint arXiv:2503.09572 (2025)
2025 arXiv
-
[17]
Jiabao Fang, Shen Gao, Pengjie Ren, Xiuying Chen, Suzan Verberne, and Zhaochun Ren. 2024. A multi-agent conversa- tional recommender system.arXiv preprint arXiv:2402.01135(2024)
2024 arXiv
-
[18]
Libo Feng, Hui Zhang, Yong Chen, and Liqi Lou. 2018. Scalable dynamic multi-agent practical byzantine fault-tolerant consensus in permissioned blockchain.Applied Sciences8, 10 (2018), 1919
2018
-
[19]
Aleksander Ficek, Jiaqi Zeng, and Oleksii Kuchaiev. 2024. GPT vs RETRO: Exploring the Intersection of Retrieval and Parameter-Efficient Fine-Tuning. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 19425–19432
2024
-
[20]
Joao Fonseca, Andrew Bell, and Julia Stoyanovich. 2025. Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs.arXiv preprint arXiv:2501.02018(2025). ACM Trans. Recomm. Syst., Vol. 1, No. 1, Article . Publication date: July 2026. The Future ...
2025 arXiv
-
[21]
Najmeh Forouzandehmehr, Reza Yousefi Maragheh, Sriram Kollipara, Kai Zhao, Topojoy Biswas, Evren Korpeoglu, and Kannan Achan. 2025. CAL-RAG: Retrieval-Augmented Multi-Agent Generation for Content-Aware Layout Design. arXiv:2506.21934 [cs.IR] https://arxiv.org/pdf/2506.21934
2025 arXiv
-
[22]
Zafeirios Fountas, Martin A Benfeghoul, Adnan Oomerjee, Fenia Christopoulou, Gerasimos Lampouras, Haitham Bou- Ammar, and Jun Wang. 2024. Human-like episodic memory for infinite context llms.arXiv preprint arXiv:2407.09450 (2024)
2024
-
[23]
Chongming Gao, Wenqiang Lei, Xiangnan He, Maarten De Rijke, and Tat-Seng Chua. 2021. Advances and challenges in conversational recommender systems: A survey.AI open2 (2021), 100–126
2021
-
[24]
Jing Guo, Nan Li, Jianchuan Qi, Hang Yang, Ruiqiao Li, Yuzhen Feng, Si Zhang, and Ming Xu. 2023. Empowering working memory for large language model agents.arXiv preprint arXiv:2312.17259(2023)
2023 arXiv
-
[25]
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. InInternational conference on machine learning. PMLR, 3929–3938
2020
-
[26]
Shanshan Han, Qifan Zhang, Yuhang Yao, Weizhao Jin, and Zhaozhuo Xu. 2024. LLM multi-agent systems: Challenges and open problems.arXiv preprint arXiv:2402.03578(2024)
2024 arXiv
-
[27]
Shibo Hao, Tianyang Liu, Zhen Wang, and Zhiting Hu. 2023. Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings.Advances in neural information processing systems36 (2023), 45870–45894
2023
-
[28]
Heidy Hazem, Ahmed Awad, and Ahmed Hassan Yousef. 2023. A distributed real-time recommender system for big data streams.Ain Shams Engineering Journal14, 8 (2023), 102026
2023
-
[29]
Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk. 2016. Session-based Recommendations with Recurrent Neural Networks. In4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedi...
2016 arXiv
- [30]
-
[31]
Xu Huang, Jianxun Lian, Yuxuan Lei, Jing Yao, Defu Lian, and Xing Xie. 2025. Recommender ai agent: Integrating large language models for interactive recommendations.ACM Transactions on Information Systems43, 4 (2025), 1–33
2025
-
[32]
Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recommendation. In2018 IEEE international conference on data mining (ICDM). IEEE, 197–206
2018
-
[33]
Yehuda Koren, Robert Bell, and Chris Volinsky. 2009. Matrix Factorization Techniques for Recommender Systems. Computer42, 8 (2009), 30–37. doi:10.1109/MC.2009.263
2009 doi
-
[34]
John E Laird, Allen Newell, and Paul S Rosenbloom. 1987. Soar: An architecture for general intelligence.Artificial intelligence33, 1 (1987), 1–64
1987
-
[35]
LangChain. 2024. LangChain Memory Types — Conceptual Guide. https://langchain-ai.github.io/langmem/concepts/ conceptual_guide/#memory-types Accessed: 2025-06-21
2024
-
[36]
Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra. 2017. Deal or No Deal? End-to-End Learning of Negotiation Dialogues. InProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. 2443–2453
2017
-
[37]
Junyou Li, Qin Zhang, Yangbin Yu, Qiang Fu, and Deheng Ye. 2024. More Agents Is All You Need.Trans. Mach. Learn. Res.2024 (2024). https://openreview.net/forum?id=bgzUSZ8aeg
2024
-
[38]
Kun Li, Xin Jing, and Chengang Jing. 2024. Vector Storage Based Long-term Memory Research on LLM.International Journal of Advanced Network, Monitoring and Controls(2024)
2024
-
[39]
Lei Li, Yongfeng Zhang, Dugang Liu, and Li Chen. 2024. Large language models for generative recommendation: A survey and visionary discussions. InProceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation (LREC-COLING...
2024
-
[40]
Zongqian Li, Yinhong Liu, Yixuan Su, and Nigel Collier. 2024. Prompt compression for large language models: A survey.arXiv preprint arXiv:2410.12388(2024)
2024 arXiv
-
[41]
Jan Malte Lichtenberg, Alexander Buchholz, and Pola Schwöbel. 2024. Large language models as recommender systems: A study of popularity bias.arXiv preprint arXiv:2406.01285(2024)
2024 arXiv
-
[42]
Jiahao Liu, Shengkang Gu, Dongsheng Li, Guangping Zhang, Mingzhe Han, Hansu Gu, Peng Zhang, Tun Lu, Li Shang, and Ning Gu. 2025. Enhancing Cross-Domain Recommendations with Memory-Optimized LLM-Based User Agents. arXiv e-prints(2025), arXiv–2502
2025
-
[43]
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the Middle: How Language Models Use Long Contexts.Transactions of the Association for Computational Linguistics11 (2024), 157–173
2024
-
[44]
Xuan Liu, Jie Zhang, Haoyang Shang, Song Guo, Chengxu Yang, and Quanyan Zhu. 2024. Exploring prosocial irrationality for llm agents: A social cognition view.arXiv preprint arXiv:2405.14744(2024)
2024 arXiv
-
[45]
Yuxing Lu and Jinzhuo Wang. 2025. KARMA: Leveraging Multi-Agent LLMs for Automated Knowledge Graph Enrichment.arXiv preprint arXiv:2502.06472(2025). ACM Trans. Recomm. Syst., Vol. 1, No. 1, Article . Publication date: July 2026. 42 Reza Yousefi Maragheh and Yashar Deldjoo
2025
-
[46]
Yu, and Ming Zhang
Junyu Luo, Weizhi Zhang, Ye Yuan, Yusheng Zhao, Junwei Yang, Yiyang Gu, Bohan Wu, Binqi Chen, Ziyue Qiao, Qingqing Long, Rongcheng Tu, Xiao Luo, Wei Ju, Zhiping Xiao, Yifan Wang, Meng Xiao, Chenwu Liu, Jingyang Yuan, Shichang Zhang, Yiqiao Jin, Fan Zhang, Xian Wu, Hanqing Zhao...
-
[47]
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023. Self-Refine: It...
2023
-
[48]
Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, and Yuwei Fang. 2024. Evaluating Very Long-Term Conversational Memory of LLM Agents. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pape...
2024
-
[49]
Potsawee Manakul, Adian Liusie, and Mark JF Gales. 2023. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models.arXiv preprint arXiv:2303.08896(2023)
2023 arXiv
-
[50]
Reza Yousefi Maragheh, Chenhao Fang, Charan Chand Irugu, Parth Parikh, Jason Cho, Jianpeng Xu, Saranyan Sukumar, Malay Patel, Evren Korpeoglu, Sushant Kumar, et al. 2023. LLM-TAKE: Theme-aware keyword extraction using large language models. In2023 IEEE International Conference...
2023
-
[51]
1990.Knapsack problems: algorithms and computer implementations
Silvano Martello and Paolo Toth. 1990.Knapsack problems: algorithms and computer implementations. John Wiley & Sons, Inc
1990
-
[52]
Chenlin Ming, Jiacheng Lin, Pangkit Fong, Han Wang, Xiaoming Duan, and Jianping He. 2023. Hicrisp: A hierarchical closed-loop robotic intelligent self-correction planner.arXiv preprint arXiv:2309.12089(2023)
2023 arXiv
-
[53]
Fangwen Mu, Junjie Wang, Lin Shi, Song Wang, Shoubin Li, and Qing Wang. 2025. EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair.arXiv preprint arXiv:2506.10484(2025)
2025 arXiv
-
[54]
Sid Nayak, Adelmo Morrison Orozco, Marina Have, Jackson Zhang, Vittal Thirumalai, Darren Chen, Aditya Kapoor, Eric Robinson, Karthik Gopalakrishnan, James Harrison, et al. 2024. Long-horizon planning for multi-agent robots in partially observable environments.Advances in Neura...
2024
-
[55]
Emil Noordeh, Roman Levin, Ruochen Jiang, and Harris Shadmany. 2020. Echo chambers in collaborative filtering based recommendation systems.arXiv preprint arXiv:2011.03890(2020)
2020 arXiv
-
[56]
Bo Pan, Jiaying Lu, Ke Wang, Li Zheng, Zhen Wen, Yingchaojie Feng, Minfeng Zhu, and Wei Chen. 2024. AgentCoord: Visually exploring coordination strategy for llm-based multi-agent collaboration.arXiv preprint arXiv:2404.11943 (2024)
2024 arXiv
-
[57]
Bhargavi Paranjape, Scott Lundberg, Sameer Singh, Hannaneh Hajishirzi, Luke Zettlemoyer, and Marco Tulio Ribeiro
-
[58]
Qiyao Peng, Hongtao Liu, Hua Huang, Qing Yang, and Minglai Shao. 2025. A Survey on LLM-powered Agents for Recommender Systems.arXiv preprint arXiv:2502.10050(2025)
2025 arXiv
-
[59]
Mathis Pink, Qinyuan Wu, Vy Ai Vo, Javier Turek, Jianing Mu, Alexander Huth, and Mariya Toneva. 2025. Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents.arXiv preprint arXiv:2502.06975(2025)
2025 arXiv
-
[60]
Stefan Poslad. 2007. Specifying protocols for multi-agent systems interaction.ACM Transactions on Autonomous and Adaptive Systems (TAAS)2, 4 (2007), 15–es
2007
-
[61]
Archiki Prasad, Alexander Koller, Mareike Hartmann, Peter Clark, Ashish Sabharwal, Mohit Bansal, and Tushar Khot
-
[62]
Preston Rasmussen, Pavlo Paliychuk, Travis Beauvais, Jack Ryan, and Daniel Chalef. 2025. Zep: A Temporal Knowledge Graph Architecture for Agent Memory.arXiv preprint arXiv:2501.13956(2025)
2025 arXiv
-
[63]
Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme. 2009. BPR: Bayesian personalized ranking from implicit feedback. InProceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence. 452–461
2009
-
[64]
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools.Advances in Neural Information Processing Systems36 (2023), 68539–68551
2023
-
[65]
Lianlei Shan, Shixian Luo, Zezhou Zhu, Yu Yuan, and Yong Wu. 2025. Cognitive memory in large language models. arXiv preprint arXiv:2504.02441(2025). ACM Trans. Recomm. Syst., Vol. 1, No. 1, Article . Publication date: July 2026. The Future is Agentic: Definitions, Perspectives...
2025 arXiv
-
[66]
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023. Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.Advances in Neural Information Processing Systems36 (2023), 38154–38180
2023
-
[67]
Adi Simhi, Itay Itzhak, Fazl Barez, Gabriel Stanovsky, and Yonatan Belinkov. 2025. Trust Me, I’m Wrong: High-Certainty Hallucinations in LLMs.arXiv preprint arXiv:2502.12964(2025)
2025 arXiv
-
[68]
frequently bought together
AP News Staff. 2017. Amazon removes “frequently bought together” items used to make explosives. https://apnews. com/article/604a73f6008846449c303ffb4b93e9d6. Accessed: 2025-06-29
2017
-
[69]
Chris Stokel-Walker. 2024. Google and Character.AI are being sued after chatbot allegedly told a 17-year-old to kill his parents. https://www.businessinsider.com/characterai-google-lawsuit-chatbot-teen-kill-parents-2024-12. Accessed: 2025-06-29
2024
-
[70]
Nicholas Sukiennik, Haoyu Wang, Zailin Zeng, Chen Gao, and Yong Li. 2025. Simulating Filter Bubble on Short-video Recommender System with Large Language Model Agents.arXiv preprint arXiv:2504.08742(2025)
2025 arXiv
-
[71]
Theodore Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas Griffiths. 2023. Cognitive architectures for language agents.Transactions on Machine Learning Research(2023)
2023
-
[72]
Haotian Sun, Yuchen Zhuang, Lingkai Kong, Bo Dai, and Chao Zhang. 2023. Adaplanner: Adaptive planning from feedback with language models.Advances in neural information processing systems36 (2023), 58202–58245
2023
-
[73]
Ding Tong, Qifeng Qiao, Ting-Po Lee, James McInerney, and Justin Basilico. 2023. Navigating the feedback loop in recommender systems: Insights and strategies from industry practice. InProceedings of the 17th ACM Conference on Recommender Systems. 1058–1061
2023
-
[74]
Mehmet Ugurbil, Dimitris Mouris, Manuel B Santos, José Cabrero-Holgueras, Miguel de Vega, and Shubho Sengupta
-
[75]
Lei Wang, Jingsen Zhang, Xiaowen Chen, Yankai Lin, Ruihua Song, Wayne Xin Zhao, and Ji-Rong Wen. 2023. Recagent: A novel simulation paradigm for recommender systems.arXiv preprint arXiv:2306.02552(2023)
2023 arXiv
-
[76]
Yancheng Wang, Ziyan Jiang, Zheng Chen, Fan Yang, Yingxue Zhou, Eunah Cho, Xing Fan, Yanbin Lu, Xiaojiang Huang, and Yingzhen Yang. 2024. RecMind: Large Language Model Powered Agent For Recommendation. InFindings of the Association for Computational Linguistics: NAACL 2024. 4351–4364
2024
-
[77]
Zhefan Wang, Yuanqing Yu, Wendi Zheng, Weizhi Ma, and Min Zhang. 2024. Macrec: A multi-agent collaboration framework for recommendation. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2760–2764
2024
-
[78]
Schaun Wheeler and Olivier Jeunen. 2025. Procedural memory is not all you need: Bridging cognitive gaps in llm-based agents. InAdjunct Proceedings of the 33rd ACM Conference on User Modeling, Adaptation and Personalization. 360–364
2025
-
[79]
Likang Wu, Zhi Zheng, Zhaopeng Qiu, Hao Wang, Hongchao Gu, Tingjia Shen, Chuan Qin, Chen Zhu, Hengshu Zhu, Qi Liu, et al. 2024. A survey on large language models for recommendation.World Wide Web27, 5 (2024), 60
2024
-
[80]
Yunjia Xi, Weiwen Liu, Jianghao Lin, Bo Chen, Ruiming Tang, Weinan Zhang, and Yong Yu. 2024. MemoCRS: Memory- enhanced Sequential Conversational Recommender Systems with Large Language Models. InProceedings of the 33rd ACM International Conference on Information and Knowledge ...
2024
-
[81]
Zidi Xiong, Yuping Lin, Wenya Xie, Pengfei He, Jiliang Tang, Himabindu Lakkaraju, and Zhen Xiang. 2025. How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior.arXiv preprint arXiv:2505.16067(2025)
2025
-
[82]
Wujiang Xu, Kai Mei, Hang Gao, Juntao Tan, Zujie Liang, and Yongfeng Zhang. 2025. A-mem: Agentic memory for llm agents.arXiv preprint arXiv:2502.12110(2025)
2025 arXiv
-
[83]
Zhengyu Yang, Danlin Jia, Stratis Ioannidis, Ningfang Mi, and Bo Sheng. 2018. Intermediate data caching optimization for multi-stage and parallel big data frameworks. In2018 IEEE 11th International Conference on Cloud Computing (CLOUD). IEEE, 277–284
2018
-
[84]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in language models. InInternational Conference on Learning Representations (ICLR)
2023
-
[85]
Hongbin Ye, Honghao Gui, Xin Xu, Xi Chen, Huajun Chen, and Ningyu Zhang. 2023. Schema-adaptable Knowledge Graph Construction. InFindings of the Association for Computational Linguistics: EMNLP 2023. 6408–6431
2023
-
[86]
Reza Yousefi Maragheh, Pratheek Vadla, Priyank Gupta, Kai Zhao, Aysenur Inan, Kehui Yao, Jianpeng Xu, Praveen Kanumala, Jason Cho, and Sushant Kumar. 2025. ARAG: Agentic Retrieval Augmented Generation for Personalized Recommendation. arXiv:2506.21931 [cs.IR] https://arxiv.org/...
2025 arXiv
-
[87]
Zhenrui Yue, Sara Rabhi, Gabriel de Souza Pereira Moreira, Dong Wang, and Even Oldridge. 2023. Llamarec: Two-stage recommendation using large language models for ranking.arXiv preprint arXiv:2311.02089(2023)
2023 arXiv
-
[88]
Ruihong Zeng, Jinyuan Fang, Siwei Liu, and Zaiqiao Meng. 2024. On the Structural Memory of LLM Agents.arXiv preprint arXiv:2412.15266(2024)
2024 arXiv
-
[89]
An Zhang, Yuxin Chen, Leheng Sheng, Xiang Wang, and Tat-Seng Chua. 2024. On generative agents in recommendation. InProceedings of the 47th international ACM SIGIR conference on research and development in Information Retrieval. ACM Trans. Recomm. Syst., Vol. 1, No. 1, Article ...
2024
- [90]
-
[91]
Yu Zhang, Shutong Qiao, Jiaqi Zhang, Tzu-Heng Lin, Chen Gao, and Yong Li. 2025. A Survey of Large Language Model Empowered Agents for Recommendation and Search: Towards Next-Generation Information Retrieval.arXiv preprint arXiv:2503.05659(2025)
2025 arXiv
-
[92]
Yuyue Zhao, Jiancan Wu, Xiang Wang, Wei Tang, Dingxian Wang, and Maarten De Rijke. 2024. Let me do it for you: Towards llm empowered recommendation via tool learning. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrie...
2024
-
[93]
Zihuai Zhao, Wenqi Fan, Jiatong Li, Yunqing Liu, Xiaowei Mei, Yiqi Wang, Zhen Wen, Fei Wang, Xiangyu Zhao, Jiliang Tang, et al. 2024. Recommender systems in the era of large language models (llms).IEEE Transactions on Knowledge and Data Engineering36, 11 (2024), 6889–6907
2024
-
[94]
Total” sums input and output tokens; “Lat. par
Jing Zhu, Chengfang Lu, Juanjuan Li, and Fei-Yue Wang. 2025. Secure Consensus Control on Multi-Agent Systems Based on Improved PBFT and Raft Blockchain Consensus Algorithms.IEEE/CAA Journal of Automatica Sinica12, 7 (2025), 1407–1417. ACM Trans. Recomm. Syst., Vol. 1, No. 1, A...
2025
-
[2023]
Art: Automatic multi-step reasoning and tool-use for large language models.arXiv preprint arXiv:2303.09014 (2023)
2023 arXiv
-
[2024]
InFindings of the Association for Computational Linguistics: NAACL 2024, Mexico City, Mexico, June 16-21, 2024 (Findings of ACL, Vol
ADaPT: As-Needed Decomposition and Planning with Language Models. InFindings of the Association for Computational Linguistics: NAACL 2024, Mexico City, Mexico, June 16-21, 2024 (Findings of ACL, Vol. NAACL 2024). Association for Computational Linguistics, 4226–4252. doi:10.186...
2024 doi
-
[2025]
Fission: Distributed Privacy-Preserving Large Language Model Inference.Cryptology ePrint Archive(2025)
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.