Pith. sign in

REVIEW 5 major objections 7 minor 167 references

A Comprehensive Survey on Integrating Large Language Models with Knowledge-Based Methods

T0 review · 5 major / 7 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This survey argues that the evidence supports integrating LLMs with structured knowledge: hybrid systems improve data contextualization, model accuracy, and use of knowledge resources, while easing hallucination and interpretability…

desk verdict A useful but citation-sloppy survey whose three-level taxonomy and recent coverage deserve referee time, though the abstract overclaims and the reference list needs a serious audit. read the letter →

arxiv 2501.13947 v3 pith:D5OMRB6B submitted 2025-01-19 cs.CL cs.AI

classification cs.CLcs.AI
keywords largelanguagemodelsknowledgebasesgraphsretrieval-augmentedgenerationpromptengineeringintegrationhybridAIsystemssurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Large language models alone are powerful but opaque, costly, and prone to outdated or invented facts. This survey argues that the research literature supports a remedy: connect the LLM to structured knowledge, through knowledge bases, knowledge graphs, and retrieval-augmented generation, so its outputs are grounded in verifiable external information. A sympathetic reading of the paper is that this integration demonstrably improves data contextualization, model accuracy, and the utilization of knowledge resources, and that the field is converging on hybrid RAG-plus-graph systems. The paper organizes the evidence into a three-level account, retrieval and integration, reasoning and utilization, and optimization and specialization, and supplies evaluation metrics, case studies, and recommendations.

What carries the argument

The carrying object is the hybrid LLM-knowledge pipeline rather than any single model. In its dominant form, retrieval-augmented generation, the pipeline couples a retriever that fetches relevant documents or graph triples with a generator that conditions its answer on that retrieved context; knowledge graphs supply the structured relations that let the system reason over entities and paths, and prompt-engineering techniques such as chain-of-thought orchestrate how the model uses the retrieved knowledge. The paper's three-level taxonomy, knowledge retrieval and integration, knowledge utilization and reasoning, and optimization and specialization, is the organizing device that turns scattered results into a progressive claim.

What would settle it

Check the paper's central claim by re-running this literature synthesis as a preregistered systematic review and independently re-extracting the performance numbers in Tables 6 and 7 from the cited papers. If many of the reported improvements, for example RAG's EM gains over DPR or KG-Agent's Hits@1 gains over GPT-4, turn out to be misreported or to wash out in a balanced meta-analysis that includes negative results, the claim that integration benefits are demonstrated would collapse.

Watch

Extended reading notes

Core claim

The central claim, stated on the paper's own terms, is that combining generative LLMs with knowledge-based systems delivers benefits that neither component achieves alone: the LLM supplies flexible language understanding and generation, while the knowledge base supplies precise, current, structured facts. The survey catalogues four integration families, basic knowledge bases, knowledge graphs, retrieval-augmented generation, and prompt augmentation, and claims each mitigates a specific LLM weakness: KBs reduce hallucination through verified facts, KGs enable multi-hop reasoning and explainability, RAG provides real-time external knowledge and traceability, and structured prompting improves step-by-step reasoning. It further claims that hybrid approaches, RAG combined with KGs, graph-based reranking, and agentic KG traversal, scale this benefit to high-stakes domains such as medicine, law, and finance.

Load-bearing premise

The load-bearing premise is that the selected literature is representative: the survey describes no search strategy, inclusion criteria, or quality assessment, so its conclusion that integration demonstrably works could fail if the paper selection is incomplete or biased.

Editorial extensions

If this is right

  • Retrieval-augmented generation can serve as an external memory that keeps LLM answers current without retraining, since the model pulls fresh documents at inference time.
  • Grounding responses in knowledge graphs supports multi-hop reasoning and gives users a traceable path from question to answer, which directly addresses the black-box concern.
  • Hybrid RAG-plus-graph systems are the most scalable option for specialized fields, because the graph can down-sample redundant data and surface rare but critical facts.
  • Prompt augmentation such as chain-of-thought makes reasoning steps inspectable, turning some of the interpretability gap into a process that can be checked.
  • Evaluation of integrated systems needs retrieval metrics, reasoning-consistency metrics, and generation-fluency metrics together, not just text-overlap scores.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the survey's synthesis is right, the next practical step is modular pipelines where retriever, reasoner, and generator can be upgraded independently, rather than retraining a monolithic model.
  • The evidence pattern implies that future benchmarks should report both retrieval quality (Recall@K, Hits@K) and reasoning consistency jointly, since the paper's own tables show dataset-dependent gains and some declines.
  • The case studies point to a testable extension: domain-specific hybrid systems in medicine, finance, and law should be compared head-to-head on the same tasks against general-purpose LLMs with and without grounding.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 7 minor

Summary. This manuscript is a survey of methods for integrating large language models (LLMs) with knowledge-based systems, including knowledge bases, knowledge graphs, retrieval-augmented generation (RAG), and prompt engineering. It presents a taxonomy of integration techniques, a three-level framework for advanced hybrid approaches, a catalog of evaluation metrics, performance comparison tables for selected RAG- and KG-based models, and four case studies (FinAgent, UMLS, Codex, BloombergGPT). The paper's central claim, stated in the Abstract and Section 6, is that the literature demonstrates that integrating generative AI with structured knowledge systems improves data contextualization, model accuracy, and utilization of knowledge resources. The survey also discusses technical, operational, and ethical challenges and offers recommendations for future work.

Significance. If its evidence base were reliable, this survey would be a useful entry point to the LLM-knowledge integration literature, with its organized taxonomy, comparative tables, and case-study coverage. The paper's strengths include a broad scope, a clear three-level framework (retrieval/integration, utilization/reasoning, optimization/specialization), and a substantial list of metrics and models. However, the central claim is a synthesis of the cited literature, and the manuscript provides no new experimental or analytical verification. The reliability of that synthesis is directly compromised by several identified citation errors, a self-citation used as an external source, and the absence of any described literature selection methodology. These issues are load-bearing because the paper's conclusions inherit their credibility entirely from its reference base.

major comments (5)
  1. [Section 5.1.2, reference [85]] The claim that knowledge graphs improve factual accuracy and knowledge probing, with the LAMA dataset as an example, is supported by reference [85]. However, [85] in the bibliography is a condensed-matter physics paper on the one-dimensional anisotropic quantum XY model (arXiv:2302.13866). This citation does not support the LAMA/knowledge-probing statement, and the same reference is also cited in Section 3.1.2 for energy consumption concerns. The factual support for a central claim in the knowledge-graph portion of the survey is therefore unverified.
  2. [Section 3.3, references [102] and [103]] The sentence 'Addressing these social issues calls for the development of new ethical standards within the generative AI sphere and the implementation of measures to reduce bias and embrace transparency in Gen AI operations' is cited to [102], which is the original generative adversarial network paper by Goodfellow et al. The following sentence, invoking 'differential privacy and federated learning' and trustworthy LLMs, is cited to [103], which is the variational autoencoder paper by Kingma and Welling. Neither reference addresses ethical standards, bias reduction, differential privacy, or federated learning. These are not borderline related citations; they are topically disjoint, and their presence indicates that the reference list has not been reliably curated.
  3. [Section 2.3, reference [42]] The text states that LLMs 'facilitate large-scale information retrieval, integrating knowledge from external sources like ERNIE and E-BERT to enhance their contextual awareness [42].' Reference [42] is the present manuscript itself (arXiv:2501.13947). This is a self-citation used as an independent source for a factual claim about ERNIE and E-BERT, which is circular. The claim should be supported by the actual ERNIE/E-BERT publications or another independent source.
  4. [Abstract and Section 1] The Abstract claims a 'comprehensive examination of the literature,' but the manuscript does not describe any search strategy, inclusion or exclusion criteria, source selection process, or quality assessment. For a survey whose central conclusions are entirely dependent on the selected references, this omission makes the synthesis non-verifiable. The authors should either add a methodology subsection describing how the literature was identified and screened, or substantially soften the claim to 'a narrative review of selected works.'
  5. [Tables 6 and 7] Tables 6 and 7 compile performance improvements for RAG-based and KG-based LLMs, and the surrounding text uses these numbers to support the claim that integration improves accuracy. However, the tables report deltas over baselines without any protocol information: no hyperparameters, evaluation setups, significance tests, or error bars are given, and the original sources are cited only as reference numbers in the 'Refs' column. Several entries are also difficult to interpret as reported, for example the repeated '↑2.6% R-L over BART' row for Open MS MARCO in Table 6 and the 'same F1 with GPT4o-mini' entry in Table 7. The quantitative evidence for the paper's central claim is therefore not reproducible from the information provided.
minor comments (7)
  1. [Section 1] The sentence 'This has landed the models in areas where they can excel in generating natural text, never reaching any peak before' is ungrammatical and should be rewritten.
  2. [Section 2.1] The parenthetical '((i.e. what a new law brings to make old predictions still work))' is confusing and appears to be an unfinished editorial note; it should be removed or clarified.
  3. [Section 5.1.1] The text contains the typo 'WiKidata' for 'Wikidata'.
  4. [Section 5.4] The sentence 'This section focuses on how integrated large language models (LLMs) effectively address the challenges encountered by traditional LLMs, as discussed in Chapter 2' should refer to Section 2, not Chapter 2.
  5. [Table 4] The entry 'MoE-RAG KAG' in the 'Adaptive Knowledge Infusion for Decision-Making' row should separate the two model names (e.g., 'MoE-RAG, KAG') for clarity.
  6. [Table 6] The 'Open MS MARCO' row lists the identical improvement '↑2.6% R-L over BART' twice; one occurrence appears to be a duplicate error.
  7. [References] References [7] and [33] are the same paper (Raiaan et al., 'A review on large language models: Architectures, applications, taxonomies, open issues and challenges'), which should be merged or cited only once.

Circularity Check

1 steps flagged · score 2.0 of 10

Minor self-citation, but the survey's central synthesis is not circular; no derivation chain reduces to its inputs.

  1. other [Section 2.3 (LLM capabilities and real-world applications) and reference [42]]
    "Furthermore, LLMs facilitate large-scale information retrieval, integrating knowledge from external sources like ERNIE and E-BERT to enhance their contextual awareness [42]."

    Reference [42] is the present manuscript itself, listed as 'Lilian Some, Wenli Yang, Michael Bain, Byeong Kang, A comprehensive survey on integrating large language models with knowledge-based methods, 2025, arXiv preprint arXiv:2501.13947'. The paper therefore cites itself as the support for a factual claim about ERNIE and E-BERT. A survey cannot independently verify a factual assertion by citing its own preprint, since the survey contains no new experiments on ERNIE/E-BERT. This is a genuine self-citation used as evidence, but it is not load-bearing for the central claim of the paper, which is a synthesis of external literature on LLM-knowledge integration.

full rationale

This paper is a survey, not a derivation. Its central claim — that integrating LLMs with knowledge bases, graphs, and retrieval improves contextualization and accuracy — is presented as a synthesis of external results, supported by cited benchmarks such as Tables 6 and 7. There are no fitted parameters, no equations, and no prediction that reduces by construction to an input. The only identifiable circularity signal is the self-citation in reference [42], used to support a peripheral statement about ERNIE and E-BERT in Section 2.3. That self-citation is not load-bearing for the survey's main conclusion. Separately, the paper's reliability is weakened by unrelated or mismatched citations: reference [85] is a condensed-matter physics paper on the quantum XY model cited for LAMA knowledge probing, reference [102] is the original GAN paper cited for ethical standards, and reference [103] is the VAE paper cited for differential privacy and federated learning. These are citation-quality problems, not circularity, and they affect correctness risk rather than the circularity score. The absence of a documented search strategy also limits verifiability of the 'comprehensive examination' claim, but again this is a methodological limitation, not a circular derivation. Accordingly, the score is 2, reflecting one minor self-citation that does not determine the paper's central finding.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey introduces no free parameters or invented entities. It depends on the trustworthiness of its references and on the representativeness of an undocumented literature selection. Both are shaky, as evidenced by unrelated citations, a self-citation, and unverifiable performance tables.

assumptions (3)
  • domain assumption Cited sources are accurately summarized and relevant to the claims they support.
    The survey's conclusions about RAG, KG, and KB integration rest entirely on the accuracy and relevance of its references. Miscitations such as [85] (physics paper for carbon footprint) and [102] (GAN paper for ethics) show this assumption is violated in places.
  • domain assumption The literature selected is representative and comprehensive for the field.
    The paper provides no systematic search protocol, inclusion criteria, or quality assessment, yet claims comprehensiveness in the Abstract and Section 1. This assumption is load-bearing for the survey's value.
  • domain assumption Performance improvements reported in Tables 6 and 7 are correctly transcribed from the cited papers.
    The tables compile numbers from external papers but give no access to evaluation artifacts or error bars; duplication and unclear formatting in Table 6 reduce confidence in transcription.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Comprehensive Survey on Integrating Large Language Models with Knowledge-Based Methods." pith.science (2026). https://pith.science/paper/D5OMRB6B

@misc{pith2026250113947,
  author       = {Pith},
  title        = {Pith review of: A Comprehensive Survey on Integrating Large Language Models with Knowledge-Based Methods},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/D5OMRB6B}},
  note         = {Machine review of arXiv:2501.13947}
}
read the original abstract

The rapid development of artificial intelligence has led to marked progress in the field. One interesting direction for research is whether Large Language Models (LLMs) can be integrated with structured knowledge-based systems. This approach aims to combine the generative language understanding of LLMs and the precise knowledge representation systems by which they are integrated. This article surveys the relationship between LLMs and knowledge bases, looks at how they can be applied in practice, and discusses related technical, operational, and ethical challenges. Utilizing a comprehensive examination of the literature, the study both identifies important issues and assesses existing solutions. It demonstrates the merits of incorporating generative AI into structured knowledge-base systems concerning data contextualization, model accuracy, and utilization of knowledge resources. The findings give a full list of the current situation of research, point out the main gaps, and propose helpful paths to take. These insights contribute to advancing AI technologies and support their practical deployment across various sectors.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

167 extracted references · 9 canonical work pages

  1. [85]

    Yun-Tong Yang, Hong-Gang Luo, Topological or not? A unified pattern descrip- tion in the one-dimensional anisotropic quantum XY model with a transverse field, 2023, arXiv:2302.13866 [cond-mat]

  2. [102]

    Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde- Farley, Sherjil Ozair, Aaron Courville, Yoshua Bengio, Generative adversarial networks, 2014

    Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde- Farley, Sherjil Ozair, Aaron Courville, Yoshua Bengio, Generative adversarial networks, 2014

  3. [103]

    Kingma, M

    D.P. Kingma, M. Welling, Auto-encoding variational bayes, 2013

  4. [42]

    Lilian Some, Wenli Yang, Michael Bain, Byeong Kang, A comprehensive survey on integrating large language models with knowledge-based methods, 2025, arXiv preprint arXiv:2501.13947

  5. [1]

    Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, Jianfeng Gao, Large language models: A survey, 2024

  6. [2]

    Zichong Wang, Zhibo Chu, Thang Viet Doan, Shiwen Ni, Min Yang, Wenbin Zhang, History, development, and principles of large language models: an introductory survey, 2024

  7. [3]

    Sanjay Kukreja, Tarun Kumar, Amit Purohit, Abhijit Dasgupta, Debashis Guha, A literature survey on open source large language models, 2024

  8. [4]

    Reyes, Andre L.C

    Teo Susnjak, Peter Hwang, Napoleon H. Reyes, Andre L.C. Barczak, Timo- thy R. McIntosh, Surangika Ranathunga, Automating research synthesis with domain-specific large language model fine-tuning, 2024

Show all 167 references
  1. [5]

    McIntosh, Teo Susnjak, Tong Liu, Paul Watters, Malka N

    Timothy R. McIntosh, Teo Susnjak, Tong Liu, Paul Watters, Malka N. Halgamuge, The inadequacy of reinforcement learning from human feedback- radicalizing large language models via semantic vulnerabilities, 2024

  2. [6]

    Nuraini Sulaiman, Farizal Hamzah, Evaluation of transfer learning and adaptability in large language models with the glue benchmark, 2024

  3. [8]

    Amanda Kau, Xuzeng He, Aishwarya Nambissan, Aland Astudillo, Hui Yin, Amir Aryani, Combining knowledge graphs and large language models, 2024, arXiv:2407.06564 [cs]

  4. [9]

    Huy Quoc To, Ming Liu, Guangyan Huang, Towards efficient large language models for scientific text: A review, 2024, arXiv:2408.10729 [cs]

  5. [10]

    Muhammad Usman Hadi, Qasem Al Tashi, Rizwan Qureshi, Abbas Shah, Amgad Muneer, Muhammad Irfan, Anas Zafar, Muhammad Bilal Shaikh, Naveed Akhtar, Jia Wu, Seyedali Mirjalili, A survey on large language models: Applications, challenges, limitations, and practical usage, 2023

  6. [11]

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al., A survey of large language models, 2023, arXiv preprint arXiv:2303.18223, 1(2)

  7. [12]

    Jean Lee, Nicholas Stevens, Soyeon Caren Han, Minseok Song, A survey of large language models in finance (FinLLMs), 2024, arXiv:2402.02315 [cs, q-fin]

  8. [13]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova, BERT: Pre-training of deep bidirectional transformers for language understanding, 2018

  9. [14]

    Enkelejda Kasneci, Kathrin Sessler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Jürgen Pfeffer, Oleksandra Poquet, Michael Sail...

  10. [15]

    Quanjun Zhang, Chunrong Fang, Yang Xie, YuXiang Ma, Weisong Sun, Yun Yang, Zhenyu Chen, A systematic literature review on large language models for automated program repair, 2024, arXiv:2405.01466 [cs]

  11. [16]

    Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, et al., DeepSeek LLM: Scaling open-source language models with longtermism, 2024, arXiv preprint arXiv:2401.02954

  12. [17]

    Li, et al., Deepseek-coder: When the large language model meets programming–the rise of code intelligence, 2024, arXiv preprint arXiv:2401.14196

    Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Yu Wu, Y.K. Li, et al., Deepseek-coder: When the large language model meets programming–the rise of code intelligence, 2024, arXiv preprint arXiv:2401.14196

  13. [18]

    Damai Dai, Chengqi Deng, Chenggang Zhao, R.X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Yu Wu, et al., DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models, 2024, arXiv preprint arXiv:2401.06066

  14. [19]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al., The llama 3 herd of models, 2024, arXiv preprint arXiv:2407.21783

  15. [20]

    Qinggang Zhang, Shengyuan Chen, Yuanchen Bei, Zheng Yuan, Huachi Zhou, Zijin Hong, Junnan Dong, Hao Chen, Yi Chang, Xiao Huang, A survey of graph retrieval-augmented generation for customized large language models, 2025, arXiv preprint arXiv:2501.13958

  16. [21]

    Han Zhang, Langshi Zhou, Hanfang Yang, Learning to retrieve and reason on knowledge graph through active self-reflection, 2025, arXiv preprint arXiv: 2502.14932

  17. [22]

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al., A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions, ACM Trans. Inf. Syst. 43 (2) ...

  18. [23]

    Tianyang Zhong, Zhengliang Liu, Yi Pan, Yutong Zhang, Yifan Zhou, Shizhe Liang, Zihao Wu, Yanjun Lyu, Peng Shu, Xiaowei Yu, et al., Evaluation of OpenAI o1: Opportunities and challenges of AGI, 2024, arXiv preprint arXiv: 2409.18486

  19. [24]

    Erik Keanius, Domain adaptation of LLMs: A study of content generation, RAG, and fine-tuning, 2024

  20. [25]

    Chengrui Wang, Qingqing Long, Meng Xiao, Xunxin Cai, Chengjun Wu, Zhen Meng, Xuezhi Wang, Yuanchun Zhou, BioRAG: A RAG-LLM framework for biological question reasoning, 2024, arXiv:2408.01107 [cs]

  21. [26]

    Rawat, Recent advances in generative AI and large language models: Current status, challenges, and perspectives, 2024, arXiv:2407.14962 [cs]

    Desta Haileselassie Hagos, Rick Battle, Danda B. Rawat, Recent advances in generative AI and large language models: Current status, challenges, and perspectives, 2024, arXiv:2407.14962 [cs]

  22. [27]

    Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, et al., Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context fi...

  23. [28]

    Tyna Eloundou, Sam Manning, Pamela Mishkin, Daniel Rock, Gpts are gpts: An early look at the labor market impact potential of large language models, 2023

  24. [29]

    Xiaonan Li, Changtai Zhu, Linyang Li, Zhangyue Yin, Tianxiang Sun, Xipeng Qiu, LLatrieval: LLM-verified retrieval for verifiable generation, 2024, arXiv: 2311.07838 [cs]

  25. [30]

    Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, Luke Zettlemoyer, Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension, 2019, arXiv preprint arXiv:1910.13461

  26. [31]

    Shuming Ma, Li Dong, Shaohan Huang, Dongdong Zhang, Alexandre Muzio, Saksham Singhal, Hany Hassan Awadalla, Xia Song, Furu Wei, Deltalm: Encoder–decoder pre-training for language generation and translation by aug- menting pretrained multilingual encoders, 2021, arXiv preprint ...

  27. [32]

    Pranjal Kumar, Large language models (llms): survey, technical frameworks, and future challenges, Artif. Intell. Rev. 57 (10) (2024) 260

  28. [33]

    Mohaimenul Azam Khan Raiaan, Md. Saddam Hossain Mukta, Kaniz Fatema, Nur Mohammad Fahad, Sadman Sakib, Most Marufatul Jannat Mim, Jubaer Ahmad, Mohammed Eunus Ali, Sami Azam, A review on large language models: Architectures, applications, taxonomies, open issues and challenges...

  29. [34]

    Kitaev, et al., Reformer: The efficient transformer, 2020

    N. Kitaev, et al., Reformer: The efficient transformer, 2020

  30. [35]

    Siyu Ding, Junyuan Shang, Shuohuan Wang, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang, Ernie-doc: A retrospective long-document modeling transformer, 2020, arXiv preprint arXiv:2012.15688

  31. [36]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, Haofen Wang, Retrieval-augmented generation for large language models: A survey, 2024, arXiv:2312.10997 [cs]

  32. [37]

    Nourhan Ibrahim, Samar Aboulela, Ahmed Ibrahim, Rasha Kashef, A survey on augmenting knowledge graphs (kgs) with large language models (llms): models, evaluation metrics, benchmarks, and challenges, Discov. Artif. Intell. 4 (1) (2024) 76

  33. [38]

    Amin Beheshti, Empowering generative ai with knowledge base 4.0: Towards linking analytical, cognitive, and generative intelligence, 2023

  34. [39]

    Andrea Matarazzo, Riccardo Torlone, A survey on large language models with some insights on their capabilities and limitations, 2025, arXiv preprint arXiv: 2501.04040

  35. [40]

    Simeone, Focus agent: LLM-powered virtual focus group, 2024, arXiv:2409.01907 [cs]

    Taiyu Zhang, Xuesong Zhang, Robbe Cools, Adalberto L. Simeone, Focus agent: LLM-powered virtual focus group, 2024, arXiv:2409.01907 [cs]

  36. [41]

    Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, Qing Li, A survey on RAG meeting LLMs: Towards retrieval-augmented large language models, 2024, arXiv:2405.06211 [cs]

  37. [43]

    Yingqing He, Zhaoyang Liu, Jingye Chen, Zeyue Tian, Hongyu Liu, Xiaowei Chi, Runtao Liu, Ruibin Yuan, Yazhou Xing, Wenhai Wang, Jifeng Dai, Yong Zhang, Wei Xue, Qifeng Liu, Yike Guo, Qifeng Chen, LLMs meet multimodal generation and editing: A survey, 2024, arXiv:2405.19334 [cs]

  38. [44]

    Ruiyao Xu, Kaize Ding, Large language models for anomaly and out-of- distribution detection: A survey, 2024, arXiv:2409.01980 [cs]

  39. [45]

    Linhao Luo, Yuan-Fang Li, Gholamreza Haffari, Shirui Pan, Reasoning on graphs: Faithful and interpretable large language model reasoning, 2024, arXiv: 2310.01061 [cs]

  40. [46]

    Sander Schulhoff, Michael Ilie, Nishant Balepur, Konstantine Kahadze, Amanda Liu, Chenglei Si, Yinheng Li, Aayush Gupta, HyoJung Han, Sevien Schulhoff, et al., The prompt report: A systematic survey of prompting techniques, 2024, arXiv preprint arXiv:2406.06608

  41. [47]

    Richard Fang, Rohan Bindu, Akul Gupta, Daniel Kang, LLM agents can autonomously exploit one-day vulnerabilities, 2024, arXiv:2404.08144 [cs]

  42. [48]

    Marissa Radensky, Daniel S. Weld, Joseph Chee Chang, Pao Siangliulue, Jonathan Bragg, Let’s get to the point: Llm-supported planning, drafting, and revising of research-paper blog posts, 2024, arXiv preprint arXiv:2406.10370

  43. [49]

    Xiaofei Sun, Xiaoya Li, Shengyu Zhang, Shuhe Wang, Fei Wu, Jiwei Li, Tianwei Zhang, Guoyin Wang, Sentiment analysis through llm negotiations, 2023, arXiv preprint arXiv:2311.01876

  44. [50]

    Rajvardhan Patil, Venkat Gudivada, A review of current trends, techniques, and challenges in large language models (llms), Appl. Sci. 14 (5) (2024) 2074

  45. [51]

    Mulvey, H

    Yuqi Nie, Yaxuan Kong, Xiaowen Dong, John M. Mulvey, H. Vincent Poor, Qingsong Wen, Stefan Zohren, A survey of large language models for finan- cial applications: Progress, prospects and challenges, 2024, arXiv:2406.11903 [q-fin]

  46. [52]

    Rui Yang, Edison Marrese-Taylor, Yuhe Ke, Lechao Cheng, Qingyu Chen, Irene Li, Integrating umls knowledge into large language models for medical question answering, 2023

  47. [53]

    11, MDPI, 2024, p

    Zabir Al Nazi, Wei Peng, Large language models in healthcare and medical domain: A review, in: Informatics, Vol. 11, MDPI, 2024, p. 57

  48. [54]

    29 (8) (2023) 1930–1940

    Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, Daniel Shu Wei Ting, Large language models in medicine, Nature Med. 29 (8) (2023) 1930–1940

  49. [55]

    Frank Fagan, A view of how language models will transform law, Tenn. Law Rev. 92 (2024) 1

  50. [56]

    Yinheng Li, Shaofei Wang, Han Ding, Hang Chen, Large language models in finance: A survey, in: Proceedings of the Fourth ACM International Conference on AI in Finance, 2023, pp. 374–382

  51. [57]

    Yanjun Gao, Ruizhe Li, Emma Croxford, John Caskey, Brian W. Patterson, Matthew Churpek, Timothy Miller, Dmitriy Dligach, Majid Afshar, Leveraging medical knowledge graphs into large language models for diagnosis prediction: Design and application study, JMIR AI 4 (2025) e58670

  52. [58]

    Huanghai Liu, Quzhe Huang, Qingjing Chen, Yiran Hu, Jiayu Ma, Yun Liu, Weixing Shen, Yansong Feng, Jurex-4e: Juridical expert-annotated four-element knowledge base for legal reasoning, 2025, arXiv preprint arXiv:2502.17166

  53. [59]

    Jean Lee, Nicholas Stevens, Soyeon Caren Han, Large language models in finance (finllms), Neural Comput. Appl. (2025) 1–15

  54. [60]

    Zibin Zheng, Kaiwen Ning, Qingyuan Zhong, Jiachi Chen, Wenqing Chen, Lianghong Guo, Weicheng Wang, Yanlin Wang, Towards an understanding of large language models in software engineering tasks, Empir. Softw. Eng. 30 (2) (2025) 50

  55. [61]

    Jingzhi Gong, Vardan Voskanyan, Paul Brookes, Fan Wu, Wei Jie, Jie Xu, Rafail Giavrimis, Mike Basios, Leslie Kanthan, Zheng Wang, Language models for code optimization: Survey, challenges and future directions, 2025, arXiv preprint arXiv:2501.01277

  56. [62]

    Yu Philip, Multi- modal large language models: A survey, in: 2023 IEEE International Conference on Big Data, BigData, IEEE, 2023, pp

    Jiayang Wu, Wensheng Gan, Zefeng Chen, Shicheng Wan, S. Yu Philip, Multi- modal large language models: A survey, in: 2023 IEEE International Conference on Big Data, BigData, IEEE, 2023, pp. 2247–2256

  57. [63]

    Dawei Huang, Chuan Yan, Qing Li, Xiaojiang Peng, From large language models to large multimodal models: A literature review, Appl. Sci. 14 (12) (2024) 5068

  58. [64]

    Zijing Liang, Yanjie Xu, Yifan Hong, Penghui Shang, Qi Wang, Qiang Fu, Ke Liu, A survey of multimodel large language models, in: Proceedings of the 3rd International Conference on Computer, Artificial Intelligence and Control Engineering, 2024, pp. 405–409

  59. [65]

    Karthik Meduri, Hari Gonaygunta, Geeta Sandeep Nadella, Priyanka Pramod Pawar, Deepak Kumar, Adaptive intelligence: Gpt-powered language models for dynamic responses to emerging healthcare challenges, IJARCCE 13 (2024) 104–109

  60. [66]

    Iker García-Ferrero, Rodrigo Agerri, Aitziber Atutxa Salazar, Elena Cabrio, Iker de la Iglesia, Alberto Lavelli, Bernardo Magnini, Benjamin Molinet, Johana Ramirez-Romero, German Rigau, et al., Medical mt5: an open-source multilingual text-to-text llm for the medical domain, 2...

  61. [67]

    Angela Fan, Beliz Gokkaya, Mark Harman, Mitya Lyubarskiy, Shubho Sengupta, Shin Yoo, Jie M. Zhang, Large language models for software engineering: Survey and open problems, in: 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering, ICSE...

  62. [68]

    Jiaqi Wang, Enze Shi, Huawen Hu, Chong Ma, Yiheng Liu, Xuhui Wang, Yincheng Yao, Xuan Liu, Bao Ge, Shu Zhang, Large language models for robotics: Opportunities, challenges, and perspectives, J. Autom. Intell. (2024)

  63. [69]

    2837–2841

    Sonal Sannigrahi, Thiago Fraga-Silva, Youssef Oualil, Christophe Van Gysel, Synthetic query generation using large language models for virtual assistants, in: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2024,...

  64. [70]

    Yetisen, Savas Tasoglu, Aydogan Ozcan, Large language model-based chatbots in higher education, Adv

    Defne Yigci, Merve Eryilmaz, Ail K. Yetisen, Savas Tasoglu, Aydogan Ozcan, Large language model-based chatbots in higher education, Adv. Intell. Syst. (2024) 2400429

  65. [71]

    Ting-Chi Chang, Yu-Jou Chen, Sheng Hung, Ning-Hsuan Chang, Chih-Hao Ku, Szu-Yin Lin, Shih-Yi Chien, A review on shaping chatbot personalities via large language models, 2025

  66. [72]

    Devon Myers, Rami Mohawesh, Venkata Ishwarya Chellaboina, Anantha Lak- shmi Sathvik, Praveen Venkatesh, Yi-Hui Ho, Hanna Henshaw, Muna Al- hawawreh, David Berdik, Yaser Jararweh, Foundation and large language models: fundamentals, challenges, opportunities, and social impacts,...

  67. [73]

    Knowledge-Based Systems 318 (2025) 113503 23 W

    Anqi Wang, Zhizhuo Yin, Yulu Hu, Yuanyuan Mao, Pan Hui, Exploring the potential of large language models in artistic creation: Collaboration and reflection on creative programming, 2024, arXiv preprint arXiv:2402.09750. Knowledge-Based Systems 318 (2025) 113503 23 W. Yang et al

  68. [74]

    Hung-Fu Chang, Tong Li, A framework for collaborating a large language model tool in brainstorming for triggering creative thoughts, Think. Ski. Creativity (2025) 101755

  69. [75]

    Transform

    Robin Qiu, Large language models: from entertainment to solutions, Digit. Transform. Soc. 3 (2) (2024) 125–126

  70. [76]

    Rose, John H

    Karthik Soman, Peter W. Rose, John H. Morris, Rabia E. Akbas, Brett Smith, Braian Peetoom, Catalina Villouta-Reyes, Gabriel Cerono, Yongmei Shi, Angela Rizk-Jackson, Sharat Israni, Charlotte A. Nelson, Sui Huang, Sergio E. Baranzini, Biomedical knowledge graph-optimized prompt...

  71. [77]

    Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, Rya...

  72. [78]

    Xin Cheng, Di Luo, Xiuying Chen, Lemao Liu, Dongyan Zhao, Rui Yan, Lift yourself up: Retrieval-augmented text generation with self memory, 2023

  73. [79]

    YunHe Su, Zhengyang Lu, Junhui Liu, Ke Pang, Haoran Dai, Sa Liu Yuxin Jia, Lujia Ge, Jing-min Yang, Applications of large models in medicine, 2025, arXiv preprint arXiv:2502.17132

  74. [80]

    Hongzhi Zhang, M. Omair Shafiq, Triple-aware reasoning: A retrieval- augmented generationapproach for enhancing question-answering tasks with- knowledge graphs and large language models, 2024, https://caiac.pubpub.org/ pub/bytcy6lo

  75. [81]

    Shengjie Ma, Chengjin Xu, Xuhui Jiang, Muzhi Li, Huaren Qu, Jian Guo, Think- on-graph 2.0: Deep and interpretable large language model reasoning with knowledge graph-guided retrieval, 2024, arXiv:2407.10805 [cs]

  76. [83]

    14511 [cs, math, stat]

    Xinyang Hu, Fengzhuo Zhang, Siyu Chen, Zhuoran Yang, Unveiling the sta- tistical foundations of chain-of-thought prompting methods, 2024, arXiv:2408. 14511 [cs, math, stat]

  77. [84]

    Alaa Abd-alrazaq, Rawan AlSaad, Dari Alhuwail, Arfan Ahmed, Padraig Mark Healy, Syed Latifi, Sarah Aziz, Rafat Damseh, Sadam Alabed Alrazak, Javaid Sheikh, Large language models in medical education: Opportunities, challenges, and future directions, JMIR Med. Educ. 9 (2023) e48291

  78. [86]

    Tailin Liang, John Glossner, Lei Wang, Shaobo Shi, Xiaotong Zhang, Pruning and quantization for deep neural network acceleration: A survey, 2021

  79. [87]

    Xiaokai Wei, Shen Wang, Dejiao Zhang, Parminder Bhatia, Andrew Arnold, Knowledge enhanced pretrained language models: A compreshensive survey, 2021, arXiv:2110.08455 [cs]

  80. [88]

    Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, Yi Zhang, Sparks of artificial general intelligence: Early experiments with gpt-4, 2023

  81. [89]

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al., Training compute-optimal large language models, 2022, arXiv preprint arXiv:2203.15556

  82. [90]

    AAAI Conf

    Jiawei Chen, Hongyu Lin, Xianpei Han, Le Sun, Benchmarking large language models in retrieval-augmented generation, Proc. AAAI Conf. Artif. Intell. 38 (16) (2024) 17754–17762

  83. [91]

    Hernandez, Mythreye Venkatesan, Paul Wang, Jason H

    Nicholas Matsumoto, Jay Moran, Hyunjun Choi, Miguel E. Hernandez, Mythreye Venkatesan, Paul Wang, Jason H. Moore, KRAGEN: a knowledge graph- enhanced RAG framework for biomedical problem solving using large language models, Bioinformatics 40 (6) (2024) btae353

  84. [92]

    Educ.: Artif

    Bodong Chen, Xinran Zhu, Fernando Díaz Del Castillo H., Integrating generative AI in knowledge building, Comput. Educ.: Artif. Intell. 5 (2023) 100184, MAG ID: 4388599367

  85. [93]

    Ahmed Menshawy, Zeeshan Nawaz, Mahmoud Fahmy, Navigating challenges and technical debt in large language models deployment, in: Proceedings of the 4th Workshop on Machine Learning and Systems, 2024, pp. 192–199

  86. [94]

    Xin Su, Tiep Le, Steven Bethard, Phillip Howard, Semi-structured chain-of- thought: Integrating multiple sources of knowledge for improved language model reasoning, 2024, arXiv:2311.08505 [cs]

  87. [95]

    Tilmann Bruckhaus, Rag does not work for enterprises, 2024

  88. [96]

    Rush, HuggingFace’s transformers: State-of-the-art natural language processing, 2019

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement De- langue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Ma...

  89. [97]

    Miao Zheng, Hao Liang, Fan Yang, Haoze Sun, Tianpeng Li, Lingchu Xiong, Yan Zhang, Youzhen Wu, Kun Li, Yanjun Shen, Mingan Lin, Tao Zhang, Guosheng Dong, Yujing Qiao, Kun Fang, Weipeng Chen, Bin Cui, Wentao Zhang, Zenan Zhou, Pas: Data-efficient plug-and-play prompt augmentati...

  90. [98]

    Wentao Zhang, Lingxuan Zhao, Haochong Xia, Shuo Sun, Jiaze Sun, Molei Qin, Xinyi Li, Yuqing Zhao, Yilei Zhao, Xinyu Cai, Longtao Zheng, Xinrun Wang, Bo An, A multimodal foundation agent for financial trading: Tool-augmented, diversified, and generalist, 2024, arXiv:2402.18485 [q-fin]

  91. [99]

    Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi- Yu, Yiming Yang, Jamie Callan, Graham Neubig, Active retrieval augmented generation, 2023, arXiv:2305.06983 [cs]

    Zhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi- Yu, Yiming Yang, Jamie Callan, Graham Neubig, Active retrieval augmented generation, 2023, arXiv:2305.06983 [cs]

  92. [100]

    Yang, Xiang Chen, Tony Q.S

    Xijun Wang, Dongshan Ye, Chenyuan Feng, Howard H. Yang, Xiang Chen, Tony Q.S. Quek, Trustworthy image semantic communication with GenAI: Explainablity, controllability, and efficiency, 2024, ARXIV_ID: 2408.03806 S2ID: 3e5473ecb44e13e6bb08b477623ab39da551943b

  93. [101]

    Fiona Fui-Hoon Nah, Ruilin Zheng, Jingyuan Cai, Keng Siau, Langtao Chen, Generative AI and ChatGPT: Applications, challenges, and AI-human collaboration, J. Inf. Technol. Case Appl. Res. 25 (3) (2023) 277–304

  94. [104]

    Rashid, Anisa Rula, Lukas Schmelzeisen, Juan Sequeda, Steffen Staab, Antoine Zimmermann, Knowledge graphs, ACM Comput

    Aidan Hogan, Eva Blomqvist, Michael Cochez, Claudia d’Amato, Gerard de Melo, Claudio Gutierrez, José Emilio Labra Gayo, Sabrina Kirrane, Sebastian Neumaier, Axel Polleres, Roberto Navigli, Axel-Cyrille Ngonga Ngomo, Sab- bir M. Rashid, Anisa Rula, Lukas Schmelzeisen, Juan Sequ...

  95. [105]

    Minbyul Jeong, Jiwoong Sohn, Mujeen Sung, Jaewoo Kang, Improving medical reasoning through retrieval and self-reflection with retrieval-augmented large language models, 2024, arXiv:2401.15269 [cs]

  96. [106]

    5 (4) (2023) 46

    Emma Yann Zhang, Adrian David Cheok, Zhigeng Pan, Jun Cai, Ying Yan, From turing to transformers: A comprehensive review and tutorial on the evolution and applications of generative transformer models, Sci. 5 (4) (2023) 46

  97. [107]

    Peizhong Gao, Ao Xie, Shaoguang Mao, Wenshan Wu, Yan Xia, Haipeng Mi, Furu Wei, Meta reasoning for large language models, 2024, arXiv:2406.11698 [cs]

  98. [108]

    Kasnesis, Augmentation of large language model capabilities with knowledge graphs, 2024

    Dr. Kasnesis, Augmentation of large language model capabilities with knowledge graphs, 2024

  99. [110]

    Sijia Chen, Baochun Li, Di Niu, Boosting of thoughts: Trial-and-error problem solving with large language models, 2024, arXiv:2402.11140 [cs]

  100. [111]

    Alonso-Moral, How to build self-explaining fuzzy systems: From interpretability to explainability [AI-eXplained], 2024, S2ID: 1d5127686d01fe1e66805c10f19284f865484632

    Ilia Stepin, Muhammad Suffian, Alejandro Catalá, J. Alonso-Moral, How to build self-explaining fuzzy systems: From interpretability to explainability [AI-eXplained], 2024, S2ID: 1d5127686d01fe1e66805c10f19284f865484632

  101. [112]

    Gonzalez, Bin Cui, Buffer of thoughts: Thought-augmented reasoning with large language models, 2024, arXiv:2406.04271 [cs]

    Ling Yang, Zhaochen Yu, Tianjun Zhang, Shiyi Cao, Minkai Xu, Wentao Zhang, Joseph E. Gonzalez, Bin Cui, Buffer of thoughts: Thought-augmented reasoning with large language models, 2024, arXiv:2406.04271 [cs]

  102. [114]

    Lukas Bahr, Christoph Wehner, Judith Wewerka, José Bittencourt, Ute Schmid, Rüdiger Daub, Daub knowledge graph enhanced retrieval-augmented generation for failure mode and effects analysis, 2024, arXiv:2406.18114 [cs]

  103. [115]

    Garima Agrawal, Tharindu Kumarage, Zeyad Alghamdi, Huan Liu, Can knowl- edge graphs reduce hallucinations in LLMs? : A survey, 2024, arXiv:2311.07914 [cs]

  104. [116]

    Hasan Abu-Rasheed, Christian Weber, Madjid Fathi, Knowledge graphs as context sources for LLM-based explanations of learning recommendations, 2024, arXiv:2403.03008 [cs]

  105. [117]

    Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, Xindong Wu, Unifying large language models and knowledge graphs: A roadmap, IEEE Trans. Knowl. Data Eng. 36 (7) (2024) 3580–3599, arXiv:2306.08302 [cs]

  106. [118]

    Julien Delile, Srayanta Mukherjee, Anton Van Pamel, Leonid Zhukov, Graph- based retriever captures the long tail of biomedical knowledge, 2024, arXiv: 2402.12352 [cs]

  107. [119]

    Paulo Finardi, Leonardo Avila, Rodrigo Castaldoni, Pedro Gengo, Celio Larcher, Marcos Piau, Pablo Costa, Vinicius Caridá, The chronicles of RAG: The retriever, the chunk and the generator, 2024, arXiv:2401.07883 [cs]

  108. [120]

    Neural Inf

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al., Retrieval-augmented generation for knowledge-intensive nlp tasks, Adv. Neural Inf. Process. Syst. 33 (2020) 9459–9474

  109. [121]

    Junde Wu, Jiayuan Zhu, Yunli Qi, Medical graph RAG: Towards safe medical large language model via graph retrieval-augmented generation, 2024, arXiv: 2408.04187 [cs]

  110. [122]

    03271 [cs]

    Yu Wang, Shiwan Zhao, Zhihu Wang, Heyuan Huang, Ming Fan, Yubo Zhang, Zhixing Wang, Haijun Wang, Ting Liu, Strategic chain-of-thought: Guiding accurate reasoning in LLMs through strategy elicitation, 2024, arXiv:2409. 03271 [cs]

  111. [123]

    Knowledge-Based Systems 318 (2025) 113503 24 W

    Jinyuan Fang, Zaiqiao Meng, Craig Macdonald, TRACE the evidence: Construct- ing knowledge-grounded reasoning chains for retrieval-augmented generation, 2024, arXiv:2406.11460 [cs]. Knowledge-Based Systems 318 (2025) 113503 24 W. Yang et al

  112. [124]

    (Accessed 25 October 2024)

    Davit Janezashvili, Rag at large enterprises, 2024, https://modulai.io/blog/rag- at-large-enterprises/. (Accessed 25 October 2024)

  113. [125]

    Haoyu Wang, Ruirui Li, Haoming Jiang, Jinjin Tian, Zhengyang Wang, Chen Luo, Xianfeng Tang, Monica Cheng, Tuo Zhao, Jing Gao, BlendFilter: Advancing retrieval-augmented large language models via query generation blending and knowledge filtering, 2024, arXiv:2402.11129 [cs]

  114. [126]

    Yujia Zhou, Zheng Liu, Jiajie Jin, Jian-Yun Nie, Zhicheng Dou, Metacognitive retrieval-augmented large language models, 2024, arXiv:2402.11626 [cs]

  115. [128]

    Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, Jonathan Larson, From local to global: A graph rag approach to query-focused summarization, 2024, arXiv preprint arXiv:2404.16130

  116. [129]

    38, 2024, pp

    Zhenyu Li, Sunqi Fan, Yu Gu, Xiuxing Li, Zhichao Duan, Bowen Dong, Ning Liu, Jianyong Wang, Flexkbqa: A flexible llm-powered framework for few-shot knowledge base question answering, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, 2024, pp. 18608–18616

  117. [130]

    Ni, Heung-Yeung Shum, Jian Guo, Think-on-graph: Deep and responsible reasoning of large language model on knowledge graph, 2023, arXiv preprint arXiv:2307.07697

    Jiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang, Chen Lin, Yeyun Gong, Lionel M. Ni, Heung-Yeung Shum, Jian Guo, Think-on-graph: Deep and responsible reasoning of large language model on knowledge graph, 2023, arXiv preprint arXiv:2307.07697

  118. [131]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao, ReAct: Synergizing reasoning and acting in language models, 2023, arXiv:2210.03629 [cs]

  119. [132]

    Shengjie Ma, Chengjin Xu, Xuhui Jiang, Muzhi Li, Huaren Qu, Jian Guo, Think- on-graph 2.0: Deep and interpretable large language model reasoning with knowledge graph-guided retrieval, 2024, arXiv e-prints, pages arXiv–2407

  120. [133]

    Yifei Zhang, Xintao Wang, Jiaqing Liang, Sirui Xia, Lida Chen, Yanghua Xiao, Chain-of-knowledge: Integrating knowledge reasoning into large language models by learning from knowledge graphs, 2024, arXiv preprint arXiv:2407. 00653

  121. [134]

    Weijian Xie, Xuefeng Liang, Yuhui Liu, Kaihua Ni, Hong Cheng, Zetian Hu, Weknow-rag: An adaptive approach for retrieval-augmented generation integrating web search and knowledge graphs, 2024, arXiv preprint arXiv: 2408.07611

  122. [135]

    Jinhao Jiang, Kun Zhou, Wayne Xin Zhao, Yang Song, Chen Zhu, Hengshu Zhu, Ji-Rong Wen, Kg-agent: An efficient autonomous agent framework for complex reasoning over knowledge graph, 2024, arXiv preprint arXiv:2402.11163

  123. [136]

    Qitan Lv, Jie Wang, Hanzhu Chen, Bin Li, Yongdong Zhang, Feng Wu, Coarse-to- fine highlighting: Reducing knowledge hallucination in large language models, 2024, arXiv preprint arXiv:2410.15116

  124. [137]

    Lei Liang, Mengshu Sun, Zhengke Gui, Zhongshu Zhu, Zhouyu Jiang, Ling Zhong, Yuan Qu, Peilong Zhao, Zhongpu Bo, Jin Yang, et al., Kag: Boosting llms in professional domains via knowledge augmented generation, 2024, arXiv preprint arXiv:2409.13731

  125. [138]

    Hanzhu Chen, Xu Shen, Qitan Lv, Jie Wang, Xiaoqi Ni, Jieping Ye, Sac-kg: Exploiting large language models as skilled automatic constructors for domain knowledge graphs, 2024, arXiv preprint arXiv:2410.02811

  126. [139]

    Jiejun Tan, Zhicheng Dou, Yutao Zhu, Peidong Guo, Kun Fang, Ji-Rong Wen, Small models, big insights: Leveraging slim proxy models to decide when and what to retrieve for LLMs, 2024, arXiv:2402.12052 [cs]

  127. [140]

    Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, Weizhu Chen, Enhancing retrieval-augmented large language models with iterative retrieval-generation synergy, 2023, arXiv:2305.15294 [cs]

  128. [141]

    Yanming Liu, Xinyue Peng, Xuhong Zhang, Weihao Liu, Jianwei Yin, Jiannan Cao, Tianyu Du, RA-ISF: Learning to answer and understand from retrieval augmentation via iterative self-feedback, 2024, arXiv:2403.06840 [cs]

  129. [142]

    17723 [cs]

    Zhentao Xu, Mark Jerome Cruz, Matthew Guevara, Tie Wang, Manasi Deshpande, Xiaofeng Wang, Zheng Li, Retrieval-augmented generation with knowledge graphs for customer service question answering, 2024, arXiv:2404. 17723 [cs]

  130. [143]

    Yang, Anton Tsitsulin, Don’t forget to connect! improving RAG with graph-based reranking, 2024, arXiv: 2405.18414 [cs]

    Jialin Dong, Bahare Fatemi, Bryan Perozzi, Lin F. Yang, Anton Tsitsulin, Don’t forget to connect! improving RAG with graph-based reranking, 2024, arXiv: 2405.18414 [cs]

  131. [144]

    Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Minlie Huang, Nan Duan, Weizhu Chen, ToRA: A tool-integrated reasoning agent for mathematical problem solving, 2024, arXiv:2309.17452 [cs]

  132. [145]

    Chawla, Olaf Wiest, Xiangliang Zhang, Large language model based multi-agents: A survey of progress and challenges, 2024

    Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V. Chawla, Olaf Wiest, Xiangliang Zhang, Large language model based multi-agents: A survey of progress and challenges, 2024

  133. [146]

    Karabacak, K

    M. Karabacak, K. Margetis, Embracing large language models for medical applications: Opportunities and challenges, Cureus 15 (5) (2023) e39305

  134. [147]

    Xiaoxi Li, Jiajie Jin, Yujia Zhou, Yuyao Zhang, Peitian Zhang, Yutao Zhu, Zhicheng Dou, From matching to generation: A survey on generative information retrieval, 2024, arXiv:2404.14851 [cs]

  135. [148]

    Zakka, R

    C. Zakka, R. Shad, A. Chaurasia, A.R. Dalal, J.L. Kim, M. Moor, R. Fong, C. Phillips, K. Alexander, E. Ashley, J. Boyd, K. Boyd, K. Hirsch, C. Langlotz, R. Lee, J. Melia, J. Nelson, K. Sallam, S. Tullis, M.A. Vogelsong, W. Hiesinger, Almanac - retrieval-augmented language mode...

  136. [149]

    Taojun Hu, Xiao-Hua Zhou, Unveiling llm evaluation focused on metrics: Challenges and solutions, 2024, arXiv preprint arXiv:2404.09135

  137. [150]

    Masoomali Fatehkia, Ji Kim Lucas, Sanjay Chawla, T-rag: lessons from the llm trenches, 2024, arXiv preprint arXiv:2402.07483

  138. [151]

    graphrag: A systematic evaluation and key insights, 2025, arXiv preprint arXiv:2502.11371

    Haoyu Han, Harry Shomer, Yu Wang, Yongjia Lei, Kai Guo, Zhigang Hua, Bo Long, Hui Liu, Jiliang Tang, Rag vs. graphrag: A systematic evaluation and key insights, 2025, arXiv preprint arXiv:2502.11371

  139. [152]

    Yinhong Liu, Zhijiang Guo, Tianya Liang, Ehsan Shareghi, Ivan Vulić, Nigel Collier, Aligning with logic: Measuring, evaluating and improving logical consistency in large language models, 2024, arXiv preprint arXiv:2410.02205

  140. [153]

    Hao Yu, Aoran Gan, Kai Zhang, Shiwei Tong, Qi Liu, Zhaofeng Liu, Evaluation of retrieval-augmented generation: A survey, in: CCF Conference on Big Data, Springer, 2024, pp. 102–120

  141. [154]

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al., A survey on evaluation of large language models, ACM Trans. Intell. Syst. Technol. 15 (3) (2024) 1–45

  142. [155]

    Sandeep Kumar, Arun Solanki, Rouge-ss: A new rouge variant for evaluation of text summarization, 2023, Authorea Preprints

  143. [156]

    Chuangtao Ma, Sriom Chakrabarti, Arijit Khan, Bálint Molnár, Knowledge graph-based retrieval-augmented generation for schema matching, 2025, arXiv preprint arXiv:2501.08686

  144. [157]

    Joel Ruben Antony Moniz, Soundarya Krishnan, Melis Ozyildirim, Prathamesh Saraf, Halim Cagri Ates, Yuan Zhang, Hong Yu, Realm: reference resolution as language modeling, 2024, arXiv preprint arXiv:2403.20329

  145. [158]

    Yinghao Zhu, Changyu Ren, Shiyun Xie, Shukai Liu, Hangyuan Ji, Zixiang Wang, Tao Sun, Long He, Zhoujun Li, Xi Zhu, et al., Realm: Rag-driven en- hancement of multimodal electronic health records analysis via large language models, 2024, arXiv preprint arXiv:2402.07016

  146. [159]

    Benjamin Reichman, Larry Heck, Dense passage retrieval: Is it retrieving? 2024, arXiv preprint arXiv:2402.11035

  147. [160]

    Xingyu Xiong, Mingliang Zheng, Merging mixture of experts and retrieval augmented generation for enhanced information retrieval and reasoning, 2024

  148. [161]

    Xinping Zhao, Yan Zhong, Zetian Sun, Xinshuo Hu, Zhenyu Liu, Dongfang Li, Baotian Hu, Min Zhang, Funnelrag: A coarse-to-fine progressive retrieval paradigm for rag, 2024, arXiv preprint arXiv:2410.10293

  149. [162]

    Majid Afshar, Yanjun Gao, Deepak Gupta, Emma Croxford, Dina Demner- Fushman, On the role of the umls in supporting diagnosis generation proposed by large language models, J. Biomed. Inform. 157 (2024) 104707

  150. [163]

    32 (suppl_1) (2004) D267–D270

    Olivier Bodenreider, The unified medical language system (umls): integrating biomedical terminology, Nucleic Acids Res. 32 (suppl_1) (2004) D267–D270

  151. [164]

    Tristan Coignion, Clément Quinton, Romain Rouvoy, A performance study of llm-generated code on leetcode, 2024

  152. [165]

    Junjielong Xu, Ziang Cui, Yuan Zhao, Xu Zhang, Shilin He, Pinjia He, Liqun Li, Yu Kang, Qingwei Lin, Yingnong Dang, Saravan Rajmohan, Dongmei Zhang, Unilog: Automatic logging via llm and in-context learning, 2024

  153. [166]

    Danie Smit, Hanlie Smuts, Paul Louw, Julia Pielmeier, Christina Eidelloth, The impact of github copilot on developer productivity from a software engineering body of knowledge perspective, 2024

  154. [167]

    Garrido-Merchan, ChatGPT is not all you need

    Roberto Gozalo-Brizuela, Eduardo C. Garrido-Merchan, ChatGPT is not all you need. A state of the art review of large generative AI models, 2023

  155. [168]

    Sen Yang, Xin Li, Leyang Cui, Lidong Bing, Wai Lam, Neuro-symbolic in- tegration brings causal and reliable reasoning proofs, 2023, arXiv preprint arXiv:2311.09802

  156. [169]

    Safoora Yousefi, Leo Betthauser, Hosein Hasanbeig, Raphaël Millière, Ida Momennejad, Decoding in-context learning: Neuroscience-inspired analysis of representations in large language models, 2023, arXiv preprint arXiv:2310. 00313

  157. [170]

    Chawla, Olaf Wiest, Xiangliang Zhang, Large language model based multi- agents: A survey of progress and challenges, 2024, arXiv preprint arXiv:2402

    Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V. Chawla, Olaf Wiest, Xiangliang Zhang, Large language model based multi- agents: A survey of progress and challenges, 2024, arXiv preprint arXiv:2402. 01680

  158. [171]

    Eric Zelikman, Eliana Lorch, Lester Mackey, Adam Tauman Kalai, Self- taught optimizer (stop): Recursively self-improving code generation, in: First Conference on Language Modeling, 2024

  159. [172]

    4027–4034

    Rameez Qureshi, Naïm Es-Sebbani, Luis Galárraga, Yvette Graham, Miguel Couceiro, Zied Bouraoui, Refine-lm: Mitigating language model stereotypes via reinforcement learning, in: ECAI 2024, IOS Press, 2024, pp. 4027–4034. Knowledge-Based Systems 318 (2025) 113503 25

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.