Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

A Survey of AIOps in the Era of Large Language Models

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper claims that large language models are reshaping every stage of AIOps and that the field now needs a single organizing map.

desk verdict Useful survey of LLM-based AIOps with a sensible taxonomy, but the corpus arithmetic doesn't add up and the 'first comprehensive' claim needs fixing before I'd cite it. read the letter →

arxiv 2507.12472 v1 pith:4CUMUO3E submitted 2025-06-23 cs.SE cs.CL

classification cs.SEcs.CL
keywords AIOpslargelanguagemodelsLLM4AIOpsfailuremanagementrootcauseanalysisanomalydetectionassistedremediationlogparsing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that large language models are not just another tool in IT operations but a force that reshapes every link of the AIOps chain, and that the field now needs a single organizing map. It presents itself as the first comprehensive survey of the entire AIOps process in the LLM era, analyzing a corpus of papers published from January 2020 to December 2024, counted as 163 in the method section and 183 in the abstract. The survey's value, if the claim holds, is that researchers and practitioners get a common taxonomy for comparing methods, choosing approaches, and spotting gaps. The central assertion is that LLMs expand AIOps data beyond logs, metrics, and traces to human-generated text, create new tasks such as root cause report generation and automatic remediation scripts, and force evaluation beyond classification accuracy into generation, execution, and human judgment.

What carries the argument

The organizing machinery is the survey's taxonomy, anchored on a three-stage operational pipeline: failure perception, root cause analysis, and assisted remediation. Within that pipeline, the paper's most distinctive instrument is a five-level automation ladder for remediation, from assisted questioning through mitigation solution generation, command recommendation, and script generation to automatic execution, which it uses to rank how much of the repair workflow LLMs are taking over. The taxonomy does the work of a map: it turns the 160-plus reviewed papers into comparable slots, making the 'first comprehensive coverage' claim checkable and letting gaps, such as the absence of effective trace-data usage, become visible.

What would settle it

Replicating the stated search string across the five databases and applying the inclusion and exclusion criteria to produce a materially different list of primary papers would falsify the 'first comprehensive' claim; so would a simple audit of the screening numbers, since the counts as printed (183 vs 163, and the subtraction in Section 2.3) do not add up.

Watch

Extended reading notes

Core claim

The paper's central claim, stated in Section 1.2, is that it provides the first comprehensive survey covering the entire AIOps process in the context of large language models. Its four research questions organize the field: RQ1 on data sources, RQ2 on task evolution, RQ3 on methods, and RQ4 on evaluation. A sympathetic reading of the findings is that LLMs have turned AIOps from a pipeline that consumes system-generated metrics, logs, and traces into one that also consumes software information, source code, question-answer pairs, and incident reports; that the task pipeline runs from failure perception to root cause analysis to assisted remediation, with new tasks appearing at each stage; that methods divide into foundation-model, fine-tuning, embedding-based, prompt-based, and knowledge-based families; and that evaluation now includes generation, execution, and manual metrics alongside classical classification and error metrics.

Load-bearing premise

The survey's whole map depends on the completeness and representativeness of the screened paper corpus, and that foundation is uncertain: the abstract says 183 papers while Section 2.3 says 163, the screening arithmetic does not reconcile (614 minus 333 does not leave the stated 395), and filters that drop models under 1 billion parameters and papers without experiments could silently exclude relevant work.

Editorial extensions

If this is right

  • Researchers can locate any new LLM-based AIOps paper in one of the four research-question slots and immediately see which data sources, methods, task stages, and evaluation metrics it combines, making comparison across papers more direct.
  • Practitioners can stage their adoption of LLM-based remediation along the five-level automation ladder, from answering operator questions to fully automatic execution, rather than attempting end-to-end automation at once.
  • The identified gaps become a research agenda: trace data is not yet effectively used by LLMs, failure perception under strict real-time constraints is unresolved, and generalizability across evolving software systems is not systematically measured.
  • The appearance of generation, execution, and manual evaluation metrics signals that accuracy-oriented benchmarks are no longer sufficient to judge AIOps systems in the LLM era.
  • The shift toward human-generated data as a first-class input implies future incident-handling pipelines will increasingly start from incident reports and source code rather than from raw logs and metrics.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One step beyond the paper: if its taxonomy becomes the field's common reference, the next bottleneck will be benchmarking, because the reviewed works mostly use idiosyncratic datasets and the survey identifies comparability as a problem without supplying a unified benchmark itself.
  • The paper's exclusion of models under 1 billion parameters means its picture of the LLM era is drawn from large models only; a complementary survey of compact and BERT-scale models would likely show a denser, cheaper landscape for log parsing and anomaly detection.
  • A testable extension: applying the same four-question taxonomy to industrial incident reports or to the Chinese-language OpsEval benchmark could reveal whether the new task categories, such as script generation and assisted questioning, hold outside the academic corpus surveyed.
  • The remediation automation ladder can be read as a maturity scale: an organization could score its own toolchain by the highest rung it reaches, turning the survey's qualitative ladder into a practical assessment instrument.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This manuscript is a systematic survey of LLM-based AIOps covering January 2020 to December 2024. It organizes the field around four research questions: data sources and preprocessing (RQ1), evolution of AIOps tasks (RQ2), LLM-based methods (RQ3), and evaluation methodologies (RQ4). The paper presents a taxonomy with three data-source families, a staged AIOps pipeline from failure perception through root cause analysis to assisted remediation, five method families, and four evaluation-metric families, and it claims to be the first comprehensive survey covering the entire AIOps process in the LLM era. The survey is based on a corpus of 163 papers selected via a five-database search with explicit inclusion and exclusion criteria.

Significance. If the corpus-construction issues are repaired, this survey would be a valuable reference map for a fast-growing field. Its strengths include the explicit four-RQ structure, the comparison with prior AIOps surveys in Table 1, the identification of genuinely new LLM-era tasks (root cause report generation, assisted questioning, command recommendation, script generation, automatic execution), the method taxonomy (foundation models, fine-tuning, embeddings, prompting, RAG/TAG), and the cataloguing of emerging evaluation metrics. The future-directions section is concrete and raises cost, trace-data, generalizability, and toolchain-integration issues that are relevant to practitioners. The central 'first comprehensive survey' claim, however, rests on the reproducibility and internal consistency of the corpus-selection process, and that process currently contains arithmetic errors and contradictory exclusion-criterion applications that must be corrected before the claim can be accepted.

major comments (4)
  1. [Section 2.3 and Figure 4] The corpus construction is arithmetically inconsistent. The main text reports 761 retrieved, 614 after deduplication, 333 excluded at screening, 395 for full-text review, 222 removed, and 163 included; however, 614-333=281 rather than 395, and 395-222=173 rather than 163. The figure's own category counts sum to 219 at screening (51+115+39+14) and 232 at full text (98+72+42+20), and 614-219=395 and 395-232=163 are consistent. The abstract, in addition, states 183 papers. Because the corpus is the empirical basis for the taxonomy and for the 'first comprehensive survey' claim in Section 1.2, the text, figure, and abstract must be reconciled before the selection process can be considered reproducible.
  2. [Section 2.2 (EC1), with Sections 5.2, 5.3, and 7.2] Exclusion criterion EC1 (models with fewer than 1 billion parameters) is contradicted by papers that the survey itself highlights as included. Section 5.2 discusses LLMParser [86] on Flan-T5-small and Flan-T5-base, Section 5.3 reviews GPT-2 embedding approaches [56,99], and Section 7.2 states that many failure-perception works use T5 and GPT-2 models 'not particularly large' [44,67,126,141]. These models are below 1B parameters, so under EC1 these papers should have been excluded. Either the criterion was applied inconsistently or the operational definition of 'LLM' used in screening differs from the one stated; in either case the selection is not reproducible.
  3. [Section 2.2 vs. Section 1.3] The mapping of inclusion criteria to research questions is inverted. Section 2.2 says a paper is selected if it addresses 'proposing novel LLM-based methods (RQ2)' and 'discussing emerging trends in AIOps tasks (RQ3)', while Section 1.3 defines RQ2 as the evolution of AIOps tasks and RQ3 as LLM-based methods. This misalignment makes it ambiguous whether a methods paper was screened against the method criterion or the task criterion, weakening the claimed systematic link between the selection process and the analysis structure.
  4. [Section 7.2 vs. Sections 3.1 and 4.2] The challenge discussion states that no current work has effectively incorporated trace data in LLM-based AIOps, but the survey itself reviews LLM-based trace generation by Kim et al. [60] in Section 3.1 and states in Section 4.2 that root cause analysis in the LLM era is often addressed using 'traces, metrics, and logs' in a single workflow. These statements are mutually inconsistent; the gap analysis should be reconciled with the corpus described in RQ1 and RQ2 before claiming a missing direction.
minor comments (5)
  1. [Section 1.2 and Table 1] Typographical errors in Section 1.2 ('This reminder of this survey') and Table 1 ('Paolop et al.' for Notaro et al.) should be corrected.
  2. [Figure 4] Figure 4 labels the third screening criterion as 'EC3: not related to FM', while the text defines EC3 as 'outside the scope of AIOps'; the labels should match the written criteria, and the criterion acronyms should be defined consistently in both places.
  3. [Section 5.4] Section 5.4 attributes chain-of-thought prompting to 'LogGPT [80]', but reference [80] is LogPrompt; the actual LogGPT work is reference [105]. Please correct the citation or the name to avoid confusion between the two systems.
  4. [Section 6.1.4] Section 6.1.4 contains typos such as 'Qualitative Assessments. .These assessments' and 'exmaple/evaluted'; these should be cleaned up.
  5. [Section 2.3 / Figure 4] The paper would benefit from a supplementary artifact listing all included and excluded papers with the criterion applied; without such a list, the 163-paper corpus cannot be audited even after the arithmetic is corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: this is a literature survey whose claims are summaries of a corpus, not results derived from fitted parameters or self-cited theorems.

full rationale

This paper is a systematic literature survey, not a derivation. Its central claims—that it is 'the first comprehensive survey that covers the entire process of AIOps in the context of large language models' and that its taxonomy organizes the field—are claims about the selected corpus and the state of the literature, not consequences derived from a fitted parameter, an equation, or a prior result. The selection criteria (IC1-IC4, EC1-EC5) are inputs to the survey process; the taxonomy and findings are summaries of the selected papers rather than predictions forced by those criteria. The authors cite some of their own papers (e.g., [153]-[159]) as examples in the reviewed literature, but the survey's arguments do not depend on accepting those papers' results; this is ordinary scholarly self-citation, not load-bearing circularity. The paper does expose internal inconsistencies: Section 2.3 reports 761 retrieved, 614 after deduplication, 333 excluded in screening, 395 for full-text review, 222 removed, and 163 included, while Figure 4 gives 232 removed and the abstract/conclusion state 183 papers. Also, EC1 excludes models under 1 billion parameters, yet the survey highlights LLMParser [86] tested on Flan-T5-small/base and GPT-2-based embedding methods [56, 99]. These are genuine reproducibility and completeness defects in the corpus construction, but they are not circularity: the central survey claim is unsupported by the stated corpus, not equivalent to it by construction. No equation, fitted parameter, or uniquely-forcing self-citation is used to derive a claimed result from its own inputs. Therefore, no circular step meets the evidentiary bar of exhibiting a specific reduction of a claimed result to its inputs.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities; the survey introduces a taxonomy and relies on the literature set plus definitional assumptions.

assumptions (3)
  • domain assumption The 163-paper corpus is representative and complete enough for the survey's taxonomy
    The survey's conclusions (e.g., which tasks are 'new', which data sources are underused) depend on the selected papers in Section 2 being representative of the whole field.
  • domain assumption Papers using models smaller than 1 billion parameters are excluded as non-LLMs (EC1)
    Section 2.2 states EC1; this definitional cutoff shapes every count and category.
  • domain assumption Title and abstract screening suffices to determine relevance
    Section 2.2: 'Judgments are primarily made based on titles and abstracts'. Full-text review only for uncertain cases; this may miss relevant work.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey of AIOps in the Era of Large Language Models." pith.science (2026). https://pith.science/paper/4CUMUO3E

@misc{pith2026250712472,
  author       = {Pith},
  title        = {Pith review of: A Survey of AIOps in the Era of Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4CUMUO3E}},
  note         = {Machine review of arXiv:2507.12472}
}
read the original abstract

As large language models (LLMs) grow increasingly sophisticated and pervasive, their application to various Artificial Intelligence for IT Operations (AIOps) tasks has garnered significant attention. However, a comprehensive understanding of the impact, potential, and limitations of LLMs in AIOps remains in its infancy. To address this gap, we conducted a detailed survey of LLM4AIOps, focusing on how LLMs can optimize processes and improve outcomes in this domain. We analyzed 183 research papers published between January 2020 and December 2024 to answer four key research questions (RQs). In RQ1, we examine the diverse failure data sources utilized, including advanced LLM-based processing techniques for legacy data and the incorporation of new data sources enabled by LLMs. RQ2 explores the evolution of AIOps tasks, highlighting the emergence of novel tasks and the publication trends across these tasks. RQ3 investigates the various LLM-based methods applied to address AIOps challenges. Finally, RQ4 reviews evaluation methodologies tailored to assess LLM-integrated AIOps approaches. Based on our findings, we discuss the state-of-the-art advancements and trends, identify gaps in existing research, and propose promising directions for future exploration.

Figures

Figures reproduced from arXiv: 2507.12472 by the authors.

Figure 1
Figure 1. Analysis of Publication Trends in LLM-Based AIOps [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Taxonomy of Research Questions To achieve this, we organize our survey around the following research questions, each targeting a critical aspect of how large language models are reshaping AIOps: • RQ1: How has the advent of LLMs transformed the sources and preprocessing methods of data in AIOps? • RQ2: How have the tasks of AIOps evolved with the advent of LLMs? J. ACM, Vol. 37, No. 4, Article 111. Publication date:… view at source ↗
Figure 3
Figure 3. Search Strategy Utilized to Identify Studies on AIOps [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Overview of Paper Selection Procedure 3.1 Advancing Preprocessing Techniques for Traditional Data Sources To begin, we review the data sources traditionally employed in ML-based and DL-based AIOps [68, 73, 122, 147, 149, 150, 152, 154, 157, 176, 178]. These studies pri…
Figure 5
Figure 5. Figure 5: Log-based Failure Perception and Root Cause Analysis: The Common Workflow [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Data Source for AIOps in the Era of LLM Software Information. Software information is generated during the software development process and includes details such as software architecture, configurations and documentation. This type of information provides valuable know…
Figure 7
Figure 7. Figure 7: AIOps Tasks in the Era of LLM (Tasks marked with * are new tasks that have emerged in the era of [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]
Figure 8
Figure 8. Figure 8: Evolution of Root Cause Analysis with the Rise of LLMs [PITH_FULL_IMAGE:figures/full_fig_p013_8.png]
Figure 9
Figure 9. Figure 9: Various Types of Auto Remediation Approaches [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: LLM-based Approaches for AIOps generated by pre-trained models to capture semantic information and improve task performance. The prompt-based approach leverages natural language prompts to guide the model’s responses, enabling it to perform specific tasks based on the…
Figure 11
Figure 11. Figure 11: Evaluation Metrics for LLM-based AIOps (Metrics marked with * are new metrics that have emerged [PITH_FULL_IMAGE:figures/full_fig_p022_11.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Progressive Crystallization: Turning Agent Exploration into Deterministic, Lower-Cost Workflows in Production

    cs.SE 2026-07 unverdicted novelty 5.0 of 10

    An evidence-based promotion/demotion lifecycle converts validated LLM agent traces into zero-token deterministic workflows, reducing per-incident cost by 70% in a production cloud-networking system.

Reference graph

Works this paper leans on

181 extracted references · 46 canonical work pages · cited by 1 Pith paper

  1. [86]

    Zeyang Ma, An Ran Chen, Dong Jae Kim, Tse-Hsun Chen, and Shaowei Wang. 2024. LLMParser: An Exploratory Study on Using Large Language Models for Log Parsing. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering. 1–13

  2. [60]

    Donghyun Kim, Sriram Ravula, Taemin Ha, Alexandros G Dimakis, Daehyeok Kim, and Aditya Akella. 2024. Large Language Models as Realistic Microservice Trace Generators. arXiv preprint arXiv:2502.17439 (2024)

  3. [1]

    Arman Ahmed, K Sadanandan Sajan, Anurag Srivastava, and Yinghui Wu. 2021. Anomaly detection, localization and classification using drifting synchrophasor data streams. IEEE Transactions on Smart Grid 12, 4 (2021), 3570–3580

  4. [2]

    Toufique Ahmed, Supriyo Ghosh, Chetan Bansal, Thomas Zimmermann, Xuchao Zhang, and Saravan Rajmohan. 2023. Recommending root-cause and mitigation steps for cloud incidents using large language models. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 1737–1749

  5. [3]

    Marco Aiello and Ilche Georgievski. 2023. Service composition in the ChatGPT era. Service Oriented Computing and Applications 17, 4 (2023), 233–238

  6. [4]

    Khalid Ayedh Alharthi, Arshad Jhumka, Sheng Di, and Franck Cappello. 2022. Clairvoyant: a log-based transformer- decoder for failure prediction in large-scale systems. In Proceedings of the 36th ACM International Conference on Supercomputing. 1–14

  7. [5]

    Khalid Ayed Alharthi, Arshad Jhumka, Sheng Di, Lin Gui, Franck Cappello, and Simon McIntosh-Smith. 2023. Time machine: generative real-time model for failure (and lead time) prediction in hpc systems. In 2023 53rd Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN) . IEEE, 508–521

  8. [6]

    Sarah Alnegheimish, Linh Nguyen, Laure Berti-Equille, and Kalyan Veeramachaneni. 2024. Large language models can be zero-shot anomaly detectors for time series? arXiv preprint arXiv:2405.14755 (2024)

Show all 181 references
  1. [7]

    Dharun Anandayuvaraj, Matthew Campbell, Arav Tewari, and James C Davis. 2024. FAIL: Analyzing Software Failures from the News Using LLMs. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering. 506–518

  2. [8]

    Bhavya Bhavya, Paulina Toro Isaza, Yu Deng, Michael Nidd, Amar Prakash Azad, Larisa Shwartz, and ChengXiang Zhai. 2023. Exploring Large Language Models for Low-Resource IT Information Extraction. In 2023 IEEE International Conference on Data Mining Workshops (ICDMW) . IEEE, 1203–1212

  3. [9]

    Charles Cao, Feiyi Wang, Lisa Lindley, and Zejiang Wang. 2024. Managing Linux servers with LLM-based AI agents: An empirical evaluation with GPT4. Machine Learning with Applications 17 (2024), 100570

  4. [10]

    Defu Cao, Furong Jia, Sercan O Arik, Tomas Pfister, Yixiang Zheng, Wen Ye, and Yan Liu. 2024. TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series Forecasting. In The Twelfth International Conference on Learning Representations

  5. [11]

    Ching Chang, Wen-Chih Peng, and Tien-Fu Chen. 2023. Llm4ts: Two-stage fine-tuning for time-series forecasting with pre-trained llms. arXiv preprint arXiv:2308.08469 (2023)

  6. [12]

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology 15, 3 (2024), 1–45

  7. [13]

    Georgios Chatzigeorgakidis, Konstantinos Lentzos, and Dimitrios Skoutas. 2024. MultiCast: Zero-Shot Multivariate Time Series Forecasting Using LLMs. In 2024 IEEE 40th International Conference on Data Engineering Workshops (ICDEW). IEEE, 119–127

  8. [14]

    Yinfang Chen, Manish Shetty, Gagan Somashekar, Minghua Ma, Yogesh Simmhan, Jonathan Mace, Chetan Bansal, Rujia Wang, and Saravan Rajmohan. [n.d.]. AIOPSLAB: AHOLISTIC FRAMEWORK TO EVALUATE AI AGENTS FOR ENABLING AUTONOMOUS CLOUDS. ([n. d.])

  9. [15]

    Yakun Chen, Xianzhi Wang, and Guandong Xu. 2023. Gatgpt: A pre-trained large language model with graph attention network for spatiotemporal imputation. arXiv preprint arXiv:2311.14332 (2023). J. ACM, Vol. 37, No. 4, Article 111. Publication date: August 2025. 111:28 Lingzhe Zh...

  10. [16]

    Yinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang, Xin Gao, Liu Shi, Yunjie Cao, Xuedong Gao, Hao Fan, Ming Wen, et al. 2024. Automatic root cause analysis via large language models for cloud incidents. In Proceedings of the Nineteenth European Conference on Computer Systems . 674–688

  11. [17]

    Qian Cheng, Doyen Sahoo, Amrita Saha, Wenzhuo Yang, Chenghao Liu, Gerald Woo, Manpreet Singh, Silvio Saverese, and Steven CH Hoi. 2023. Ai for it operations (aiops) on cloud platforms: Reviews, opportunities and challenges. arXiv preprint arXiv:2304.04661 (2023)

  12. [18]

    Tianyu Cui, Shiyu Ma, Ziang Chen, Tong Xiao, Shimin Tao, Yilun Liu, Shenglin Zhang, Duoming Lin, Changchang Liu, Yuzhe Cai, et al. 2024. Logeval: A comprehensive benchmark suite for large language models in log analysis. arXiv preprint arXiv:2407.01896 (2024)

  13. [19]

    Abhimanyu Das, Weihao Kong, Rajat Sen, and Yichen Zhou. 2024. A decoder-only foundation model for time-series forecasting. In Forty-first International Conference on Machine Learning

  14. [20]

    Josu Diaz-de Arcaya, Juan López-de Armentia, Gorka Zárate, and Ana I Torre-Bastida. 2024. Towards the self- healing of Infrastructure as Code projects using constrained LLM technologies. In Proceedings of the 5th ACM/IEEE International Workshop on Automated Program Repair . 22–25

  15. [21]

    Josu Diaz-De-Arcaya, Ana I Torre-Bastida, Gorka Zarate, Raul Minon, and Aitor Almeida. 2023. A joint study of the challenges, opportunities, and roadmap of mlops and aiops: A systematic survey. Comput. Surveys 56, 4 (2023), 1–30

  16. [22]

    Jiaxiang Dong, Haixu Wu, Haoran Zhang, Li Zhang, Jianmin Wang, and Mingsheng Long. 2024. Simmtm: A simple pre-training framework for masked time-series modeling. Advances in Neural Information Processing Systems 36 (2024)

  17. [23]

    Chiming Duan, Tong Jia, Yong Yang, Guiyang Liu, Jinbu Liu, Huxing Zhang, Qi Zhou, Ying Li, and Gang Huang. 2025. EagerLog: Active Learning Enhanced Retrieval Augmented Generation for Log-based Anomaly Detection. In ICASSP 2025-2025 IEEE International Conference on Acoustics, S...

  18. [24]

    Chris Egersdoerfer, Di Zhang, and Dong Dai. 2023. Early exploration of using chatgpt for log-based anomaly detection on parallel file systems logs. In Proceedings of the 32nd International Symposium on High-Performance Parallel and Distributed Computing. 315–316

  19. [25]

    Vijay Ekambaram, Arindam Jati, Nam H Nguyen, Pankaj Dayama, Chandra Reddy, Wesley M Gifford, and Jayant Kalagnanam. 2024. TTMs: Fast Multi-level Tiny Time Mixers for Improved Zero-shot and Few-shot Forecasting of Multivariate Time Series. arXiv preprint arXiv:2401.03955 (2024)

  20. [26]

    Stephen Elliot. 2014. DevOps and the cost of downtime: Fortune 1000 best practice metrics quantified. International Data Corporation (IDC) (2014)

  21. [27]

    Asma Fariha, Vida Gharavian, Masoud Makrehchi, Shahryar Rahnamayan, Sanaa Alwidian, and Akramul Azim. 2024. Log Anomaly Detection by Leveraging LLM-Based Parsing and Embedding with Attention Mechanism. In 2024 IEEE Canadian Conference on Electrical and Computer Engineering (CC...

  22. [28]

    Azul Garza and Max Mergenthaler-Canseco. 2023. TimeGPT-1. arXiv preprint arXiv:2310.03589 (2023)

  23. [29]

    Drishti Goel, Fiza Husain, Aditya Singh, Supriyo Ghosh, Anjaly Parayil, Chetan Bansal, Xuchao Zhang, and Saravan Rajmohan. 2024. X-lifecycle learning for cloud incident management using llms. In Companion Proceedings of the 32nd ACM International Conference on the Foundations ...

  24. [30]

    Nate Gruver, Marc Finzi, Shikai Qiu, and Andrew G Wilson. 2024. Large language models are zero-shot time series forecasters. Advances in Neural Information Processing Systems 36 (2024)

  25. [31]

    Jindong Gu, Zhen Han, Shuo Chen, Ahmad Beirami, Bailan He, Gengyuan Zhang, Ruotong Liao, Yao Qin, Volker Tresp, and Philip Torr. 2023. A systematic survey of prompt engineering on vision-language foundation models. arXiv preprint arXiv:2307.12980 (2023)

  26. [32]

    Hongcheng Guo, Jian Yang, Jiaheng Liu, Liqun Yang, Linzheng Chai, Jiaqi Bai, Junran Peng, Xiaorong Hu, Chao Chen, Dongfeng Zhang, et al. 2024. OWL: A Large Language Model for IT Operations. In The Twelfth International Conference on Learning Representations

  27. [33]

    Hongcheng Guo, Wei Zhang, Anjie Le, Jian Yang, Jiaheng Liu, Zhoujun Li, Tieqiao Zheng, Shi Xu, Runqiang Zang, Liangfan Zheng, et al. 2024. Lemur: Log Parsing with Entropy Sampling and Chain-of-Thought Merging. arXiv preprint arXiv:2402.18205 (2024)

  28. [34]

    Pranjal Gupta, Harshit Kumar, Debanjana Kar, Karan Bhukar, Pooja Aggarwal, and Prateeti Mohapatra. 2023. Learning Representations on Logs for AIOps. In 2023 IEEE 16th International Conference on Cloud Computing (CLOUD) . IEEE, 155–166

  29. [35]

    Fatemeh Hadadi, Qinghua Xu, Domenico Bianculli, and Lionel Briand. 2024. Anomaly Detection on Unstable Logs with GPT Models. arXiv preprint arXiv:2406.07467 (2024)

  30. [36]

    Pouya Hamadanian, Behnaz Arzani, Sadjad Fouladi, Siva Kesava Reddy Kakarla, Rodrigo Fonseca, Denizcan Billor, Ahmad Cheema, Edet Nkposong, and Ranveer Chandra. 2023. A Holistic View of AI-driven Network Incident Management. In Proceedings of the 22nd ACM Workshop on Hot Topics...

  31. [37]

    Shangbin Han, Qianhong Wu, Han Zhang, Bo Qin, Jiankun Hu, Xingang Shi, Linfeng Liu, and Xia Yin. 2021. Log-based anomaly detection with robust feature extraction and online learning. IEEE Transactions on Information Forensics and Security 16 (2021), 2300–2311

  32. [38]

    Yongqi Han, Qingfeng Du, Ying Huang, Jiaqi Wu, Fulong Tian, and Cheng He. 2024. The Potential of One-Shot Failure Root Cause Analysis: Collaboration of the Large Language Model and Small Classifier. In Proceedings of the 39th IEEE/ACM International Conference on Automated Soft...

  33. [39]

    Ahatsham Hayat and Mohammad R Hasan. 2025. A Context-Aware Approach for Enhancing Data Imputation with Pre-trained Language Models. In Proceedings of the 31st International Conference on Computational Linguistics . 5668–5685

  34. [40]

    Minghua He, Tong Jia, Chiming Duan, Huaqian Cai, Ying Li, and Gang Huang. 2024. LLMeLog: An Approach for Anomaly Detection based on LLM-enriched Log Events. In 2024 IEEE 35th International Symposium on Software Reliability Engineering (ISSRE). IEEE, 132–143

  35. [41]

    Shilin He, Jieming Zhu, Pinjia He, and Michael R Lyu. 2023. Loghub: A large collection of system log datasets towards automated log analytics. (2023)

  36. [42]

    Stefan Hegselmann, Alejandro Buendia, Hunter Lang, Monica Agrawal, Xiaoyi Jiang, and David Sontag. 2023. Tabllm: Few-shot classification of tabular data with large language models. InInternational Conference on Artificial Intelligence and Statistics. PMLR, 5549–5581

  37. [43]

    Adha Hrusto, Per Runeson, and Magnus C Ohlsson. 2024. Autonomous monitors for detecting failures early and reporting interpretable alerts in cloud operations. In Proceedings of the 46th International Conference on Software Engineering: Software Engineering in Practice . 47–57

  38. [44]

    Junjie Huang, Zhihan Jiang, Jinyang Liu, Yintong Huo, Jiazhen Gu, Zhuangbin Chen, Cong Feng, Hui Dong, Zengyin Yang, and Michael R Lyu. 2024. Demystifying and Extracting Fault-indicating Information from Logs for Failure Diagnosis. In 2024 IEEE 35th International Symposium on ...

  39. [45]

    Shaohan Huang, Yi Liu, Jiaxing Qi, Jing Shang, Zhiwen Xiao, Carol Fung, Zhihui Wu, Hailong Yang, Zhongzhi Luan, and Depei Qian. 2024. Gloss: Guiding Large Language Models to Answer Questions from System Logs. In 2024 IEEE International Conference on Software Analysis, Evolutio...

  40. [46]

    Azam Ikram, Sarthak Chakraborty, Subrata Mitra, Shiv Saini, Saurabh Bagchi, and Murat Kocaoglu. 2022. Root cause analysis of failures in microservices through causal discovery. Advances in Neural Information Processing Systems 35 (2022), 31158–31170

  41. [47]

    Michel Jacobsen and Marina Tropmann-Frick. 2024. Imputation Strategies in Time Series Based on Language Models. Datenbank-Spektrum 24, 3 (2024), 197–207

  42. [48]

    Yuhe JI, Jing HAN, Yongxin ZHAO, Shenglin ZHANG, and Zican GONG. 2023. Log Anomaly Detection Through GPT-2 for Large Scale Systems. ZTE Communications 21, 3 (2023), 70

  43. [49]

    Tong Jia, Yifan Wu, Chuanjia Hou, and Ying Li. 2021. Logflash: Real-time streaming anomaly detection and diagnosis from system logs for large-scale software systems. In 2021 IEEE 32nd International Symposium on Software Reliability Engineering (ISSRE). IEEE, 80–90

  44. [50]

    Yuxuan Jiang, Chaoyun Zhang, Shilin He, Zhihao Yang, Minghua Ma, Si Qin, Yu Kang, Yingnong Dang, Saravan Rajmohan, Qingwei Lin, et al. 2024. Xpert: Empowering incident management with query recommendations via large language models. In Proceedings of the IEEE/ACM 46th Internat...

  45. [51]

    Zhihan Jiang, Jinyang Liu, Zhuangbin Chen, Yichen Li, Junjie Huang, Yintong Huo, Pinjia He, Jiazhen Gu, and Michael R Lyu. 2024. Lilac: Log parsing using llms with adaptive parsing cache. Proceedings of the ACM on Software Engineering 1, FSE (2024), 137–160

  46. [52]

    Hongwei Jin, George Papadimitriou, Krishnan Raghavan, Pawel Zuk, Prasanna Balaprakash, Cong Wang, Anirban Mandal, and Ewa Deelman. 2024. Large Language Models for Anomaly Detection in Computational Workflows: From Supervised Fine-Tuning to In-Context Learning. In 2024 SC24: In...

  47. [53]

    Ming Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu, James Y Zhang, Xiaoming Shi, Pin-Yu Chen, Yuxuan Liang, Yuan-Fang Li, Shirui Pan, et al. 2023. Time-llm: Time series forecasting by reprogramming large language models. arXiv preprint arXiv:2310.01728 (2023)

  48. [54]

    Pengxiang Jin, Shenglin Zhang, Minghua Ma, Haozhe Li, Yu Kang, Liqun Li, Yudong Liu, Bo Qiao, Chaoyun Zhang, Pu Zhao, et al. 2023. Assess and summarize: Improve outage understanding with large language models. In Proceedings of the 31st ACM Joint European Software Engineering ...

  49. [55]

    Yuyuan Kang, Xiangdong Huang, Shaoxu Song, Lingzhe Zhang, Jialin Qiao, Chen Wang, Jianmin Wang, and Julian Feinauer. 2022. Separation or not: On handing out-of-order time-series data in leveled lsm-tree. In 2022 IEEE 38th International Conference on Data Engineering (ICDE) . I...

  50. [56]

    Egil Karlsen, Xiao Luo, Nur Zincir-Heywood, and Malcolm Heywood. 2024. Large language models and unsupervised feature learning: implications for log analysis. Annals of Telecommunications (2024), 1–19

  51. [57]

    Khalid S Khan, Regina Kunz, Jos Kleijnen, and Gerd Antes. 2003. Five steps to conducting a systematic review.Journal of the royal society of medicine 96, 3 (2003), 118–121

  52. [58]

    Subina Khanal, Seshu Tirupathi, Giulio Zizzo, Ambrish Rawat, and Torben Bach Pedersen. 2024. Domain Adaptation for Time series Transformers using One-step fine-tuning. arXiv preprint arXiv:2401.06524 (2024)

  53. [59]

    Pitikorn Khlaisamniang, Prachaya Khomduean, Kriangkrai Saetan, and Supasin Wonglapsuwan. 2023. Generative AI for self-healing systems. In 2023 18th International Joint Symposium on Artificial Intelligence and Natural Language Processing (iSAI-NLP). IEEE, 1–6

  54. [61]

    Barbara Kitchenham, O Pearl Brereton, David Budgen, Mark Turner, John Bailey, and Stephen Linkman. 2009. Systematic literature reviews in software engineering–a systematic literature review. Information and software technology 51, 1 (2009), 7–15

  55. [62]

    Jinxi Kuang, Jinyang Liu, Junjie Huang, Renyi Zhong, Jiazhen Gu, Lan Yu, Rui Tan, Zengyin Yang, and Michael R Lyu. 2024. Knowledge-aware Alert Aggregation in Large-scale Cloud Systems: a Hybrid Approach. arXiv preprint arXiv:2403.06485 (2024)

  56. [63]

    Max Landauer, Florian Skopik, and Markus Wurzenberger. 2023. A Critical Review of Common Log Data Sets Used for Evaluation of Sequence-based Anomaly Detection Techniques. arXiv preprint arXiv:2309.02854 (2023)

  57. [64]

    Pedro Las-Casas, Alok Gautum Kumbhare, Rodrigo Fonseca, and Sharad Agarwal. 2024. LLexus: an AI agent system for incident management. ACM SIGOPS Operating Systems Review 58, 1 (2024), 23–36

  58. [65]

    Van-Hoang Le and Hongyu Zhang. 2021. Log-based anomaly detection without log parsing. In 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 492–504

  59. [66]

    Van-Hoang Le and Hongyu Zhang. 2023. Log Parsing: How Far Can ChatGPT Go?. In2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 1699–1704

  60. [67]

    Van-Hoang Le and Hongyu Zhang. 2024. PreLog: A Pre-trained Model for Log Analytics. Proceedings of the ACM on Management of Data 2, 3 (2024), 1–28

  61. [68]

    Guoliang Li, Xuanhe Zhou, Ji Sun, Xiang Yu, Yue Han, Lianyuan Jin, Wenbo Li, Tianqing Wang, and Shifu Li. 2021. opengauss: An autonomous database system. Proceedings of the VLDB Endowment 14, 12 (2021), 3028–3042

  62. [69]

    Hongbo Li, Wenli Zheng, Feilong Tang, Yanmin Zhu, and Jielong Huang. 2023. Few-shot time-series anomaly detection with unsupervised domain adaptation. Information Sciences 649 (2023), 119610

  63. [70]

    Junyi Li, Tianyi Tang, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2024. Pre-trained language models for text generation: A survey. Comput. Surveys 56, 9 (2024), 1–39

  64. [71]

    Peiwen Li, Xin Wang, Zeyang Zhang, Yuan Meng, Fang Shen, Yue Li, Jialong Wang, Yang Li, and Wenweu Zhu. 2024. LLM-Enhanced Causal Discovery in Temporal Domain from Interventional Data. arXiv preprint arXiv:2404.14786 (2024)

  65. [72]

    Wenlong Liao, Fernando Porte-Agel, Jiannong Fang, Christian Rehtanz, Shouxiang Wang, Dechang Yang, and Zhe Yang. 2024. TimeGPT in Load Forecasting: A Large Time Series Model Perspective. arXiv preprint arXiv:2404.04885 (2024)

  66. [73]

    Qingwei Lin, Hongyu Zhang, Jian-Guang Lou, Yu Zhang, and Xuewei Chen. 2016. Log clustering based problem identification for online service systems. In Proceedings of the 38th International Conference on Software Engineering Companion. 102–111

  67. [74]

    Yifei Lin, Hanqiu Deng, and Xingyu Li. 2024. FastLogAD: Log Anomaly Detection with Mask-Guided Pseudo Anomaly Generation and Discrimination. arXiv preprint arXiv:2404.08750 (2024)

  68. [75]

    Haoxin Liu, Zhiyuan Zhao, Jindong Wang, Harshavardhan Kamarthi, and B Aditya Prakash. 2024. LSTPrompt: Large Language Models as Zero-Shot Time Series Forecasters by Long-Short-Term Prompting. arXiv preprint arXiv:2402.16132 (2024)

  69. [76]

    Shuo Liu, Di Yao, Lanting Fang, Zhetao Li, Wenbin Li, Kaiyu Feng, XiaoWen Ji, and Jingping Bi. 2024. AnomalyLLM: Few-shot Anomaly Edge Detection for Dynamic Graphs using Large Language Models.arXiv preprint arXiv:2405.07626 (2024)

  70. [77]

    Xu Liu, Junfeng Hu, Yuan Li, Shizhe Diao, Yuxuan Liang, Bryan Hooi, and Roger Zimmermann. 2024. Unitime: A language-empowered unified model for cross-domain time series forecasting. In Proceedings of the ACM on Web Conference 2024. 4095–4106

  71. [78]

    Yilun Liu, Yuhe Ji, Shimin Tao, Minggui He, Weibin Meng, Shenglin Zhang, Yongqian Sun, Yuming Xie, Boxing Chen, and Hao Yang. 2024. LogLM: From Task-based to Instruction-based Automated Log Analysis. arXiv preprint arXiv:2410.09352 (2024). J. ACM, Vol. 37, No. 4, Article 111. ...

  72. [79]

    Yuhe Liu, Changhua Pei, Longlong Xu, Bohan Chen, Mingze Sun, Zhirui Zhang, Yongqian Sun, Shenglin Zhang, Kun Wang, Haiming Zhang, et al. 2023. OpsEval: A Comprehensive Task-Oriented AIOps Benchmark for Large Language Models. arXiv preprint arXiv:2310.07637 (2023)

  73. [80]

    Yilun Liu, Shimin Tao, Weibin Meng, Jingyu Wang, Wenbing Ma, Yuhang Chen, Yanqing Zhao, Hao Yang, and Yanfei Jiang. 2024. Interpretable online log analysis using large language models with prompt strategies. In Proceedings of the 32nd IEEE/ACM International Conference on Progr...

  74. [81]

    Yong Liu, Haoran Zhang, Chenyu Li, Xiangdong Huang, Jianmin Wang, and Mingsheng Long. 2024. Timer: Generative Pre-trained Transformers Are Large Time Series Models. In Forty-first International Conference on Machine Learning

  75. [82]

    Yingzhe Lyu, Heng Li, Zhen Ming, Ahmed E Hassan, et al . 2023. Assessing the maturity of model maintenance techniques for AIOps solutions. arXiv preprint arXiv:2311.03213 (2023)

  76. [83]

    A Hamou-Lhadj M Mehrabi and H Moosavi. 2024. The Effectiveness of Compact Fine-Tuned LLMs in Log Parsing. In 2024 IEEE International Conference on Software Maintenance and Evolution (ICSME) . IEEE, 97–109

  77. [84]

    Lipeng Ma, Weidong Yang, Sihang Jiang, Ben Fei, Mingjie Zhou, Shuhao Li, Bo Xu, and Yanghua Xiao. 2024. LUK: Empowering Log Understanding with Expert Knowledge from Large Language Models.arXiv preprint arXiv:2409.01909 (2024)

  78. [85]

    Lipeng Ma, Weidong Yang, Bo Xu, Sihang Jiang, Ben Fei, Jiaqing Liang, Mingjie Zhou, and Yanghua Xiao. 2024. Knowlog: Knowledge enhanced pre-trained language model for log understanding. InProceedings of the 46th IEEE/ACM International Conference on Software Engineering . 1–13

  79. [87]

    Ehud Malul, Yair Meidan, Dudu Mimran, Yuval Elovici, and Asaf Shabtai. 2024. GenKubeSec: LLM-Based Kubernetes Misconfiguration Detection, Localization, Reasoning, and Remediation. arXiv preprint arXiv:2405.19954 (2024)

  80. [88]

    Maryam Mehrabi, Abdelwahab Hamou-Lhadj, and Hossein Moosavi. 2024. The Effectiveness of Compact Fine-Tuned LLMs in Log Parsing. In 2024 IEEE International Conference on Software Maintenance and Evolution (ICSME) . IEEE, 1–12

  81. [89]

    Weibin Meng, Federico Zaiter, Yuzhe Zhang, Ying Liu, Shenglin Zhang, Shimin Tao, Yichen Zhu, Tao Han, Yongpeng Zhao, En Wang, et al. 2023. Logsummary: Unstructured log summarization for software systems. IEEE Transactions on Network and Service Management 20, 3 (2023), 3803–3815

  82. [90]

    Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2024. Large language models: A survey. arXiv preprint arXiv:2402.06196 (2024)

  83. [91]

    Francesco Minna, Fabio Massacci, and Katja Tuma. 2024. Analyzing and Mitigating (with LLMs) the Security Misconfigurations of Helm Charts from Artifact Hub. arXiv preprint arXiv:2403.09537 (2024)

  84. [92]

    Panagiotis Misiakos, Chris Wendler, and Markus Püschel. 2024. Learning DAGs from data with few root causes. Advances in Neural Information Processing Systems 36 (2024)

  85. [93]

    Priyanka Mudgal, Bijan Arbab, and Swaathi Sampath Kumar. 2024. CrashEventLLM: Predicting System Crashes with Large Language Models. In 2024 International Conference on Information Technology and Computing (ICITCOM) . IEEE, 72–76

  86. [94]

    Priyanka Mudgal and Rita Wouhaybi. 2023. An Assessment of ChatGPT on Log Data. In International Conference on AI-generated Content. Springer, 148–169

  87. [95]

    Zakeya Namrud, Komal Sarda, Marin Litoiu, Larisa Shwartz, and Ian Watts. 2024. KubePlaybook: A Repository of Ansible Playbooks for Kubernetes Auto-Remediation with LLMs. In Companion of the 15th ACM/SPEC International Conference on Performance Engineering . 57–61

  88. [96]

    Giang Nguyen, Stefan Dlugolinsky, Viet Tran, and Álvaro López García. 2024. Network security AIOps for online stream data monitoring. Neural Computing and Applications 36, 24 (2024), 14925–14949

  89. [97]

    Paolo Notaro, Jorge Cardoso, and Michael Gerndt. 2021. A survey of aiops methods for failure management. ACM Transactions on Intelligent Systems and Technology (TIST) 12, 6 (2021), 1–45

  90. [98]

    Achraf Othman, Amira Dhouib, and Aljazi Nasser Al Jabor. 2023. Fostering websites accessibility: A case study on the use of the Large Language Models ChatGPT for automatic remediation. In Proceedings of the 16th International Conference on PErvasive Technologies Related to Ass...

  91. [99]

    Harold Ott, Jasmin Bogatinovski, Alexander Acker, Sasho Nedelkoski, and Odej Kao. 2021. Robust and transferable anomaly detection in log data using pre-trained language models. In 2021 IEEE/ACM international workshop on cloud intelligence (CloudIntelligence). IEEE, 19–24

  92. [100]

    Jonathan Pan, Wong Swee Liang, and Yuan Yidi. 2024. Raglog: Log anomaly detection using retrieval augmented generation. In 2024 IEEE World Forum on Public Safety Technology (WFPST) . IEEE, 169–174

  93. [101]

    Guansong Pang, Chunhua Shen, Longbing Cao, and Anton Van Den Hengel. 2021. Deep learning for anomaly detection: A review. ACM computing surveys (CSUR) 54, 2 (2021), 1–38. J. ACM, Vol. 37, No. 4, Article 111. Publication date: August 2025. 111:32 Lingzhe Zhang et al

  94. [102]

    Gijun Park and Dohoon Kim. 2023. Formulating an Korean LLM-Based Interactive Assistant for Enhanced IT Collaboration in Microservice Environments. In 2023 IEEE 8th International Conference on Smart Cloud (SmartCloud) . IEEE, 176–181

  95. [103]

    Robin D Pesl, Miles Stötzner, Ilche Georgievski, and Marco Aiello. 2023. Uncovering LLMs for Service-Composition: Challenges and Opportunities. In International Conference on Service-Oriented Computing . Springer, 39–48

  96. [104]

    Pankaj Prasad and Charley Rich. 2018. Market guide for aiops platforms. Retrieved March 12, 2020 (2018), 2–9

  97. [105]

    Jiaxing Qi, Shaohan Huang, Zhongzhi Luan, Shu Yang, Carol Fung, Hailong Yang, Depei Qian, Jing Shang, Zhiwen Xiao, and Zhihui Wu. 2023. Loggpt: Exploring chatgpt for log-based anomaly detection. In 2023 IEEE International Conference on High Performance Computing & Communicatio...

  98. [106]

    Andres Quan, Leah Howell, and Hugh Greenberg. 2023. Heterogeneous Syslog Analysis: There Is Hope. InProceedings of the SC’23 Workshops of The International Conference on High Performance Computing, Network, Storage, and Analysis . 581–587

  99. [107]

    Kashif Rasul, Arjun Ashok, Andrew Robert Williams, Arian Khorasani, George Adamopoulos, Rishika Bhagwatkar, Marin Biloš, Hena Ghonia, Nadhir Hassen, Anderson Schneider, et al. 2023. Lag-llama: Towards foundation models for time series forecasting. In R0-FoMo: Robustness of Few...

  100. [108]

    Youcef Remil, Anes Bendimerad, Romain Mathonat, and Mehdi Kaytoue. 2024. Aiops solutions for incident manage- ment: Technical guidelines and a comprehensive literature review. arXiv preprint arXiv:2404.01363 (2024)

  101. [109]

    Devjeet Roy, Xuchao Zhang, Rashi Bhave, Chetan Bansal, Pedro Las-Casas, Rodrigo Fonseca, and Saravan Rajmohan

  102. [110]

    Alicia Russell-Gilbert, Alexander Sommers, Andrew Thompson, Logan Cummins, Sudip Mittal, Shahram Rahimi, Maria Seale, Joseph Jaboure, Thomas Arnold, and Joshua Church. 2024. AAD-LLM: Adaptive Anomaly Detection Using Large Language Models. arXiv preprint arXiv:2411.00914 (2024)

  103. [111]

    Pranab Sahoo, Ayush Kumar Singh, Sriparna Saha, Vinija Jain, Samrat Mondal, and Aman Chadha. 2024. A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.arXiv preprint arXiv:2402.07927 (2024)

  104. [112]

    Komal Sarda. 2023. Leveraging Large Language Models for Auto-remediation in Microservices Architecture. In 2023 IEEE International Conference on Autonomic Computing and Self-Organizing Systems Companion (ACSOS-C) . IEEE, 16–18

  105. [113]

    Komal Sarda, Zakeya Namrud, Marin Litoiu, Larisa Shwartz, and Ian Watts. 2024. Leveraging Large Language Models for the Auto-remediation of Microservice Applications: An Experimental Study. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of...

  106. [114]

    Komal Sarda, Zakeya Namrud, Raphael Rouf, Harit Ahuja, Mohammadreza Rasolroveicy, Marin Litoiu, Larisa Shwartz, and Ian Watts. 2023. Adarma auto-detection and auto-remediation of microservice anomalies by leveraging large language models. In Proceedings of the 33rd Annual Inte...

  107. [115]

    Jahanggir Hossain Setu, Md Shazzad Hossain, Nabarun Halder, Ashraful Islam, and M Ashraful Amin. 2024. Optimizing Software Release Management with GPT-Enabled Log Anomaly Detection. In International Conference on Pattern Recognition. Springer, 351–365

  108. [116]

    Shiwen Shan, Yintong Huo, Yuxin Su, Yichen Li, Dan Li, and Zibin Zheng. 2024. Face It Yourselves: An LLM-Based Two-Stage Strategy to Localize Configuration Errors via Logs. arXiv preprint arXiv:2404.00640 (2024)

  109. [117]

    Manish Shetty, Yinfang Chen, Gagan Somashekar, Minghua Ma, Yogesh Simmhan, Xuchao Zhang, Jonathan Mace, Dax Vandevoorde, Pedro Las-Casas, Shachee Mishra Gupta, et al. 2024. Building AI Agents for Autonomous Clouds: Challenges and Design Principles. In Proceedings of the 2024 A...

  110. [118]

    Honghao Shi, Longkai Cheng, Wenli Wu, Yuhang Wang, Xuan Liu, Shaokai Nie, Weixv Wang, Xuebin Min, Chunlei Men, and Yonghua Lin. 2024. Enhancing Cluster Resilience: LLM-agent Based Autonomous Intelligent Cluster Diagnosis System and Evaluation Framework. arXiv preprint arXiv:24...

  111. [119]

    Jie Shi, Sihang Jiang, Bo Xu, Jiaqing Liang, Yanghua Xiao, and Wei Wang. 2023. ShellGPT: Generative Pre-trained Transformer Model for Shell Language Understanding. In 2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE). IEEE, 671–682

  112. [120]

    Jacopo Soldani and Antonio Brogi. 2022. Anomaly detection and failure root cause analysis in (micro) service-based cloud applications: A survey. ACM Computing Surveys (CSUR) 55, 3 (2022), 1–39

  113. [121]

    Jing Su, Chufeng Jiang, Xin Jin, Yuxin Qiao, Tingsong Xiao, Hongda Ma, Rong Wei, Zhi Jing, Jiajun Xu, and Junhong Lin. 2024. Large Language Models for Forecasting and Anomaly Detection: A Systematic Literature Review. arXiv preprint arXiv:2402.10350 (2024)

  114. [122]

    Yicheng Sui, Yuzhe Zhang, Jianjun Sun, Ting Xu, Shenglin Zhang, Zhengdan Li, Yongqian Sun, Fangrui Guo, Junyu Shen, Yuzhi Zhang, et al . 2023. LogKG: Log Failure Diagnosis through Knowledge Graph. IEEE Transactions on J. ACM, Vol. 37, No. 4, Article 111. Publication date: Augu...

  115. [123]

    Chenxi Sun, Hongyan Li, Yaliang Li, and Shenda Hong. 2024. TEST: Text Prototype Aligned Embedding to Activate LLM’s Ability for Time Series. In The Twelfth International Conference on Learning Representations

  116. [124]

    Yuchen Sun, Yanpiao Chen, Haotian Zhao, and Shan Peng. 2023. Design and Development of a Log Management System Based on Cloud Native Architecture. In 2023 9th International Conference on Systems and Informatics (ICSAI) . IEEE, 1–6

  117. [125]

    Yicheng SUN, Jacky Keung, Zhen Yang, Shuo Liu, and Hi Kuen Yu. 2024. Semirald: A Semi-Supervised Hybrid Language Model for Robust Anomalous Log Detection. A vailable at SSRN 4927951(2024)

  118. [126]

    Yongqian Sun, Binpeng Shi, Mingyu Mao, Minghua Ma, Sibo Xia, Shenglin Zhang, and Dan Pei. 2024. ART: A Unified Unsupervised Framework for Incident Management in Microservice Systems. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering...

  119. [127]

    Łukasz Tulczyjew, Kinan Jarrah, Charles Abondo, Dina Bennett, and Nathanael Weill. 2024. LLMcap: Large Language Model for Unsupervised PCAP Failure Detection. In 2024 IEEE International Conference on Communications Workshops (ICC Workshops). IEEE, 1559–1565

  120. [128]

    Jiabo Wang, Guojun Chu, Jingyu Wang, Haifeng Sun, Qi Qi, Yuanyi Wang, Ji Qi, and Jianxin Liao. 2024. LogExpert: Log-based Recommended Resolutions Generation using Large Language Model. In Proceedings of the 2024 ACM/IEEE 44th International Conference on Software Engineering: N...

  121. [129]

    Jingyu Wang, Lei Zhang, Yiran Yang, Zirui Zhuang, Qi Qi, Haifeng Sun, Lu Lu, Junlan Feng, and Jianxin Liao. 2023. Network Meets ChatGPT: Intent Autonomous Management, Control and Operation. Journal of Communications and Information Networks 8, 3 (2023), 239–255

  122. [130]

    Lu Wang, Chaoyun Zhang, Ruomeng Ding, Yong Xu, Qihang Chen, Wentao Zou, Qingjun Chen, Meng Zhang, Xuedong Gao, Hao Fan, et al. 2023. Root cause analysis for microservice systems via hierarchical reinforcement learning from human feedback. In Proceedings of the 29th ACM SIGKDD ...

  123. [131]

    Qing Wang, Wubai Zhou, Chunqiu Zeng, Tao Li, Larisa Shwartz, and Genady Ya Grabarnik. 2017. Constructing the knowledge base for cognitive it service management. In 2017 IEEE International Conference on Services Computing (SCC). IEEE, 410–417

  124. [132]

    Sheng-Kai Wang, Shang-Pin Ma, Chen-Hao Chao, and Guan-Hong Lai. 2023. Low-code ChatOps for Microservices Systems Using Service Composition. In 2023 IEEE International Conference on e-Business Engineering (ICEBE) . IEEE, 55–62

  125. [133]

    Zefan Wang, Zichuan Liu, Yingying Zhang, Aoxiao Zhong, Jihong Wang, Fengbin Yin, Lunting Fan, Lingfei Wu, and Qingsong Wen. 2024. Rcagent: Cloud root cause analysis by autonomous agents with tool-augmented large language models. In Proceedings of the 33rd ACM International Con...

  126. [134]

    Xinjie Wei, Jie Wang, Chang-ai Sun, Dave Towey, Shoufeng Zhang, Wanqing Zuo, Yiming Yu, Ruoyi Ruan, and Guyang Song. 2024. Log-based anomaly detection for distributed systems: State of the art, industry experience, and open issues. Journal of Software: Evolution and Process (2024)

  127. [135]

    Jinfeng Wen, Zhenpeng Chen, Federica Sarro, Zixi Zhu, Yi Liu, Haodi Ping, and Shangguang Wang. 2024. LLM-Based Misconfiguration Detection for AWS Serverless Computing. arXiv preprint arXiv:2411.00642 (2024)

  128. [136]

    Yi Xiao, Van-Hoang Le, and Hongyu Zhang. 2024. Free: Towards More Practical Log Parsing with Large Language Models. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering . 153–165

  129. [137]

    Yi Xiao, Van-Hoang Le, and Hongyu Zhang. 2024. Stronger, Faster, and Cheaper Log Parsing with LLMs. arXiv preprint arXiv:2406.06156 (2024)

  130. [138]

    Zhiqiang Xie, Yujia Zheng, Lizi Ottens, Kun Zhang, Christos Kozyrakis, and Jonathan Mace. 2024. Cloud Atlas: Efficient Fault Localization for Cloud Systems using Language Models and Causal Insight. arXiv preprint arXiv:2407.08694 (2024)

  131. [139]

    Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi. 2023. Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms. arXiv preprint arXiv:2306.13063 (2023)

  132. [140]

    Junjielong Xu, Ruichun Yang, Yintong Huo, Chengyu Zhang, and Pinjia He. 2024. DivLog: Log Parsing with Prompt Enhanced In-Context Learning. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . 1–12

  133. [141]

    Hao Xue and Flora D Salim. 2023. Promptcast: A new prompt-based learning paradigm for time series forecasting. IEEE Transactions on Knowledge and Data Engineering (2023)

  134. [142]

    Senming Yan, Lei Shi, Jing Ren, Wei Wang, Yaxin Liu, Limin Sun, Xiong Wang, and Wei Zhang. 2024. Log-Based Anomaly Detection with Transformers Pre-Trained on Large-Scale Unlabeled Data. In ICC 2024-IEEE International Conference on Communications. IEEE, 1–6. J. ACM, Vol. 37, No...

  135. [143]

    Fangkai Yang, Wenjie Yin, Lu Wang, Tianci Li, Pu Zhao, Bo Liu, Paul Wang, Bo Qiao, Yudong Liu, Mårten Björkman, et al. 2023. Diffusion-Based Time Series Data Imputation for Cloud Failure Prediction at Microsoft 365. In Proceedings of the 31st ACM Joint European Software Engine...

  136. [144]

    Jiexia Ye, Weiqi Zhang, Ke Yi, Yongzi Yu, Ziyue Li, Jia Li, and Fugee Tsung. 2024. A Survey of Time Series Foundation Models: Generalizing Time Series Representation with Large Language Model. ArXiv abs/2405.02358 (2024). https: //api.semanticscholar.org/CorpusID:269605992

  137. [145]

    Zhaoyang Yu, Minghua Ma, Chaoyun Zhang, Si Qin, Yu Kang, Chetan Bansal, Saravan Rajmohan, Yingnong Dang, Changhua Pei, Dan Pei, et al. 2024. Monitorassistant: Simplifying cloud service monitoring via large language models. In Companion Proceedings of the 32nd ACM International...

  138. [146]

    Zhaoyang Yu, Changhua Pei, Xin Wang, Minghua Ma, Chetan Bansal, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang, Xidao Wen, Jianhui Li, et al. 2024. Pre-trained kpi anomaly detection model through disentangled transformer. In Proceedings of the 30th ACM SIGKDD Conference on Knowl...

  139. [147]

    Yue Yuan, Wenchang Shi, Bin Liang, and Bo Qin. 2019. An approach to cloud execution failure diagnosis based on exception logs in openstack. In 2019 IEEE 12th International Conference on Cloud Computing (CLOUD) . IEEE, 124–131

  140. [148]

    Chunqiu Zeng, Wubai Zhou, Tao Li, Larisa Shwartz, and Genady Ya Grabarnik. 2017. Knowledge guided hierarchical multi-label classification over ticket data. IEEE Transactions on Network and Service Management 14, 2 (2017), 246–260

  141. [149]

    Zhengran Zeng, Yuqun Zhang, Yong Xu, Minghua Ma, Bo Qiao, Wentao Zou, Qingjun Chen, Meng Zhang, Xu Zhang, Hongyu Zhang, et al. 2023. Traceark: Towards actionable performance anomaly alerting for online service systems. In 2023 IEEE/ACM 45th International Conference on Software...

  142. [150]

    Chenxi Zhang, Xin Peng, Chaofeng Sha, Ke Zhang, Zhenqing Fu, Xiya Wu, Qingwei Lin, and Dongmei Zhang. 2022. Deeptralog: Trace-log combined microservice anomaly detection through graph-based deep learning. In Proceedings of the 44th International Conference on Software Engineer...

  143. [151]

    Dylan Zhang, Xuchao Zhang, Chetan Bansal, Pedro Las-Casas, Rodrigo Fonseca, and Saravan Rajmohan. 2024. LM- PACE: Confidence estimation by large language models for effective root causing of cloud incidents. In Companion Proceedings of the 32nd ACM International Conference on ...

  144. [152]

    Ke Zhang, Chenxi Zhang, Xin Peng, and Chaofeng Sha. 2022. Putracead: Trace anomaly detection with partial labels based on gnn and pu learning. In 2022 IEEE 33rd International Symposium on Software Reliability Engineering (ISSRE) . IEEE, 239–250

  145. [153]

    Lingzhe Zhang, Tong Jia, Mengxi Jia, Ying Li, Yong Yang, and Zhonghai Wu. 2024. Multivariate Log-based Anomaly Detection for Distributed Database. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 4256–4267

  146. [154]

    Lingzhe Zhang, Tong Jia, Mengxi Jia, Hongyi Liu, Yong Yang, Zhonghai Wu, and Ying Li. 2024. Towards Close-To-Zero Runtime Collection Overhead: Raft-Based Anomaly Diagnosis on System Faults for Distributed Storage System. IEEE Transactions on Services Computing (2024)

  147. [155]

    Lingzhe Zhang, Tong Jia, Mengxi Jia, Yifan Wu, Hongyi Liu, and Ying Li. 2025. ScalaLog: Scalable Log-Based Failure Diagnosis Using LLM. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5

  148. [156]

    Lingzhe Zhang, Tong Jia, Mengxi Jia, Yifan Wu, Hongyi Liu, and Ying Li. 2025. XRAGLog: A Resource-Efficient and Context-Aware Log-Based Anomaly Detection Method Using Retrieval-Augmented Generation. InAAAI 2025 Workshop on Preventing and Detecting LLM Misinformation (PDLM)

  149. [157]

    Lingzhe Zhang, Tong Jia, Kangjin Wang, Mengxi Jia, Yong Yang, and Ying Li. 2024. Reducing Events to Augment Log- based Anomaly Detection Models: An Empirical Study. In Proceedings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement . 538–548

  150. [158]

    Lingzhe Zhang, Yunpeng Zhai, Tong Jia, Chiming Duan, Siyu Yu, Jinyang Gao, Bolin Ding, Zhonghai Wu, and Ying Li. 2025. ThinkFL: Self-Refining Failure Localization for Microservice Systems via Reinforcement Fine-Tuning. arXiv preprint arXiv:2504.18776 (2025)

  151. [159]

    Lingzhe Zhang, Yunpeng Zhai, Tong Jia, Xiaosong Huang, Chiming Duan, and Ying Li. 2025. AgentFM: Role-Aware Failure Management for Distributed Databases with LLM-Driven Multi-Agents. arXiv preprint arXiv:2504.06614 (2025)

  152. [160]

    Ling-Zhe Zhang, Xiang-Dong Huang, Yan-Kai Wang, Jia-Lin Qiao, Shao-Xu Song, and Jian-Min Wang. 2024. Time- tired compaction: An elastic compaction scheme for LSM-tree based time-series database. Advanced Engineering Informatics 59 (2024), 102224

  153. [161]

    Mingyang Zhang, Jianfei Chen, Jianyi Liu, Jingchu Wang, Rui Shi, and Hua Sheng. 2022. LogST: Log semi-supervised anomaly detection based on sentence-BERT. In 2022 7th International Conference on Signal and Image Processing (ICSIP). IEEE, 356–361. J. ACM, Vol. 37, No. 4, Articl...

  154. [162]

    Shenglin Zhang, Yuhe Ji, Jiaqi Luan, Xiaohui Nie, Ziang Chen, Minghua Ma, Yongqian Sun, and Dan Pei. 2024. End- to-end automl for unsupervised log anomaly detection. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering. 1680–1692

  155. [163]

    Shenglin Zhang, Sibo Xia, Wenzhao Fan, Binpeng Shi, Xiao Xiong, Zhenyu Zhong, Minghua Ma, Yongqian Sun, and Dan Pei. 2024. Failure Diagnosis in Microservice Systems: A Comprehensive Survey and Analysis. arXiv preprint arXiv:2407.01710 (2024)

  156. [164]

    Shenglin Zhang, Pengtian Zhu, Minghua Ma, Jiagang Wang, Yongqian Sun, Dongwen Li, Jingyu Wang, Qianying Guo, Xiaolei Hua, Lin Zhu, et al. 2024. Enhanced Fine-Tuning of Lightweight Domain-Specific Q&A Model Based on Large Language Models. In 2024 IEEE 35th International Symposi...

  157. [165]

    Ting Zhang, Xin Huang, Wen Zhao, Shaohuang Bian, and Peng Du. 2023. LogPrompt: A Log-based Anomaly Detection Framework Using Prompts. In 2023 International Joint Conference on Neural Networks (IJCNN) . IEEE, 1–8

  158. [166]

    Tianyang Zhang, Zhuoxuan Jiang, Shengguang Bai, Tianrui Zhang, Lin Lin, Yang Liu, and Jiawei Ren. 2024. RAG4ITOps: A Supervised Fine-Tunable and Comprehensive RAG Framework for IT Operations and Maintenance. In Proceedings of the 2024 Conference on Empirical Methods in Natural...

  159. [167]

    Wei Zhang, Xianfu Cheng, Yi Zhang, Jian Yang, Hongcheng Guo, Zhoujun Li, Xiaolin Yin, Xiangyuan Guan, Xu Shi, Liangfan Zheng, et al. 2024. ECLIPSE: Semantic Entropy-LCS for Cross-Lingual Industrial Log Parsing. arXiv preprint arXiv:2405.13548 (2024)

  160. [168]

    Wei Zhang, Hongcheng Guo, Jian Yang, Yi Zhang, Chaoran Yan, Zhoujin Tian, Hangyuan Ji, Zhoujun Li, Tongliang Li, Tieqiao Zheng, et al . 2024. mABC: multi-Agent Blockchain-Inspired Collaboration for root cause analysis in micro-services architecture. arXiv preprint arXiv:2404.1...

  161. [169]

    Wanhao Zhang, Qianli Zhang, Enyu Yu, Yuxiang Ren, Yeqing Meng, Mingxi Qiu, and Jilong Wang. 2024. LogRAG: Semi-Supervised Log-based Anomaly Detection with Retrieval-Augmented Generation. In 2024 IEEE International Conference on Web Services (ICWS) . IEEE, 1100–1102

  162. [170]

    Xuchao Zhang, Supriyo Ghosh, Chetan Bansal, Rujia Wang, Minghua Ma, Yu Kang, and Saravan Rajmohan. 2024. Automated root causing of cloud incidents using in-context learning with gpt-4. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Soft...

  163. [171]

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223 (2023)

  164. [172]

    Aoxiao Zhong, Dengyao Mo, Guiyang Liu, Jinbu Liu, Qingda Lu, Qi Zhou, Jiesheng Wu, Quanzheng Li, and Qingsong Wen. 2024. Logparser-llm: Advancing efficient log parsing with large language models. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data ...

  165. [173]

    Huazhen Zhong, Jibin Wang, Xuejian Wang, Xin Wang, Wenjie Xiao, Xuehai Tang, and Liangjun Zang. 2024. DBPrompt: A Database Anomaly Operation Detection and Analysis via Prompt Learning. In International Conference on Intelligent Computing. Springer, 357–368

  166. [174]

    Tian Zhou, Peisong Niu, Liang Sun, Rong Jin, et al. 2023. One fits all: Power general time series analysis by pretrained lm. Advances in neural information processing systems 36 (2023), 43322–43355

  167. [175]

    Wubai Zhou, Liang Tang, Chunqiu Zeng, Tao Li, Larisa Shwartz, and Genady Ya Grabarnik. 2016. Resolution recommendation for event tickets in service management. IEEE Transactions on Network and Service Management 13, 4 (2016), 954–967

  168. [176]

    Xuanhe Zhou, Lianyuan Jin, Ji Sun, Xinyang Zhao, Xiang Yu, Jianhua Feng, Shifu Li, Tianqing Wang, Kun Li, and Luyang Liu. 2021. Dbmind: A self-driving platform in opengauss. Proceedings of the VLDB Endowment 14, 12 (2021), 2743–2746

  169. [177]

    Xuanhe Zhou, Guoliang Li, Zhaoyan Sun, Zhiyuan Liu, Weize Chen, Jianming Wu, Jiesi Liu, Ruohang Feng, and Guoyang Zeng. 2023. D-bot: Database diagnosis system using large language models. arXiv preprint arXiv:2312.01454 (2023)

  170. [178]

    Xiang Zhou, Xin Peng, Tao Xie, Jun Sun, Chao Ji, Wenhai Li, and Dan Ding. 2018. Fault analysis and debugging of microservice systems: Industrial survey, benchmark system, and empirical study. IEEE Transactions on Software Engineering 47, 2 (2018), 243–260

  171. [179]

    Xuanhe Zhou, Zhaoyan Sun, and Guoliang Li. 2024. Db-gpt: Large language model meets database. Data Science and Engineering 9, 1 (2024), 102–111

  172. [180]

    Xuanhe Zhou, Xinyang Zhao, and Guoliang Li. 2024. LLM-Enhanced Data Management.arXiv preprint arXiv:2402.02643 (2024). J. ACM, Vol. 37, No. 4, Article 111. Publication date: August 2025

  173. [2024]

    arXiv preprint arXiv:2403.04123 (2024)

    Exploring LLM-based Agents for Root Cause Analysis. arXiv preprint arXiv:2403.04123 (2024)

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.