Pith. sign in

REVIEW 2 cited by

RCAgent: Cloud Root Cause Analysis by Autonomous Agents with Tool-Augmented Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.16340 v3 pith:ANEI2YAX submitted 2023-10-25 cs.SE cs.CL

classification cs.SEcs.CL
keywords rcagentanalysiscloudrootautonomousbeencausecurrent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language model (LLM) applications in cloud root cause analysis (RCA) have been actively explored recently. However, current methods are still reliant on manual workflow settings and do not unleash LLMs' decision-making and environment interaction capabilities. We present RCAgent, a tool-augmented LLM autonomous agent framework for practical and privacy-aware industrial RCA usage. Running on an internally deployed model rather than GPT families, RCAgent is capable of free-form data collection and comprehensive analysis with tools. Our framework combines a variety of enhancements, including a unique Self-Consistency for action trajectories, and a suite of methods for context management, stabilization, and importing domain knowledge. Our experiments show RCAgent's evident and consistent superiority over ReAct across all aspects of RCA -- predicting root causes, solutions, evidence, and responsibilities -- and tasks covered or uncovered by current rules, as validated by both automated metrics and human evaluations. Furthermore, RCAgent has already been integrated into the diagnosis and issue discovery workflow of the Real-time Compute Platform for Apache Flink of Alibaba Cloud.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning

    cs.SE 2026-07 conditional novelty 6.0 of 10

    Even when RAG-based LLMs identify the right root-cause service 91–99% of the time, their recovery-action validity stays only 37–60% on a 302-incident Kubernetes benchmark.

  2. Simplifying Root Cause Analysis in Kubernetes with StateGraph and LLM

    cs.DC 2025-06 conditional novelty 6.0 of 10

    SynergyRCA uses GPT-4o and a graph database of Kubernetes entity states to identify root causes of cluster incidents, reporting about 90 percent precision on two production clusters.

Pith tools