REVIEW 5 cited by
Cloud Atlas: Efficient Fault Localization for Cloud Systems using Language Models and Causal Insight
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Runtime failure and performance degradation is commonplace in modern cloud systems. For cloud providers, automatically determining the root cause of incidents is paramount to ensuring high reliability and availability as prompt fault localization can enable faster diagnosis and triage for timely resolution. A compelling solution explored in recent work is causal reasoning using causal graphs to capture relationships between varied cloud system performance metrics. To be effective, however, systems developers must correctly define the causal graph of their system, which is a time-consuming, brittle, and challenging task that increases in difficulty for large and dynamic systems and requires domain expertise. Alternatively, automated data-driven approaches have limited efficacy for cloud systems due to the inherent rarity of incidents. In this work, we present Atlas, a novel approach to automatically synthesizing causal graphs for cloud systems. Atlas leverages large language models (LLMs) to generate causal graphs using system documentation, telemetry, and deployment feedback. Atlas is complementary to data-driven causal discovery techniques, and we further enhance Atlas with a data-driven validation step. We evaluate Atlas across a range of fault localization scenarios and demonstrate that Atlas is capable of generating causal graphs in a scalable and generalizable manner, with performance that far surpasses that of data-driven algorithms and is commensurate to the ground-truth baseline.
Forward citations
Cited by 5 Pith papers
-
Enabling Multi-Dimensional Distributed Trace Comparison with Contrast
A mergeable trace-summary representation (TPO) lets operators dynamically compare arbitrary groups of traces across structural, timing, critical-path, and semantic dimensions, powering a visual and an LLM-based compar...
-
Root Cause Analysis of Outliers in Unknown Cyclic Graphs
Applying the inverse covariance (precision) matrix of normal data to one anomalous sample exposes the root cause plus its cycle-involved parents in linear cyclic systems, without knowing the graph.
-
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training
Mycroft adds collective-communication-level tracing to NCCL so that slow or stuck data transfers in LLM training can be detected and traced to likely faulty ranks in seconds.
-
Generating representative macrobenchmark microservice systems from distributed traces with Palette
Palette generates representative microservice benchmark systems from distributed traces using a topology built from a directed graph, a probabilistic automaton, and a graphical causal model.
-
A Survey of AIOps in the Era of Large Language Models
A systematic survey that categorizes LLM-based AIOps research into four dimensions: data sources, tasks, methods, and evaluation, claiming to be the first comprehensive such overview.
Discussion (0). Continue with ORCID to comment.