Pith. sign in

REVIEW 2 cited by

Constructing Large-Scale Real-World Benchmark Datasets for AIOps

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.03938 v1 pith:CNJH6ECM submitted 2022-08-08 cs.SE cs.PF

classification cs.SEcs.PF
keywords aiopsdatasetslarge-scalereal-worldanomalybeencausecompetitions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, AIOps (Artificial Intelligence for IT Operations) has been well studied in academia and industry to enable automated and effective software service management. Plenty of efforts have been dedicated to AIOps, including anomaly detection, root cause localization, incident management, etc. However, most existing works are evaluated on private datasets, so their generality and real performance cannot be guaranteed. The lack of public large-scale real-world datasets has prevented researchers and engineers from enhancing the development of AIOps. To tackle this dilemma, in this work, we introduce three public real-world, large-scale datasets about AIOps, mainly aiming at KPI anomaly detection, root cause localization on multi-dimensional data, and failure discovery and diagnosis. More importantly, we held three competitions in 2018/2019/2020 based on these datasets, attracting thousands of teams to participate. In the future, we will continue to publish more datasets and hold competitions to promote the development of AIOps further.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Argos: Agentic Time-Series Anomaly Detection with Autonomous Rule Generation via Large Language Models

    cs.LG 2025-01 reject novelty 7.0 of 10

    ARGOS uses LLM agents to generate explainable, reproducible anomaly detection rules and fuses them with a base detector, reporting higher F1 than deep-learning and LLM baselines on KPI, Yahoo, and a Microsoft internal...

  2. RCAEval: A Benchmark for Root Cause Analysis of Microservice Systems with Telemetry Data

    cs.SE 2024-12 conditional novelty 6.0 of 10

    RCAEval provides three telemetry datasets with 735 microservice failure cases and an evaluation framework with 15 baselines for metric-based, trace-based, and multi-source root cause analysis.

Pith tools