REVIEW 4 cited by
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Reliability is a fundamental challenge in operating large-scale machine learning (ML) infrastructures, particularly as the scale of ML models and training clusters continues to grow. Despite decades of research on infrastructure failures, the impact of job failures across different scales remains unclear. This paper presents a view of managing two large, multi-tenant ML clusters, providing quantitative analysis, operational experience, and our own perspective in understanding and addressing reliability concerns at scale. Our analysis reveals that while large jobs are most vulnerable to failures, smaller jobs make up the majority of jobs in the clusters and should be incorporated into optimization objectives. We identify key workload properties, compare them across clusters, and demonstrate essential reliability requirements for pushing the boundaries of ML training at scale. We hereby introduce a taxonomy of failures and key reliability metrics, analyze 11 months of data from two state-of-the-art ML environments with 4 million jobs and over 150 million A100 GPU hours. Building on our data, we fit a failure model to project Mean Time to Failure for various GPU scales. We further propose a method to estimate a related metric, Effective Training Time Ratio, as a function of job parameters, and we use this model to gauge the efficacy of potential software mitigations at scale. Our work provides valuable insights and future research directions for improving the reliability of AI supercomputer clusters, emphasizing the need for flexible, workload-agnostic, and reliability-aware infrastructure, system software, and algorithms.
Forward citations
Cited by 4 Pith papers
-
On Topology's Role in ML Training Performance
For ML training collectives, a Clos network usually beats a torus except for AllGather, where the torus's higher per-node access bandwidth wins.
-
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training
Mycroft adds collective-communication-level tracing to NCCL so that slow or stuck data transfers in LLM training can be detected and traced to likely faulty ranks in seconds.
-
NIXT: A NCCL Inspector Exporter Tool for Observability of Collective Communication in Large Model Training
NIXT is a relational exporter and analysis layer for NCCL Inspector data that turns high-volume profiler logs into summary statistics and correlation queries, validated on Nemotron-4 training up to 2,048 GPUs.
-
Evolving HPC services to enable ML workloads on HPE Cray EX
CSCS proposes seven service enhancements for ML workloads on the Alps supercomputer without quantitative validation.
Discussion (0). Continue with ORCID to comment.