Pith. sign in

REVIEW 7 cited by

MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.15627 v1 pith:XR2UHLWN submitted 2024-02-23 cs.LG cs.DC

classification cs.LGcs.DC
keywords trainingmodelexperiencegpuslargemegascalescalestability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present the design, implementation and engineering experience in building and deploying MegaScale, a production system for training large language models (LLMs) at the scale of more than 10,000 GPUs. Training LLMs at this scale brings unprecedented challenges to training efficiency and stability. We take a full-stack approach that co-designs the algorithmic and system components across model block and optimizer design, computation and communication overlapping, operator optimization, data pipeline, and network performance tuning. Maintaining high efficiency throughout the training process (i.e., stability) is an important consideration in production given the long extent of LLM training jobs. Many hard stability issues only emerge at large scale, and in-depth observability is the key to address them. We develop a set of diagnosis tools to monitor system components and events deep in the stack, identify root causes, and derive effective techniques to achieve fault tolerance and mitigate stragglers. MegaScale achieves 55.2% Model FLOPs Utilization (MFU) when training a 175B LLM model on 12,288 GPUs, improving the MFU by 1.34x compared to Megatron-LM. We share our operational experience in identifying and fixing failures and stragglers. We hope by articulating the problems and sharing our experience from a systems perspective, this work can inspire future LLM systems research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 24 citations worldwide. Full citation record

  1. InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching Transceivers

    cs.NI 2025-02 conditional novelty 8.0 of 10

    InfiniteHBD embeds optical circuit switching inside each transceiver to build reconfigurable ring networks for GPU clusters, claiming node-level fault isolation at roughly one-third the cost of NVL-72.

  2. The Cost and Network Limits of Space-Based AI Compute

    cs.DC 2026-07 conditional novelty 6.0 of 10

    Orbital laser-mesh networks have ~10,000x less bisection bandwidth than terrestrial Clos networks, making LEO training of frontier LLMs 100x+ more expensive while single-satellite inference remains plausible.

  3. BlueLM-2.5-3B Technical Report

    cs.AI 2025-07 conditional novelty 5.0 of 10

    BlueLM-2.5-3B is a small multimodal model with a switchable thinking mode that reportedly matches larger models like Qwen3-4B and comes close to Kimi-VL-A3B-16B on many benchmarks.

  4. DeepCEE: Efficient Cross-Region Model Distributed Training System under Heterogeneous GPUs and Networks

    eess.SY 2025-05 conditional novelty 5.0 of 10

    DeepCEE groups heterogeneous GPUs by network and compute speed, schedules a compact zero-bubble pipeline across regions, and adapts micro-batch sizes to network fluctuations, reporting 1.3-2.8x higher training through...

  5. Goku: Flow Based Video Generative Foundation Models

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A joint image-video generation model family reports state-of-the-art benchmark scores using rectified flow transformers, with all key evidence self-reported and no artifacts released.

  6. Towards Experiment Execution in Support of Community Benchmark Workflows for HPC

    cs.DC 2025-07 reject novelty 4.0 of 10

    The paper proposes workflow templates and experiment management as key to simpler HPC benchmarking, but validates this only through the authors' own two tools.

  7. Evolving HPC services to enable ML workloads on HPE Cray EX

    cs.DC 2025-07 unverdicted novelty 4.0 of 10

    CSCS proposes seven service enhancements for ML workloads on the Alps supercomputer without quantitative validation.

Pith tools