Pith. sign in

REVIEW 8 cited by

Efficient Training of Large Language Models on Distributed Infrastructures: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.20018 v1 pith:XHNNPWFN submitted 2024-07-29 cs.DC

classification cs.DC
keywords trainingsurveycomputingllmsmodelssystemschallengesdistributed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) like GPT and LLaMA are revolutionizing the AI industry with their sophisticated capabilities. Training these models requires vast GPU clusters and significant computing time, posing major challenges in terms of scalability, efficiency, and reliability. This survey explores recent advancements in training systems for LLMs, including innovations in training infrastructure with AI accelerators, networking, storage, and scheduling. Additionally, the survey covers parallelism strategies, as well as optimizations for computation, communication, and memory in distributed LLM training. It also includes approaches of maintaining system reliability over extended training periods. By examining current innovations and future directions, this survey aims to provide valuable insights towards improving LLM training systems and tackling ongoing challenges. Furthermore, traditional digital circuit-based computing systems face significant constraints in meeting the computational demands of LLMs, highlighting the need for innovative solutions such as optical computing and optical networks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall

    cs.DC 2025-09 conditional novelty 7.0 of 10

    MLP-Offload accelerates LLM pre-training on memory-constrained GPUs by mixing local NVMe and remote PFS offloading with cache-aware subgroup reordering, achieving up to 2.5x faster iterations than DeepSpeed ZeRO-3.

  2. CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training

    cs.DC 2025-06 conditional novelty 7.0 of 10

    CrossPipe generates latency- and bandwidth-aware pipeline schedules that cut emulated cross-datacenter LLM training time by up to 33.6%.

  3. Design-CP: Context Parallelism for Design of Protein Nanoparticles

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Context-parallel inference for RFdiffusion 3 enables end-to-end all-atom design of large symmetric protein nanoparticles on multi-GPU hardware without retraining.

  4. SAKURAONE: Empowering Transparent and Open AI Platforms through Private-Sector HPC Investment in Japan

    cs.DC 2025-07 conditional novelty 6.0 of 10

    SAKURAONE, a 100-node, 800-GPU cluster with an open 800GbE SONiC network, reached rank 49 on TOP500 and shows that open Ethernet can scale to a major supercomputer.

  5. The Foundation Cracks: A Comprehensive Study on Bugs and Testing Practices in LLM Libraries

    cs.SE 2025-06 conditional novelty 6.0 of 10

    API misuse is the dominant root cause of bugs in HuggingFace Transformers and vLLM, and most bugs are missed by existing tests because of missing drivers, missing cases, and weak oracles.

  6. NoLoCo: No-all-reduce Low Communication Training Method for Large Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    NoLoCo trains large language models without any all-to-all synchronization by using pairwise weight averaging and random pipeline routing, matching or slightly beating DiLoCo in experiments.

  7. Joint Partitioning and Placement of Foundation Models for Real-Time Edge AI

    cs.DC 2025-11 reject novelty 4.0 of 10

    A framework for runtime re-splitting and re-placement of foundation model layers across edge nodes is proposed, but its claimed latency gains are inherited from prior work rather than measured.

  8. Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead

    cs.SE 2025-06 conditional novelty 4.0 of 10

    A literature review organizes LLM development into a six-phase software engineering lifecycle and identifies challenges and research directions for each phase.

Pith tools