Pith. sign in

REVIEW 1 cited by

Resource Allocation and Workload Scheduling for Large-Scale Distributed Deep Learning: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08115 v1 pith:NQOSFHZY submitted 2024-06-12 cs.DC cs.AI

classification cs.DCcs.AI
keywords distributeddeeplearningresourceschedulinglarge-scaleworkloadallocation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With rapidly increasing distributed deep learning workloads in large-scale data centers, efficient distributed deep learning framework strategies for resource allocation and workload scheduling have become the key to high-performance deep learning. The large-scale environment with large volumes of datasets, models, and computational and communication resources raises various unique challenges for resource allocation and workload scheduling in distributed deep learning, such as scheduling complexity, resource and workload heterogeneity, and fault tolerance. To uncover these challenges and corresponding solutions, this survey reviews the literature, mainly from 2019 to 2024, on efficient resource allocation and workload scheduling strategies for large-scale distributed DL. We explore these strategies by focusing on various resource types, scheduling granularity levels, and performance goals during distributed training and inference processes. We highlight critical challenges for each topic and discuss key insights of existing technologies. To illustrate practical large-scale resource allocation and workload scheduling in real distributed deep learning scenarios, we use a case study of training large language models. This survey aims to encourage computer science, artificial intelligence, and communications researchers to understand recent advances and explore future research directions for efficient framework strategies for large-scale distributed deep learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Secure Resource Allocation via Constrained Deep Reinforcement Learning

    cs.LG 2025-01 reject novelty 3.0 of 10

    A deep Q-network with a fixed deadline penalty is claimed to cut simulated system cost by up to 40% and energy use by 41.5% in serverless multi-cloud offloading.

Pith tools