REVIEW 2 cited by
Cephalo: Harnessing Heterogeneous GPU Clusters for Training Transformer Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Training transformer models requires substantial GPU compute and memory resources. In homogeneous clusters, distributed strategies allocate resources evenly, but this approach is inefficient for heterogeneous clusters, where GPUs differ in power and memory. As high-end GPUs are costly and limited in availability, heterogeneous clusters with diverse GPU types are becoming more common. Existing methods attempt to balance compute across GPUs based on capacity but often underutilize compute due to memory constraints. We present Cephalo, a system that optimizes compute and memory usage by decoupling compute distribution from training state assignment. Cephalo outperforms state-of-the-art methods by achieving significantly higher training throughput while supporting larger models and batch sizes.
Forward citations
Cited by 2 Pith papers
-
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
Zorse integrates interleaved pipeline parallelism, ZeRO-2 data parallelism, and CPU offloading to accelerate LLM training on heterogeneous GPU clusters by up to 4x.
-
ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing
ViFusion combines dynamic tensor fusion with hierarchical AllReduce to speed up distributed video feature indexing, but the 8-22x throughput claim is an overstatement of bandwidth gains over a self-defined baseline.
Discussion (0). Continue with ORCID to comment.