Pith. sign in

REVIEW 2 cited by

A Survey of Methods for Collective Communication Optimization and Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1611.06334 v1 pith:2WJZDKOX submitted 2016-11-19 cs.DC

classification cs.DC
keywords collectivecommunicationoptimizationtuningmanybecomecomputingmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

New developments in HPC technology in terms of increasing computing power on multi/many core processors, high-bandwidth memory/IO subsystems and communication interconnects, pose a direct impact on software and runtime system development. These advancements have become useful in producing high-performance collective communication interfaces that integrate efficiently on a wide variety of platforms and environments. However, number of optimization options that shows up with each new technology or software framework has resulted in a \emph{combinatorial explosion} in feature space for tuning collective parameters such that finding the optimal set has become a nearly impossible task. Applicability of algorithmic choices available for optimizing collective communication depends largely on the scalability requirement for a particular usecase. This problem can be further exasperated by any requirement to run collective problems at very large scales such as in the case of exascale computing, at which impractical tuning by brute force may require many months of resources. Therefore application of statistical, data mining and artificial Intelligence or more general hybrid learning models seems essential in many collectives parameter optimization problems. We hope to explore current and the cutting edge of collective communication optimization and tuning methods and culminate with possible future directions towards this problem.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On Topology's Role in ML Training Performance

    cs.NI 2026-08 conditional novelty 6.0 of 10

    For ML training collectives, a Clos network usually beats a torus except for AllGather, where the torus's higher per-node access bandwidth wins.

  2. Terabyte-Scale Analytics in the Blink of an Eye

    cs.DB 2025-06 conditional novelty 6.0 of 10

    Distributed TQP, a GPU-accelerated SQL engine using NCCL and RCCL collectives, runs the full TPC-H 1TB workload in 0.53s on 40 H100 GPUs, more than 60x faster than a high-end CPU server.

Pith tools