Pith. sign in

REVIEW 2 cited by

Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.00801 v1 pith:HEMCW3YW submitted 2024-10-01 cs.DC

classification cs.DC
keywords gpuscommunicationmemorysystemsmulti-gpuapplicationsdataevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern GPU systems are constantly evolving to meet the needs of computing-intensive applications in scientific and machine learning domains. However, there is typically a gap between the hardware capacity and the achievable application performance. This work aims to provide a better understanding of the Infinity Fabric interconnects on AMD GPUs and CPUs. We propose a test and evaluation methodology for characterizing the performance of data movements on multi-GPU systems, stressing different communication options on AMD MI250X GPUs, including point-to-point and collective communication, and memory allocation strategies between GPUs, as well as the host CPU. In a single-node setup with four GPUs, we show that direct peer-to-peer memory accesses between GPUs and utilization of the RCCL library outperform MPI-based solutions in terms of memory/communication latency and bandwidth. Our test and evaluation method serves as a base for validating memory and communication strategies on a system and improving applications on AMD multi-GPU computing systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications

    cs.DC 2025-07 conditional novelty 5.0 of 10

    Overlapping computation and communication in distributed GPU training slows compute kernels by up to 40% and raises power use, while remaining faster than sequential execution.

  2. Apple vs. Oranges: Evaluating the Apple Silicon M-Series SoCs for HPC Performance and Efficiency

    cs.AR 2025-02 conditional novelty 5.0 of 10

    Apple Silicon M-Series chips achieve up to 2.9 FP32 TFLOPS and over 200 GFLOPS per watt, making them energy-efficient but low-absolute-performance HPC options.

Pith tools