REVIEW 3 major objections 5 minor 69 references
Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
T0 review · 3 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read By sharding communication one level deeper than existing shard-based overlap, this paper shows that data-dependent GPU communication can be overlapped with computation in an all-to-all pattern, delivering up to 1.6x speedup and a heuristic
desk verdict Finer-grain decomposition for compute–communication overlap is a solid experimental idea with real measured speedups on MI300X, but the schedule-picking heuristic is under-validated and sloppily specified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
FiCCO (Finer-grain Compute-Communication Overlap): communication is re-sharded by the number of GPUs, so in an eight-GPU system each transfer is one-eighth the size of a shard-level transfer. This turns communication into an all-to-all pattern and enables a design space of schedules distinguished by computation uniformity (uniform vs. heterogeneous), computation granularity (fused vs. unfused GEMM kernels), and communication shape (1D vs. 2D). The selection heuristic uses a GEMM's OTB and MT, forms their product, compares it against a machine-level threshold with a 5x multiplier, and picks among the schedules.
What would settle it
Run FiCCO on a ring or torus topology, or with GEMM shapes outside the 16 synthetic scenarios, and measure whether the heuristic still picks the winning schedule in roughly 81% of cases; if accuracy drops well below that, the static OTB*MT signature is not portable across topologies.
Extended reading notes
Core claim
The paper's central claim is that finer-grain decomposition of data-dependent communication, one level below shard granularity, converts peer-to-peer transfers into all-to-all transfers that keep direct-connection GPU interconnects busy. The resulting overheads, decomposition and contention inefficiencies, can be characterized using two static GEMM properties: operations per byte (OTB) and memory traffic (MT). From this characterization the paper derives four concrete schedules and a heuristic that selects among them, and offloading communication to GPU DMA engines reduces contention further. This combination is what delivers the reported speedups.
Load-bearing premise
The heuristic assumes that a GEMM's static arithmetic intensity (OTB) and memory traffic (MT), combined as a product and compared with a hand-chosen 5x machine-level threshold, reliably predicts which FiCCO schedule minimizes total inefficiency on unseen workloads and hardware.
Editorial extensions
If this is right
- FiCCO attains up to 1.6x speedup over serial execution across realistic GEMM and all-gather scenarios.
- The proposed heuristic picks the optimal schedule in 81% of unseen synthetic scenarios, with mispredictions losing about 14% of speedup.
- DMA-based communication offload reduces contention inefficiency versus GPU-core-driven communication in all measured cases.
- Shard-based overlap can be slower than serial execution on direct-connection topologies, while FiCCO avoids that degradation.
- The design-space framing gives frameworks and runtimes a concrete mechanism for choosing overlap schedules based on operation shapes.
Reading between the lines
- The heuristic's 5x threshold and OTB*MT product are likely calibrated to the specific GPU and topology studied; on other hardware the demarcation may shift, so the heuristic may need recalibration rather than re-derivation.
- If DMA engines gain support for 2D copies, the emulated 2D schedules should be re-measured; real 2D DMA may close part of the remaining gap to ideal speedup.
- The same one-level-deeper decomposition idea could extend to reduce-scatter-based parallelism, such as tensor parallelism with gradients, once DMA engines support arithmetic operations, which the paper explicitly leaves out.
- The correlation between static GEMM properties and inefficiency signatures suggests that a similar OTB/MT-based selector could be applied to other overlap schemes, not just the four schedules studied here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes FiCCO, a finer-granularity compute-communication overlap scheme that decomposes communication one level deeper than shard-based overlap (e.g., transfer sizes one-eighth of shard-based in an 8-GPU system). This decomposition is argued to unlock all-to-all communication, better utilize direct-connection topologies such as AMD MI300X, and enable a richer space of execution schedules. The paper characterizes two inefficiencies—decomposition inefficiency loss (DIL) and contention inefficiency loss (CIL)—and maps them to static GEMM features (OTB and MT). It then defines four FiCCO schedules, proposes a heuristic based on OTB×MT and a 5× machine-dependent threshold to select among them, and offloads communication to GPU DMA engines. Experiments on 16 real GEMM scenarios report up to 1.6× speedup over serial execution, and an additional 16 synthetic scenarios are used to claim 81% heuristic accuracy.
Significance. If the heuristic generalizes, the paper offers a practical, software-only method for improving dependent compute-communication overlap on full-mesh GPU systems, a setting where existing shard-based overlap degrades. The strengths are the concrete measurements on a real MI300X system, the use of DMA offload to reduce contention, and the attempt to ground schedule selection in static operator properties. The core speedup result is plausible and independent of the heuristic. However, the heuristic's central claim of 81% accuracy on unseen scenarios rests on a small, author-generated validation set, a hand-picked threshold, and a unit-inconsistent decision variable. The 2D communication schedule is also emulated rather than measured. These issues limit the current evidence for transferability, which is the main practical contribution.
major comments (3)
- [V-C and VI-C] The heuristic's decision boundary is not well-founded. Section V-C defines combined OTB×MT for the GEMM and compares it to 'machine-level' OTB×MT defined as 'peak compute FLOPs' (FLOPs/s). This is a unit mismatch: OTB (FLOPs/byte) times MT (bytes) yields FLOPs, so the comparison implicitly imposes a 1-second timescale. The 5× threshold is hand-chosen and appears tuned to the 15 real scenarios in Table I, on which the heuristic then achieves 100% accuracy. The only out-of-sample evaluation is 16 synthetic scenarios with no cross-validation, no sensitivity analysis for the 5× value, and no confidence intervals. With one free parameter and 16 test points, 13/16 correct is weak evidence for the '81% of unseen scenarios' claim. Please provide a principled derivation of the threshold, report accuracy as a function of the threshold, or validate on independently collected workloads/hardware.
- [VI-B and Figure 12b] The 2D schedule is emulated using 1D memory copies of the same size because '2D memory copies with DMAs are not supported today.' The paper still reports 'with emulated 2D schedules we attain as high as 1.7× speedup' and includes uniform-fused-2D in the headline results and heuristic evaluation. A 1D copy does not reproduce the buffer layout, gather/scatter behavior, or link-level traffic pattern of a true 2D communication shape. This means the 2D arm of the design space is not actually evaluated, and the 1.7× number is a best-case estimate, not a measured result. Please either implement true 2D transfers, use a simulator validated against the 1D results, or clearly label these as estimates throughout the abstract, Section V, and Section VI.
- [IV-B and VI] The paper reports speedups and DIL/CIL values as point averages of 5 runs (after 10 warmups) but gives no error bars, standard deviations, or statistical significance tests. This matters because the heuristic is selecting among schedules whose speedups can be close (e.g., in Figure 12b the difference between schedules is often small relative to the reported 6% operator variation). Without variance information, the reader cannot tell whether the 81% accuracy or the 1.6× speedup is robust, or whether the heuristic's choices are within measurement noise. Please report error bars or variance for the headline numbers and for the per-schedule speedups in Figure 12b.
minor comments (5)
- [VI-C] The 16 synthetic scenarios are described only as 'wide ranging OTB and MT combinations.' Please specify how they were generated (parameter ranges, GEMM shapes, whether they include the M<K case), and list them so the out-of-sample claim is reproducible.
- [IV-C] The phrase 'op-to-byte' is used inconsistently; consider using 'OTB' consistently after first definition. Also, in Figure 7 the data labels are small and hard to read; enlarge them or report the values in a table.
- [VI-A] The text says 'we observe 7× communication slowdown' for shard-based overlap on MI300X, but Figure 13 shows speedup, not communication slowdown. Clarify where the 7× comes from or add a direct measurement.
- [IV-B.2] The omitted scenarios (e.g., tensor parallelism with reduce-scatter) are excluded because DMA engines do not support math today. This is a reasonable limitation, but it should be stated earlier and in the abstract or conclusion, since it bounds the applicability of FiCCO.
- [VII] The related work discussion is brief for a design-space paper. In particular, the comparison to Triton-Distributed is reported as out-of-memory; consider adding a small-scale comparison or at least a qualitative discussion of how FiCCO differs in kernel-authoring burden.
Circularity Check
No significant circularity: FiCCO speedups are measured and the heuristic, though in-sample validated, is not defined by the predicted schedule labels.
full rationale
The paper's central speedup claim is an empirical result: FiCCO is implemented and measured on MI300X against serial execution, shard-overlap, RCCL-based FiCCO, and an ideal roofline, with the up-to-1.6x speedup coming from measured execution times (Sections VI-B and VI-D). The DIL/CIL characterization is also measured, and the heuristic in Section V-C selects schedules from static GEMM features (M vs K, OTB, MT) using a decision rule; the rule is not defined in terms of the measured winning schedule, so no predicted quantity is equal to an input by construction. The 5x machine-level threshold is hand-picked rather than derived, and evaluating the heuristic on the same 15 real scenarios is in-sample validation—a generalization limitation, not a definitional circularity. The 16-scenario synthetic test provides a separate, though small, out-of-sample check. Self-citations such as ConCCL [1], T3 [42], and Tale of Two Cs [41] provide background on DMA offload and motivation, but the paper implements DMA via hipMemcpyDtoDAsync and measures its effect directly, so the citations are not load-bearing. No step in the derivation chain reduces a predicted result to its own inputs.
Assumptions & free parameters
free parameters (2)
- 5x threshold in FiCCO heuristic =
5
- OTB*MT product as combined static metric
assumptions (4)
- domain assumption 8-GPU MI300X full-mesh topology is representative of modern GPU systems
- domain assumption DMA-based async memory copies can overlap with GEMM kernels and reduce interference
- domain assumption Static GEMM dimensions (M,N,K) determine OTB and MT, which predict DIL and CIL
- domain assumption The 16 synthetic scenarios are representative of unseen real deployments
Cite this review
Pith. "Pith review of Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap." pith.science (2026). https://pith.science/paper/LXHGMB5Q
@misc{pith2026251210236,
author = {Pith},
title = {Pith review of: Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap},
year = {2026},
howpublished = {\url{https://pith.science/paper/LXHGMB5Q}},
note = {Machine review of arXiv:2512.10236}
}
read the original abstract
Modern ML workloads demand distributing training and inference across multiple GPUs. However, these parallelization techniques often suffer from exposed critical-path communication, leaving a potential 1.7x speedup on the table through compute-communication overlap. Prior overlapping methods harness the fact that ML model state and inputs are already sharded into the number of GPUs, and overlap the compute and communication at shard granularity. However, such coarse-grained overlap suffers from limited network topology support, and suboptimal dataflows. In this work, we instead make a case for finer-grain compute-communication overlap which we term FiCCO. FiCCO operates one level deeper than traditional sharding, and unlocks overlap for a wider set of network topologies and enables finer-grain dataflow. We show that FiCCO opens up a wider design space of execution schedules than possible at shard-level alone. To walk the design space of schedules, we study and characterize the performance inefficiencies on doing overlap and overlay the schedules with the associated inefficiency signatures. Our characterization reveals decomposition and contention based slowdowns to be the major performance limiters, and we correlate the slowdown factors with the static compute/communication operator sizes. This helps us design heuristics (that frameworks and runtimes can harness) to select bespoke FiCCO schedules based on the nature of underlying ML operations. Finally, to further minimize contention inefficiencies inherent with operation overlap, we offload communication to GPU DMA engines. We evaluate several scenarios from realistic ML deployments and demonstrate that our proposed heuristics driven bespoke schedules deliver up to 1.6x speedup. Further, our heuristics provide accurate guidance to pick the optimal schedule in 81% of unseen scenarios.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Conccl: Optimizing ML concurrent computation and communication with GPU DMA engines,
A. Agrawal, S. Aga, S. Pati, and M. Islam, “Conccl: Optimizing ML concurrent computation and communication with GPU DMA engines,” inIEEE International Symposium on Performance Analysis of Systems and Software, ISPASS 2025, Ghent, Belgium, May 11-13, 2025. IEEE, 2025, pp. 1–11. [Online]. Available: https: //doi.org/10.1109/ISPASS64960.2025.00018
arXiv 2025
-
[2]
[Distributed GEMM: A novel CUTLASS-based implementation of Tensor Parallelism for NVLink-enabled systems,
Ali Hassani, Michael Isaev, Nic McDonald, Jie Ren, Vijay Thakkar, Haicheng Wu, and Humphrey Shi, “[Distributed GEMM: A novel CUTLASS-based implementation of Tensor Parallelism for NVLink-enabled systems,” https://blog.shi-labs. com/distributed-gemm-88be6a481e2b, December 2024
2024
-
[3]
(2023) Amd instinct™ mi300x accelerators
AMD. (2023) Amd instinct™ mi300x accelerators. [On- line]. Available: https://www.amd.com/en/products/accelerators/instinct/ mi300/mi300x.html
2023
-
[4]
HIP: C++ Heterogeneous-Compute Interface for Portability,
AMD, “HIP: C++ Heterogeneous-Compute Interface for Portability,” https://github.com/ROCm/HIP, 2024
2024
-
[5]
ROCm Communication Collectives Library (RCCL),
——, “ROCm Communication Collectives Library (RCCL),” https: //github.com/ROCm/rccl, 2024
2024
-
[6]
ROCm: HIPStream,
——, “ROCm: HIPStream,” https://rocm.docs.amd.com/projects/HIP/ en/latest/reference/hip runtime api/modules/stream management.html, 2024
2024
-
[7]
ROCm/rocBLAS: Next generation BLAS implementation for ROCm platform,
——, “ROCm/rocBLAS: Next generation BLAS implementation for ROCm platform,” https://github.com/ROCm/rocBLAS, 2024
2024
-
[8]
(2025) Hip graphs
AMD. (2025) Hip graphs. [Online]. Avail- able: https://rocm.docs.amd.com/projects/HIP/en/docs-develop/how-to/ hip runtime api/hipgraph.html
2025
Show all 69 references
-
[9]
(2025) hipblaslt
——. (2025) hipblaslt. [Online]. Available: https://github.com/ROCm/ rocm-libraries
2025
-
[10]
FLUX: fast software-based communication overlap on gpus through kernel fusion,
L. Chang, W. Bao, Q. Hou, C. Jiang, N. Zheng, Y . Zhong, X. Zhang, Z. Song, Z. Jiang, H. Lin, X. Jin, and X. Liu, “FLUX: fast software-based communication overlap on gpus through kernel fusion,”CoRR, vol. abs/2406.06858, 2024. [Online]. Available: https://doi.org/10.48550/arXi...
-
[11]
Centauri: Enabling efficient scheduling for communication- computation overlap in large model training via communication partitioning,
C. Chen, X. Li, Q. Zhu, J. Duan, P. Sun, X. Zhang, and C. Yang, “Centauri: Enabling efficient scheduling for communication- computation overlap in large model training via communication partitioning,” inProceedings of the 29th ACM International Conference on Architectural Supp...
2024
-
[12]
Evaluating large language models trained on code,
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y . Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter,...
2021 arXiv
-
[13]
Revisiting scaling laws for language models: The role of data quality and training strategies,
Z. Chen, S. Wang, T. Xiao, Y . Wang, S. Chen, X. Cai, J. He, and J. Wang, “Revisiting scaling laws for language models: The role of data quality and training strategies,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long...
2025
-
[14]
Concerto: Automatic communication optimization and scheduling for large-scale deep learning,
S. Cheng, S. Lin, L. Diao, H. Wu, S. Wang, C. Si, Z. Liu, X. Zhao, J. Du, W. Lin, and Y . You, “Concerto: Automatic communication optimization and scheduling for large-scale deep learning,” inProceedings of the 30th ACM International Conference on Architectural Support for Pro...
2025
-
[15]
Scaling llama 3 training with efficient parallelism strategies,
W. Chu, X. Xie, J. Yu, J. Wang, A. Phanishayee, C. Tang, Y . Hao, J. Huang, M. Ozdal, J. Wang, V . Goswami, N. Goyal, A. Kadian, A. Gu, C. Cai, F. Tian, X. Wang, M. Si, P. Balaji, C.-H. Chu, and J. Park, “Scaling llama 3 training with efficient parallelism strategies,” inProce...
2025
-
[16]
Transformations to parallel codes for communication-computation overlap,
A. Danalis, K.-Y . Kim, L. Pollock, and M. Swany, “Transformations to parallel codes for communication-computation overlap,” inSC ’05: Proceedings of the 2005 ACM/IEEE Conference on Supercomputing, 2005, pp. 58–58
2005
-
[17]
Mpi-aware compiler optimizations for improving communication-computation overlap,
A. Danalis, L. Pollock, M. Swany, and J. Cavazos, “Mpi-aware compiler optimizations for improving communication-computation overlap,” inProceedings of the 23rd International Conference on Supercomputing, ser. ICS ’09. New York, NY , USA: Association for Computing Machinery, 20...
2009
-
[18]
Flashattention: Fast and memory-efficient exact attention with io-awareness,
T. Dao, D. Y . Fu, S. Ermon, A. Rudra, and C. R ´e, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” inAdvances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans...
2022
-
[19]
Deepseek-v3 technical report,
DeepSeek-AI, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, ...
-
[20]
The llama 3 herd of models,
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al- Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru...
-
[21]
Compiler-assisted overlapping of communication and computation in mpi applications,
J. Guo, Q. Yi, J. Meng, J. Zhang, and P. Balaji, “Compiler-assisted overlapping of communication and computation in mpi applications,” in 2016 IEEE International Conference on Cluster Computing (CLUSTER), 2016, pp. 60–69
2016
-
[22]
Bandwidth characterization of deepspeed on distributed large language model training,
B. Hanindhito, B. Patel, and L. K. John, “Bandwidth characterization of deepspeed on distributed large language model training,” in IEEE International Symposium on Performance Analysis of Systems and Software, ISPASS 2024, Indianapolis, IN, USA, May 5- 7, 2024. IEEE, 2024, pp....
2024
-
[23]
Efficient and adaptable overlapping for computation and communication via signaling and reordering,
K. Hong, X. Li, M. Liu, Q. Mao, T. Wu, Z. Huang, L. Chen, Z. Wang, Y . Zhang, Z. Zhu, G. Dai, and Y . Wang, “Efficient and adaptable overlapping for computation and communication via signaling and reordering,” 2025
2025
-
[24]
[Distributed w/ TorchTitan] Introducing Async Tensor Parallelism in PyTorch,
Horace He, Less Wright, Luca Wehrstedt, Tianyu Liu, Wanchao Liang, “[Distributed w/ TorchTitan] Introducing Async Tensor Parallelism in PyTorch,” https://discuss.pytorch.org/t/ distributed-w-torchtitan-introducing-async-tensor-parallelism-in-pytorch/ 209487, September 2024
2024
-
[25]
Demystifying nccl: An in-depth analysis of gpu communication protocols and algorithms,
Z. Hu, S. Shen, T. Bonato, S. Jeaugey, C. Alexander, E. Spada, J. Dinan, J. Hammond, and T. Hoefler, “Demystifying nccl: An in-depth analysis of gpu communication protocols and algorithms,”
-
[26]
Gpipe: Efficient training of giant neural networks using pipeline parallelism,
Y . Huang, Y . Cheng, A. Bapna, O. Firat, D. Chen, M. X. Chen, H. Lee, J. Ngiam, Q. V . Le, Y . Wu, and Z. Chen, “Gpipe: Efficient training of giant neural networks using pipeline parallelism,” inAdvances in Neural Information Processing Systems 32: Annual Conference on Neural...
2019
-
[27]
A loop transformation algorithm for communication overlapping,
K. Ishizaki, H. Komatsu, and T. Nakatani, “A loop transformation algorithm for communication overlapping,”Int. J. Parallel Program., vol. 28, no. 2, p. 135–154, Apr. 2000. [Online]. Available: https://doi.org/10.1023/A:1007554715418
-
[28]
Breaking the computation and communication abstraction barrier in distributed machine learning workloads,
A. Jangda, J. Huang, G. Liu, A. H. N. Sabet, S. Maleki, Y . Miao, M. Musuvathi, T. Mytkowicz, and O. Saarikivi, “Breaking the computation and communication abstraction barrier in distributed machine learning workloads,” inProceedings of the 27th ACM International Conference on...
2022
-
[29]
Available: https://arxiv.org/abs/2507.04786
[Online]. Available: https://arxiv.org/abs/2507.04786
-
[30]
A survey on large language models for code generation,
J. Jiang, F. Wang, J. Shen, S. Kim, and S. Kim, “A survey on large language models for code generation,”CoRR, vol. abs/2406.00515,
-
[31]
Reducing activation recomputation in large transformer models,
V . A. Korthikanti, J. Casper, S. Lym, L. McAfee, M. Andersch, M. Shoeybi, and B. Catanzaro, “Reducing activation recomputation in large transformer models,” inProceedings of the Sixth Conference on Machine Learning and Systems, MLSys 2023, Miami, FL, USA, June 4-8, 2023, D. S...
2023
-
[32]
Lit silicon: A case where thermal imbalance couples concurrent execution in multiple gpus,
M. Kurzynski, S. Aga, and D. Wu, “Lit silicon: A case where thermal imbalance couples concurrent execution in multiple gpus,”
-
[33]
Mixtral of experts,
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de Las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T. L...
-
[34]
Pytorch 11 distributed: Experiences on accelerating data parallel training,
S. Li, Y . Zhao, R. Varma, O. Salpekar, P. Noordhuis, T. Li, A. Paszke, J. Smith, B. Vaughan, P. Damania, and S. Chintala, “Pytorch 11 distributed: Experiences on accelerating data parallel training,”Proc. VLDB Endow., vol. 13, no. 12, pp. 3005–3018, 2020. [Online]. Available:...
2020
- [35]
-
[36]
Ring attention with blockwise transformers for near-infinite context,
——, “Ring attention with blockwise transformers for near-infinite context,” 2023. [Online]. Available: https://arxiv.org/abs/2310.01889
2023 arXiv
-
[37]
MLPerf Inference Results v5.0,
MLCommons, “MLPerf Inference Results v5.0,” https://github.com/ mlcommons/inference results v5.0, 2025, accessed: 2025-11-16
2025
-
[38]
Available: https://arxiv.org/abs/2511.09861
[Online]. Available: https://arxiv.org/abs/2511.09861
-
[39]
Gshard: Scaling giant models with conditional computation and automatic sharding,
D. Lepikhin, H. Lee, Y . Xu, D. Chen, O. Firat, Y . Huang, M. Krikun, N. Shazeer, and Z. Chen, “Gshard: Scaling giant models with conditional computation and automatic sharding,” in9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May...
2021
-
[40]
Pytorch: An imperative style, high-performance deep learning library,
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. K ¨opf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-...
2019 arXiv
-
[42]
T3: Transparent tracking & triggering for fine-grained overlap of compute & collectives,
——, “T3: Transparent tracking & triggering for fine-grained overlap of compute & collectives,” inProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, ser. ASPLOS ’24. New York, NY , USA: Associ...
2024
-
[43]
Exact dependence analysis for increased communication overlap,
S. Pellegrini, T. Hoefler, and T. Fahringer, “Exact dependence analysis for increased communication overlap,” inRecent Advances in the Message Passing Interface, J. L. Tr ¨aff, S. Benkner, and J. J. Dongarra, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2012, pp. 89–99
2012
-
[44]
Automatic gen- eration of software pipelines for heterogeneous parallel systems,
J. A. Pienaar, S. Chakradhar, and A. Raghunathan, “Automatic gen- eration of software pipelines for heterogeneous parallel systems,” in Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, ser. SC ’12. Washington, DC, USA...
2012
-
[45]
Efficient large-scale language model training on gpu clusters using megatron-lm,
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V . Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro, A. Phanishayee, and M. Zaharia, “Efficient large-scale language model training on gpu clusters using megatron-lm,” inProceedings of the Inte...
2021
-
[46]
Stream-k: Work-centric parallel decomposition for dense matrix- matrix multiplication on the GPU,
M. Osama, D. Merrill, C. Cecka, M. Garland, and J. D. Owens, “Stream-k: Work-centric parallel decomposition for dense matrix- matrix multiplication on the GPU,” inProceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming, PPoPP 202...
2023
-
[47]
Enabling compute-communication overlap in distributed deep learning training platforms,
S. Rashidi, M. Denton, S. Sridharan, S. Srinivasan, A. Suresh, J. Nie, and T. Krishna, “Enabling compute-communication overlap in distributed deep learning training platforms,” in48th ACM/IEEE Annual International Symposium on Computer Architecture, ISCA 2021, Virtual Event / ...
2021
-
[48]
Tale of two cs: Computation vs. communication scaling for future transformers on future hardware,
S. Pati, S. Aga, M. Islam, N. Jayasena, and M. D. Sinclair, “Tale of two cs: Computation vs. communication scaling for future transformers on future hardware,” inIEEE International Symposium on Workload Characterization, IISWC 2023, Ghent, Belgium, October 1-3, 2023. IEEE, 202...
2023
-
[49]
Slechta, N
B. Slechta, N. Comly, A. Eassa, J. DeLaere, and S. Raj. (2024, Aug) Nvidia nvlink and nvidia nvswitch supercharge large language model inference. NVIDIA Technical Blog. [Online]. Available: https://developer.nvidia.com/blog/ nvidia-nvlink-and-nvidia-nvswitch-supercharge-large-...
2024
-
[50]
Triton: an intermediate language and compiler for tiled neural network computations,
P. Tillet, H. Kung, and D. D. Cox, “Triton: an intermediate language and compiler for tiled neural network computations,” in Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, MAPL@PLDI 2019, Phoenix, AZ, USA, June 22, 2019...
2019
-
[51]
Llama: Open and efficient foundation language models,
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” 2023. [Online]. Available: https: //arxiv.org/abs/2...
2023 arXiv
-
[52]
Optimizing distributed ML communication with fused computation-collective operations,
K. Punniyamurthy, K. Hamidouche, and B. M. Beckmann, “Optimizing distributed ML communication with fused computation-collective operations,” inProceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis, SC 2024, Atlanta, GA, U...
2024 arXiv
-
[53]
Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation AI scale,
S. Rajbhandari, C. Li, Z. Yao, M. Zhang, R. Y . Aminabadi, A. A. Awan, J. Rasley, and Y . He, “Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation AI scale,” in International Conference on Machine Learning, ICML 2022, 17-23 July 2022, B...
2022
-
[54]
Domino: Eliminating communication in LLM training via generic tensor slicing and overlapping,
G. Wang, C. Zhang, Z. Shen, A. Li, and O. Ruwase, “Domino: Eliminating communication in LLM training via generic tensor slicing and overlapping,”CoRR, vol. abs/2409.15241, 2024. [Online]. Available: https://doi.org/10.48550/arXiv.2409.15241
-
[55]
A hybrid tensor-expert-data parallelism approach to optimize mixture-of-experts training,
S. Singh, O. Ruwase, A. A. Awan, S. Rajbhandari, Y . He, and A. Bhatele, “A hybrid tensor-expert-data parallelism approach to optimize mixture-of-experts training,” inProceedings of the 37th International Conference on Supercomputing, ICS 2023, Orlando, FL, USA, June 21-23, 20...
2023
-
[56]
Pytorch symmetricmemory: Harnessing nvlink programmability with ease,
Y . Wang, H. He, and L. Wehrstedt, “Pytorch symmetricmemory: Harnessing nvlink programmability with ease,” Feb 2025, pyTorch Developer Forum. [Online]. Available: https://dev-discuss.pytorch.org/t/ pytorch-symmetricmemory-harnessing-nvlink-programmability-with-ease/ 2798
2025
-
[57]
Petuum: A new platform for distributed machine learning on big data,
E. P. Xing, Q. Ho, W. Dai, J. K. Kim, J. Wei, S. Lee, X. Zheng, P. Xie, A. Kumar, and Y . Yu, “Petuum: A new platform for distributed machine learning on big data,”IEEE Trans. Big Data, vol. 1, no. 2, pp. 49–67, 2015. [Online]. Available: https://doi.org/10.1109/TBDATA.2015.2472014
2015
-
[58]
Context parallelism for scalable million-token inference,
A. Yang, J. Yang, A. Ibrahim, X. Xie, B. Tang, G. Sizov, J. Reizenstein, J. Park, and J. Huang, “Context parallelism for scalable million-token inference,”CoRR, vol. abs/2411.01783, 2024. [Online]. Available: https://doi.org/10.48550/arXiv.2411.01783 12
-
[59]
Llama 2: Open foundation and fine-tuned chat models,
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V . Goswami, N. Goyal, A. Hartshorn, S. Ho...
2023 arXiv
-
[60]
Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models,
P. Vaithilingam, T. Zhang, and E. L. Glassman, “Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models,” inCHI ’22: CHI Conference on Human Factors in Computing Systems, New Orleans, LA, USA, 29 April 2022 - 5 May 2022, E...
2022
-
[61]
Triton-distributed: Programming overlapping kernels on distributed ai systems with the triton compiler,
S. Zheng, W. Bao, Q. Hou, X. Zheng, J. Fang, C. Huang, T. Li, H. Duanmu, R. Chen, R. Xu, Y . Guo, N. Zheng, Z. Jiang, X. Di, D. Wang, J. Ye, H. Lin, L.-W. Chang, L. Lu, Y . Liang, J. Zhai, and X. Liu, “Triton-distributed: Programming overlapping kernels on distributed ai syste...
2025 arXiv
-
[62]
Overlap communication with dependent computation via decomposition in large deep learning models,
S. Wang, J. Wei, A. Sabne, A. Davis, B. Ilbeyi, B. Hechtman, D. Chen, K. S. Murthy, M. Maggioni, Q. Zhang, S. Kumar, T. Guo, Y . Xu, and Z. Zhou, “Overlap communication with dependent computation via decomposition in large deep learning models,” inProceedings of the 28th ACM I...
2023
-
[66]
Comet: Fine-grained computation-communication overlapping for mixture- of-experts,
S. Zhang, N. Zheng, H. Lin, Z. Jiang, W. Bao, C. Jiang, Q. Hou, W. Cui, S. Zheng, L. Chang, Q. Chen, and X. Liu, “Comet: Fine-grained computation-communication overlapping for mixture- of-experts,”CoRR, vol. abs/2502.19811, 2025. [Online]. Available: https://doi.org/10.48550/a...
-
[67]
Pytorch FSDP: experiences on scaling fully sharded data parallel,
Y . Zhao, A. Gu, R. Varma, L. Luo, C. Huang, M. Xu, L. Wright, H. Shojanazeri, M. Ott, S. Shleifer, A. Desmaison, C. Balioglu, P. Damania, B. Nguyen, G. Chauhan, Y . Hao, A. Mathews, and S. Li, “Pytorch FSDP: experiences on scaling fully sharded data parallel,” Proc. VLDB Endo...
2023
-
[69]
Tilelink: Generating efficient compute-communication overlapping kernels using tile-centric primitives,
S. Zheng, J. Fang, X. Zheng, Q. Hou, W. Bao, N. Zheng, Z. Jiang, D. Wang, J. Ye, H. Lin, L.-W. Chang, and X. Liu, “Tilelink: Generating efficient compute-communication overlapping kernels using tile-centric primitives,” inEighth Conference on Machine Learning and Systems,
-
[70]
Available: https://openreview.net/forum?id=ccjvBkTRRe 13
[Online]. Available: https://openreview.net/forum?id=ccjvBkTRRe 13
-
[2022]
Available: http://papers.nips.cc/paper files/paper/2022/ hash/67d57c32e20fd0a7a302cb81d36e40d5-Abstract-Conference.html
[Online]. Available: http://papers.nips.cc/paper files/paper/2022/ hash/67d57c32e20fd0a7a302cb81d36e40d5-Abstract-Conference.html
2022
- [2023]
-
[2024]
Available: https://arxiv.org/abs/2407.21783
[Online]. Available: https://arxiv.org/abs/2407.21783
-
[2025]
Available: https://arxiv.org/abs/2412.19437
[Online]. Available: https://arxiv.org/abs/2412.19437
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.