REVIEW 4 major objections 5 minor 1 cited by
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Full-stack search of distributed ML systems beats isolated-stack tuning by up to 48x.
desk verdict A useful framework with a misleading headline: the 1.50–48.41x claim is measured against a per-bandwidth objective that full-stack search trivially wins by cutting bandwidth to the floor. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing abstraction is Parameter Set Architecture (PSA), a schema analogous to an instruction set architecture in that it defines the interface between a search agent and the underlying system. PSA has three parts: a parameter set covering workload, collective, network, and compute knobs; value ranges for each parameter; and constraints capturing cross-parameter dependencies such as the product of parallelization dimensions being bounded by the number of NPUs. A Parameter Set Scheduler reads this schema and automatically configures both the simulator environment (action space, observation space, constraints) and the search agent (parameter types, step sizes, reward function), so that a domain expert never touches agent internals. This decoupling is what allows the same framework to run workload-only, collective-only, network-only, or full-stack searches and to swap in different ML agents without code changes.
What would settle it
Run the paper's best full-stack configuration and its best single-stack configuration for the same model on a real 512-, 1024-, or 2048-NPU cluster, or at least on full-model simulation without layer truncation, and measure end-to-end training time; if the full-stack configuration does not maintain a comparable advantage over the best single-stack configuration, the central claim fails.
Extended reading notes
Core claim
The discovery is that cross-stack interactions dominate distributed-ML performance: the configurations found by full-stack optimization are qualitatively different from, and reliably better than, the ones found by optimizing any single stack. In the paper's experiments on GPT3-175B, restricting the search to workload-only, collective-only, or network-only knobs yields training runtimes 1.50-48.41x worse per bandwidth/NPU and 3.94-127.17x worse per network dollar cost than the full-stack result on 512- and 1024-NPU systems. On a 2048-NPU cluster, full-stack optimization beats workload-only optimization by 1.71-5.05x across global batch sizes from 1024 to 16384, with larger gains for the larger model. The full-stack optimum is objective-dependent: when the reward is runtime per bandwidth, the agent selects a topology and workload split that differ sharply from those chosen under a runtime-per-network-cost reward, and the collective algorithm choices change accordingly. Across four search algorithms, agents converge to the same optimal reward but reach it through different parameter settings, yielding eight non-obvious high-performance design points.
Load-bearing premise
The reported speedups assume the simulator's runtime estimates are faithful for the full models on real clusters, because the paper simulates only four layers per model and re-scales latency and memory, and no discovered configuration was run on real hardware.
Editorial extensions
If this is right
- Design-space exploration for distributed ML should be treated as a joint search over workload, collective, network, and compute knobs, not as sequential or isolated tuning passes.
- The optimal distributed system depends on the objective: runtime-per-bandwidth and runtime-per-network-cost reward functions lead to different topologies, workload splits, and collective algorithms.
- The advantage of full-stack co-design grows with model scale: GPT3-175B shows larger relative improvements than ViT-Large on the same 2048-NPU system.
- Because multiple agents converge to equal-reward but structurally different configurations, system designers can choose among near-optimal points by secondary criteria such as cost, complexity, or deployment constraints.
- The PSA abstraction makes agent-agnostic search possible, so adding a new search algorithm family does not require reconfiguring the simulator interface.
Reading between the lines
- If the simulation-to-hardware transfer holds, the PSA abstraction could become a reusable interface for co-design across other simulators and real testbeds, letting different teams share a common design-space definition.
- The equal-reward diversity of discovered configurations suggests a Pareto surface; a multi-objective version of the search could map the cost-latency tradeoff explicitly instead of collapsing it into a single reward scalar.
- The layer-truncation approximation (four simulated layers re-scaled to the full model) is the most likely place for simulator error to enter; running one full-model validation on a real cluster would bound that error and is the natural next test.
- Dynamic factors such as stragglers, faults, and varying batch compositions are not modeled, so cross-layer interactions may matter even more in production settings than in the static results reported here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces COSMIC, a framework that couples ASTRA-sim with ArchGym through a Parameter Set Architecture (PSA) abstraction, and uses ML search agents to co-optimize workload parallelization, collective algorithms, network topology/bandwidth, and compute parameters for distributed transformer training and inference. The authors report that full-stack optimization outperforms single-stack optimization by 1.50-48.41x on 'runtime per BW/NPU' and by 3.40-127.17x on 'runtime per network cost' across three simulated systems, and they present a scalability study plus two co-design use cases. The central claim is that jointly searching multiple design stacks yields substantially better configurations and reveals non-obvious design points.
Significance. If the results held under a correctly specified objective and at full model scale, COSMIC would be a useful addition to the distributed-ML design-space exploration literature: the PSA abstraction is a clean separation-of-concerns idea, the agent-agnostic integration is demonstrated with four search algorithms, and the observation that multiple distinct configurations achieve similar reward is practically valuable. These strengths are real but are currently overshadowed by two problems: the printed reward function does not match the stated objective, and the headline speedups are measured in a metric that full-stack search can improve by reducing bandwidth alone. The paper also provides no artifact, no error bars, and no hardware validation to support the simulation-based speedups, so the quantitative central claim is not yet supported.
major comments (4)
- [Section 5.4] The reward function for 'Perf per BW/NPU' is printed as reward = 1/sqrt((Latency x sum(BW per Dim) - 1)^2). Maximizing this reward drives Latency x sum(BW) toward 1, not toward 0; the prose says COSMIC minimizes the product of latency and bandwidth. The 'minus-one offset' does not prevent divide-by-zero; it introduces a reward singularity at product = 1. Because every result in Section 6 uses this reward, the formula and the prose cannot both be correct, and the quantitative claims cannot be interpreted until this discrepancy is resolved.
- [Section 6.1, Tables 5 and 6] The full-stack versus single-stack comparison is not apples-to-apples: full-stack search controls the bandwidth knobs, while workload-only, collective-only, and network-only baselines use the fixed bandwidths in Table 3. All reported full-stack configurations in Tables 5 and 6 select bandwidth 50 per dimension. For System 1 the fixed baseline sums to 650 units versus 200 for the full-stack configuration, and for System 2 the fixed baseline sums to 800 versus 200, giving full-stack a 3.25x or 4x reward advantage before any latency difference is considered. Thus the reported 1.50-48.41x improvements do not establish raw latency improvements; the paper must report latency or training time separately, or hold the network configuration fixed.
- [Section 4.4 and Table 2 footnote] The paper asserts that ASTRA-sim 'has been cross-validated against real system measurements' but gives no citation and no validation experiment. The evaluations simulate only 4 layers per model and then re-scale latency and memory to the full model, and no discovered configuration is validated on real hardware. The headline speedups and the claimed 'non-obvious configurations' therefore rest on the unverified assumption that 4-layer simulated rankings, after re-scaling, match full-model behavior. Please provide validation evidence, or explicitly characterize the sensitivity of the conclusions to the 4-layer approximation and the re-scaling procedure.
- [Abstract, Section 6.1, Conclusion] The central claim is stated without its metric qualifier. The abstract and conclusion say full-stack optimization delivers '1.50-48.41x higher performance', but Section 6.1 actually measures 'runtime per BW/NPU' or 'runtime per network dollar cost', not raw runtime. This overstates the finding even if the reward formula were corrected, because the denominator includes bandwidth or cost terms that full-stack search can deliberately reduce. The authors should either re-state the claim as 'performance per bandwidth/cost' or provide raw-latency results that support the stronger phrasing.
minor comments (5)
- [Section 6.4] The text refers to 'experiment 4' and to GPT3-175B training, but Table 6 defines Experiment 2 as inference (Chat and QA) and no Experiment 4 exists anywhere in the paper.
- [Table 5 caption] Table 5 reuses the caption 'System configurations used for evaluation' from Table 3; it should instead describe these as the discovered full-stack configurations, since the table reports search outputs rather than fixed baselines.
- [Table 1 and Section 3.2] The design-space count is not self-consistent: the constraint is written as product(DP, SP, PP) <= 1024, but the stated 286 options appear to count tuples with product exactly 1024, and multiplying the listed per-knob point counts does not give 7.69e13. Please recompute or clarify the counting procedure.
- [Section 5.4] The notation 'BW per Dim' is used for what Table 3 calls 'Bandwidth per Dim' and the reward expression sums these values, but the units and the 'per NPU' normalization are never defined; please clarify whether the sum is total network bandwidth or bandwidth per link, and why dividing by NPU count is omitted.
- [Throughout] There are several typographical and consistency issues: 'Cerberas' should be 'Cerebras', the framework name appears as both COSMIC and Cosmic, and Figure 9's parameter labels (a through o) are not mapped to a legend in the figure itself, making the discussion of which parameters vary hard to follow.
Circularity Check
The headline '1.50-48.41x higher performance' is partially self-definitional: 'performance' is measured as a per-bandwidth reward whose denominator the full-stack search can shrink, while the workload-only and collective-only baselines cannot.
-
self definitional
[Section 5.4 (Optimization Objectives), Section 6.1 (Full-Stack Optimizations), Tables 3, 5 and 6]
"reward(Perf per BW/NPU) = 1 / sqrt((Latency × Σ(BW per Dim)−1)^2). ... Cosmic aims to maximize ML performance while limiting the bandwidth allocated per NPU. To achieve this, Cosmic minimizes the product of ML execution latency and the total network bandwidth consumed per NPU. ... Bandwidth per Dim [50, 50, 50, 50]"
The headline improvement is measured under the 'Perf per BW/NPU' reward, whose denominator contains Σ(BW per Dim). In Section 6.1, workload-only and collective-only searches keep the fixed Table 3 networks (System 1 bandwidth sums to 650; System 2 to 800), while full-stack search is free to choose bandwidth knobs, and every reported full-stack result (Tables 5 and 6) selects [50,50,50,50], summing to 200. Thus the full-stack reward receives a 3.25x (System 1) or 4x (System 2) advantage purely from the metric's denominator, before any latency difference. The quoted 1.50-48.41x 'higher performance' is therefore in part a restatement of the objective's definition rather than an emergent co-design gain.
full rationale
The paper's central result is a simulation measurement, not a theorem derived from assumptions that include the speedup. I found no load-bearing self-citation chain: ASTRA-sim [58] and ArchGym [23] are prior systems from overlapping author groups, but the paper does not invoke a uniqueness theorem, and these are open, externally checkable artifacts. Section 4.4's assertion that 'ASTRA-sim has been cross-validated against real system measurements' is, however, made without an in-paper citation and is an evidence gap rather than a circular step. The Table 2 footnote that only 4 layers are simulated and then re-scaled is likewise a validity limitation, not circularity. The significant circularity concern is the reward construction: because the performance metric is defined per total bandwidth and full-stack alone can reduce that bandwidth in the workload-only and collective-only comparisons, a substantial part of the reported improvement is forced by the metric's definition. This makes the headline claim partially self-definitional, though the full-stack search also has genuine independent content in jointly tuning parallelization, collectives, and topology. A more rigorous evaluation would report raw latency separately and use a reward that actually minimizes the stated product, not |Latency×ΣBW − 1|.
Assumptions & free parameters
free parameters (2)
- Simulated layers per model =
4 layers, then re-scaled
- Memory capacity constraint =
24 GB per NPU
assumptions (5)
- domain assumption ASTRA-sim accurately models runtime of distributed training at scale
- domain assumption Roofline compute model with peak-perf, local-mem-bw, and memory-capacity captures device behavior
- ad hoc to paper 4-layer trace templates with symbolic shapes represent full transformer workloads after scaling
- domain assumption Network topology as stacking of Ring, Switch, and FC dimensions captures real cluster fabrics
- ad hoc to paper Reward functions in Section 5.4 encode the intended performance-per-resource objective
invented entities (1)
-
Parameter Set Architecture (PSA)
Cite this review
Pith. "Pith review of COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems." pith.science (2026). https://pith.science/paper/WFGPR6V7
@misc{pith2026250515020,
author = {Pith},
title = {Pith review of: COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/WFGPR6V7}},
note = {Machine review of arXiv:2505.15020}
}
read the original abstract
Large-scale machine learning models necessitate distributed systems, posing significant design challenges due to the large parameter space across distinct design stacks. Existing studies often focus on optimizing individual system aspects in isolation. This work challenges this limitation and introduces COSMIC, a full-stack distributed machine learning systems environment enabling end-to-end simulation and agent-based design space exploration. To facilitate efficient exploration and optimization across the entire stack, we introduce Parameter Set Architecture-an abstraction concept analogous to the instruction set architecture-abstracting away configuration complexities of agent-based search methods. Case studies demonstrate COSMIC's ability to consolidate parameters across multiple layers of design abstraction, discovering eight non-obvious high-performance system configurations across four transformer-based models with up to 175 billion parameters. By optimizing across the stack, COSMIC full-stack optimization delivers 1.50-48.41x higher performance compared to the isolated single-stack optimization.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 1 Pith paper
-
Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference
An RL agent that co-optimizes parallelism degrees and per-operator sharding dimensions finds distributed inference strategies that beat random search and simulated annealing, and slightly outperform Megatron-LM heuris...
Reference graph
Works this paper leans on
-
[1]
Amazon. 2020. AWS Trainium. https://aws.amazon.com/machine-learning/ trainium
work page 2020
-
[2]
AMD. 2023. AMD Instinct MI300 Series Accelerators. https://www.amd.com/en/ products/accelerators/instinct/mi300.html
work page 2023
-
[3]
Anthropic. 2024. Introducing the next generation of Claude. https://www. anthropic.com/news/claude-3-family
work page 2024
-
[4]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...
work page 2020
-
[5]
Cerebras. 2021. Cerebras Systems: Achieving Industry Best AI Performance Through A Systems Approach. https://cerebras.net/wp-content/uploads/2021/ 04/Cerebras-CS-2-Whitepaper.pdf
work page 2021
-
[6]
Ernie Chan, Robert van de Geijn, William Gropp, and Rajeev Thakur. 2006. Collective Communication on Architectures that Support Simultaneous Com- munication over Multiple Links. In PPoPP
work page 2006
-
[7]
M. Cho, U. Finkler, M. Serrano, D. Kung, and H. Hunter. 2019. BlueConnect: De- composing All-Reduce for Deep Learning on Heterogeneous Network Hierarchy. In MLSys
work page 2019
-
[8]
Sajal Dash, Isaac Lyngaas, Junqi Yin, Xiao Wang, Romain Egele, Guojing Cong, Feiyi Wang, and Prasanna Balaprakash. 2023. Optimizing Distributed Training on Frontier for Large Language Models. In arXiv:2312.12705 [cs.DC]
arXiv 2023
Show all 63 references
-
[9]
Dorigo and G
M. Dorigo and G. Di Caro. 1999. Ant Colony Optimization: A New Meta-Heuristic. In CEC
1999
-
[10]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognit...
2021 arXiv
-
[11]
Hadi Esmaeilzadeh, Soroush Ghodrati, Andrew Kahng, Joon Kyung Kim, Sean Kinzer, Sayak Kundu, Rohan Mahapatra, Susmita Dey Manasi, Sachin Sapatnekar, Zhiang Wang, and Ziqing Zeng. 2024. An Open-Source ML-Based Full-Stack Optimization Framework for Machine Learning Accelerators....
2024
-
[12]
Clement, Neel Sun- daresan, and Chen Wu
Spandan Garg, Roshanak Zilouchian Moghaddam, Colin B. Clement, Neel Sun- daresan, and Chen Wu. 2022. DeepPERF: A Deep Learning-Based Approach For Improving Software Performance. In arXiv:2206.13619 [cs.SE]
2022 arXiv
-
[13]
Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei, Irfan Essa, Lu Jiang, and José Lezama. 2023. Photorealistic Video Generation with Diffusion Models. In arXiv:2312.06662 [cs.CV]
2023 arXiv
-
[14]
Aaron Harlap, Deepak Narayanan, Amar Phanishayee, Vivek Seshadri, Nikhil Devanur, Greg Ganger, and Phil Gibbons. 2018. PipeDream: Fast and Efficient Pipeline Parallel DNN Training. In arXiv:1806.03377 [cs.DC]
2018 arXiv
-
[15]
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, and zhifeng Chen. 2019. GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism. In NeurIPS
2019
-
[16]
Sylvain Jeaugey. 2019. Massively Scale Your Deep Learning Training with NCCL 2.4. https://developer.nvidia.com/blog/massively-scale-deep-learning-training- nccl-2-4
2019
-
[17]
Norm Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, Clifford Young, Xiang Zhou, Zongwei Zhou, and David A Patterson. 2023. TPU v4: An Optically Reconfigurable Supercomputer for Machine Learn...
2023
-
[18]
Jouppi, Doe Hyun Yoon, George Kurian, Sheng Li, Nishant Patil, James Laudon, Cliff Young, and David Patterson
Norman P. Jouppi, Doe Hyun Yoon, George Kurian, Sheng Li, Nishant Patil, James Laudon, Cliff Young, and David Patterson. 2020. A Domain-Specific Supercomputer for Training Deep neural networks. Commun. ACM (2020)
2020
-
[19]
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al. 2021. Highly Accurate Protein Structure Prediction with Al- phaFold. Nature (2021)
2021
-
[20]
Sheng-Chun Kao and Tushar Krishna. 2020. GAMMA: Automating the HW Mapping of DNN Models on Accelerators via Genetic Algorithm. In ICCAD
2020
-
[21]
Sourabh Katoch, Sumit Singh Chauhan, and Vijay Kumar. 2021. A Review on Genetic Algorithm: Past, Present, and Future. Multimedia Tools Appl. (2021)
2021
-
[22]
Benjamin Klenk, Nan Jiang, Greg Thorson, and Larry Dennison. 2020. An In-Network Architecture for Accelerating Shared-memory Multiprocessor Col- lectives. In ISCA
2020
-
[23]
Srivatsan Krishnan, Amir Yazdanbakhsh, Shvetank Prakash, Jason Jabbour, Ikechukwu Uchendu, Susobhan Ghosh, Behzad Boroujerdian, Daniel Richins, Devashree Tripathy, Aleksandra Faust, and Vijay Janapa Reddi. 2023. Arch- Gym: An Open-Source Gymnasium for Machine Learning Assisted...
2023
-
[24]
H. J. Kushner. 1964. A New Method of Locating the Maximum Point of an Arbitrary Multipeak Curve in the Presence of Noise. Journal of Basic Engineering (1964)
1964
-
[25]
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding. In arXiv:2006.16668 [cs.CL]
2020 arXiv
-
[26]
Chuan Li. 2020. OpenAI’s GPT-3 Language Model: A Technical Overview. https: //lambdalabs.com/blog/demystifying-gpt-3
2020
-
[27]
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, and Soumith Chintala
-
[28]
Hao Lin, Ke Wu, Jie Li, Jun Li, and Wu-Jun Li. 2024. UniAP: Unifying Inter- and Intra-Layer Automatic Parallelism by Mixed Integer Quadratic Programming. In arXiv:2307.16375 [cs.LG]
2024 arXiv
-
[29]
Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jianfeng Gao, Lifang He, and Lichao Sun. 2024. Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models. In arXiv:2402.17177 [cs.CV]
2024 arXiv
-
[30]
Wenyan Lu, Guihai Yan, Jiajun Li, Shijun Gong, Yinhe Han, and Xiaowei Li
-
[31]
Meta. 2023. Introducing LLaMA: A Foundational, 65-billion-Parameter Large Language Model. https://ai.meta.com/blog/large-language-model-llama-meta- ai
2023
-
[32]
J. Močkus. 1975. On Bayesian Methods for Seeking the Extremum. Optimization Techniques IFIP Technical Conference (1975)
1975
-
[33]
NVIDIA. 2020. NVIDIA A100 TENSOR CORE GPU. https://www.nvidia.com/en- us/data-center/a100
2020
-
[34]
NVIDIA. 2023. NVIDIA H100 Tensor Core GPU. https://www.nvidia.com/en- us/data-center/h100
2023
-
[35]
NVIDIA. 2024. NVIDIA Collective Communications Library (NCCL). https: //developer.nvidia.com/nccl
2024
-
[36]
NVIDIA. 2024. NVLink and NVLink Switch. https://www.nvidia.com/en-us/data- center/nvlink/
2024
-
[37]
NVIDIA. 2024. Parallelisms - NVIDIA Docs. https://docs.nvidia.com/nemo- framework/user-guide/latest/nemotoolkit/features/parallelisms.html
2024
-
[38]
OpenAI. 2022. Introducing ChatGPT. https://openai.com/blog/chatgpt
2022
-
[39]
KARL PEARSON. 1905. The Problem of the Random Walk. Nature (1905)
1905
-
[40]
Sundar Pichai and Demis Hassabis. 2023. Introducing Gemini: Our Largest and Most Capable AI Model. https://blog.google/technology/ai/google-gemini-ai
2023
-
[41]
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models. In arXiv:1910.02054 [cs.LG]
2020 arXiv
-
[42]
Saeed Rashidi, Srinivas Sridharan, Sudarshan Srinivasan, and Tushar Krishna
-
[43]
Saeed Rashidi, William Won, Sudarshan Srinivasan, Srinivas Sridharan, and Tushar Krishna. 2022. Themis: A Network Bandwidth-Aware Collective Sched- uling Policy for Distributed Training of DL Models. In ISCA
2022
-
[44]
Aashaka Shah, Vijay Chidambaram, Meghan Cowan, Saeed Maleki, Madan Musu- vathi, Todd Mytkowicz, Jacob Nelson, Olli Saarikivi, and Rachee Singh. 2023. TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches. In NSDI
2023
-
[45]
In ISPASS
ASTRA-SIM: Enabling SW/HW Co-Design Exploration for Distributed DL Training Platforms. In ISPASS
-
[46]
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2020. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. In arXiv:1909.08053 [cs.CL]
2020 arXiv
-
[47]
Gardner, Yiming Yang, Milad Hashemi, Graham Neubig, Parthasarathy Ranganathan, Osbert Bastani, and Amir Yazdanbakhsh
Alexander Shypula, Aman Madaan, Yimeng Zeng, Uri Alon, Jacob R. Gardner, Yiming Yang, Milad Hashemi, Graham Neubig, Parthasarathy Ranganathan, Osbert Bastani, and Amir Yazdanbakhsh. 2024. Learning Performance-Improving Code Edits. In ICLR
2024
-
[48]
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. In arXiv:1701.06538 [cs.LG]
2017 arXiv
-
[49]
Siddharth Singh, Olatunji Ruwase, Ammar Ahmad Awan, Samyam Rajbhan- dari, Yuxiong He, and Abhinav Bhatele. 2023. A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training. In ICS
2023
-
[50]
Mellanox Technologies. 2003. Introduction to InfiniBand. https://network.nvidia. com/pdf/whitepapers/IB_Intro_WP_190.pdf
2003
-
[51]
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. 2022. Make-A-Video: Text-to-Video Generation without Text-Video Data. In arXiv:2209.14792 [cs.CV]. A. Raju et al
2022 arXiv
-
[52]
Rajeev Thakur, Rolf Rabenseifner, and William Gropp. 2005. Optimization of Collective Communication Operations in MPICH. Int. J. High Perform. Comput. Appl. 19, 1 (2005), 49–66. doi:10.1177/1094342005051521
2005 doi
-
[53]
Trieu H Trinh, Yuhuai Wu, Quoc V Le, He He, and Thang Luong. 2024. Solving Olympiad Geometry without Human Demonstrations. Nature (2024)
2024
-
[54]
Tesla. 2021. Tesla Dojo Technology. https://cdn.motor1.com/pdf-files/535242876- tesla-dojo-technology.pdf
2021
-
[55]
Leslie G. Valiant. 1990. A Bridging Model for Parallel Computation. Commun. ACM 33, 8 (1990), 103–111. doi:10.1145/79173.79181
1990
-
[56]
Weiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi, Zhihao Jia, Dheevatsa Mudigere, Ying Zhang, and Anthony Kewitsch. 2023. TopoOpt: Co- optimizing Network Topology and Parallelization Strategy for Distributed Train- ing Jobs. In NSDI
2023
-
[57]
Amin Vahdat and Mark Lohmeyer. 2023. Enabling next-generation AI workloads: Announcing TPU v5p and AI Hypercomputer. https: //cloud.google.com/blog/products/ai-machine-learning/introducing-cloud- tpu-v5p-and-ai-hypercomputer
2023
-
[58]
William Won, Taekyung Heo, Saeed Rashidi, Srinivas Sridharan, Sudarshan Srinivasan, and Tushar Krishna. 2023. ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale. In Proceedings of the 2023 IEEE International Symposium on Pe...
2023
-
[59]
William Won, Saeed Rashidi, Sudarshan Srinivasan, and Tushar Krishna. 2024. LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Opti- mization for Distributed Training of Large AI Models. In ISPASS
2024
-
[60]
William Won, Midhilesh Elavazhagan, Sudarshan Srinivasan, Ajaya Durg, Samvit Kaul, Swati Gupta, and Tushar Krishna. 2024. TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning. In arXiv:2304.05301 [cs.DC]
2024 arXiv
-
[63]
Liangyu Zhao, Siddharth Pal, Tapan Chugh, Weiyang Wang, Jason Fantl, Prith- wish Basu, Joud Khoury, and Arvind Krishnamurthy. 2024. Efficient Direct- Connect Topologies for Collective Communications. In arXiv:2202.03356 [cs.NI]
2024 arXiv
-
[2017]
FlexFlow: A Flexible Dataflow Accelerator Architecture for Convolutional Neural Networks. In HPCA
-
[2020]
In arXiv:2006.15704 [cs.DC]
PyTorch Distributed: Experiences on Accelerating Data Parallel Training. In arXiv:2006.15704 [cs.DC]
2006 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.