REVIEW 3 major objections 5 minor 1 cited by
ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read ATLAHS claims that tracing real AI, HPC, and storage applications into GOAL graphs lets a flexible simulator predict runtimes within 5% error.
desk verdict A genuinely useful open-source simulator toolchain with solid HPC validation, but the AI <5% claim rests on unmeasured LogGOPS parameters and the storage support is unvalidated, so treat the headline accuracy claim as conditional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the GOAL directed acyclic graph, a format in which every workload is a set of send, receive, and calc (computation) vertices with dependency edges, assigned to compute streams to model concurrency. ATLAHS's tracer front ends convert real executions into these graphs: NCCL collectives are decomposed into point-to-point send/receive schedules according to algorithm, channel, and protocol settings; MPI operations are converted through PMPI tracing; block I/O commands are converted through a bpftrace-based tracer. Dummy zero-cost vertices synchronize parallel streams, and merging graphs from multiple jobs models multi-tenancy. The GOAL graphs are then scheduled by ATLAHS and executed on pluggable backends—LogGOPSim for message-level speed and htsim for packet-level fidelity—through a minimal interface of send, recv, calc, and eventOver. This combination is what lets one unified representation span AI, HPC, and storage.
What would settle it
On the same cluster used for tracing, directly measure the interconnect's latency, overhead, and per-byte transfer time and rerun the Llama 7B 128-GPU validation with those measured values; if the predicted iteration time deviates from the measured runtime by more than 5%, the central accuracy claim depends on the borrowed parameters rather than on the toolchain itself.
Extended reading notes
Core claim
The core discovery is that a single, compact intermediate representation—the GOAL format, with only send, receive, and computation vertices and edges expressing dependencies—is expressive enough to capture the communication and computation patterns of LLM training, MPI scientific codes, and distributed storage traffic, and to reproduce measured application runtimes within 5% error. The toolchain obtains these graphs by tracing NCCL through a GPU profiler with added annotations, tracing MPI through the PMPI interface, and tracing block I/O through an eBPF-based tracer; it then decomposes NCCL collectives into point-to-point schedules, merges per-GPU graphs into per-node graphs, and replaces intra-node communication with computation. On the two configurations where AstraSim ran successfully, ATLAHS was both more accurate and faster, and its GOAL trace files were consistently smaller than AstraSim's Chakra traces. The paper further shows that the message-level and packet-level backends agree within 1-2% on fully provisioned symmetric networks, while diverging by over 120% when oversubscription causes packet drops that only the packet-level backend can see.
Load-bearing premise
The AI workload predictions rest on network-cost parameters (latency, overhead, per-byte transfer time) borrowed from benchmarks of similar hardware rather than measured on the target cluster, and the paper does not show how much the predictions would change if those parameters were different.
Editorial extensions
If this is right
- Application traces, not synthetic microbenchmarks, become the default workload for evaluating network designs, since they expose issues such as Swift's multi-hop congestion weakness that microbenchmarks hide.
- Network architects can safely use the fast message-level backend for fully provisioned symmetric topologies, but must switch to a packet-level backend when oversubscription, packet drops, or queue dynamics matter.
- Multi-job and multi-tenant performance questions—such as how job placement affects a shared cluster—can be answered by merging GOAL graphs of different applications.
- The storage case study implies that congestion control choice (sender-based vs receiver-based) significantly changes storage request completion under oversubscribed topologies.
- Because GOAL trace files are several times smaller than Chakra traces, sharing reproducible workload traces at large scale becomes cheaper.
Reading between the lines
- If the 5%-error claim is robust, a natural next step is to use ATLAHS itself as a fast screening tool for congestion-control and topology changes before packet-level simulation, since the two backends agree when the network is well provisioned.
- The acknowledged static-DAG limitation suggests the approach would under-predict runtime for dynamically scheduled communication, such as fault-tolerant storage protocols or data-dependent GNN training; a testable extension is to add dynamic vertices to GOAL and compare against real runs.
- The lack of a sensitivity analysis for the AI LogGOPS parameters leaves open whether the reported accuracy is stable; a direct measurement of those parameters on the target hardware would settle it.
- Combining traces from different domains in one graph (as done for Llama and LULESH) could be extended to model shared infrastructure contention more generally, but the paper only demonstrates network-level contention, not memory or cache interference.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents ATLAHS, an application-centric network simulation toolchain that translates real-world traces from AI (NCCL), HPC (MPI), and distributed storage (block I/O) into GOAL-style DAGs and simulates them through multiple backends (LogGOPSim, htsim, NS-3). It validates the toolchain on six LLM/MoE training configurations on the Alps cluster and fifteen HPC configurations on a CSCS testbed, reporting prediction errors mostly within 5% for both ATLAHS LGS and ATLAHS htsim, while also comparing against AstraSim for two configurations and reporting significantly smaller trace files than Chakra. Case studies illustrate congestion-control effects on storage traffic, differences between message-level and packet-level backends, and job-placement effects in a shared cluster.
Significance. ATLAHS is a potentially valuable open-source infrastructure: it unifies diverse workload formats under GOAL, releases a substantial trace dataset, supports multi-job and multi-tenant scenarios, and demonstrates that a message-level backend can match packet-level accuracy in fully provisioned, symmetric topologies. The HPC validation is broad (15 configurations) and uses measured LogGOPS parameters, and the trace-size and simulation-speed comparisons are concrete and reproducible. If the AI-accuracy claim can be supported against parameter uncertainty, this work would be a solid contribution to network-simulation practice for AI/HPC convergence.
major comments (3)
- [§5.2, Fig. 8] The AI-validation accuracy claim depends entirely on LogGOPS parameters (L=3700 ns, o=200 ns, g=5 ns, O=0, G=0.04 ns/B, S=0) that are estimated from external benchmarking works rather than measured on the Alps cluster where the traces were collected. No sensitivity analysis is reported, so the asserted consistent <5% error for AI workloads is a point estimate on six configurations, not a demonstrated robustness property. Since several workloads are communication-dominated (MoE 8x70B has only 5.3% non-overlapped computation; Llama 70B has 9.9%), modest errors in G or L translate almost directly into per-iteration time errors. Please add a sensitivity analysis over plausible parameter ranges (especially G and L) for both ATLAHS backends, or measure the parameters directly on the validation system, and also isolate the effect of the static-GOAL-DAG simplification acknowledged in §7.
- [Fig. 10, §5.3] The text states that prediction error 'remains consistently below 5% across all cases and applications,' but Fig. 10 includes a reported error of -5.2% (for the htsim backend on one of the HPC configurations). This contradicts the stated claim. Please correct the text or the figure, or clarify whether the -5.2% value is a typo or an excluded outlier.
- [§3.1.3, §6.1] The distributed-storage component is described as a first-class part of the toolchain and appears in the title and abstract, but it is never validated against measured storage-system runtimes or packet-level ground truth. The storage case study compares two congestion-control algorithms using traces generated under assumptions about Azure Direct Drive from public documentation; the results are relative and have no accuracy check. Please add a validation experiment for the storage path, or explicitly scope the accuracy claims in the abstract and conclusion to AI and HPC workloads.
minor comments (5)
- [Abstract / §5.3] The abstract claims 'consistently less than 5% error' without noting that the storage component is unvalidated; please qualify the claim to the validated AI and HPC domains.
- [§5.1] The sentence 'From our testing, the runtime of complex traces is reduced from 10× to100 × the after the improvements' contains a typo and should be reworded to state the speedup range clearly.
- [Fig. 8 / Fig. 10] Backend naming is inconsistent ('HTSIM' in Fig. 10 caption vs 'htsim' in the text); also, in Fig. 10 the two sets of red percentages are not explicitly labelled in the visualization, making it hard to map errors to backends.
- [§5.2] The claim that ATLAHS 'consistently outperforms AstraSim' in accuracy is based on only two configurations where AstraSim completed; consider phrasing this as 'in the two successful AstraSim comparisons' to match the evidence.
- [References] Reference [82] is incomplete; it ends with 'embedded-checkout=true' and appears to be a truncated Bloomberg URL.
Circularity Check
No significant circularity; ATLAHS accuracy is validated against measured runtimes with parameters taken from external benchmarks, not fitted to the predictions.
full rationale
The paper's central claim is that ATLAHS predicts runtimes of real AI and HPC workloads with consistently less than 5% error. This claim is supported by direct comparison against measured application runtimes on real clusters (Alps for AI, a CSCS testbed for HPC). The AI LogGOPS parameters (L=3700 ns, o=200 ns, g=5 ns, O=0, G=0.04 ns/B, S=0) are explicitly estimated from external benchmarking works rather than fitted to the six validation points, and the HPC parameters are measured with Netgauge. No equation in the paper reduces to the measured runtime by construction: the GOAL representation of computation and communication is independent of the network backends, and the two backends (LogGOPSim and htsim) are cross-checked against one another and against AstraSim. The paper does cite its own prior work on GOAL and LogGOPSim, but these citations are used as background for the trace format choice and do not carry the load of the accuracy validation, which rests on measured ground truth. The acknowledged simplifications (static GOAL DAG, no CUDA cross-stream dependencies, unmeasured AI LogGOPS parameters) are correctness-robustness concerns, not circularity: they could make the predictions wrong, but they do not make the predictions equal to the inputs. No fitted input is renamed as a prediction, and no self-citation chain forces the result. Therefore the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (1)
- AI LogGOPS parameter set (L, o, g, G) =
L=3700, o=200, g=5, O=0, G=0.04, S=0 (ns)
assumptions (5)
- domain assumption GOAL DAG abstraction is sufficient to model communication and computation for AI, HPC, and storage workloads.
- domain assumption The NCCL collective-to-P2P decomposition in Stage 3 is accurate across NCCL algorithms, protocols, and channel counts.
- domain assumption Replacing intra-node sends/receives with calc vertices preserves simulation fidelity.
- domain assumption A static GOAL DAG without CUDA inter-stream dependencies or dynamic scheduling is sufficient for the validated workloads.
- domain assumption The Azure Direct Drive model, based on public documentation, captures the storage architecture's network behavior.
Cite this review
Pith. "Pith review of ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage." pith.science (2026). https://pith.science/paper/WJ3QXMIH
@misc{pith2026250508936,
author = {Pith},
title = {Pith review of: ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage},
year = {2026},
howpublished = {\url{https://pith.science/paper/WJ3QXMIH}},
note = {Machine review of arXiv:2505.08936}
}
read the original abstract
Network simulators play a crucial role in evaluating the performance of large-scale systems. However, existing simulators rely heavily on synthetic microbenchmarks or narrowly focus on specific domains, limiting their ability to provide comprehensive performance insights. In this work, we introduce ATLAHS, a flexible, extensible, and open-source toolchain designed to trace real-world applications and accurately simulate their workloads. ATLAHS leverages the GOAL format to model communication and computation patterns in AI, HPC, and distributed storage applications. It supports multiple network simulation backends and handles multi-job and multi-tenant scenarios. Through extensive validation, we demonstrate that ATLAHS achieves high accuracy in simulating realistic workloads (consistently less than 5% error), while significantly outperforming AstraSim, the current state-of-the-art AI systems simulator, in terms of simulation runtime and trace size efficiency. We further illustrate ATLAHS's utility via detailed case studies, highlighting the impact of congestion control algorithms on the performance of distributed storage systems, as well as the influence of job-placement strategies on application runtimes.
Figures
Figures from the paper (10 more)
Forward citations
Cited by 1 Pith paper
-
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
This survey classifies distributed DNN training simulators into analytical, profiling-based, and execution-driven categories, and compares them alongside TCO and carbon-emission models.
Reference graph
Works this paper leans on
-
[1]
[n. d.]. Broadcom htsim repository. https://github.com/Broadcom/csg-htsim. Accessed: 2025-02-13
2025
-
[2]
[n. d.]. Ultra Ethernet Consortium. https://ultraethernet.org/. Accessed: 2025- 01-13
2025
-
[3]
Evensky, Joseph P
Helgi Adalsteinsson, Scott Cranford, David A. Evensky, Joseph P. Kenny, Jackson Mayo, Ali Pinar, and Curtis L. Janssen. 2010. A Simulator for Large-Scale Parallel Computer Architectures.Int. J. Distrib. Syst. Technol.1, 2 (April 2010), 57–73
2010
-
[4]
2025.ROCm Communication Collectives Library (RCCL) Documentation
Advanced Micro Devices, Inc. 2025.ROCm Communication Collectives Library (RCCL) Documentation. https://rocm.docs.amd.com/projects/rccl/en/latest/what- is-rccl.html Accessed: 2025-01-28
2025
-
[5]
Ionescu, Klaus E
Albert Alexandrov, Mihai F. Ionescu, Klaus E. Schauser, and Chris Scheiman
-
[6]
Jens Axboe. 2005. blktrace: A Block I/O Tracing Mechanism for Linux. https: //www.kernel.org/doc/Documentation/block/blktrace.txt. Accessed: February 04, 2025; an open-source tool for tracing block I/O events in Linux
2005
-
[7]
Tal Ben-Nun and Torsten Hoefler. 2019. Demystifying Parallel and Distributed Deep Learning: An In-depth Concurrency Analysis.ACM Comput. Surv.52, 4, Article 65 (Aug. 2019), 43 pages
2019
-
[8]
Maciej Besta, Julia Barth, Eric Schreiber, Ales Kubicek, Afonso Catarino, Robert Gerstenberger, Piotr Nyczyk, Patrick Iff, Yueling Li, Sam Houliston, Tomasz Sternal, Marcin Copik, Grzegorz Kwaśniewski, Jürgen Müller, Łukasz Flis, Hannes Eberhard, Hubert Niewiadomski, and Torsten Hoefler. 2025. Reasoning Language Models: A Blueprint. arXiv:2501.11223 [cs.A...
arXiv 2025
Show all 98 references
-
[9]
Maciej Besta and Torsten Hoefler. 2023. Parallel and Distributed Graph Neural Networks: An In-Depth Concurrency Analysis. arXiv:2205.09702 [cs.LG] https: //arxiv.org/abs/2205.09702
2023 arXiv
-
[10]
Maciej Besta, Marcel Schneider, Salvatore Di Girolamo, Ankit Singla, and Torsten Hoefler. 2021. Towards Million-Server Network Simulations on Just a Laptop. arXiv:2105.12663 [cs.NI] https://arxiv.org/abs/2105.12663
2021 arXiv
-
[11]
Tommaso Bonato, Abdul Kabbani, Ahmad Ghalayini, Michael Papamichael, Mohammad Dohadwala, Lukas Gianinazzi, Mikhail Khalilov, Elias Acher- mann, Daniele De Sensi, and Torsten Hoefler. 2025. REPS: Recycled En- tropy Packet Spraying for Adaptive Load Balancing and Failure Mitigat...
2025
-
[12]
Tommaso Bonato, Abdul Kabbani, Daniele De Sensi, Rong Pan, Yanfang Le, Costin Raiciu, Mark Handley, Timo Schneider, Nils Blach, Ahmad Ghalayini, Daniel Alves, Michael Papamichael, Adrian Caulfield, and Torsten Hoefler. 2024. FASTFLOW: Flexible Adaptive Congestion Control for H...
2024
-
[13]
Carothers, D
C.D. Carothers, D. Bauer, and S. Pearce. 2000. ROSS: a high-performance, low memory, modular time warp system. InProceedings Fourteenth Workshop on Parallel and Distributed Simulation. 53–60. https://doi.org/10.1109/PADS.2000. 847144
2000 doi
-
[14]
Henri Casanova, Arnaud Giersch, Arnaud Legrand, Martin Quinson, and Frédéric Suter. 2025. Lowering entry barriers to developing custom simulators of dis- tributed applications and platforms with SimGrid.Parallel Comput.123 (2025), 103–125. https://doi.org/10.1016/j.parco.2025.103125
2025
-
[15]
Henri Casanova, Arnaud Legrand, and Martin Quinson. 2008. SimGrid: A Generic Framework for Large-Scale Distributed Experiments. InTenth Inter- national Conference on Computer Modeling and Simulation (uksim 2008). 126–131. https://doi.org/10.1109/UKSIM.2008.28
2008 doi
-
[16]
Abhishek Chard, R. G. Dreslinski, Thomas F. Wenisch, Greg Ganger, and An- drew A. Chien. 2018. On the Diversity of Cluster Workloads and Its Impact on Research Results. In2018 USENIX Annual Technical Conference (USENIX ATC 18). 533–546. https://www.pdl.cmu.edu/ATLAS/ Includes ...
2018
-
[17]
Jaehong Cho, Minsu Kim, Hyunmin Choi, Guseul Heo, and Jongse Park. 2024. LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale. In2024 IEEE International Symposium on Workload Characteri- zation (IISWC). 15–29. https://doi.org/10.1109/IISWC6309...
2024
-
[18]
Pierre-Nicolas Clauss, Mark Stillwell, Stephane Genaud, Frederic Suter, Henri Casanova, and Martin Quinson. 2011. Single Node On-Line Simulation of MPI Ap- plications with SMPI. In2011 IEEE International Parallel & Distributed Processing Symposium. 664–675. https://doi.org/10....
2011 doi
-
[19]
Storage Performance Council. 2025. SPC Trace File Format Specification. https: //skulddata.cs.umass.edu/traces/storage/SPC-Traces.pdf. Accessed: 2025-01-22
2025
-
[20]
CSCS. [n. d.]. New Research Infrastructure: ’Alps’ Supercomputer Inaugurated. Swiss National Supercomputing Center([n. d.]). https://www.cscs.ch/publications/ news/2024/new-research-infrastructure-alps-supercomputer-inaugurated
2024
-
[21]
Daniele De Sensi, Tiziano De Matteis, Konstantin Taranov, Salvatore Di Girolamo, Tobias Rahn, and Torsten Hoefler. 2022. Noise in the Clouds: Influence of Net- work Performance Variability on Application Scalability.Proc. ACM Meas. Anal. Comput. Syst.6, 3, Article 49 (dec 2022...
2022 doi
-
[22]
Daniele De Sensi, Lorenzo Pichetti, Flavio Vella, Tiziano De Matteis, Zebin Ren, Luigi Fusco, Matteo Turisini, Daniele Cesarini, Kurt Lust, Animesh Trivedi, Duncan Roweth, Filippo Spiga, Salvatore Di Girolamo, and Torsten Hoefler
-
[23]
Ewa Deelman, Rafael Ferreira da Silva, Gideon Juve, Mats Rynge, Karan Vahi, and Miron Livny. 2022. WfCommons: A Framework for Enabling Scientific Workflow Research and Development.Future Generation Computer Systems129 (2022), 166–182. https://doi.org/10.1016/j.future.2022.01.011
2022 doi
-
[24]
Denzel, Jian Li, Peter Walker, and Yuho Jin
Wolfgang E. Denzel, Jian Li, Peter Walker, and Yuho Jin. 2010. A Framework for End-to-End Simulation of High-performance Computing Systems.SIM- ULATION86, 5-6 (2010), 331–350. https://doi.org/10.1177/0037549709340840 arXiv:https://doi.org/10.1177/0037549709340840
2010 doi
-
[25]
Jiangfei Duan, Xiuhong Li, Ping Xu, Xingcheng Zhang, Shengen Yan, Yun Liang, and Dahua Lin. 2024. Proteus: Simulating the Performance of Distributed DNN Training.IEEE Transactions on Parallel and Distributed Systems35, 10 (2024), 1867–1878. https://doi.org/10.1109/TPDS.2024.3443255
2024
-
[26]
Jiangfei Duan, Shuo Zhang, Zerui Wang, Lijuan Jiang, Wenwen Qu, Qinghao Hu, Guoteng Wang, Qizhen Weng, Hang Yan, Xingcheng Zhang, Xipeng Qiu, Dahua Lin, Yonggang Wen, Xin Jin, Tianwei Zhang, and Peng Sun. 2024. Efficient Training of Large Language Models on Distributed Infrast...
2024 arXiv
-
[27]
Feitelson
Dror G. Feitelson. [n. d.]. The Parallel Workloads Archive. https://www.cs.huji. ac.il/labs/parallel/workload/. Accessed: 2025-04-12
2025
-
[28]
Yinxiao Feng, Yuchen Wei, Dong Xiang, and Kaisheng Ma. 2024. Evaluating Chiplet-based Large-Scale Interconnection Networks via Cycle-Accurate Packet- Parallel Simulation. In2024 USENIX Annual Technical Conference (USENIX ATC 24). USENIX Association, Santa Clara, CA, 731–747. h...
2024
-
[29]
Mackenzie Ferguson. 2025. Nvidia’s Unstoppable Rise: Dominating the AI Chip Market.OpenTools.ai(2025). https://opentools.ai/news/nvidias-unstoppable- rise-dominating-the-ai-chip-market Accessed: 2025-01-28
2025
-
[30]
Luigi Fusco, Mikhail Khalilov, Marcin Chrapek, Giridhar Chukkapalli, Thomas Schulthess, and Torsten Hoefler. 2024. Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip. arXiv:2408.11556 [cs.DC] https://arxiv.org/abs...
2024 arXiv
-
[31]
Markus Geimer, Felix Wolf, Brian J. N. Wylie, Erika Ábrahám, Daniel Becker, and Bernd Mohr. 2010. The Scalasca performance toolset architecture.Concurrency and Computation: Practice and Experience22 (4 2010), 702–719. Issue 6. https: //doi.org/10.1002/cpe.1556 Shen et al
2010 doi
-
[32]
Ibrahim, Lenny Oliker, Nicholas J
Taylor Groves, Ben Brock, Yuxin Chen, Khaled Z. Ibrahim, Lenny Oliker, Nicholas J. Wright, Samuel Williams, and Katherine Yelick. 2020. Performance Trade-offs in GPU Communication: A Study of Host and Device-initiated Ap- proaches. In2020 IEEE/ACM Performance Modeling, Benchma...
2020
-
[33]
Shin, Yibo Zhu, Myeongjae Jeon, Junjie Qian, Hongqiang Liu, and Chuanxiong Guo
Juncheng Gu, Mosharaf Chowdhury, Kang G. Shin, Yibo Zhu, Myeongjae Jeon, Junjie Qian, Hongqiang Liu, and Chuanxiong Guo. 2019. Tiresias: A GPU Cluster Manager for Distributed Deep Learning. In16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19). USENI...
2019
-
[34]
Moore, Gianni Antichi, and Marcin Wójcik
Mark Handley, Costin Raiciu, Alexandru Agache, Andrei Voinescu, Andrew W. Moore, Gianni Antichi, and Marcin Wójcik. 2017. Re-architecting datacenter networks and stacks for low latency and high performance. InProceedings of the Conference of the ACM Special Interest Group on D...
2017
-
[35]
Thomas Henderson, Sally Floyd, and George Riley. 2006. ns3 Project Goals. Workshop on NS-2: the IP Network Simulator.(01 2006). https://doi.org/10.1145/ 1190455.1190468
2006
-
[36]
Torsten Hoefler, Tommaso Bonato, Daniele De Sensi, Salvatore Di Girolamo, Shigang Li, Marco Heddes, Jon Belk, Deepak Goel, Miguel Castro, and Steve Scott. 2022. HammingMesh: A Network Topology for Large-Scale Deep Learning. arXiv:2209.01346 [cs.DC] https://arxiv.org/abs/2209.01346
2022 arXiv
-
[37]
Torsten Hoefler, Torsten Mehlan, Andrew Lumsdaine, and Wolfgang Rehm. 2007. Netgauge: A Network Performance Measurement Framework. InProceedings of High Performance Computing and Communications, HPCC’07(Houston, USA), Vol. 4782. Springer, 659–671
2007
-
[38]
Torsten Hoefler, Timo Schneider, and Andrew Lumsdaine. 2010. Character- izing the Influence of System Noise on Large-Scale Applications by Simula- tion. InSC ’10: Proceedings of the 2010 ACM/IEEE International Conference for High Performance Computing, Networking, Storage and ...
2010 doi
-
[39]
Torsten Hoefler, Timo Schneider, and Andrew Lumsdaine. 2010. LogGOPSim: simulating large-scale applications in the LogGOPS model. InProceedings of the 19th ACM International Symposium on High Performance Distributed Computing (Chicago, Illinois)(HPDC ’10). Association for Comp...
2010
-
[40]
Torsten Hoefler, Christian Siebert, and Andrew Lumsdaine. 2009. Group Operation Assembly Language - A Flexible Way to Express Collective Com- munication. In2009 International Conference on Parallel Processing. 574–581. https://doi.org/10.1109/ICPP.2009.70
2009 doi
-
[41]
2024.Intel®oneAPI Collective Communications Li- brary (oneCCL)
Intel Corporation. 2024.Intel®oneAPI Collective Communications Li- brary (oneCCL). https://www.intel.com/content/www/us/en/docs/oneapi/ programming-guide/2024-1/intel-oneapi-collective-communications- library.html Accessed: 2025-01-28
2024
-
[42]
Nikhil Jain, Abhinav Bhatele, Sam White, Todd Gamblin, and Laxmikant V. Kale
-
[43]
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne...
2024 arXiv
-
[44]
Kale and Sanjeev Krishnan
Laxmikant V. Kale and Sanjeev Krishnan. 1993. CHARM++: a portable concurrent object oriented system based on C++.SIGPLAN Not.28, 10 (oct 1993), 91–108. https://doi.org/10.1145/167962.165874
1993
-
[45]
Chamberlain, Jonathan Cohen, Zachary Devito, Riyaz Haque, Dan Laney, Edward Luke, Felix Wang, David Richards, Martin Schulz, and Charles H
Ian Karlin, Abhinav Bhatele, Jeff Keasler, Bradford L. Chamberlain, Jonathan Cohen, Zachary Devito, Riyaz Haque, Dan Laney, Edward Luke, Felix Wang, David Richards, Martin Schulz, and Charles H. Still. 2013. Exploring Traditional and Emerging Parallel Programming Models Using ...
2013 doi
-
[46]
Andreas Knüpfer, Ronny Brendel, Holger Brunst, Hartmut Mix, and Wolfgang E. Nagel. 2006. Introducing the open trace format (OTF). InProceedings of the 6th International Conference on Computational Science - Volume Part II(Reading, UK) (ICCS’06). Springer-Verlag, Berlin, Heidel...
2006
-
[47]
Müller, and Wolfgang E
Andreas Knüpfer, Holger Brunst, Jens Doleschal, Matthias Jurenz, Matthias Lieber, Holger Mickler, Matthias S. Müller, and Wolfgang E. Nagel. 2008. The Vampir Performance Analysis Tool-Set. InTools for High Performance Computing, Michael Resch, Rainer Keller, Valentin Himmler, ...
2008
-
[48]
Andreas Knüpfer, Markus Geimer, Johannes Spazier, Joseph Schuchart, Michael Wagner, Dominic Eschweiler, and Matthias S. Müller. 2010. A generic attribute extension to OTF and its use for MPI replay.Procedia Computer Science1, 1 (2010), 2115–2124. https://doi.org/10.1016/j.proc...
2010 doi
-
[49]
Nagel, Yury Oleynik, Peter Philippen, Pavel Saviankou, Dirk Schmidl, Sameer Shende, Ronny Tschüter, Michael Wagner, Bert Wesarg, and Felix Wolf
Andreas Knüpfer, Christian Rössel, Dieter an Mey, Scott Biersdorff, Kai Diethelm, Dominic Eschweiler, Markus Geimer, Michael Gerndt, Daniel Lorenz, Allen Mal- ony, Wolfgang E. Nagel, Yury Oleynik, Peter Philippen, Pavel Saviankou, Dirk Schmidl, Sameer Shende, Ronny Tschüter, M...
2012
-
[50]
Andreas Knüpfer, Ronny Brendel, Holger Brunst, Hartmut Mix, and Wolfgang E Nagel. 2006. LNCS 3992 - Introducing the Open Trace Format (OTF). https: //doi.org/doi:10.3233/978-1-61499-041-3-481
2006 doi
-
[51]
Greg Kramer. 2023. Direct Drive - Azure’s next-generation block storage ar- chitecture. https://storagedeveloper.org/events/agenda/session/347. [Online]. Accessed: 2024-02-13
2023
-
[52]
Wetherall, and Amin Vahdat
Gautam Kumar, Nandita Dukkipati, Keon Jang, Hassan Wassel, Xian Wu, Behnam Montazeri, Yaogong Wang, Kevin Springborn, Christopher Alfeld, Mike Ryan, David J. Wetherall, and Amin Vahdat. 2020. Swift: Delay is Simple and Effective for Congestion Control in the Datacenter. https:...
2020
-
[53]
Lawrence Livermore National Laboratory. [n. d.]. Lawrence Livermore Na- tional Laboratory’s El Capitan verified as world’s fastest supercomputer. LLNL([n. d.]). https://www.llnl.gov/article/52061/lawrence-livermore-national- laboratorys-el-capitan-verified-worlds-fastest-supercomputer
-
[54]
Tian Li, Jie Zhong, Ji Liu, Wentao Wu, and Ce Zhang. 2018. Ease.ml: towards multi-tenant resource sharing for machine learning workloads.Proc. VLDB Endow.11, 5 (Oct. 2018), 607–620. https://doi.org/10.1145/3177732.3177737
2018
-
[55]
Yuliang Li, Rui Miao, Hongqiang Harry Liu, Yan Zhuang, Fei Feng, Lingbo Tang, Zheng Cao, Ming Zhang, Frank Kelly, Mohammad Alizadeh, and Minlan Yu
-
[56]
Ruofan Liang, Bingsheng He, Shengen Yan, and Peng Sun. 2022. A Simulation Platform for Multi-tenant Machine Learning Services on Thousands of GPUs. arXiv:2201.03175 [cs.DC] https://arxiv.org/abs/2201.03175
2022 arXiv
-
[57]
Linux Kernel Documentation. 2025. Extended Berkeley Packet Filter (eBPF). https://www.kernel.org/doc/html/latest/bpf/index.html. Accessed: February 04,
2025
-
[58]
Yuanwei Lu, Guo Chen, Bojie Li, Kun Tan, Yongqiang Xiong, Peng Cheng, Jian- song Zhang, Enhong Chen, and Thomas Moscibroda. 2018. Multi-Path Transport for RDMA in Datacenters. In15th USENIX Symposium on Networked Systems Design and Implementation (NSDI 18). USENIX Association,...
2018
-
[59]
Dorian Maillard. 2024.Nvidia’s AI market dominance: Can anyone mount a serious challenge?https://www.techradar.com/pro/nvidias-ai-market-dominance-can- anyone-mount-a-serious-challenge Accessed: 2025-01-28
2024
-
[60]
Carothers, Robert B
Misbah Mubarak, Christopher D. Carothers, Robert B. Ross, and Philip Carns
-
[61]
2025.NVIDIA Collective Communications Library (NCCL) Documentation
NVIDIA Corporation. 2025.NVIDIA Collective Communications Library (NCCL) Documentation. https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/ index.html Accessed: 2025-01-28
2025
-
[62]
2025.NVIDIA Nsight Systems
NVIDIA Corporation. 2025.NVIDIA Nsight Systems. https://developer.nvidia. com/nsight-systems
2025
-
[63]
Vladimir Olteanu, Haggai Eran, Dragos Dumitrescu, Adrian Popa, Cristi Baciu, Mark Silberstein, Georgios Nikolaidis, Mark Handley, and Costin Raiciu. 2022. An edge-queued datagram service for all datacenter traffic. In19th USENIX Symposium on Networked Systems Design and Implem...
2022
-
[64]
T. V. Pham, C. Steger, B. Rockel, K. Keuler, I. Kirchner, M. Mertens, D. Rieger, G. Zängl, and B. Früh. 2021. ICON in Climate Limited-area Mode (ICON release version 2.6.1): a new regional climate model.Geoscientific Model Development14, 2 (2021), 985–1005. https://doi.org/10....
2021 doi
-
[65]
Costin Raiciu, Christopher Pluntke, Sebastien Barre, Adam Greenhalgh, Damon Wischik, and Mark Handley. 2010. Data Center Networking with Multipath TCP. InProceedings of the 9th ACM SIGCOMM Workshop on Hot Topics in Networks (Monterey, California)(Hotnets-IX). Association for C...
2010
-
[66]
Saeed Rashidi, Joongun Park, Abhilash Kolluri, and Taekyung Heo. 2024. Chakra Execution Trace Collection – A Comprehensive Guide on Merging PyTorch and ATLAHS: An Application-centric Network Simulator T oolchain for A I, HPC, and Distributed S torage Kineto Traces. https://git...
2024
-
[67]
Thomas Rausch, Waldemar Hummer, and Vinod Muthusamy. 2020. PipeSim: Trace-driven Simulation of Large-Scale AI Operations Platforms. arXiv:2006.12587 [cs.DC] https://arxiv.org/abs/2006.12587
2020 arXiv
-
[68]
Ganger, Randy H
Charles Reiss, Alexey Tumanov, Gregory R. Ganger, Randy H. Katz, and Michael A. Kozuch. 2012. Heterogeneity and dynamicity of clouds at scale: Google trace analysis. InProceedings of the Third ACM Symposium on Cloud Computing (SoCC). 7:1–7:13. https://doi.org/10.1145/2391229.2391236
2012
-
[69]
George F Riley and Thomas R Henderson. 2010. The ns-3 network simulator. In Modeling and tools for network simulation. Springer, 15–34
2010
-
[70]
Alastair Robertson and contributors. 2025. bpftrace: A High-Level Tracing Language for Linux. https://github.com/iovisor/bpftrace Version 0.22.1; accessed February 04, 2025; licensed under Apache-2.0
2025
-
[71]
Hongzhang Shan, Filip Blagojević, Seung-Jai Min, Paul Hargrove, Haoqiang Jin, Karl Fuerlinger, Alice Koniges, and Nicholas J. Wright. 2010. A programming model performance study using the NAS parallel benchmarks.Sci. Program.18, 3–4 (Aug. 2010), 153–167. https://doi.org/10.115...
2010 doi
-
[72]
Siyuan Shen, Langwen Huang, Marcin Chrapek, Timo Schneider, Jai Dayal, Manisha Gajbe, Robert Wisniewski, and Torsten Hoefler. 2024. LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming. In SC24: International Conference for High Performance Co...
2024 arXiv
-
[73]
Shende and Allen D
Sameer S. Shende and Allen D. Malony. 2006. The Tau Parallel Performance System.The International Journal of High Performance Computing Applications 20 (5 2006), 287–311. Issue 2. https://doi.org/10.1177/1094342006064482
2006 doi
-
[74]
Srinivas Sridharan, Taekyung Heo, Louis Feng, Zhaodong Wang, Matt Bergeron, Wenyin Fu, Shengbao Zheng, Brian Coutinho, Saeed Rashidi, Changhai Man, and Tushar Krishna. 2023. Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces. arXiv:230...
2023 arXiv
-
[75]
sstsimulator. 2025. SST-DUMPI Trace Library. https://github.com/sstsimulator/ sst-dumpi. Accessed: 2025-02-14
2025
-
[76]
A. P. Thompson, H. M. Aktulga, R. Berger, D. S. Bolintineanu, W. M. Brown, P. S. Crozier, P. J. in ’t Veld, A. Kohlmeyer, S. G. Moore, T. D. Nguyen, R. Shan, M. J. Stevens, J. Tranchida, C. Trott, and S. J. Plimpton. 2022. LAMMPS - a flexible simulation tool for particle-based...
2022
-
[77]
Tikir, Michael A
Mustafa M. Tikir, Michael A. Laurenzano, Laura Carrington, and Allan Snavely
-
[78]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guil- laume Lample. 2023. LLaMA: Open and Efficient Foundation ...
2023 arXiv
-
[79]
University of Massachusetts Amherst. 2016. UMass Trace Repository: Storage. https://traces.cs.umass.edu/docs/traces/storage/. Accessed: 4 February 2025. Copyright©2016–2024 University of Massachusetts Amherst
2016
-
[80]
Erico Vanini, Rong Pan, Mohammad Alizadeh, Parvin Taheri, and Tom Edsall
-
[81]
András Varga and Rudolf Hornig. 2008. An overview of the OMNeT++ simulation environment. InProceedings of the 1st International Conference on Simulation Tools and Techniques for Communications, Networks and Systems & Workshops (Marseille, France)(Simutools ’08). ICST (Institut...
2008
-
[82]
Kurt Wagner. [n. d.]. Meta Is Building New $800 Million AI-Focused Data Center in Indiana.Bloomberg([n. d.]). https://www.bloomberg.com/news/articles/2024- 01-25/meta-building-new-800-million-ai-focused-data-center-in- indiana?embedded-checkout=true
2024
-
[83]
Xizheng Wang, Qingxu Li, Yichi Xu, Gang Lu, Dan Li, Chen Li, Heyang Zhou, Linkang Zheng, Sen Zhang, Yikai Zhu, Yang Liu, Pengcheng Zhang, Kun Qian, and Kunling He. 2025. SimAI: Unifying Architecture Design and Performance Tunning for Large-Scale Large Language Model Training w...
2025
-
[84]
William Won, Taekyung Heo, Saeed Rashidi, Srinivas Sridharan, Sudarshan Srinivasan, and Tushar Krishna. 2023. ASTRA-sim2.0: Modeling Hierarchi- cal Networks and Disaggregated Systems for Large-model Training at Scale. arXiv:2303.14006 [cs.DC] https://arxiv.org/abs/2303.14006
2023 arXiv
-
[85]
Feroz Zahid, Ernst Gunnar Gran, Bartosz Bogdański, Bjørn Dag Johnsen, and Tor Skeie. 2017. Efficient network isolation and load balancing in multi-tenant HPC clusters.Future Generation Computer Systems72 (2017), 145–162. https: //doi.org/10.1016/j.future.2016.04.003
2017 doi
-
[86]
Ben Zaitlen. 2021. NVIDIA Tools Extension API: An Annotation Tool for Profil- ing Code in Python and C/C++. https://developer.nvidia.com/blog/nvidia-tools- extension-api-nvtx-annotation-tool-for-profiling-code-in-python-and-c-c/ Ac- cessed: 2025-01-28
2021
-
[87]
Jidong Zhai, Wenguang Chen, and Weimin Zheng. 2010. PHANTOM: predicting performance of parallel applications on large-scale parallel machines using a single node. InProceedings of the 15th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming(Bangalore, Indi...
2010
-
[88]
In14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17)
Let It Flow: Resilient Asymmetric Load Balancing with Flowlet Switching. In14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). USENIX Association, Boston, MA, 407–420. https://www.usenix.org/ conference/nsdi17/technical-sessions/presentation/vanini
-
[89]
Zheng, Gunavardhan Kakulapati, and L.V
G. Zheng, Gunavardhan Kakulapati, and L.V. Kale. 2004. BigSim: a parallel simu- lator for performance prediction of extremely large parallel machines. In18th International Parallel and Distributed Processing Symposium, 2004. Proceedings. 78–. https://doi.org/10.1109/IPDPS.2004.1303013
2004 arXiv
-
[96]
Tianqi Zhang, Yanqi Zhang, Yuandong Tian, Lin Ma, Wei Lin, and Bin Cui
-
[1995]
InProceedings of the Seventh Annual ACM Symposium on Parallel Algorithms and Architectures(Santa Barbara, California, USA)(SPAA ’95)
LogGP: incorporating long messages into the LogP model—one step closer towards a realistic model for parallel computation. InProceedings of the Seventh Annual ACM Symposium on Parallel Algorithms and Architectures(Santa Barbara, California, USA)(SPAA ’95). Association for Comp...
-
[2009]
In2009 DoD High Performance Computing Modernization Program Users Group Conference
PSINS: An Open Source Event Tracer and Execution Simulator. In2009 DoD High Performance Computing Modernization Program Users Group Conference. 444–449. https://doi.org/10.1109/HPCMP-UGC.2009.73
2009 doi
-
[2016]
InSC16: International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
Evaluating HPC Networks via Simulation of Parallel Workloads . InSC16: International Conference for High Performance Computing, Networking, Storage and Analysis (SC). IEEE Computer Society, Los Alamitos, CA, USA, 154–165. https://doi.org/10.1109/SC.2016.13
-
[2017]
https: //doi.org/10.1109/TPDS.2016.2543725
Enabling Parallel Simulation of Large-Scale HPC Network Systems.IEEE Transactions on Parallel and Distributed Systems28, 1 (2017), 87–100. https: //doi.org/10.1109/TPDS.2016.2543725
2017
-
[2019]
InProceedings of the ACM Special Interest Group on Data Communication(Beijing, China)(SIGCOMM ’19)
HPCC: high precision congestion control. InProceedings of the ACM Special Interest Group on Data Communication(Beijing, China)(SIGCOMM ’19). Association for Computing Machinery, New York, NY, USA, 44–58. https: //doi.org/10.1145/3341302.3342085
-
[2022]
In19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22)
MLaaS in the Wild: Workload Analysis and Scheduling in Large-scale Heterogeneous GPU Clusters. In19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22). 827–841. https://github.com/alibaba/ clusterdata/blob/master/cluster-trace-gpu-v2020/README.md
-
[2024]
InSC24: International Conference for High Performance Computing, Networking, Storage and Analysis
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects. InSC24: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1–15. https://doi.org/10.1109/sc41406. 2024.00039
-
[2025]
eBPF extends the classic BPF mechanism to run sandboxed programs in the Linux kernel for tracing, networking, and more
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.