Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read ATLAHS claims that tracing real AI, HPC, and storage applications into GOAL graphs lets a flexible simulator predict runtimes within 5% error.

desk verdict A genuinely useful open-source simulator toolchain with solid HPC validation, but the AI <5% claim rests on unmeasured LogGOPS parameters and the storage support is unvalidated, so treat the headline accuracy claim as conditional. read the letter →

arxiv 2505.08936 v1 pith:WJ3QXMIH submitted 2025-05-13 cs.DC

classification cs.DC
keywords application-centricnetworksimulationGOALformatexecutiontracingLLMtrainingMPIworkloadsdistributedstorageLogGOPSimcongestioncontrol
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

ATLAHS is an open-source toolchain that traces real applications from AI, HPC, and distributed storage into GOAL directed acyclic graphs of send, receive, and compute tasks, then replays those graphs on message-level or packet-level network simulators. The paper's central claim is that this application-centric approach predicts the runtimes of real workloads—training iterations for Llama and Mixture-of-Experts models and runs of five HPC codes—with errors consistently below 5% across the validated configurations. The paper also claims that ATLAHS is faster than the current AI-focused simulator AstraSim, produces smaller trace files, and supports multi-job and multi-tenant scenarios by merging GOAL graphs. A sympathetic reader would care because simulators that rely on synthetic microbenchmarks can miss congestion behavior that real traces expose, as demonstrated in the paper's congestion-control case study.

What carries the argument

The load-bearing mechanism is the GOAL directed acyclic graph, a format in which every workload is a set of send, receive, and calc (computation) vertices with dependency edges, assigned to compute streams to model concurrency. ATLAHS's tracer front ends convert real executions into these graphs: NCCL collectives are decomposed into point-to-point send/receive schedules according to algorithm, channel, and protocol settings; MPI operations are converted through PMPI tracing; block I/O commands are converted through a bpftrace-based tracer. Dummy zero-cost vertices synchronize parallel streams, and merging graphs from multiple jobs models multi-tenancy. The GOAL graphs are then scheduled by ATLAHS and executed on pluggable backends—LogGOPSim for message-level speed and htsim for packet-level fidelity—through a minimal interface of send, recv, calc, and eventOver. This combination is what lets one unified representation span AI, HPC, and storage.

What would settle it

On the same cluster used for tracing, directly measure the interconnect's latency, overhead, and per-byte transfer time and rerun the Llama 7B 128-GPU validation with those measured values; if the predicted iteration time deviates from the measured runtime by more than 5%, the central accuracy claim depends on the borrowed parameters rather than on the toolchain itself.

Watch

Extended reading notes

Core claim

The core discovery is that a single, compact intermediate representation—the GOAL format, with only send, receive, and computation vertices and edges expressing dependencies—is expressive enough to capture the communication and computation patterns of LLM training, MPI scientific codes, and distributed storage traffic, and to reproduce measured application runtimes within 5% error. The toolchain obtains these graphs by tracing NCCL through a GPU profiler with added annotations, tracing MPI through the PMPI interface, and tracing block I/O through an eBPF-based tracer; it then decomposes NCCL collectives into point-to-point schedules, merges per-GPU graphs into per-node graphs, and replaces intra-node communication with computation. On the two configurations where AstraSim ran successfully, ATLAHS was both more accurate and faster, and its GOAL trace files were consistently smaller than AstraSim's Chakra traces. The paper further shows that the message-level and packet-level backends agree within 1-2% on fully provisioned symmetric networks, while diverging by over 120% when oversubscription causes packet drops that only the packet-level backend can see.

Load-bearing premise

The AI workload predictions rest on network-cost parameters (latency, overhead, per-byte transfer time) borrowed from benchmarks of similar hardware rather than measured on the target cluster, and the paper does not show how much the predictions would change if those parameters were different.

Editorial extensions

If this is right

  • Application traces, not synthetic microbenchmarks, become the default workload for evaluating network designs, since they expose issues such as Swift's multi-hop congestion weakness that microbenchmarks hide.
  • Network architects can safely use the fast message-level backend for fully provisioned symmetric topologies, but must switch to a packet-level backend when oversubscription, packet drops, or queue dynamics matter.
  • Multi-job and multi-tenant performance questions—such as how job placement affects a shared cluster—can be answered by merging GOAL graphs of different applications.
  • The storage case study implies that congestion control choice (sender-based vs receiver-based) significantly changes storage request completion under oversubscribed topologies.
  • Because GOAL trace files are several times smaller than Chakra traces, sharing reproducible workload traces at large scale becomes cheaper.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the 5%-error claim is robust, a natural next step is to use ATLAHS itself as a fast screening tool for congestion-control and topology changes before packet-level simulation, since the two backends agree when the network is well provisioned.
  • The acknowledged static-DAG limitation suggests the approach would under-predict runtime for dynamically scheduled communication, such as fault-tolerant storage protocols or data-dependent GNN training; a testable extension is to add dynamic vertices to GOAL and compare against real runs.
  • The lack of a sensitivity analysis for the AI LogGOPS parameters leaves open whether the reported accuracy is stable; a direct measurement of those parameters on the target hardware would settle it.
  • Combining traces from different domains in one graph (as done for Llama and LULESH) could be extended to model shared infrastructure contention more generally, but the paper only demonstrates network-level contention, not memory or cache interference.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents ATLAHS, an application-centric network simulation toolchain that translates real-world traces from AI (NCCL), HPC (MPI), and distributed storage (block I/O) into GOAL-style DAGs and simulates them through multiple backends (LogGOPSim, htsim, NS-3). It validates the toolchain on six LLM/MoE training configurations on the Alps cluster and fifteen HPC configurations on a CSCS testbed, reporting prediction errors mostly within 5% for both ATLAHS LGS and ATLAHS htsim, while also comparing against AstraSim for two configurations and reporting significantly smaller trace files than Chakra. Case studies illustrate congestion-control effects on storage traffic, differences between message-level and packet-level backends, and job-placement effects in a shared cluster.

Significance. ATLAHS is a potentially valuable open-source infrastructure: it unifies diverse workload formats under GOAL, releases a substantial trace dataset, supports multi-job and multi-tenant scenarios, and demonstrates that a message-level backend can match packet-level accuracy in fully provisioned, symmetric topologies. The HPC validation is broad (15 configurations) and uses measured LogGOPS parameters, and the trace-size and simulation-speed comparisons are concrete and reproducible. If the AI-accuracy claim can be supported against parameter uncertainty, this work would be a solid contribution to network-simulation practice for AI/HPC convergence.

major comments (3)
  1. [§5.2, Fig. 8] The AI-validation accuracy claim depends entirely on LogGOPS parameters (L=3700 ns, o=200 ns, g=5 ns, O=0, G=0.04 ns/B, S=0) that are estimated from external benchmarking works rather than measured on the Alps cluster where the traces were collected. No sensitivity analysis is reported, so the asserted consistent <5% error for AI workloads is a point estimate on six configurations, not a demonstrated robustness property. Since several workloads are communication-dominated (MoE 8x70B has only 5.3% non-overlapped computation; Llama 70B has 9.9%), modest errors in G or L translate almost directly into per-iteration time errors. Please add a sensitivity analysis over plausible parameter ranges (especially G and L) for both ATLAHS backends, or measure the parameters directly on the validation system, and also isolate the effect of the static-GOAL-DAG simplification acknowledged in §7.
  2. [Fig. 10, §5.3] The text states that prediction error 'remains consistently below 5% across all cases and applications,' but Fig. 10 includes a reported error of -5.2% (for the htsim backend on one of the HPC configurations). This contradicts the stated claim. Please correct the text or the figure, or clarify whether the -5.2% value is a typo or an excluded outlier.
  3. [§3.1.3, §6.1] The distributed-storage component is described as a first-class part of the toolchain and appears in the title and abstract, but it is never validated against measured storage-system runtimes or packet-level ground truth. The storage case study compares two congestion-control algorithms using traces generated under assumptions about Azure Direct Drive from public documentation; the results are relative and have no accuracy check. Please add a validation experiment for the storage path, or explicitly scope the accuracy claims in the abstract and conclusion to AI and HPC workloads.
minor comments (5)
  1. [Abstract / §5.3] The abstract claims 'consistently less than 5% error' without noting that the storage component is unvalidated; please qualify the claim to the validated AI and HPC domains.
  2. [§5.1] The sentence 'From our testing, the runtime of complex traces is reduced from 10× to100 × the after the improvements' contains a typo and should be reworded to state the speedup range clearly.
  3. [Fig. 8 / Fig. 10] Backend naming is inconsistent ('HTSIM' in Fig. 10 caption vs 'htsim' in the text); also, in Fig. 10 the two sets of red percentages are not explicitly labelled in the visualization, making it hard to map errors to backends.
  4. [§5.2] The claim that ATLAHS 'consistently outperforms AstraSim' in accuracy is based on only two configurations where AstraSim completed; consider phrasing this as 'in the two successful AstraSim comparisons' to match the evidence.
  5. [References] Reference [82] is incomplete; it ends with 'embedded-checkout=true' and appears to be a truncated Bloomberg URL.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; ATLAHS accuracy is validated against measured runtimes with parameters taken from external benchmarks, not fitted to the predictions.

full rationale

The paper's central claim is that ATLAHS predicts runtimes of real AI and HPC workloads with consistently less than 5% error. This claim is supported by direct comparison against measured application runtimes on real clusters (Alps for AI, a CSCS testbed for HPC). The AI LogGOPS parameters (L=3700 ns, o=200 ns, g=5 ns, O=0, G=0.04 ns/B, S=0) are explicitly estimated from external benchmarking works rather than fitted to the six validation points, and the HPC parameters are measured with Netgauge. No equation in the paper reduces to the measured runtime by construction: the GOAL representation of computation and communication is independent of the network backends, and the two backends (LogGOPSim and htsim) are cross-checked against one another and against AstraSim. The paper does cite its own prior work on GOAL and LogGOPSim, but these citations are used as background for the trace format choice and do not carry the load of the accuracy validation, which rests on measured ground truth. The acknowledged simplifications (static GOAL DAG, no CUDA cross-stream dependencies, unmeasured AI LogGOPS parameters) are correctness-robustness concerns, not circularity: they could make the predictions wrong, but they do not make the predictions equal to the inputs. No fitted input is renamed as a prediction, and no self-citation chain forces the result. Therefore the appropriate finding is no significant circularity.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The central accuracy claim rests on the estimated AI LogGOPS parameters, the assumed fidelity of NCCL collective decomposition, the static-DAG simplification, and the unvalidated Direct Drive model. No new physical entities are introduced.

free parameters (1)
  • AI LogGOPS parameter set (L, o, g, G) = L=3700, o=200, g=5, O=0, G=0.04, S=0 (ns)
    Estimated from published benchmarks of Grace Hopper and Dragonfly interconnects (Fusco et al., De Sensi et al., Groves et al.), not measured on the Alps cluster; used by both the LGS and htsim backends in Section 5.2. No sensitivity analysis is reported.
assumptions (5)
  • domain assumption GOAL DAG abstraction is sufficient to model communication and computation for AI, HPC, and storage workloads.
    The paper cites prior validation of GOAL and LogGOPSim [21, 38, 39, 72] and uses GOAL as the universal intermediate format throughout.
  • domain assumption The NCCL collective-to-P2P decomposition in Stage 3 is accurate across NCCL algorithms, protocols, and channel counts.
    The paper omits detailed breakdowns ('full implementations are available in the source code') and relies on manual analysis of NCCL configurations.
  • domain assumption Replacing intra-node sends/receives with calc vertices preserves simulation fidelity.
    Section 3.1.2 Stage 4; the cost is inferred from profiling data, but no independent validation of intra-node behavior is given.
  • domain assumption A static GOAL DAG without CUDA inter-stream dependencies or dynamic scheduling is sufficient for the validated workloads.
    Stated as a limitation in Section 7; the authors argue it does not significantly impact accuracy for NCCL workloads.
  • domain assumption The Azure Direct Drive model, based on public documentation, captures the storage architecture's network behavior.
    Section 3.1.3: 'As Direct Drive is proprietary, we made assumptions based on public documentation.' The storage case study is not validated against real measurements.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage." pith.science (2026). https://pith.science/paper/WJ3QXMIH

@misc{pith2026250508936,
  author       = {Pith},
  title        = {Pith review of: ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WJ3QXMIH}},
  note         = {Machine review of arXiv:2505.08936}
}
read the original abstract

Network simulators play a crucial role in evaluating the performance of large-scale systems. However, existing simulators rely heavily on synthetic microbenchmarks or narrowly focus on specific domains, limiting their ability to provide comprehensive performance insights. In this work, we introduce ATLAHS, a flexible, extensible, and open-source toolchain designed to trace real-world applications and accurately simulate their workloads. ATLAHS leverages the GOAL format to model communication and computation patterns in AI, HPC, and distributed storage applications. It supports multiple network simulation backends and handles multi-job and multi-tenant scenarios. Through extensive validation, we demonstrate that ATLAHS achieves high accuracy in simulating realistic workloads (consistently less than 5% error), while significantly outperforming AstraSim, the current state-of-the-art AI systems simulator, in terms of simulation runtime and trace size efficiency. We further illustrate ATLAHS's utility via detailed case studies, highlighting the impact of congestion control algorithms on the performance of distributed storage systems, as well as the influence of job-placement strategies on application runtimes.

Figures

Figures reproduced from arXiv: 2505.08936 by the authors.

Figure 1
Figure 1. A illustrates a space-time diagram of a realistic training scenario for Large Language Models (LLMs), showing overlapping communication from data parallelism (DP) and pipeline parallelism (PP). B depicts a network-level view demonstrating how PP victim flows become congested due to simultaneous DP ring allreduce communications within a two-level fat tree topology. C compares the performance of Swift and MPRDMA conge… view at source ↗
Figure 2
Figure 2. Overview of the ATLAHS toolchain. Application and hardware components are represented in shades of [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. A shows an example GOAL schedule of node 0 in its textual format, while B shows the visualization of the same schedule as a DAG. Vertices in green are assigned to compute stream 0 to execute while vertex𝑙3 in teal is assigned to be executed on compute stream 1. storage and execution efficiency, GOAL schedules are stored and executed in a compact binary format. We chose GOAL as the intermediate trace format for ATLAH… view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Example of an AI application with 2 nodes and 4 [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: This structured approach ensures a modular design, making [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]
Figure 5
Figure 5. Figure 5: An example showing the 4 stages GOAL file generation for large-scale distributed AI applications. [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: A provides an overview of how applications in￾teract with Azure Direct Drive. B presents a space-time diagram illustrating the sequence of operations involved in a read request. Sends and receives are depicted as dotted green blocks, while computation is shown as paste…
Figure 7
Figure 7. Figure 7: ALTAHS APIs and an overview of its code. In the [PITH_FULL_IMAGE:figures/full_fig_p006_7.png]
Figure 8
Figure 8. Figure 8: Comparison of measured runtimes against predicted runtimes from ATLAHS and AstraSim for various AI training [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 9
Figure 9. Figure 9: Trace size comparison of GOAL and Chakra. The [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]
Figure 10
Figure 10. Figure 10: Comparison of measured and ATLAHS-predicted runtimes for HPC applications from various domains. The second [PITH_FULL_IMAGE:figures/full_fig_p009_10.png]
Figure 11
Figure 11. Figure 11: Comparison of the Message Completion Time [PITH_FULL_IMAGE:figures/full_fig_p009_11.png]
Figure 12
Figure 12. Figure 12: Comparison of the runtime of ATLAHS LGS and [PITH_FULL_IMAGE:figures/full_fig_p010_12.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO

    cs.DC 2025-06 conditional novelty 2.0 of 10

    This survey classifies distributed DNN training simulators into analytical, profiling-based, and execution-driven categories, and compares them alongside TCO and carbon-emission models.

Reference graph

Works this paper leans on

98 extracted references · 48 canonical work pages · cited by 1 Pith paper

  1. [1]

    [n. d.]. Broadcom htsim repository. https://github.com/Broadcom/csg-htsim. Accessed: 2025-02-13

  2. [2]

    [n. d.]. Ultra Ethernet Consortium. https://ultraethernet.org/. Accessed: 2025- 01-13

  3. [3]

    Evensky, Joseph P

    Helgi Adalsteinsson, Scott Cranford, David A. Evensky, Joseph P. Kenny, Jackson Mayo, Ali Pinar, and Curtis L. Janssen. 2010. A Simulator for Large-Scale Parallel Computer Architectures.Int. J. Distrib. Syst. Technol.1, 2 (April 2010), 57–73

  4. [4]

    2025.ROCm Communication Collectives Library (RCCL) Documentation

    Advanced Micro Devices, Inc. 2025.ROCm Communication Collectives Library (RCCL) Documentation. https://rocm.docs.amd.com/projects/rccl/en/latest/what- is-rccl.html Accessed: 2025-01-28

  5. [5]

    Ionescu, Klaus E

    Albert Alexandrov, Mihai F. Ionescu, Klaus E. Schauser, and Chris Scheiman

  6. [6]

    Jens Axboe. 2005. blktrace: A Block I/O Tracing Mechanism for Linux. https: //www.kernel.org/doc/Documentation/block/blktrace.txt. Accessed: February 04, 2025; an open-source tool for tracing block I/O events in Linux

  7. [7]

    Tal Ben-Nun and Torsten Hoefler. 2019. Demystifying Parallel and Distributed Deep Learning: An In-depth Concurrency Analysis.ACM Comput. Surv.52, 4, Article 65 (Aug. 2019), 43 pages

  8. [8]

    Maciej Besta, Julia Barth, Eric Schreiber, Ales Kubicek, Afonso Catarino, Robert Gerstenberger, Piotr Nyczyk, Patrick Iff, Yueling Li, Sam Houliston, Tomasz Sternal, Marcin Copik, Grzegorz Kwaśniewski, Jürgen Müller, Łukasz Flis, Hannes Eberhard, Hubert Niewiadomski, and Torsten Hoefler. 2025. Reasoning Language Models: A Blueprint. arXiv:2501.11223 [cs.A...

Show all 98 references
  1. [9]

    Maciej Besta and Torsten Hoefler. 2023. Parallel and Distributed Graph Neural Networks: An In-Depth Concurrency Analysis. arXiv:2205.09702 [cs.LG] https: //arxiv.org/abs/2205.09702

  2. [10]

    Maciej Besta, Marcel Schneider, Salvatore Di Girolamo, Ankit Singla, and Torsten Hoefler. 2021. Towards Million-Server Network Simulations on Just a Laptop. arXiv:2105.12663 [cs.NI] https://arxiv.org/abs/2105.12663

  3. [11]

    Tommaso Bonato, Abdul Kabbani, Ahmad Ghalayini, Michael Papamichael, Mohammad Dohadwala, Lukas Gianinazzi, Mikhail Khalilov, Elias Acher- mann, Daniele De Sensi, and Torsten Hoefler. 2025. REPS: Recycled En- tropy Packet Spraying for Adaptive Load Balancing and Failure Mitigat...

  4. [12]

    Tommaso Bonato, Abdul Kabbani, Daniele De Sensi, Rong Pan, Yanfang Le, Costin Raiciu, Mark Handley, Timo Schneider, Nils Blach, Ahmad Ghalayini, Daniel Alves, Michael Papamichael, Adrian Caulfield, and Torsten Hoefler. 2024. FASTFLOW: Flexible Adaptive Congestion Control for H...

  5. [13]

    Carothers, D

    C.D. Carothers, D. Bauer, and S. Pearce. 2000. ROSS: a high-performance, low memory, modular time warp system. InProceedings Fourteenth Workshop on Parallel and Distributed Simulation. 53–60. https://doi.org/10.1109/PADS.2000. 847144

  6. [14]

    Henri Casanova, Arnaud Giersch, Arnaud Legrand, Martin Quinson, and Frédéric Suter. 2025. Lowering entry barriers to developing custom simulators of dis- tributed applications and platforms with SimGrid.Parallel Comput.123 (2025), 103–125. https://doi.org/10.1016/j.parco.2025.103125

  7. [15]

    Henri Casanova, Arnaud Legrand, and Martin Quinson. 2008. SimGrid: A Generic Framework for Large-Scale Distributed Experiments. InTenth Inter- national Conference on Computer Modeling and Simulation (uksim 2008). 126–131. https://doi.org/10.1109/UKSIM.2008.28

  8. [16]

    Abhishek Chard, R. G. Dreslinski, Thomas F. Wenisch, Greg Ganger, and An- drew A. Chien. 2018. On the Diversity of Cluster Workloads and Its Impact on Research Results. In2018 USENIX Annual Technical Conference (USENIX ATC 18). 533–546. https://www.pdl.cmu.edu/ATLAS/ Includes ...

  9. [17]

    Jaehong Cho, Minsu Kim, Hyunmin Choi, Guseul Heo, and Jongse Park. 2024. LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale. In2024 IEEE International Symposium on Workload Characteri- zation (IISWC). 15–29. https://doi.org/10.1109/IISWC6309...

  10. [18]

    Pierre-Nicolas Clauss, Mark Stillwell, Stephane Genaud, Frederic Suter, Henri Casanova, and Martin Quinson. 2011. Single Node On-Line Simulation of MPI Ap- plications with SMPI. In2011 IEEE International Parallel & Distributed Processing Symposium. 664–675. https://doi.org/10....

  11. [19]

    Storage Performance Council. 2025. SPC Trace File Format Specification. https: //skulddata.cs.umass.edu/traces/storage/SPC-Traces.pdf. Accessed: 2025-01-22

  12. [20]

    CSCS. [n. d.]. New Research Infrastructure: ’Alps’ Supercomputer Inaugurated. Swiss National Supercomputing Center([n. d.]). https://www.cscs.ch/publications/ news/2024/new-research-infrastructure-alps-supercomputer-inaugurated

  13. [21]

    Daniele De Sensi, Tiziano De Matteis, Konstantin Taranov, Salvatore Di Girolamo, Tobias Rahn, and Torsten Hoefler. 2022. Noise in the Clouds: Influence of Net- work Performance Variability on Application Scalability.Proc. ACM Meas. Anal. Comput. Syst.6, 3, Article 49 (dec 2022...

  14. [22]

    Daniele De Sensi, Lorenzo Pichetti, Flavio Vella, Tiziano De Matteis, Zebin Ren, Luigi Fusco, Matteo Turisini, Daniele Cesarini, Kurt Lust, Animesh Trivedi, Duncan Roweth, Filippo Spiga, Salvatore Di Girolamo, and Torsten Hoefler

  15. [23]

    Ewa Deelman, Rafael Ferreira da Silva, Gideon Juve, Mats Rynge, Karan Vahi, and Miron Livny. 2022. WfCommons: A Framework for Enabling Scientific Workflow Research and Development.Future Generation Computer Systems129 (2022), 166–182. https://doi.org/10.1016/j.future.2022.01.011

  16. [24]

    Denzel, Jian Li, Peter Walker, and Yuho Jin

    Wolfgang E. Denzel, Jian Li, Peter Walker, and Yuho Jin. 2010. A Framework for End-to-End Simulation of High-performance Computing Systems.SIM- ULATION86, 5-6 (2010), 331–350. https://doi.org/10.1177/0037549709340840 arXiv:https://doi.org/10.1177/0037549709340840

  17. [25]

    Jiangfei Duan, Xiuhong Li, Ping Xu, Xingcheng Zhang, Shengen Yan, Yun Liang, and Dahua Lin. 2024. Proteus: Simulating the Performance of Distributed DNN Training.IEEE Transactions on Parallel and Distributed Systems35, 10 (2024), 1867–1878. https://doi.org/10.1109/TPDS.2024.3443255

  18. [26]

    Jiangfei Duan, Shuo Zhang, Zerui Wang, Lijuan Jiang, Wenwen Qu, Qinghao Hu, Guoteng Wang, Qizhen Weng, Hang Yan, Xingcheng Zhang, Xipeng Qiu, Dahua Lin, Yonggang Wen, Xin Jin, Tianwei Zhang, and Peng Sun. 2024. Efficient Training of Large Language Models on Distributed Infrast...

  19. [27]

    Feitelson

    Dror G. Feitelson. [n. d.]. The Parallel Workloads Archive. https://www.cs.huji. ac.il/labs/parallel/workload/. Accessed: 2025-04-12

  20. [28]

    Yinxiao Feng, Yuchen Wei, Dong Xiang, and Kaisheng Ma. 2024. Evaluating Chiplet-based Large-Scale Interconnection Networks via Cycle-Accurate Packet- Parallel Simulation. In2024 USENIX Annual Technical Conference (USENIX ATC 24). USENIX Association, Santa Clara, CA, 731–747. h...

  21. [29]

    Mackenzie Ferguson. 2025. Nvidia’s Unstoppable Rise: Dominating the AI Chip Market.OpenTools.ai(2025). https://opentools.ai/news/nvidias-unstoppable- rise-dominating-the-ai-chip-market Accessed: 2025-01-28

  22. [30]

    Luigi Fusco, Mikhail Khalilov, Marcin Chrapek, Giridhar Chukkapalli, Thomas Schulthess, and Torsten Hoefler. 2024. Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip. arXiv:2408.11556 [cs.DC] https://arxiv.org/abs...

  23. [31]

    Markus Geimer, Felix Wolf, Brian J. N. Wylie, Erika Ábrahám, Daniel Becker, and Bernd Mohr. 2010. The Scalasca performance toolset architecture.Concurrency and Computation: Practice and Experience22 (4 2010), 702–719. Issue 6. https: //doi.org/10.1002/cpe.1556 Shen et al

  24. [32]

    Ibrahim, Lenny Oliker, Nicholas J

    Taylor Groves, Ben Brock, Yuxin Chen, Khaled Z. Ibrahim, Lenny Oliker, Nicholas J. Wright, Samuel Williams, and Katherine Yelick. 2020. Performance Trade-offs in GPU Communication: A Study of Host and Device-initiated Ap- proaches. In2020 IEEE/ACM Performance Modeling, Benchma...

  25. [33]

    Shin, Yibo Zhu, Myeongjae Jeon, Junjie Qian, Hongqiang Liu, and Chuanxiong Guo

    Juncheng Gu, Mosharaf Chowdhury, Kang G. Shin, Yibo Zhu, Myeongjae Jeon, Junjie Qian, Hongqiang Liu, and Chuanxiong Guo. 2019. Tiresias: A GPU Cluster Manager for Distributed Deep Learning. In16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19). USENI...

  26. [34]

    Moore, Gianni Antichi, and Marcin Wójcik

    Mark Handley, Costin Raiciu, Alexandru Agache, Andrei Voinescu, Andrew W. Moore, Gianni Antichi, and Marcin Wójcik. 2017. Re-architecting datacenter networks and stacks for low latency and high performance. InProceedings of the Conference of the ACM Special Interest Group on D...

  27. [35]

    Thomas Henderson, Sally Floyd, and George Riley. 2006. ns3 Project Goals. Workshop on NS-2: the IP Network Simulator.(01 2006). https://doi.org/10.1145/ 1190455.1190468

  28. [36]

    Torsten Hoefler, Tommaso Bonato, Daniele De Sensi, Salvatore Di Girolamo, Shigang Li, Marco Heddes, Jon Belk, Deepak Goel, Miguel Castro, and Steve Scott. 2022. HammingMesh: A Network Topology for Large-Scale Deep Learning. arXiv:2209.01346 [cs.DC] https://arxiv.org/abs/2209.01346

  29. [37]

    Torsten Hoefler, Torsten Mehlan, Andrew Lumsdaine, and Wolfgang Rehm. 2007. Netgauge: A Network Performance Measurement Framework. InProceedings of High Performance Computing and Communications, HPCC’07(Houston, USA), Vol. 4782. Springer, 659–671

  30. [38]

    Torsten Hoefler, Timo Schneider, and Andrew Lumsdaine. 2010. Character- izing the Influence of System Noise on Large-Scale Applications by Simula- tion. InSC ’10: Proceedings of the 2010 ACM/IEEE International Conference for High Performance Computing, Networking, Storage and ...

  31. [39]

    Torsten Hoefler, Timo Schneider, and Andrew Lumsdaine. 2010. LogGOPSim: simulating large-scale applications in the LogGOPS model. InProceedings of the 19th ACM International Symposium on High Performance Distributed Computing (Chicago, Illinois)(HPDC ’10). Association for Comp...

  32. [40]

    Torsten Hoefler, Christian Siebert, and Andrew Lumsdaine. 2009. Group Operation Assembly Language - A Flexible Way to Express Collective Com- munication. In2009 International Conference on Parallel Processing. 574–581. https://doi.org/10.1109/ICPP.2009.70

  33. [41]

    2024.Intel®oneAPI Collective Communications Li- brary (oneCCL)

    Intel Corporation. 2024.Intel®oneAPI Collective Communications Li- brary (oneCCL). https://www.intel.com/content/www/us/en/docs/oneapi/ programming-guide/2024-1/intel-oneapi-collective-communications- library.html Accessed: 2025-01-28

  34. [42]

    Nikhil Jain, Abhinav Bhatele, Sam White, Todd Gamblin, and Laxmikant V. Kale

  35. [43]

    Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne...

  36. [44]

    Kale and Sanjeev Krishnan

    Laxmikant V. Kale and Sanjeev Krishnan. 1993. CHARM++: a portable concurrent object oriented system based on C++.SIGPLAN Not.28, 10 (oct 1993), 91–108. https://doi.org/10.1145/167962.165874

  37. [45]

    Chamberlain, Jonathan Cohen, Zachary Devito, Riyaz Haque, Dan Laney, Edward Luke, Felix Wang, David Richards, Martin Schulz, and Charles H

    Ian Karlin, Abhinav Bhatele, Jeff Keasler, Bradford L. Chamberlain, Jonathan Cohen, Zachary Devito, Riyaz Haque, Dan Laney, Edward Luke, Felix Wang, David Richards, Martin Schulz, and Charles H. Still. 2013. Exploring Traditional and Emerging Parallel Programming Models Using ...

  38. [46]

    Andreas Knüpfer, Ronny Brendel, Holger Brunst, Hartmut Mix, and Wolfgang E. Nagel. 2006. Introducing the open trace format (OTF). InProceedings of the 6th International Conference on Computational Science - Volume Part II(Reading, UK) (ICCS’06). Springer-Verlag, Berlin, Heidel...

  39. [47]

    Müller, and Wolfgang E

    Andreas Knüpfer, Holger Brunst, Jens Doleschal, Matthias Jurenz, Matthias Lieber, Holger Mickler, Matthias S. Müller, and Wolfgang E. Nagel. 2008. The Vampir Performance Analysis Tool-Set. InTools for High Performance Computing, Michael Resch, Rainer Keller, Valentin Himmler, ...

  40. [48]

    Andreas Knüpfer, Markus Geimer, Johannes Spazier, Joseph Schuchart, Michael Wagner, Dominic Eschweiler, and Matthias S. Müller. 2010. A generic attribute extension to OTF and its use for MPI replay.Procedia Computer Science1, 1 (2010), 2115–2124. https://doi.org/10.1016/j.proc...

  41. [49]

    Nagel, Yury Oleynik, Peter Philippen, Pavel Saviankou, Dirk Schmidl, Sameer Shende, Ronny Tschüter, Michael Wagner, Bert Wesarg, and Felix Wolf

    Andreas Knüpfer, Christian Rössel, Dieter an Mey, Scott Biersdorff, Kai Diethelm, Dominic Eschweiler, Markus Geimer, Michael Gerndt, Daniel Lorenz, Allen Mal- ony, Wolfgang E. Nagel, Yury Oleynik, Peter Philippen, Pavel Saviankou, Dirk Schmidl, Sameer Shende, Ronny Tschüter, M...

  42. [50]

    Andreas Knüpfer, Ronny Brendel, Holger Brunst, Hartmut Mix, and Wolfgang E Nagel. 2006. LNCS 3992 - Introducing the Open Trace Format (OTF). https: //doi.org/doi:10.3233/978-1-61499-041-3-481

  43. [51]

    Greg Kramer. 2023. Direct Drive - Azure’s next-generation block storage ar- chitecture. https://storagedeveloper.org/events/agenda/session/347. [Online]. Accessed: 2024-02-13

  44. [52]

    Wetherall, and Amin Vahdat

    Gautam Kumar, Nandita Dukkipati, Keon Jang, Hassan Wassel, Xian Wu, Behnam Montazeri, Yaogong Wang, Kevin Springborn, Christopher Alfeld, Mike Ryan, David J. Wetherall, and Amin Vahdat. 2020. Swift: Delay is Simple and Effective for Congestion Control in the Datacenter. https:...

  45. [53]

    Lawrence Livermore National Laboratory. [n. d.]. Lawrence Livermore Na- tional Laboratory’s El Capitan verified as world’s fastest supercomputer. LLNL([n. d.]). https://www.llnl.gov/article/52061/lawrence-livermore-national- laboratorys-el-capitan-verified-worlds-fastest-supercomputer

  46. [54]

    Tian Li, Jie Zhong, Ji Liu, Wentao Wu, and Ce Zhang. 2018. Ease.ml: towards multi-tenant resource sharing for machine learning workloads.Proc. VLDB Endow.11, 5 (Oct. 2018), 607–620. https://doi.org/10.1145/3177732.3177737

  47. [55]

    Yuliang Li, Rui Miao, Hongqiang Harry Liu, Yan Zhuang, Fei Feng, Lingbo Tang, Zheng Cao, Ming Zhang, Frank Kelly, Mohammad Alizadeh, and Minlan Yu

  48. [56]

    Ruofan Liang, Bingsheng He, Shengen Yan, and Peng Sun. 2022. A Simulation Platform for Multi-tenant Machine Learning Services on Thousands of GPUs. arXiv:2201.03175 [cs.DC] https://arxiv.org/abs/2201.03175

  49. [57]

    Linux Kernel Documentation. 2025. Extended Berkeley Packet Filter (eBPF). https://www.kernel.org/doc/html/latest/bpf/index.html. Accessed: February 04,

  50. [58]

    Yuanwei Lu, Guo Chen, Bojie Li, Kun Tan, Yongqiang Xiong, Peng Cheng, Jian- song Zhang, Enhong Chen, and Thomas Moscibroda. 2018. Multi-Path Transport for RDMA in Datacenters. In15th USENIX Symposium on Networked Systems Design and Implementation (NSDI 18). USENIX Association,...

  51. [59]

    Dorian Maillard. 2024.Nvidia’s AI market dominance: Can anyone mount a serious challenge?https://www.techradar.com/pro/nvidias-ai-market-dominance-can- anyone-mount-a-serious-challenge Accessed: 2025-01-28

  52. [60]

    Carothers, Robert B

    Misbah Mubarak, Christopher D. Carothers, Robert B. Ross, and Philip Carns

  53. [61]

    2025.NVIDIA Collective Communications Library (NCCL) Documentation

    NVIDIA Corporation. 2025.NVIDIA Collective Communications Library (NCCL) Documentation. https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/ index.html Accessed: 2025-01-28

  54. [62]

    2025.NVIDIA Nsight Systems

    NVIDIA Corporation. 2025.NVIDIA Nsight Systems. https://developer.nvidia. com/nsight-systems

  55. [63]

    Vladimir Olteanu, Haggai Eran, Dragos Dumitrescu, Adrian Popa, Cristi Baciu, Mark Silberstein, Georgios Nikolaidis, Mark Handley, and Costin Raiciu. 2022. An edge-queued datagram service for all datacenter traffic. In19th USENIX Symposium on Networked Systems Design and Implem...

  56. [64]

    T. V. Pham, C. Steger, B. Rockel, K. Keuler, I. Kirchner, M. Mertens, D. Rieger, G. Zängl, and B. Früh. 2021. ICON in Climate Limited-area Mode (ICON release version 2.6.1): a new regional climate model.Geoscientific Model Development14, 2 (2021), 985–1005. https://doi.org/10....

  57. [65]

    Costin Raiciu, Christopher Pluntke, Sebastien Barre, Adam Greenhalgh, Damon Wischik, and Mark Handley. 2010. Data Center Networking with Multipath TCP. InProceedings of the 9th ACM SIGCOMM Workshop on Hot Topics in Networks (Monterey, California)(Hotnets-IX). Association for C...

  58. [66]

    Saeed Rashidi, Joongun Park, Abhilash Kolluri, and Taekyung Heo. 2024. Chakra Execution Trace Collection – A Comprehensive Guide on Merging PyTorch and ATLAHS: An Application-centric Network Simulator T oolchain for A I, HPC, and Distributed S torage Kineto Traces. https://git...

  59. [67]

    Thomas Rausch, Waldemar Hummer, and Vinod Muthusamy. 2020. PipeSim: Trace-driven Simulation of Large-Scale AI Operations Platforms. arXiv:2006.12587 [cs.DC] https://arxiv.org/abs/2006.12587

  60. [68]

    Ganger, Randy H

    Charles Reiss, Alexey Tumanov, Gregory R. Ganger, Randy H. Katz, and Michael A. Kozuch. 2012. Heterogeneity and dynamicity of clouds at scale: Google trace analysis. InProceedings of the Third ACM Symposium on Cloud Computing (SoCC). 7:1–7:13. https://doi.org/10.1145/2391229.2391236

  61. [69]

    George F Riley and Thomas R Henderson. 2010. The ns-3 network simulator. In Modeling and tools for network simulation. Springer, 15–34

  62. [70]

    Alastair Robertson and contributors. 2025. bpftrace: A High-Level Tracing Language for Linux. https://github.com/iovisor/bpftrace Version 0.22.1; accessed February 04, 2025; licensed under Apache-2.0

  63. [71]

    Hongzhang Shan, Filip Blagojević, Seung-Jai Min, Paul Hargrove, Haoqiang Jin, Karl Fuerlinger, Alice Koniges, and Nicholas J. Wright. 2010. A programming model performance study using the NAS parallel benchmarks.Sci. Program.18, 3–4 (Aug. 2010), 153–167. https://doi.org/10.115...

  64. [72]

    Siyuan Shen, Langwen Huang, Marcin Chrapek, Timo Schneider, Jai Dayal, Manisha Gajbe, Robert Wisniewski, and Torsten Hoefler. 2024. LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming. In SC24: International Conference for High Performance Co...

  65. [73]

    Shende and Allen D

    Sameer S. Shende and Allen D. Malony. 2006. The Tau Parallel Performance System.The International Journal of High Performance Computing Applications 20 (5 2006), 287–311. Issue 2. https://doi.org/10.1177/1094342006064482

  66. [74]

    Srinivas Sridharan, Taekyung Heo, Louis Feng, Zhaodong Wang, Matt Bergeron, Wenyin Fu, Shengbao Zheng, Brian Coutinho, Saeed Rashidi, Changhai Man, and Tushar Krishna. 2023. Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces. arXiv:230...

  67. [75]

    sstsimulator. 2025. SST-DUMPI Trace Library. https://github.com/sstsimulator/ sst-dumpi. Accessed: 2025-02-14

  68. [76]

    A. P. Thompson, H. M. Aktulga, R. Berger, D. S. Bolintineanu, W. M. Brown, P. S. Crozier, P. J. in ’t Veld, A. Kohlmeyer, S. G. Moore, T. D. Nguyen, R. Shan, M. J. Stevens, J. Tranchida, C. Trott, and S. J. Plimpton. 2022. LAMMPS - a flexible simulation tool for particle-based...

  69. [77]

    Tikir, Michael A

    Mustafa M. Tikir, Michael A. Laurenzano, Laura Carrington, and Allan Snavely

  70. [78]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guil- laume Lample. 2023. LLaMA: Open and Efficient Foundation ...

  71. [79]

    University of Massachusetts Amherst. 2016. UMass Trace Repository: Storage. https://traces.cs.umass.edu/docs/traces/storage/. Accessed: 4 February 2025. Copyright©2016–2024 University of Massachusetts Amherst

  72. [80]

    Erico Vanini, Rong Pan, Mohammad Alizadeh, Parvin Taheri, and Tom Edsall

  73. [81]

    András Varga and Rudolf Hornig. 2008. An overview of the OMNeT++ simulation environment. InProceedings of the 1st International Conference on Simulation Tools and Techniques for Communications, Networks and Systems & Workshops (Marseille, France)(Simutools ’08). ICST (Institut...

  74. [82]

    Kurt Wagner. [n. d.]. Meta Is Building New $800 Million AI-Focused Data Center in Indiana.Bloomberg([n. d.]). https://www.bloomberg.com/news/articles/2024- 01-25/meta-building-new-800-million-ai-focused-data-center-in- indiana?embedded-checkout=true

  75. [83]

    Xizheng Wang, Qingxu Li, Yichi Xu, Gang Lu, Dan Li, Chen Li, Heyang Zhou, Linkang Zheng, Sen Zhang, Yikai Zhu, Yang Liu, Pengcheng Zhang, Kun Qian, and Kunling He. 2025. SimAI: Unifying Architecture Design and Performance Tunning for Large-Scale Large Language Model Training w...

  76. [84]

    William Won, Taekyung Heo, Saeed Rashidi, Srinivas Sridharan, Sudarshan Srinivasan, and Tushar Krishna. 2023. ASTRA-sim2.0: Modeling Hierarchi- cal Networks and Disaggregated Systems for Large-model Training at Scale. arXiv:2303.14006 [cs.DC] https://arxiv.org/abs/2303.14006

  77. [85]

    Feroz Zahid, Ernst Gunnar Gran, Bartosz Bogdański, Bjørn Dag Johnsen, and Tor Skeie. 2017. Efficient network isolation and load balancing in multi-tenant HPC clusters.Future Generation Computer Systems72 (2017), 145–162. https: //doi.org/10.1016/j.future.2016.04.003

  78. [86]

    Ben Zaitlen. 2021. NVIDIA Tools Extension API: An Annotation Tool for Profil- ing Code in Python and C/C++. https://developer.nvidia.com/blog/nvidia-tools- extension-api-nvtx-annotation-tool-for-profiling-code-in-python-and-c-c/ Ac- cessed: 2025-01-28

  79. [87]

    Jidong Zhai, Wenguang Chen, and Weimin Zheng. 2010. PHANTOM: predicting performance of parallel applications on large-scale parallel machines using a single node. InProceedings of the 15th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming(Bangalore, Indi...

  80. [88]

    In14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17)

    Let It Flow: Resilient Asymmetric Load Balancing with Flowlet Switching. In14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). USENIX Association, Boston, MA, 407–420. https://www.usenix.org/ conference/nsdi17/technical-sessions/presentation/vanini

  81. [89]

    Zheng, Gunavardhan Kakulapati, and L.V

    G. Zheng, Gunavardhan Kakulapati, and L.V. Kale. 2004. BigSim: a parallel simu- lator for performance prediction of extremely large parallel machines. In18th International Parallel and Distributed Processing Symposium, 2004. Proceedings. 78–. https://doi.org/10.1109/IPDPS.2004.1303013

  82. [96]

    Tianqi Zhang, Yanqi Zhang, Yuandong Tian, Lin Ma, Wei Lin, and Bin Cui

  83. [1995]

    InProceedings of the Seventh Annual ACM Symposium on Parallel Algorithms and Architectures(Santa Barbara, California, USA)(SPAA ’95)

    LogGP: incorporating long messages into the LogP model—one step closer towards a realistic model for parallel computation. InProceedings of the Seventh Annual ACM Symposium on Parallel Algorithms and Architectures(Santa Barbara, California, USA)(SPAA ’95). Association for Comp...

  84. [2009]

    In2009 DoD High Performance Computing Modernization Program Users Group Conference

    PSINS: An Open Source Event Tracer and Execution Simulator. In2009 DoD High Performance Computing Modernization Program Users Group Conference. 444–449. https://doi.org/10.1109/HPCMP-UGC.2009.73

  85. [2016]

    InSC16: International Conference for High Performance Computing, Networking, Storage and Analysis (SC)

    Evaluating HPC Networks via Simulation of Parallel Workloads . InSC16: International Conference for High Performance Computing, Networking, Storage and Analysis (SC). IEEE Computer Society, Los Alamitos, CA, USA, 154–165. https://doi.org/10.1109/SC.2016.13

  86. [2017]

    https: //doi.org/10.1109/TPDS.2016.2543725

    Enabling Parallel Simulation of Large-Scale HPC Network Systems.IEEE Transactions on Parallel and Distributed Systems28, 1 (2017), 87–100. https: //doi.org/10.1109/TPDS.2016.2543725

  87. [2019]

    InProceedings of the ACM Special Interest Group on Data Communication(Beijing, China)(SIGCOMM ’19)

    HPCC: high precision congestion control. InProceedings of the ACM Special Interest Group on Data Communication(Beijing, China)(SIGCOMM ’19). Association for Computing Machinery, New York, NY, USA, 44–58. https: //doi.org/10.1145/3341302.3342085

  88. [2022]

    In19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22)

    MLaaS in the Wild: Workload Analysis and Scheduling in Large-scale Heterogeneous GPU Clusters. In19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22). 827–841. https://github.com/alibaba/ clusterdata/blob/master/cluster-trace-gpu-v2020/README.md

  89. [2024]

    InSC24: International Conference for High Performance Computing, Networking, Storage and Analysis

    Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects. InSC24: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1–15. https://doi.org/10.1109/sc41406. 2024.00039

  90. [2025]

    eBPF extends the classic BPF mechanism to run sandboxed programs in the Linux kernel for tracing, networking, and more

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.