Pith. sign in

REVIEW 5 major objections 5 minor 3 cited by

Machine Learning Fleet Efficiency: Analyzing and Optimizing Large-Scale Google TPU Systems with ML Productivity Goodput

T0 review · 5 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read ML Productivity Goodput splits ML fleet efficiency into scheduling, runtime, and program components, exposing optimization opportunities across the stack.

desk verdict MPG decomposition is a useful lens for ML fleets, but Program Goodput's compute-only roofline conflates communication-boundness with program inefficiency, so the bottleneck-localization claims need qualification. read the letter →

arxiv 2502.06982 v2 pith:33COF5SI submitted 2025-02-10 cs.LG

classification cs.LG
keywords MLfleetsTPUgoodputschedulingruntimeefficiencycompileroptimizationwarehouse-scalecomputingperformancemetrics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the efficiency of a large-scale ML fleet—thousands of accelerators running diverse training and serving jobs—cannot be measured by traditional utilization metrics like occupancy or duty cycle, because those metrics count busy devices rather than useful progress. To fix this, the authors introduce ML Productivity Goodput (MPG), which multiplies three sub-metrics: Scheduling Goodput (fraction of time all requested chips are simultaneously available), Runtime Goodput (fraction of allocated time spent on progress that survives checkpointing), and Program Goodput (ratio of an idealized FLOP-based execution time to actual execution time). They show, on Google's production TPU fleet, that segmenting MPG by workload type, hardware generation, framework, and job size exposes bottlenecks at specific layers of the system stack, and that acting on those signals—such as compiler optimizations, runtime changes, and scheduler tuning—improves fleet efficiency. The pay-off, if the approach is right, is a general playbook for operating and optimizing ML fleets at warehouse scale.

What carries the argument

The central object is the ML Productivity Goodput (MPG) metric, defined as $MPG = SG \times RG \times PG$, where $SG$ (Scheduling Goodput) is allocated chip-time over fleet capacity, $RG$ (Runtime Goodput) is checkpointed productive chip-time over allocated chip-time, and $PG$ (Program Goodput) is the ratio of an ideal compute-based execution time—FLOPs of the unoptimized HLO graph divided by theoretical peak FLOPS—to the actual execution time. The decomposition is designed as the functional analog of the 'iron law' of processor performance, and it does the argument's work because each factor is computable from production telemetry, is attributable to a distinct stack layer (scheduler, runtime, compiler/program), and can be measured before and after an optimization to attribute the gain or loss to that layer.

What would settle it

Run a fixed-FLOP communication-bound training job with and without a network-bandwidth increase; if the runtime drops but Program Goodput remains unchanged while Runtime Goodput rises, the compute-only roofline is attributing the communication loss to the wrong layer, contradicting the claim that each MPG component isolates one stack layer.

Watch

Extended reading notes

Core claim

The central claim is that ML fleet efficiency is the product of three independent, measurable terms—Scheduling Goodput, Runtime Goodput, and Program Goodput—and that this decomposition is both actionable and validatable at fleet scale. Scheduling Goodput is the fraction of time an application has all its requested accelerators simultaneously available; Runtime Goodput is the fraction of allocated time during which the workload makes progress that would survive a failure or preemption (that is, progress that has been checkpointed); Program Goodput is the ratio of an ideal execution time, estimated from the FLOP count of the unoptimized HLO graph divided by theoretical peak FLOPS, to the actual execution time. The paper demonstrates on a production Google TPU fleet that these components can be tracked over time and segmented by workload characteristics, and that optimizations aimed at each component—communication-computation overlap and compiler autotuning for Program Goodput, asynchronous checkpointing and ahead-of-time compilation for Runtime Goodput, scheduler defragmentation and pre-emption policies for Scheduling Goodput—move the corresponding sub-metric in the expected direction. That agreement between targeted interventions and component-wise changes is what validates the decomposition as a guide for fleet management, rather than merely a reporting exercise.

Load-bearing premise

The central claim rests on the assumption that an ML workload's ideal execution time is an intrinsic property—its FLOP count divided by the accelerator's peak FLOPS, taken from the unoptimized computation graph—so Program Goodput's baseline stays fixed even for communication-bound workloads and across compiler transformations.

Editorial extensions

If this is right

  • Fleet operators can replace utilization-based dashboards with MPG to quantify actual forward progress and set efficiency targets per stack layer.
  • The decomposition lets a team attribute a fleet-level efficiency change to a specific layer, so a compiler change or scheduler tweak can be validated in production before a large rollout.
  • Segmenting MPG by model architecture, workload phase, framework, and hardware generation reveals hidden losses that aggregate utilization metrics miss, such as low Program Goodput on newly introduced accelerators.
  • Because the metric is expressed as a product of three rates, improvements at different layers multiply, giving operators a way to prioritize which layer to attack first.
  • The methodology is not TPU-specific and should transfer to GPU or other DSA fleets, provided the telemetry for all-allocated time, checkpointed progress, and graph FLOPs is available.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The checkpoint-based definition of Runtime Goodput may unfairly penalize workloads that checkpoint rarely; a normalization by checkpoint interval would make cross-workload RG comparisons fairer, but would also weaken the metric's simplicity.
  • A compute-only roofline for Program Goodput cannot distinguish 'poorly optimized program' from 'communication-bound algorithm'; extending PG with a communication-aware roofline could split those two causes.
  • The lifecycle pattern in the notional PG-versus-allocation curve (low PG when a chip debuts, rising as software matures, falling before decommission) suggests that fleet operators could use PG to time when to invest in compiler work for a new accelerator.
  • The multiplicative form of MPG implies that an optimization that improves one layer can mask regressions in another; reporting the three components side by side is therefore essential, not optional.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. This paper proposes ML Productivity Goodput (MPG), defined as the product of Scheduling Goodput (SG), Runtime Goodput (RG), and Program Goodput (PG), to measure and decompose the efficiency of large-scale ML fleets. Using anonymized internal Google TPU fleet data, the authors apply MPG to analyze fleet-wide trends across hardware, scheduler, runtime/compiler, framework, and model layers, and report using MPG segmentation to guide scheduling, runtime, and compiler optimizations, including a claimed fleet-wide impact of an XLA algebraic simplification.

Significance. If the metric and its decomposition are sound, MPG is a useful, interpretable alternative to occupancy and utilization metrics for warehouse-scale ML systems, with the practical advantage of targeting optimization efforts to specific stack layers. The paper's strengths are that MPG is conceptually simple, has no fitted parameters, and is tied to concrete optimization interventions such as communication-computation overlap, XLA changes, and scheduling policies. The main limitations are the lack of precise formal definitions of the sub-metrics, the internal tension in the PG definition for communication-bound workloads, and the heavy reliance on anonymized or explicitly notional data, which limits independent verification of the quantitative claims. Overall, this is a promising industrial framework rather than a fully validated scientific metric; the decomposition is partly definitional, since SG by construction reflects scheduling losses, but the reported external improvements give it practical value.

major comments (5)
  1. [Section 4.2, Figures 8-9] The definition of Program Goodput is internally inconsistent with the paper's own workload characterization. PG is defined as T_ideal/T_actual with T_ideal = FLOPs(unoptimized HLO)/peak FLOPS, and the text calls this 'intrinsic' and 'agnostic to compiler decisions.' Yet Section 2.1 states that ML workloads 'tend to be more communication-bound rather than compute-bound,' and Section 5.3 uses low PG to conclude that 'many high-cost workloads are communication-bound.' For a communication-bound workload, T_actual is dominated by collective communication and memory stalls, so PG is low even for a perfectly optimized program. PG is therefore a compute-utilization proxy, not a measure of 'how well optimized the ML program is,' and the MPG decomposition does not cleanly isolate the program/compiler layer. The denominator should either include communication time explicitly, or the interpretation of PG should be restricted.
  2. [Section 4.2, Figure 10] The definitions of SG, RG, and PG are not formal enough to be reproducible. For SG, the numerator is 'simultaneous uptime of all tasks' and the denominator is 'fleet capacity, expressed as chip-time,' but it is not specified whether capacity means all fleet chips during the window, only allocated chips, or something else, nor how partially allocated intervals are counted. For RG, the numerator is 'productive chip-time ... saved in checkpoints,' but the text does not define how productive time is measured between two checkpoints when no failure or preemption occurs; presumably all such time is productive, but this must be stated. For PG, the ideal-time computation from HLO FLOPs lacks a precise specification of which HLO stage is counted as 'unoptimized' and which peak FLOPS number is used per accelerator generation. Equations and a time-window definition are needed before the metric can be evaluated.
  3. [Section 5, Table 2] The predicted effects of compiler, runtime, and scheduler optimizations are presented as qualitative assertions without a formal model or empirical validation. For example, the compiler row claims PG 'Increases,' SG and MPG 'Decreases if device-bound' but 'No change if host-bound,' yet no derivation in the text connects a change in step time to these component changes. Since the surrounding sections claim to use MPG to validate optimizations, Table 2 should be either derived from the definitions in Section 4.2 or accompanied by measured before-and-after MPG decomposition data.
  4. [Section 5.3, Figure 11] The claim that an XLA algebraic simplification improves fleet-wide Program Goodput is based on a benchmark of the 'top 150 most costly workloads,' but no evidence is given that this set is representative of the fleet for evaluating compiler changes. Cost-based selection likely overweights compute-heavy, long-running workloads, so the reported PG improvement may not generalize to the rest of the fleet. The paper should justify representativeness or report the effect on the full fleet.
  5. [Section 5.2, Figures 13-15] Several central quantitative results are reported with anonymized or explicitly notional data, including 'segments in Figure 13 are not explicitly identified,' a 'notional slice' in Figure 14, and a 'notional example' in Figure 15. This prevents independent verification of the claimed trends and of the general conclusion that MPG 'precisely pinpoint[s] optimization opportunities.' The authors should either declassify the data, describe the anonymization procedure, or clearly separate notional illustrations from measured results and state which claims depend on which category.
minor comments (5)
  1. [Section 4.1] The statement that 'the MPG metric must be a clearly defined and accurate measure of forward progress' reads as a design goal, but the paper never formally defines 'forward progress'; consider defining it or restricting the claim to the three sub-metrics.
  2. [Section 3.1 and Section 4.2] Notation is inconsistent: 'FLOPS' and 'FLOPs' are used interchangeably, and 'TOPs/Watt' appears without a definition; please standardize the units and their capitalization.
  3. [Section 5.1, Figure 12] The claim that 'the overall SG is already close to optimal' is not supported by a numeric baseline or a definition of 'optimal'; please provide a reference value or a formal optimality criterion.
  4. [Section 5.2, Figure 14] The sentence 'This has resulted in an temporary decrease of RG for the bulk inference segment' contains a grammar error ('an temporary' should be 'a temporary').
  5. [Section 2.1] The claim that ML workloads 'tend to be more communication-bound rather than compute-bound' cites a parameter-server paper [36]; a more direct citation on collective communication in distributed DNN training, such as Wang et al. [63], would better support the statement.

Circularity Check

1 steps flagged · score 3.0 of 10

MPG decomposition is an accounting identity by construction; PG's compute-only roofline is a validity concern, but no fitted-input or self-citation circularity found.

  1. self definitional [Section 1 (contributions) and Section 4.2, Figure 8: definition of Scheduling/Runtime/Program Goodput and MPG decomposition.]
    "By breaking down the metric into subcomponents—Scheduling Goodput (SG), Runtime Goodput (RG), and Program Goodput (PG)—and examining it along the axes of fleet characteristics, we can precisely pinpoint optimization opportunities. ... Scheduling Goodput (SG) quantifies the efficiency of resource allocation in an ML fleet. It measures the fraction of time that an application has all the required resources simultaneously available to make progress."

    The paper builds MPG as the product of SG, RG, and PG, and defines SG as the scheduling-layer term, RG as the runtime-layer term, and PG as the program-layer term. Consequently, 'low SG identifies a scheduling bottleneck' is true by definition: the decomposition is an accounting identity over the authors' chosen layer taxonomy, not an empirical derivation of where losses come from. The 'precisely pinpoint' claim therefore restates the metric's construction rather than an independent finding. This is mild: no data are fitted to make the components line up, and the Section 5 optimization outcomes are validated with measured throughput/speedup improvements rather than by MPG alone.

full rationale

I found no fitted-input-as-prediction, no load-bearing self-citation, and no uniqueness theorem. Self-citations to XTAT [48] and Kumar et al. [34] are examples of deployed optimizations, not the basis of the metric. The decomposition's 'pinpointing' language is definitional because each goodput is constructed as a layer-specific ratio; this is the only circular element. The PG roofline concern—FLOP-only ideal time for workloads the paper itself characterizes as communication-bound—is a measurement-validity and attribution risk rather than a result-in-implies-result-out circularity, so it does not raise the score above 3. The central contribution is a proposed metric plus empirical fleet observations and externally measured optimizations, which carry independent content beyond the definitions.

Assumptions & free parameters 0 free parameters · 5 assumptions · 1 invented entities

The central claim depends on MPG being a valid and measurable definition of fleet efficiency. No numeric free parameters are fitted, but meaningfulness relies on several domain assumptions: all-or-nothing chip availability, checkpoint-based productivity, a compute-only roofline, and homogeneous chip-time aggregation. These are plausible for Google's TPU training workloads, but they are not validated for heterogeneous serving or inference workloads.

assumptions (5)
  • domain assumption ML workloads make progress only when all allocated chips are available simultaneously (bulk-synchronous assumption).
    Scheduling Goodput's definition in Section 4.2 treats fully allocated time as the prerequisite for progress. This is typical for distributed training, but less accurate for serving or asynchronous workloads.
  • domain assumption Only checkpointed work counts as productive runtime progress.
    Runtime Goodput's numerator in Section 4.2 excludes work since the last checkpoint, making the metric dependent on checkpoint policy and failure or preemption rates.
  • ad hoc to paper The ideal execution time for Program Goodput is the unoptimized HLO graph's FLOP count divided by theoretical peak FLOPS.
    This compute-only roofline is introduced in Section 4.2 without validation against memory-bound or communication-bound workloads, and it is sensitive to HLO graph representation.
  • domain assumption Fleet capacity can be expressed as a homogeneous chip-time denominator across heterogeneous accelerators.
    Scheduling Goodput in Section 4.2 divides all-allocated time by fleet capacity in chip-time, implicitly treating different accelerator types and generations as fungible.
  • ad hoc to paper The top 150 most costly workloads are a representative benchmark for evaluating fleet-wide compiler changes.
    Section 5.3 uses this set to measure Program Goodput changes, but no selection criteria or representativeness analysis is provided.
invented entities (1)
  • ML Productivity Goodput (MPG) and its sub-metrics SG, RG, PG
    purpose: Quantify fleet-wide ML efficiency and decompose losses into scheduling, runtime, and program layers.
    MPG is defined by this paper. The paper provides no external benchmark, release artifact, or independent validation, so the metric's usefulness rests on the assumptions listed above.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Machine Learning Fleet Efficiency: Analyzing and Optimizing Large-Scale Google TPU Systems with ML Productivity Goodput." pith.science (2026). https://pith.science/paper/33COF5SI

@misc{pith2026250206982,
  author       = {Pith},
  title        = {Pith review of: Machine Learning Fleet Efficiency: Analyzing and Optimizing Large-Scale Google TPU Systems with ML Productivity Goodput},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/33COF5SI}},
  note         = {Machine review of arXiv:2502.06982}
}
read the original abstract

Recent years have seen the emergence of machine learning (ML) workloads deployed in warehouse-scale computing (WSC) settings, also known as ML fleets. As the computational demands placed on ML fleets have increased due to the rise of large models and growing demand for ML applications, it has become increasingly critical to measure and improve the efficiency of such systems. However, there is not yet an established methodology to characterize ML fleet performance and identify potential performance optimizations accordingly. This paper presents a large-scale analysis of an ML fleet based on Google's TPUs, introducing a framework to capture fleet-wide efficiency, systematically evaluate performance characteristics, and identify optimization strategies for the fleet. We begin by defining an ML fleet, outlining its components, and analyzing an example Google ML fleet in production comprising thousands of accelerators running diverse workloads. Our study reveals several critical insights: first, ML fleets extend beyond the hardware layer, with model, data, framework, compiler, and scheduling layers significantly impacting performance; second, the heterogeneous nature of ML fleets poses challenges in characterizing individual workload performance; and third, traditional utilization-based metrics prove insufficient for ML fleet characterization. To address these challenges, we present the "ML Productivity Goodput" (MPG) metric to measure ML fleet efficiency. We show how to leverage this metric to characterize the fleet across the ML system stack. We also present methods to identify and optimize performance bottlenecks using MPG, providing strategies for managing warehouse-scale ML systems in general. Lastly, we demonstrate quantitative evaluations from applying these methods to a real ML fleet for internal-facing Google TPU workloads, where we observed tangible improvements.

Figures

Figures reproduced from arXiv: 2502.06982 by the authors.

Figure 1
Figure 1. Five-year historical ML fleet breakdown by accelerator type. The [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. ML Productivity Goodput (MPG) and its components. [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. The ML fleet system stack of a production system at Google. The [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (8 more)
Figure 5
Figure 5. Figure 5: An ML workload requires all requested TPUs to be allocated before [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]
Figure 6
Figure 6. Figure 6: The prevalence of fleet-wide workloads using the Pathways runtime [PITH_FULL_IMAGE:figures/full_fig_p005_6.png]
Figure 8
Figure 8. Figure 8: ML Productivity Goodput (MPG) and its components. [PITH_FULL_IMAGE:figures/full_fig_p006_8.png]
Figure 10
Figure 10. Figure 10: The scheduling goodput for training workloads measures the per [PITH_FULL_IMAGE:figures/full_fig_p007_10.png]
Figure 12
Figure 12. Figure 12: Scheduling goodput by job size. Extra-large and small jobs tend to [PITH_FULL_IMAGE:figures/full_fig_p008_12.png]
Figure 14
Figure 14. Figure 14: Runtime Goodput trends for a notional slice of a sample ML fleet [PITH_FULL_IMAGE:figures/full_fig_p009_14.png]
Figure 15
Figure 15. Figure 15: Tracking the Program Goodput (PG) versus allocation trends for a [PITH_FULL_IMAGE:figures/full_fig_p009_15.png]
Figure 16
Figure 16. Figure 16: Traditional utilization-based metrics. We replace these using good [PITH_FULL_IMAGE:figures/full_fig_p010_16.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MLSYSIM: First-Principles Infrastructure Modeling for Machine Learning Systems

    cs.DC 2026-06 accept novelty 6.5 of 10

    A dimensionally strict analytical framework codifies 22 ML systems walls into 28 composable resolvers for sub-second full-stack design-space exploration and hardware synthesis.

  2. A Taxonomy of Performance Metrics for the Distributed Computing Continuum

    cs.DC 2026-07 conditional novelty 4.0 of 10

    A three-layer taxonomy (compute, network, application) plus novel continuum-specific metrics and acquisition requirements for evaluating distributed computing continuum systems.

  3. A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO

    cs.DC 2025-06 conditional novelty 2.0 of 10

    This survey classifies distributed DNN training simulators into analytical, profiling-based, and execution-driven categories, and compares them alongside TCO and carbon-emission models.

Reference graph

Works this paper leans on

76 extracted references · 52 canonical work pages · cited by 3 Pith papers

  1. [1]

    [n. d.]. XLA. https://openxla.org/xla

  2. [2]

    Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, San- jay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Leven- berg, Dan Mane, Rajat Monga, Sherry Moore, Derek Murray,...

  3. [3]

    Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng

    Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Man- junath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016. TensorFlow: A system f...

  4. [4]

    Banning, Sumeer Bhola, Rick Buskens, Ming Chen, Xi Chen, Yoo Chung, Qin Jia, Nick Sakharov, George T

    Colin Adams, Luis Alonso, Ben Atkin, John P. Banning, Sumeer Bhola, Rick Buskens, Ming Chen, Xi Chen, Yoo Chung, Qin Jia, Nick Sakharov, George T. Talbot, Adam Jacob Tart, and Nick Taylor (Eds.). 2020. Monarch: Google’s Planet- Scale In-Memory Time Series Database

  5. [5]

    Anthropic AI. 2024. The Claude 3 Model Family: Opus, Sonnet, Haiku

  6. [6]

    Krste Asanovic, Rastislav Bodik, James Demmel, Tony Keaveny, Kurt Keutzer, John Kubiatowicz, Nelson Morgan, David Patterson, Koushik Sen, John Wawrzynek, David Wessel, and Katherine Yelick. 2009. A view of the paral- lel computing landscape. Commun. ACM 52, 10 (Oct. 2009), 56–67. https: //doi.org/10.1145/1562764.1562783

  7. [7]

    Thekkath, and Yonghui Wu

    Paul Barham, Aakanksha Chowdhery, Jeff Dean, Sanjay Ghemawat, Steven Hand, Dan Hurt, Michael Isard, Hyeontaek Lim, Ruoming Pang, Sudip Roy, Brennan Saeta, Parker Schuh, Ryan Sepassi, Laurent El Shafey, Chandramohan A. Thekkath, and Yonghui Wu. 2022. Pathways: Asynchronous Distributed Dataflow for ML. arXiv preprint arXiv:2203.12533 (2022). https://arxiv.o...

  8. [8]

    Luiz André Barroso and Urs Hölzle. 2009. The Datacenter as a Computer: An Introduction to the Design of Warehouse-Scale Machines . http://dx.doi.org/10. 2200/S00193ED1V01Y200905CAC006

Show all 76 references
  1. [9]

    Nathan Bell and Michael Garland. 2008. Efficient sparse matrix-vector multiplica- tion on CUDA. Technical Report. Nvidia Technical Report NVR-2008-004, Nvidia Corporation

  2. [10]

    Cooper, and Linda Torczon

    Preston Briggs, Keith D. Cooper, and Linda Torczon. 1992. Rematerializa- tion. In Proceedings of the ACM SIGPLAN 1992 Conference on Programming Language Design and Implementation (San Francisco, California, USA) (PLDI ’92). Association for Computing Machinery, New York, NY, US...

  3. [11]

    Sergey Brin and Lawrence Page. 1998. The Anatomy of a Large-Scale Hypertex- tual Web Search Engine. Computer Networks 30 (1998), 107–117. http://www- db.stanford.edu/~backrub/google.html

  4. [12]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...

  5. [13]

    Brendan Burns, Brian Grant, David Oppenheimer, Eric Brewer, and John Wilkes

  6. [14]

    Arnab Choudhury, Yang Wang, Tuomas Pelkonen, Kutta Srinivasan, Abha Jain, Shenghao Lin, Delia David, Siavash Soleimanifard, Michael Chen, Abhishek Yadav, Ritesh Tijoriwala, Denis Samoylov, and Chunqiang Tang. 2024. MAST: Global Scheduling of ML Training across Geo-Distributed ...

  7. [15]

    Borg, Omega, and Kubernetes. Commun. ACM 59, 5 (apr 2016), 50–57. https://doi.org/10.1145/2890784

  8. [16]

    Lieven Eeckhout. 2010. Computer architecture performance evaluation methods . Morgan & Claypool Publishers

  9. [17]

    James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost, JJ Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Pe- ter Hochschild, Wilson Hsieh, Sebastian Kanthak, Eugene Kogan, Hongyi Li, Alexander Lloyd, Sergey Melnik, David Mwaura, David N...

  10. [18]

    Andre Esteva, Katherine Chou, Serena Yeung, Nikhil Naik, Ali Madani, Ali Mot- taghi, Yun Liu, Eric Topol, Jeff Dean, and Richard Socher. 2021. Deep learning- enabled medical computer vision. NPJ digital medicine 4, 1 (2021), 5

  11. [19]

    Emer and Douglas W

    Joel S. Emer and Douglas W. Clark. 1984. A Characterization of Processor Performance in the vax-11/780. SIGARCH Comput. Archit. News 12, 3 (jan 1984), 301–310. https://doi.org/10.1145/773453.808199

  12. [20]

    Roy Frostig, Matthew Johnson, and Chris Leary. 2018. Compiling machine learning programs via high-level tracing. https://mlsys.org/Conferences/doc/ 2018/146.pdf

  13. [21]

    Kayvon Fatahalian, Jeremy Sugerman, and Pat Hanrahan. 2004. Understanding the efficiency of GPU algorithms for matrix-matrix multiplication. In Proceedings of the ACM SIGGRAPH/EUROGRAPHICS conference on Graphics hardware . 133– 137

  14. [22]

    James Hamilton. 2007. On Designing and Deploying Internet-Scale Services. In 21st Large Installation System Administration Conference (LISA 07) . USENIX Association, Dallas, TX. https://www.usenix.org/conference/lisa-07/designing- and-deploying-internet-scale-services

  15. [23]

    Sanjay Ghemawat, Howard Gobioff, and Shun-Tak Leung. 2003. The Google File System. In Proceedings of the 19th ACM Symposium on Operating Systems Principles. Bolton Landing, NY, 20–43

  16. [24]

    Dean Hildebrand and Denis Serenyi. 2021. Colossus under the hood: a peek into Google’s scalable storage system . https://cloud.google.com/blog/products/ storage-data-transfer/a-peek-behind-colossus-googles-file-system

  17. [25]

    Hennessy and David A

    John L. Hennessy and David A. Patterson. 2019. A new golden age for computer architecture. Commun. ACM 62, 2 (jan 2019), 48–60. https://doi.org/10.1145/ 3282307

  18. [26]

    JAX. [n. d.]. Ahead-of-time lowering and compilation. https://jax.readthedocs.io/ en/latest/aot.html

  19. [27]

    Joel Janai, Fatma Güney, Aseem Behl, Andreas Geiger, et al. 2020. Computer vision for autonomous vehicles: Problems, datasets and state of the art. Foundations and Trends® in Computer Graphics and Vision 12, 1–3 (2020), 1–308

  20. [28]

    Norman P. Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, Cliff Young, Xiang Zhou, Zongwei Zhou, and David Patterson. 2023. TPU v4: An Optically Reconfigurable Supercomputer for Machine Learn...

  21. [29]

    Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B

    Norman P. Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B. Jablin, George Kurian, James Laudon, Sheng Li, Peter Ma, Xiaoyu Ma, Thomas Norrie, Nishant Patil, Sushma Prasad, Cliff Young, Zongwei Zhou, and David Patterson. 2021. Ten Lessons From Three Generations...

  22. [30]

    Shoaib Kamil, John Shalf, and Erich Strohmaier. 2008. Power efficiency in high performance computing. In 2008 IEEE International Symposium on Parallel and Distributed Processing. 1–8. https://doi.org/10.1109/IPDPS.2008.4536223

  23. [31]

    Jouppi, Cliff Young, Nishant Patil, David A

    Norman P. Jouppi, Cliff Young, Nishant Patil, David A. Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, 11 Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tar...

  24. [32]

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. ImageNet Clas- sification with Deep Convolutional Neural Networks. In Advances in Neural Information Processing Systems, F. Pereira, C.J. Burges, L. Bottou, and K.Q. Wein- berger (Eds.), Vol. 25. Curran Associates, ...

  25. [33]

    Michael Kuchnik, Ana Klimovic, Jiri Simsa, Virginia Smith, and George Amvrosiadis. 2022. Plumber: Diagnosing and Removing Performance Bottlenecks in Machine Learning Data Pipelines.Proceedings of Machine Learning and Systems 4 (2022), 33–51

  26. [34]

    Svilen Kanev, Juan Pablo Darago, Kim Hazelwood, Parthasarathy Ranganathan, Tipp Moseley, Gu-Yeon Wei, and David Brooks. 2015. Profiling a warehouse- scale computer. In Proceedings of the 42nd Annual International Symposium on Computer Architecture. 158–169

  27. [35]

    Jie Li, George Michelogiannakis, Brandon Cook, Dulanya Cooray, and Yong Chen

  28. [36]

    Mu Li, David G Andersen, Alexander J Smola, and Kai Yu. 2014. Communication efficient distributed machine learning with the parameter server. Advances in Neural Information Processing Systems 27 (2014)

  29. [37]

    Sameer Kumar, Yu Wang, Cliff Young, James Bradbury, Naveen Kumar, Dehao Chen, and Andy Swing. 2021. Exploring the limits of Concurrency in ML Training on Google TPUs. Proceedings of Machine Learning and Systems 3 (2021), 81–92

  30. [38]

    Peter Mattson, Christine Cheng, Gregory Diamos, Cody Coleman, Paulius Micike- vicius, David Patterson, Hanlin Tang, Gu-Yeon Wei, Peter Bailis, Victor Bittorf, et al. 2020. Mlperf training benchmark. Proceedings of Machine Learning and Systems 2 (2020), 336–349

  31. [39]

    Mustafa Rafique, Franck Cappello, and Bogdan Nicolae

    Avinash Maurya, Robert Underwood, M. Mustafa Rafique, Franck Cappello, and Bogdan Nicolae. 2024. DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models. In Proceedings of the 33rd International Symposium on High-Performance Parallel and Distributed Computing...

  32. [40]

    D Menemenlis, C Hill, A Adcrocft, J-M Campin, B Cheng, B Ciotti, I Fukumori, P Heimbach, C Henze, A Kohl, et al . 2005. NASA supercomputer improves prospects for ocean climate research. Eos, Transactions American Geophysical Union 86, 9 (2005), 89–96

  33. [41]

    Jason Mars, Lingjia Tang, Robert Hundt, Kevin Skadron, and Mary Lou Soffa

  34. [42]

    Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G. Azzolini, Dmytro Dzhulgakov, Andrey Mallevich, Ilia Cherni- avskii, Yinghai Lu, Raghuraman Krishnamoorthi, Ansha Yu, V...

  35. [43]

    Wozniak, George Bosilca, Matthieu Dorier, and Franck Cappello

    Bogdan Nicolae, Jiali Li, Justin M. Wozniak, George Bosilca, Matthieu Dorier, and Franck Cappello. 2020. DeepFreeze: Towards Scalable Asynchronous Check- pointing of Deep Learning Models. In 2020 20th IEEE/ACM International Sym- posium on Cluster, Cloud and Internet Computing ...

  36. [44]

    Li, Ryan McElroy, Mike Paleczny, Daniel Peek, Paul Saab, David Stafford, Tony Tung, and Venkateshwaran Venkataramani

    Rajesh Nishtala, Hans Fugal, Steven Grimm, Marc Kwiatkowski, Herman Lee, Harry C. Li, Ryan McElroy, Mike Paleczny, Daniel Peek, Paul Saab, David Stafford, Tony Tung, and Venkateshwaran Venkataramani. 2013. Scaling Memcache at Facebook. In 10th USENIX Symposium on Networked Sys...

  37. [45]

    Tetsuya Odajima, Yuetsu Kodama, Miwako Tsuji, Motohiko Matsuda, Yutaka Maruyama, and Mitsuhisa Sato. 2020. Preliminary Performance Evaluation of the Fujitsu A64FX Using HPC Applications. In 2020 IEEE International Conference on Cluster Computing (CLUSTER). 523–530. https://doi...

  38. [46]

    Murray, Jiri Simsa, Ana Klimovic, and Ihor Indyk

    Derek G. Murray, Jiri Simsa, Ana Klimovic, and Ihor Indyk. 2021. tf.data: A Machine Learning Data Processing Framework. arXiv:2101.12127 [cs.LG] https: //arxiv.org/abs/2101.12127

  39. [47]

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fan...

  40. [48]

    Phitchaya Mangpo Phothilimthana, Amit Sabne, Nikhil Sarda, Karthik Srini- vasa Murthy, Yanqi Zhou, Christof Angermueller, Mike Burrows, Sudip Roy, Ketan Mandke, Rezsa Farahani, et al. 2021. A Flexible Approach to Autotuning Multi-Pass Machine Learning Compilers. In 2021 30th I...

  41. [49]

    Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean

  42. [50]

    Parthasarathy Ranganathan and Urs Holzle. 2024. Twenty Five Years of Warehouse-Scale Computing . IEEE Micro 44, 05 (Sept. 2024), 11–22. https: //doi.org/10.1109/MM.2024.3409469

  43. [51]

    OpenXLA. [n. d.]. Using AOT compilation. https://openxla.org/xla/tf2xla/ tfcompile

  44. [52]

    Paul Menage Sanjay Ghemawat. [n. d.]. TCMalloc : Thread-Caching Malloc . https://goog-perftools.sourceforge.net/doc/tcmalloc.html

  45. [53]

    Roland Schulz, Benjamin Lindner, Loukas Petridis, and Jeremy C Smith. 2009. Scaling of multimillion-atom biological molecular dynamics simulation on a petascale supercomputer. Journal of Chemical Theory and Computation 5, 10 (2009), 2798–2808

  46. [54]

    Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. arXiv:1701.06538 [cs.LG] https: //arxiv.org/abs/1701.06538

  47. [55]

    In Proceedings of Ma- chine Learning and Systems , D

    Efficiently Scaling Transformer Inference. In Proceedings of Ma- chine Learning and Systems , D. Song, M. Carbin, and T. Chen (Eds.), Vol. 5. Curan, 606–624. https://proceedings.mlsys.org/paper_files/paper/2023/file/ c4be71ab8d24cdfb45e3d06dbfca2780-Paper-mlsys2023.pdf

  48. [56]

    Vladimir Stegailov, Ekaterina Dlinnova, Timur Ismagilov, Mikhail Khalilov, Niko- lay Kondratyuk, Dmitry Makagon, Alexander Semenov, Alexei Simonov, Grigory Smirnov, and Alexey Timofeev. 2019. Angara interconnect makes GPU-based Desmos supercomputer an efficient tool for molecu...

  49. [57]

    David K. Rensin. 2015. Kubernetes - Scheduling the Future at Cloud Scale . 1005 Gravenstein Highway North Sebastopol, CA 95472. All pages. http://www.oreilly. com/webops-perf/free/kubernetes.csp

  50. [58]

    Gemini Team. 2024. Gemini: A Family of Highly Capable Multimodal Models. arXiv:2312.11805 [cs.CL] https://arxiv.org/abs/2312.11805

  51. [59]

    Amin Vahdat and Mark Lohmeyer. 2023. Enabling next-generation AI workloads: Announcing TPU v5p and AI Hypercomputer. https: //cloud.google.com/blog/products/ai-machine-learning/introducing-cloud-tpu- v5p-and-ai-hypercomputer

  52. [60]

    Kenton Varda. 2008. Protocol Buffers: Google’s Data Interchange Format . https: //opensource.googleblog.com/2008/07/protocol-buffers-googles-data.html

  53. [61]

    Zhan Shi, Chirag Sakhuja, Milad Hashemi, Kevin Swersky, and Calvin Lin

  54. [62]

    Korupolu, David Oppenheimer, Eric Tune, and John Wilkes

    Abhishek Verma, Luis Pedrosa, Madhukar R. Korupolu, David Oppenheimer, Eric Tune, and John Wilkes. 2015. Large-scale cluster management at Google with Borg. In Proceedings of the European Conference on Computer Systems (EuroSys) . Bordeaux, France

  55. [63]

    Shibo Wang, Jinliang Wei, Amit Sabne, Andy Davis, Berkin Ilbeyi, Blake Hecht- man, Dehao Chen, Karthik Srinivasa Murthy, Marcello Maggioni, Qiao Zhang, et al. 2022. Overlap Communication with Dependent Computation via Decompo- sition in Large Deep Learning Models. InProceeding...

  56. [64]

    Varun Talwar. 2016. gRPC: a true internet-scale RPC framework is now 1.0 and ready for production deployments . https://cloud.google.com/blog/ products/gcp/grpc-a-true-internet-scale-rpc-framework-is-now-1-and-ready- for-production-deployments

  57. [65]

    Amir Yazdanbakhsh, Kiran Seshadri, Berkin Akin, James Laudon, and Ravi Narayanaswami. 2021. An Evaluation of Edge TPU Accelerators for Con- volutional Neural Networks. CoRR abs/2102.10423 (2021). arXiv:2102.10423 https://arxiv.org/abs/2102.10423

  58. [66]

    Yoo, Morris A

    Andy B. Yoo, Morris A. Jette, and Mark Grondona. 2003. SLURM: Simple Linux Utility for Resource Management. In Job Scheduling Strategies for Parallel Pro- cessing, Dror Feitelson, Larry Rudolph, and Uwe Schwiegelshohn (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 44–60

  59. [67]

    Mark Zhao, Niket Agarwal, Aarti Basant, Bugra Gedik, Satadru Pan, Mustafa Ozdal, Rakesh Komuravelli, Jerry Pan, Tianshu Bao, Haowei Lu, Sundaram Narayanan, Jack Langman, Kevin Wilfong, Harsha Rastogi, Carole-Jean Wu, Christos Kozyrakis, and Parik Pol. 2021. Understanding and C...

  60. [68]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems , I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. ...

  61. [69]

    Yazhou Zu, Alireza Ghaffarkhah, Hoang-Vu Dang, Brian Towles, Steven Hand, Safeen Huda, Adekunle Bello, Alexander Kolbasov, Arash Rezaei, Dayou Du, Steve Lacy, Hang Wang, Aaron Wisner, Chris Lewis, and Henri Bahini. 2024. Resiliency at Scale: Managing Google’s TPUv4 Machine Lea...

  62. [71]

    Samuel Williams, Andrew Waterman, and David Patterson. 2009. Roofline: an insightful visual performance model for multicore architectures. Commun. ACM 52, 4 (apr 2009), 65–76. https://doi.org/10.1145/1498765.1498785

  63. [75]

    Mark Zhao, Niket Agarwal, Aarti Basant, Buğra Gedik, Satadru Pan, Mustafa Ozdal, Rakesh Komuravelli, Jerry Pan, Tianshu Bao, Haowei Lu, Sundaram Narayanan, Jack Langman, Kevin Wilfong, Harsha Rastogi, Carole-Jean Wu, Christos Kozyrakis, and Parik Pol. 2022. Understanding data ...

  64. [2011]

    In Proceedings of the 44th annual IEEE/ACM International Symposium on Microarchitecture

    Bubble-up: Increasing utilization in modern warehouse scale computers via sensible co-locations. In Proceedings of the 44th annual IEEE/ACM International Symposium on Microarchitecture. 248–259

  65. [2016]

    arXiv:1603.04467 [cs.DC] https://arxiv.org/abs/1603.04467

    TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems. arXiv:1603.04467 [cs.DC] https://arxiv.org/abs/1603.04467

  66. [2017]

    CoRR abs/1704.04760 (2017)

    In-Datacenter Performance Analysis of a Tensor Processing Unit. CoRR abs/1704.04760 (2017). arXiv:1704.04760 http://arxiv.org/abs/1704.04760

  67. [2020]

    CoRR abs/2010.02075 (2020)

    Learned Hardware/Software Co-Design of Neural Accelerators. CoRR abs/2010.02075 (2020). arXiv:2010.02075 https://arxiv.org/abs/2010.02075

  68. [2023]

    In International Conference on High Performance Computing

    Analyzing resource utilization in an HPC system: A case study of NERSC’s Perlmutter. In International Conference on High Performance Computing . Springer, 297–316

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.