REVIEW 5 major objections 5 minor 3 cited by
Machine Learning Fleet Efficiency: Analyzing and Optimizing Large-Scale Google TPU Systems with ML Productivity Goodput
T0 review · 5 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read ML Productivity Goodput splits ML fleet efficiency into scheduling, runtime, and program components, exposing optimization opportunities across the stack.
desk verdict MPG decomposition is a useful lens for ML fleets, but Program Goodput's compute-only roofline conflates communication-boundness with program inefficiency, so the bottleneck-localization claims need qualification. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the ML Productivity Goodput (MPG) metric, defined as $MPG = SG \times RG \times PG$, where $SG$ (Scheduling Goodput) is allocated chip-time over fleet capacity, $RG$ (Runtime Goodput) is checkpointed productive chip-time over allocated chip-time, and $PG$ (Program Goodput) is the ratio of an ideal compute-based execution time—FLOPs of the unoptimized HLO graph divided by theoretical peak FLOPS—to the actual execution time. The decomposition is designed as the functional analog of the 'iron law' of processor performance, and it does the argument's work because each factor is computable from production telemetry, is attributable to a distinct stack layer (scheduler, runtime, compiler/program), and can be measured before and after an optimization to attribute the gain or loss to that layer.
What would settle it
Run a fixed-FLOP communication-bound training job with and without a network-bandwidth increase; if the runtime drops but Program Goodput remains unchanged while Runtime Goodput rises, the compute-only roofline is attributing the communication loss to the wrong layer, contradicting the claim that each MPG component isolates one stack layer.
Extended reading notes
Core claim
The central claim is that ML fleet efficiency is the product of three independent, measurable terms—Scheduling Goodput, Runtime Goodput, and Program Goodput—and that this decomposition is both actionable and validatable at fleet scale. Scheduling Goodput is the fraction of time an application has all its requested accelerators simultaneously available; Runtime Goodput is the fraction of allocated time during which the workload makes progress that would survive a failure or preemption (that is, progress that has been checkpointed); Program Goodput is the ratio of an ideal execution time, estimated from the FLOP count of the unoptimized HLO graph divided by theoretical peak FLOPS, to the actual execution time. The paper demonstrates on a production Google TPU fleet that these components can be tracked over time and segmented by workload characteristics, and that optimizations aimed at each component—communication-computation overlap and compiler autotuning for Program Goodput, asynchronous checkpointing and ahead-of-time compilation for Runtime Goodput, scheduler defragmentation and pre-emption policies for Scheduling Goodput—move the corresponding sub-metric in the expected direction. That agreement between targeted interventions and component-wise changes is what validates the decomposition as a guide for fleet management, rather than merely a reporting exercise.
Load-bearing premise
The central claim rests on the assumption that an ML workload's ideal execution time is an intrinsic property—its FLOP count divided by the accelerator's peak FLOPS, taken from the unoptimized computation graph—so Program Goodput's baseline stays fixed even for communication-bound workloads and across compiler transformations.
Editorial extensions
If this is right
- Fleet operators can replace utilization-based dashboards with MPG to quantify actual forward progress and set efficiency targets per stack layer.
- The decomposition lets a team attribute a fleet-level efficiency change to a specific layer, so a compiler change or scheduler tweak can be validated in production before a large rollout.
- Segmenting MPG by model architecture, workload phase, framework, and hardware generation reveals hidden losses that aggregate utilization metrics miss, such as low Program Goodput on newly introduced accelerators.
- Because the metric is expressed as a product of three rates, improvements at different layers multiply, giving operators a way to prioritize which layer to attack first.
- The methodology is not TPU-specific and should transfer to GPU or other DSA fleets, provided the telemetry for all-allocated time, checkpointed progress, and graph FLOPs is available.
Reading between the lines
- The checkpoint-based definition of Runtime Goodput may unfairly penalize workloads that checkpoint rarely; a normalization by checkpoint interval would make cross-workload RG comparisons fairer, but would also weaken the metric's simplicity.
- A compute-only roofline for Program Goodput cannot distinguish 'poorly optimized program' from 'communication-bound algorithm'; extending PG with a communication-aware roofline could split those two causes.
- The lifecycle pattern in the notional PG-versus-allocation curve (low PG when a chip debuts, rising as software matures, falling before decommission) suggests that fleet operators could use PG to time when to invest in compiler work for a new accelerator.
- The multiplicative form of MPG implies that an optimization that improves one layer can mask regressions in another; reporting the three components side by side is therefore essential, not optional.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes ML Productivity Goodput (MPG), defined as the product of Scheduling Goodput (SG), Runtime Goodput (RG), and Program Goodput (PG), to measure and decompose the efficiency of large-scale ML fleets. Using anonymized internal Google TPU fleet data, the authors apply MPG to analyze fleet-wide trends across hardware, scheduler, runtime/compiler, framework, and model layers, and report using MPG segmentation to guide scheduling, runtime, and compiler optimizations, including a claimed fleet-wide impact of an XLA algebraic simplification.
Significance. If the metric and its decomposition are sound, MPG is a useful, interpretable alternative to occupancy and utilization metrics for warehouse-scale ML systems, with the practical advantage of targeting optimization efforts to specific stack layers. The paper's strengths are that MPG is conceptually simple, has no fitted parameters, and is tied to concrete optimization interventions such as communication-computation overlap, XLA changes, and scheduling policies. The main limitations are the lack of precise formal definitions of the sub-metrics, the internal tension in the PG definition for communication-bound workloads, and the heavy reliance on anonymized or explicitly notional data, which limits independent verification of the quantitative claims. Overall, this is a promising industrial framework rather than a fully validated scientific metric; the decomposition is partly definitional, since SG by construction reflects scheduling losses, but the reported external improvements give it practical value.
major comments (5)
- [Section 4.2, Figures 8-9] The definition of Program Goodput is internally inconsistent with the paper's own workload characterization. PG is defined as T_ideal/T_actual with T_ideal = FLOPs(unoptimized HLO)/peak FLOPS, and the text calls this 'intrinsic' and 'agnostic to compiler decisions.' Yet Section 2.1 states that ML workloads 'tend to be more communication-bound rather than compute-bound,' and Section 5.3 uses low PG to conclude that 'many high-cost workloads are communication-bound.' For a communication-bound workload, T_actual is dominated by collective communication and memory stalls, so PG is low even for a perfectly optimized program. PG is therefore a compute-utilization proxy, not a measure of 'how well optimized the ML program is,' and the MPG decomposition does not cleanly isolate the program/compiler layer. The denominator should either include communication time explicitly, or the interpretation of PG should be restricted.
- [Section 4.2, Figure 10] The definitions of SG, RG, and PG are not formal enough to be reproducible. For SG, the numerator is 'simultaneous uptime of all tasks' and the denominator is 'fleet capacity, expressed as chip-time,' but it is not specified whether capacity means all fleet chips during the window, only allocated chips, or something else, nor how partially allocated intervals are counted. For RG, the numerator is 'productive chip-time ... saved in checkpoints,' but the text does not define how productive time is measured between two checkpoints when no failure or preemption occurs; presumably all such time is productive, but this must be stated. For PG, the ideal-time computation from HLO FLOPs lacks a precise specification of which HLO stage is counted as 'unoptimized' and which peak FLOPS number is used per accelerator generation. Equations and a time-window definition are needed before the metric can be evaluated.
- [Section 5, Table 2] The predicted effects of compiler, runtime, and scheduler optimizations are presented as qualitative assertions without a formal model or empirical validation. For example, the compiler row claims PG 'Increases,' SG and MPG 'Decreases if device-bound' but 'No change if host-bound,' yet no derivation in the text connects a change in step time to these component changes. Since the surrounding sections claim to use MPG to validate optimizations, Table 2 should be either derived from the definitions in Section 4.2 or accompanied by measured before-and-after MPG decomposition data.
- [Section 5.3, Figure 11] The claim that an XLA algebraic simplification improves fleet-wide Program Goodput is based on a benchmark of the 'top 150 most costly workloads,' but no evidence is given that this set is representative of the fleet for evaluating compiler changes. Cost-based selection likely overweights compute-heavy, long-running workloads, so the reported PG improvement may not generalize to the rest of the fleet. The paper should justify representativeness or report the effect on the full fleet.
- [Section 5.2, Figures 13-15] Several central quantitative results are reported with anonymized or explicitly notional data, including 'segments in Figure 13 are not explicitly identified,' a 'notional slice' in Figure 14, and a 'notional example' in Figure 15. This prevents independent verification of the claimed trends and of the general conclusion that MPG 'precisely pinpoint[s] optimization opportunities.' The authors should either declassify the data, describe the anonymization procedure, or clearly separate notional illustrations from measured results and state which claims depend on which category.
minor comments (5)
- [Section 4.1] The statement that 'the MPG metric must be a clearly defined and accurate measure of forward progress' reads as a design goal, but the paper never formally defines 'forward progress'; consider defining it or restricting the claim to the three sub-metrics.
- [Section 3.1 and Section 4.2] Notation is inconsistent: 'FLOPS' and 'FLOPs' are used interchangeably, and 'TOPs/Watt' appears without a definition; please standardize the units and their capitalization.
- [Section 5.1, Figure 12] The claim that 'the overall SG is already close to optimal' is not supported by a numeric baseline or a definition of 'optimal'; please provide a reference value or a formal optimality criterion.
- [Section 5.2, Figure 14] The sentence 'This has resulted in an temporary decrease of RG for the bulk inference segment' contains a grammar error ('an temporary' should be 'a temporary').
- [Section 2.1] The claim that ML workloads 'tend to be more communication-bound rather than compute-bound' cites a parameter-server paper [36]; a more direct citation on collective communication in distributed DNN training, such as Wang et al. [63], would better support the statement.
Circularity Check
MPG decomposition is an accounting identity by construction; PG's compute-only roofline is a validity concern, but no fitted-input or self-citation circularity found.
-
self definitional
[Section 1 (contributions) and Section 4.2, Figure 8: definition of Scheduling/Runtime/Program Goodput and MPG decomposition.]
"By breaking down the metric into subcomponents—Scheduling Goodput (SG), Runtime Goodput (RG), and Program Goodput (PG)—and examining it along the axes of fleet characteristics, we can precisely pinpoint optimization opportunities. ... Scheduling Goodput (SG) quantifies the efficiency of resource allocation in an ML fleet. It measures the fraction of time that an application has all the required resources simultaneously available to make progress."
The paper builds MPG as the product of SG, RG, and PG, and defines SG as the scheduling-layer term, RG as the runtime-layer term, and PG as the program-layer term. Consequently, 'low SG identifies a scheduling bottleneck' is true by definition: the decomposition is an accounting identity over the authors' chosen layer taxonomy, not an empirical derivation of where losses come from. The 'precisely pinpoint' claim therefore restates the metric's construction rather than an independent finding. This is mild: no data are fitted to make the components line up, and the Section 5 optimization outcomes are validated with measured throughput/speedup improvements rather than by MPG alone.
full rationale
I found no fitted-input-as-prediction, no load-bearing self-citation, and no uniqueness theorem. Self-citations to XTAT [48] and Kumar et al. [34] are examples of deployed optimizations, not the basis of the metric. The decomposition's 'pinpointing' language is definitional because each goodput is constructed as a layer-specific ratio; this is the only circular element. The PG roofline concern—FLOP-only ideal time for workloads the paper itself characterizes as communication-bound—is a measurement-validity and attribution risk rather than a result-in-implies-result-out circularity, so it does not raise the score above 3. The central contribution is a proposed metric plus empirical fleet observations and externally measured optimizations, which carry independent content beyond the definitions.
Assumptions & free parameters
assumptions (5)
- domain assumption ML workloads make progress only when all allocated chips are available simultaneously (bulk-synchronous assumption).
- domain assumption Only checkpointed work counts as productive runtime progress.
- ad hoc to paper The ideal execution time for Program Goodput is the unoptimized HLO graph's FLOP count divided by theoretical peak FLOPS.
- domain assumption Fleet capacity can be expressed as a homogeneous chip-time denominator across heterogeneous accelerators.
- ad hoc to paper The top 150 most costly workloads are a representative benchmark for evaluating fleet-wide compiler changes.
invented entities (1)
-
ML Productivity Goodput (MPG) and its sub-metrics SG, RG, PG
Cite this review
Pith. "Pith review of Machine Learning Fleet Efficiency: Analyzing and Optimizing Large-Scale Google TPU Systems with ML Productivity Goodput." pith.science (2026). https://pith.science/paper/33COF5SI
@misc{pith2026250206982,
author = {Pith},
title = {Pith review of: Machine Learning Fleet Efficiency: Analyzing and Optimizing Large-Scale Google TPU Systems with ML Productivity Goodput},
year = {2026},
howpublished = {\url{https://pith.science/paper/33COF5SI}},
note = {Machine review of arXiv:2502.06982}
}
read the original abstract
Recent years have seen the emergence of machine learning (ML) workloads deployed in warehouse-scale computing (WSC) settings, also known as ML fleets. As the computational demands placed on ML fleets have increased due to the rise of large models and growing demand for ML applications, it has become increasingly critical to measure and improve the efficiency of such systems. However, there is not yet an established methodology to characterize ML fleet performance and identify potential performance optimizations accordingly. This paper presents a large-scale analysis of an ML fleet based on Google's TPUs, introducing a framework to capture fleet-wide efficiency, systematically evaluate performance characteristics, and identify optimization strategies for the fleet. We begin by defining an ML fleet, outlining its components, and analyzing an example Google ML fleet in production comprising thousands of accelerators running diverse workloads. Our study reveals several critical insights: first, ML fleets extend beyond the hardware layer, with model, data, framework, compiler, and scheduling layers significantly impacting performance; second, the heterogeneous nature of ML fleets poses challenges in characterizing individual workload performance; and third, traditional utilization-based metrics prove insufficient for ML fleet characterization. To address these challenges, we present the "ML Productivity Goodput" (MPG) metric to measure ML fleet efficiency. We show how to leverage this metric to characterize the fleet across the ML system stack. We also present methods to identify and optimize performance bottlenecks using MPG, providing strategies for managing warehouse-scale ML systems in general. Lastly, we demonstrate quantitative evaluations from applying these methods to a real ML fleet for internal-facing Google TPU workloads, where we observed tangible improvements.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 3 Pith papers
-
MLSYSIM: First-Principles Infrastructure Modeling for Machine Learning Systems
A dimensionally strict analytical framework codifies 22 ML systems walls into 28 composable resolvers for sub-second full-stack design-space exploration and hardware synthesis.
-
A Taxonomy of Performance Metrics for the Distributed Computing Continuum
A three-layer taxonomy (compute, network, application) plus novel continuum-specific metrics and acquisition requirements for evaluating distributed computing continuum systems.
-
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
This survey classifies distributed DNN training simulators into analytical, profiling-based, and execution-driven categories, and compares them alongside TCO and carbon-emission models.
Reference graph
Works this paper leans on
-
[1]
[n. d.]. XLA. https://openxla.org/xla
-
[2]
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, San- jay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Leven- berg, Dan Mane, Rajat Monga, Sherry Moore, Derek Murray,...
-
[3]
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Man- junath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016. TensorFlow: A system f...
arXiv 2016
-
[4]
Banning, Sumeer Bhola, Rick Buskens, Ming Chen, Xi Chen, Yoo Chung, Qin Jia, Nick Sakharov, George T
Colin Adams, Luis Alonso, Ben Atkin, John P. Banning, Sumeer Bhola, Rick Buskens, Ming Chen, Xi Chen, Yoo Chung, Qin Jia, Nick Sakharov, George T. Talbot, Adam Jacob Tart, and Nick Taylor (Eds.). 2020. Monarch: Google’s Planet- Scale In-Memory Time Series Database
work page 2020
-
[5]
Anthropic AI. 2024. The Claude 3 Model Family: Opus, Sonnet, Haiku
work page 2024
-
[6]
Krste Asanovic, Rastislav Bodik, James Demmel, Tony Keaveny, Kurt Keutzer, John Kubiatowicz, Nelson Morgan, David Patterson, Koushik Sen, John Wawrzynek, David Wessel, and Katherine Yelick. 2009. A view of the paral- lel computing landscape. Commun. ACM 52, 10 (Oct. 2009), 56–67. https: //doi.org/10.1145/1562764.1562783
arXiv 2009
-
[7]
Paul Barham, Aakanksha Chowdhery, Jeff Dean, Sanjay Ghemawat, Steven Hand, Dan Hurt, Michael Isard, Hyeontaek Lim, Ruoming Pang, Sudip Roy, Brennan Saeta, Parker Schuh, Ryan Sepassi, Laurent El Shafey, Chandramohan A. Thekkath, and Yonghui Wu. 2022. Pathways: Asynchronous Distributed Dataflow for ML. arXiv preprint arXiv:2203.12533 (2022). https://arxiv.o...
arXiv 2022
-
[8]
Luiz André Barroso and Urs Hölzle. 2009. The Datacenter as a Computer: An Introduction to the Design of Warehouse-Scale Machines . http://dx.doi.org/10. 2200/S00193ED1V01Y200905CAC006
work page 2009
Show all 76 references
-
[9]
Nathan Bell and Michael Garland. 2008. Efficient sparse matrix-vector multiplica- tion on CUDA. Technical Report. Nvidia Technical Report NVR-2008-004, Nvidia Corporation
2008
-
[10]
Cooper, and Linda Torczon
Preston Briggs, Keith D. Cooper, and Linda Torczon. 1992. Rematerializa- tion. In Proceedings of the ACM SIGPLAN 1992 Conference on Programming Language Design and Implementation (San Francisco, California, USA) (PLDI ’92). Association for Computing Machinery, New York, NY, US...
1992
-
[11]
Sergey Brin and Lawrence Page. 1998. The Anatomy of a Large-Scale Hypertex- tual Web Search Engine. Computer Networks 30 (1998), 107–117. http://www- db.stanford.edu/~backrub/google.html
1998
-
[12]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...
2020 arXiv
-
[13]
Brendan Burns, Brian Grant, David Oppenheimer, Eric Brewer, and John Wilkes
-
[14]
Arnab Choudhury, Yang Wang, Tuomas Pelkonen, Kutta Srinivasan, Abha Jain, Shenghao Lin, Delia David, Siavash Soleimanifard, Michael Chen, Abhishek Yadav, Ritesh Tijoriwala, Denis Samoylov, and Chunqiang Tang. 2024. MAST: Global Scheduling of ML Training across Geo-Distributed ...
2024
-
[15]
Borg, Omega, and Kubernetes. Commun. ACM 59, 5 (apr 2016), 50–57. https://doi.org/10.1145/2890784
2016 doi
-
[16]
Lieven Eeckhout. 2010. Computer architecture performance evaluation methods . Morgan & Claypool Publishers
2010
-
[17]
James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost, JJ Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Pe- ter Hochschild, Wilson Hsieh, Sebastian Kanthak, Eugene Kogan, Hongyi Li, Alexander Lloyd, Sergey Melnik, David Mwaura, David N...
2012
-
[18]
Andre Esteva, Katherine Chou, Serena Yeung, Nikhil Naik, Ali Madani, Ali Mot- taghi, Yun Liu, Eric Topol, Jeff Dean, and Richard Socher. 2021. Deep learning- enabled medical computer vision. NPJ digital medicine 4, 1 (2021), 5
2021
-
[19]
Emer and Douglas W
Joel S. Emer and Douglas W. Clark. 1984. A Characterization of Processor Performance in the vax-11/780. SIGARCH Comput. Archit. News 12, 3 (jan 1984), 301–310. https://doi.org/10.1145/773453.808199
1984
-
[20]
Roy Frostig, Matthew Johnson, and Chris Leary. 2018. Compiling machine learning programs via high-level tracing. https://mlsys.org/Conferences/doc/ 2018/146.pdf
2018
-
[21]
Kayvon Fatahalian, Jeremy Sugerman, and Pat Hanrahan. 2004. Understanding the efficiency of GPU algorithms for matrix-matrix multiplication. In Proceedings of the ACM SIGGRAPH/EUROGRAPHICS conference on Graphics hardware . 133– 137
2004
-
[22]
James Hamilton. 2007. On Designing and Deploying Internet-Scale Services. In 21st Large Installation System Administration Conference (LISA 07) . USENIX Association, Dallas, TX. https://www.usenix.org/conference/lisa-07/designing- and-deploying-internet-scale-services
2007
-
[23]
Sanjay Ghemawat, Howard Gobioff, and Shun-Tak Leung. 2003. The Google File System. In Proceedings of the 19th ACM Symposium on Operating Systems Principles. Bolton Landing, NY, 20–43
2003
-
[24]
Dean Hildebrand and Denis Serenyi. 2021. Colossus under the hood: a peek into Google’s scalable storage system . https://cloud.google.com/blog/products/ storage-data-transfer/a-peek-behind-colossus-googles-file-system
2021
-
[25]
Hennessy and David A
John L. Hennessy and David A. Patterson. 2019. A new golden age for computer architecture. Commun. ACM 62, 2 (jan 2019), 48–60. https://doi.org/10.1145/ 3282307
2019
-
[26]
JAX. [n. d.]. Ahead-of-time lowering and compilation. https://jax.readthedocs.io/ en/latest/aot.html
-
[27]
Joel Janai, Fatma Güney, Aseem Behl, Andreas Geiger, et al. 2020. Computer vision for autonomous vehicles: Problems, datasets and state of the art. Foundations and Trends® in Computer Graphics and Vision 12, 1–3 (2020), 1–308
2020
-
[28]
Norman P. Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, Cliff Young, Xiang Zhou, Zongwei Zhou, and David Patterson. 2023. TPU v4: An Optically Reconfigurable Supercomputer for Machine Learn...
2023 arXiv
-
[29]
Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B
Norman P. Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B. Jablin, George Kurian, James Laudon, Sheng Li, Peter Ma, Xiaoyu Ma, Thomas Norrie, Nishant Patil, Sushma Prasad, Cliff Young, Zongwei Zhou, and David Patterson. 2021. Ten Lessons From Three Generations...
2021
-
[30]
Shoaib Kamil, John Shalf, and Erich Strohmaier. 2008. Power efficiency in high performance computing. In 2008 IEEE International Symposium on Parallel and Distributed Processing. 1–8. https://doi.org/10.1109/IPDPS.2008.4536223
2008
-
[31]
Jouppi, Cliff Young, Nishant Patil, David A
Norman P. Jouppi, Cliff Young, Nishant Patil, David A. Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, 11 Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tar...
-
[32]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. ImageNet Clas- sification with Deep Convolutional Neural Networks. In Advances in Neural Information Processing Systems, F. Pereira, C.J. Burges, L. Bottou, and K.Q. Wein- berger (Eds.), Vol. 25. Curran Associates, ...
2012
-
[33]
Michael Kuchnik, Ana Klimovic, Jiri Simsa, Virginia Smith, and George Amvrosiadis. 2022. Plumber: Diagnosing and Removing Performance Bottlenecks in Machine Learning Data Pipelines.Proceedings of Machine Learning and Systems 4 (2022), 33–51
2022
-
[34]
Svilen Kanev, Juan Pablo Darago, Kim Hazelwood, Parthasarathy Ranganathan, Tipp Moseley, Gu-Yeon Wei, and David Brooks. 2015. Profiling a warehouse- scale computer. In Proceedings of the 42nd Annual International Symposium on Computer Architecture. 158–169
2015
-
[35]
Jie Li, George Michelogiannakis, Brandon Cook, Dulanya Cooray, and Yong Chen
-
[36]
Mu Li, David G Andersen, Alexander J Smola, and Kai Yu. 2014. Communication efficient distributed machine learning with the parameter server. Advances in Neural Information Processing Systems 27 (2014)
2014
-
[37]
Sameer Kumar, Yu Wang, Cliff Young, James Bradbury, Naveen Kumar, Dehao Chen, and Andy Swing. 2021. Exploring the limits of Concurrency in ML Training on Google TPUs. Proceedings of Machine Learning and Systems 3 (2021), 81–92
2021
-
[38]
Peter Mattson, Christine Cheng, Gregory Diamos, Cody Coleman, Paulius Micike- vicius, David Patterson, Hanlin Tang, Gu-Yeon Wei, Peter Bailis, Victor Bittorf, et al. 2020. Mlperf training benchmark. Proceedings of Machine Learning and Systems 2 (2020), 336–349
2020
-
[39]
Mustafa Rafique, Franck Cappello, and Bogdan Nicolae
Avinash Maurya, Robert Underwood, M. Mustafa Rafique, Franck Cappello, and Bogdan Nicolae. 2024. DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models. In Proceedings of the 33rd International Symposium on High-Performance Parallel and Distributed Computing...
2024
-
[40]
D Menemenlis, C Hill, A Adcrocft, J-M Campin, B Cheng, B Ciotti, I Fukumori, P Heimbach, C Henze, A Kohl, et al . 2005. NASA supercomputer improves prospects for ocean climate research. Eos, Transactions American Geophysical Union 86, 9 (2005), 89–96
2005
-
[41]
Jason Mars, Lingjia Tang, Robert Hundt, Kevin Skadron, and Mary Lou Soffa
-
[42]
Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G. Azzolini, Dmytro Dzhulgakov, Andrey Mallevich, Ilia Cherni- avskii, Yinghai Lu, Raghuraman Krishnamoorthi, Ansha Yu, V...
2019 arXiv
-
[43]
Wozniak, George Bosilca, Matthieu Dorier, and Franck Cappello
Bogdan Nicolae, Jiali Li, Justin M. Wozniak, George Bosilca, Matthieu Dorier, and Franck Cappello. 2020. DeepFreeze: Towards Scalable Asynchronous Check- pointing of Deep Learning Models. In 2020 20th IEEE/ACM International Sym- posium on Cluster, Cloud and Internet Computing ...
2020
-
[44]
Li, Ryan McElroy, Mike Paleczny, Daniel Peek, Paul Saab, David Stafford, Tony Tung, and Venkateshwaran Venkataramani
Rajesh Nishtala, Hans Fugal, Steven Grimm, Marc Kwiatkowski, Herman Lee, Harry C. Li, Ryan McElroy, Mike Paleczny, Daniel Peek, Paul Saab, David Stafford, Tony Tung, and Venkateshwaran Venkataramani. 2013. Scaling Memcache at Facebook. In 10th USENIX Symposium on Networked Sys...
2013
-
[45]
Tetsuya Odajima, Yuetsu Kodama, Miwako Tsuji, Motohiko Matsuda, Yutaka Maruyama, and Mitsuhisa Sato. 2020. Preliminary Performance Evaluation of the Fujitsu A64FX Using HPC Applications. In 2020 IEEE International Conference on Cluster Computing (CLUSTER). 523–530. https://doi...
2020
-
[46]
Murray, Jiri Simsa, Ana Klimovic, and Ihor Indyk
Derek G. Murray, Jiri Simsa, Ana Klimovic, and Ihor Indyk. 2021. tf.data: A Machine Learning Data Processing Framework. arXiv:2101.12127 [cs.LG] https: //arxiv.org/abs/2101.12127
2021 arXiv
-
[47]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fan...
2019 arXiv
-
[48]
Phitchaya Mangpo Phothilimthana, Amit Sabne, Nikhil Sarda, Karthik Srini- vasa Murthy, Yanqi Zhou, Christof Angermueller, Mike Burrows, Sudip Roy, Ketan Mandke, Rezsa Farahani, et al. 2021. A Flexible Approach to Autotuning Multi-Pass Machine Learning Compilers. In 2021 30th I...
2021
-
[49]
Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean
-
[50]
Parthasarathy Ranganathan and Urs Holzle. 2024. Twenty Five Years of Warehouse-Scale Computing . IEEE Micro 44, 05 (Sept. 2024), 11–22. https: //doi.org/10.1109/MM.2024.3409469
2024
-
[51]
OpenXLA. [n. d.]. Using AOT compilation. https://openxla.org/xla/tf2xla/ tfcompile
-
[52]
Paul Menage Sanjay Ghemawat. [n. d.]. TCMalloc : Thread-Caching Malloc . https://goog-perftools.sourceforge.net/doc/tcmalloc.html
-
[53]
Roland Schulz, Benjamin Lindner, Loukas Petridis, and Jeremy C Smith. 2009. Scaling of multimillion-atom biological molecular dynamics simulation on a petascale supercomputer. Journal of Chemical Theory and Computation 5, 10 (2009), 2798–2808
2009
-
[54]
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. arXiv:1701.06538 [cs.LG] https: //arxiv.org/abs/1701.06538
2017 arXiv
-
[55]
In Proceedings of Ma- chine Learning and Systems , D
Efficiently Scaling Transformer Inference. In Proceedings of Ma- chine Learning and Systems , D. Song, M. Carbin, and T. Chen (Eds.), Vol. 5. Curan, 606–624. https://proceedings.mlsys.org/paper_files/paper/2023/file/ c4be71ab8d24cdfb45e3d06dbfca2780-Paper-mlsys2023.pdf
2023
-
[56]
Vladimir Stegailov, Ekaterina Dlinnova, Timur Ismagilov, Mikhail Khalilov, Niko- lay Kondratyuk, Dmitry Makagon, Alexander Semenov, Alexei Simonov, Grigory Smirnov, and Alexey Timofeev. 2019. Angara interconnect makes GPU-based Desmos supercomputer an efficient tool for molecu...
2019
-
[57]
David K. Rensin. 2015. Kubernetes - Scheduling the Future at Cloud Scale . 1005 Gravenstein Highway North Sebastopol, CA 95472. All pages. http://www.oreilly. com/webops-perf/free/kubernetes.csp
2015
-
[58]
Gemini Team. 2024. Gemini: A Family of Highly Capable Multimodal Models. arXiv:2312.11805 [cs.CL] https://arxiv.org/abs/2312.11805
2024 arXiv
-
[59]
Amin Vahdat and Mark Lohmeyer. 2023. Enabling next-generation AI workloads: Announcing TPU v5p and AI Hypercomputer. https: //cloud.google.com/blog/products/ai-machine-learning/introducing-cloud-tpu- v5p-and-ai-hypercomputer
2023
-
[60]
Kenton Varda. 2008. Protocol Buffers: Google’s Data Interchange Format . https: //opensource.googleblog.com/2008/07/protocol-buffers-googles-data.html
2008
-
[61]
Zhan Shi, Chirag Sakhuja, Milad Hashemi, Kevin Swersky, and Calvin Lin
-
[62]
Korupolu, David Oppenheimer, Eric Tune, and John Wilkes
Abhishek Verma, Luis Pedrosa, Madhukar R. Korupolu, David Oppenheimer, Eric Tune, and John Wilkes. 2015. Large-scale cluster management at Google with Borg. In Proceedings of the European Conference on Computer Systems (EuroSys) . Bordeaux, France
2015
-
[63]
Shibo Wang, Jinliang Wei, Amit Sabne, Andy Davis, Berkin Ilbeyi, Blake Hecht- man, Dehao Chen, Karthik Srinivasa Murthy, Marcello Maggioni, Qiao Zhang, et al. 2022. Overlap Communication with Dependent Computation via Decompo- sition in Large Deep Learning Models. InProceeding...
2022
-
[64]
Varun Talwar. 2016. gRPC: a true internet-scale RPC framework is now 1.0 and ready for production deployments . https://cloud.google.com/blog/ products/gcp/grpc-a-true-internet-scale-rpc-framework-is-now-1-and-ready- for-production-deployments
2016
-
[65]
Amir Yazdanbakhsh, Kiran Seshadri, Berkin Akin, James Laudon, and Ravi Narayanaswami. 2021. An Evaluation of Edge TPU Accelerators for Con- volutional Neural Networks. CoRR abs/2102.10423 (2021). arXiv:2102.10423 https://arxiv.org/abs/2102.10423
2021 arXiv
-
[66]
Yoo, Morris A
Andy B. Yoo, Morris A. Jette, and Mark Grondona. 2003. SLURM: Simple Linux Utility for Resource Management. In Job Scheduling Strategies for Parallel Pro- cessing, Dror Feitelson, Larry Rudolph, and Uwe Schwiegelshohn (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 44–60
2003
-
[67]
Mark Zhao, Niket Agarwal, Aarti Basant, Bugra Gedik, Satadru Pan, Mustafa Ozdal, Rakesh Komuravelli, Jerry Pan, Tianshu Bao, Haowei Lu, Sundaram Narayanan, Jack Langman, Kevin Wilfong, Harsha Rastogi, Carole-Jean Wu, Christos Kozyrakis, and Parik Pol. 2021. Understanding and C...
2021 arXiv
-
[68]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems , I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. ...
2017
-
[69]
Yazhou Zu, Alireza Ghaffarkhah, Hoang-Vu Dang, Brian Towles, Steven Hand, Safeen Huda, Adekunle Bello, Alexander Kolbasov, Arash Rezaei, Dayou Du, Steve Lacy, Hang Wang, Aaron Wisner, Chris Lewis, and Henri Bahini. 2024. Resiliency at Scale: Managing Google’s TPUv4 Machine Lea...
2024
-
[71]
Samuel Williams, Andrew Waterman, and David Patterson. 2009. Roofline: an insightful visual performance model for multicore architectures. Commun. ACM 52, 4 (apr 2009), 65–76. https://doi.org/10.1145/1498765.1498785
2009
-
[75]
Mark Zhao, Niket Agarwal, Aarti Basant, Buğra Gedik, Satadru Pan, Mustafa Ozdal, Rakesh Komuravelli, Jerry Pan, Tianshu Bao, Haowei Lu, Sundaram Narayanan, Jack Langman, Kevin Wilfong, Harsha Rastogi, Carole-Jean Wu, Christos Kozyrakis, and Parik Pol. 2022. Understanding data ...
2022
-
[2011]
In Proceedings of the 44th annual IEEE/ACM International Symposium on Microarchitecture
Bubble-up: Increasing utilization in modern warehouse scale computers via sensible co-locations. In Proceedings of the 44th annual IEEE/ACM International Symposium on Microarchitecture. 248–259
-
[2016]
arXiv:1603.04467 [cs.DC] https://arxiv.org/abs/1603.04467
TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems. arXiv:1603.04467 [cs.DC] https://arxiv.org/abs/1603.04467
-
[2017]
CoRR abs/1704.04760 (2017)
In-Datacenter Performance Analysis of a Tensor Processing Unit. CoRR abs/1704.04760 (2017). arXiv:1704.04760 http://arxiv.org/abs/1704.04760
2017 arXiv
-
[2020]
CoRR abs/2010.02075 (2020)
Learned Hardware/Software Co-Design of Neural Accelerators. CoRR abs/2010.02075 (2020). arXiv:2010.02075 https://arxiv.org/abs/2010.02075
2020 arXiv
-
[2023]
In International Conference on High Performance Computing
Analyzing resource utilization in an HPC system: A case study of NERSC’s Perlmutter. In International Conference on High Performance Computing . Springer, 297–316
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.