REVIEW 3 major objections 5 minor 3 cited by
Hardware Trends Impacting Floating-Point Computations In Scientific Applications
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper argues that AI hardware's low-precision engines can be emulated up to double-precision accuracy, making the fastest chips also the most efficient for science.
desk verdict A competent survey with one load-bearing preliminary benchmark table that needs accuracy validation before its emulation speedup claim is taken at face value. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is emulation of high-precision arithmetic from many low-precision 'slices,' combined with mixed-precision iterative refinement. In the $s=7$ case, double-precision data are represented and multiplied as seven 8-bit integer slices on the tensor-core matrix-multiply units that AI workloads use, so a chip's cheap integer throughput becomes FP64-grade arithmetic; the paper relies on a recently published scheme for making integer matrix-multiplication units deliver that emulation efficiently. The companion mechanism is iterative refinement: a low-precision factorization does the heavy lifting, and a higher-precision residual correction repeats until the solution meets double-precision accuracy. Stochastic rounding is also cited as an error-control tool that keeps low-precision steps from polluting the final result.
What would settle it
On one GPU, solve the same dense system with native FP64 and with INT8 $s=7$ emulation at matched problem size, and compare the residual norms of the computed solutions; if the emulated residual is meaningfully worse, the reported 2x speedup is a precision trade-off, not a pure emulation gain.
Extended reading notes
Core claim
The paper's central claim, stated on its own terms, is that the AI-driven adoption of reduced-precision floating-point types is not merely a challenge for scientific computing but an opportunity, because emulation and mixed-precision algorithms can convert abundant low-precision throughput into double-precision-quality results. The load-bearing demonstration is an HPL measurement on a B200-class GPU: with emulation using $s=7$ eight-bit integer data elements, the run reached about 68 TFLOP/s compared with 34.5 TFLOP/s for native FP64 at maximum performance, and about 53 TFLOP/s compared with 23 TFLOP/s at maximum efficiency, with energy efficiency improving by roughly 60 to 70 percent. For a 32,000-by-32,000 complex double-precision system, the paper's mixed-precision iterative refinement solver reached 124 TFLOP/s and 529 GFLOP/s/Watt on an H200 GPU, against 42.6 and 78 for native FP64. The broader assertion is that dynamically switching precision during a computation, and emulating high precision on low-precision hardware, will define how scientific applications stay accurate while riding the performance curve of AI hardware.
Load-bearing premise
The emulated INT8 HPL run is assumed to be as accurate as native FP64, but the paper reports only speed and power efficiency, not residuals or error.
Editorial extensions
If this is right
- A machine's scientific throughput becomes tied to its low-precision throughput, so the correlation between the standard FP64 ranking and application-relevant performance weakens.
- Dense solvers can be restructured into a low-precision bulk phase plus a high-precision correction phase, with the reported payoff of 4.4x speed and 5.8x energy efficiency on data-center GPUs.
- Energy per useful operation, not peak FLOPS, becomes the binding design constraint as power budgets approach 40 megawatts.
- Emulation becomes a deliberate design feature, not a stopgap, letting a single system cover both AI and double-precision scientific workloads.
Reading between the lines
- If the accuracy of the emulated HPL run is confirmed, the same seven-slice technique should transfer to other dense linear algebra kernels, giving near-2x speedups on existing AI accelerators.
- The paper stops short of showing the emulated HPL run is as accurate as native FP64; a reader who wants the speedup should check residuals first.
- For ill-conditioned systems, iterative refinement will need extra passes, so the emulation advantage should shrink as the condition number grows; that is a testable prediction.
- If the historical widening of the gap between low-precision and high-precision throughput continues, the emulation speedup should grow across future hardware generations.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a perspective/review article on how hardware trends, particularly the adoption of reduced-precision floating-point formats driven by AI, are reshaping scientific computing. It surveys the historical evolution of floating-point support from software emulation and coprocessors to integrated FPUs and GPUs, discusses community benchmarks (HPL, HPCG, Green500, HPL-MxP), and synthesizes recent developments in mixed-precision algorithms and emulation techniques. The original contributions are quantitative: Table I charts throughput and memory bandwidth across four NVIDIA GPU generations; Table II reports performance and power efficiency for a mixed-precision iterative-refinement solver; Table III reports preliminary HPL results on a Blackwell B200 comparing native FP64 with INT8-slice emulation (s=7), claiming a 2.0-2.3x speedup and 60-70% power-efficiency improvement. The paper argues that these trends reflect a broader shift toward flexibility in precision, where emulation and mixed-precision allow lower-precision hardware to serve high-precision scientific workloads.
Significance. If the new measurements are trustworthy, the paper provides a valuable and credible synthesis of the current state and near-term trajectory of floating-point computing, with an authoritative author list. The survey of historical developments and benchmark evolution is generally accurate and well organized, and the emphasis on INT8-slice emulation as a means to leverage AI-oriented tensor hardware for FP64-class computation is a timely and potentially important observation. The paper would be strengthened as a reference point for the community if its quantitative claims were backed by reproducible methodology. The main significance currently rests on the two self-reported benchmark tables (II and III), which are exactly the parts that lack detail; this limits the paper's value as more than an opinionated literature review.
major comments (3)
- [Section IV-C, Table III] The central quantitative claim, that INT8 emulation with s=7 roughly doubles HPL performance and improves power efficiency by 60-70% versus native FP64 on a B200, assumes that both configurations are solving the same problem to the same numerical accuracy. The table and surrounding text report only TFLOP/s and GFLOP/s/Watt, with no residual, validation metric, or reference to the HPL acceptance criterion. The table is explicitly labeled 'preliminary' and cites reference [38], but the paper does not state whether that reference contains an accuracy validation, nor does Section IV-C provide any error analysis. As written, the comparison is only a throughput/efficiency comparison, not a capability-equivalent one. Please either provide the scaled residuals or other acceptance data demonstrating FP64-comparable accuracy, or revise the claim to be explicitly about raw speed without implying equivalent solution quality.
- [Section IV-B, Table II / Figure 4] The mixed-precision iterative-refinement results in Table II (and the A100 power curve in Figure 4) lack the experimental detail needed to assess the claimed 4.4x speedup and 5.8x power-efficiency gain. The table reports only the matrix size (32K complex) and the labels 'FP16+FP64 MxP', but does not specify the number of iterative-refinement iterations, the convergence threshold, the achieved residual, or whether the final solution is returned in FP64. Since this table is used as evidence that mixed-precision solvers preserve accuracy while improving efficiency, the authors should cite a public reproducible implementation, provide typical residual values, or state the accuracy target. Without this, the reader cannot tell whether the speedup is achieved while meeting the same numerical quality as the FP64 baseline.
- [Section XI-A, Figure 5] The dashed line in Figure 5, described as 'Tensor Core accelerated DGEMM performance using integer based emulation with 7 slices (INT8 data storage elements) [38]', is presented as if it were a direct point of comparison with the Bytes/FLOP curves of Table I, but the figure does not show the data points or the accuracy of the emulated DGEMM. If this line represents a single measurement or an extrapolation, that should be stated. The same accuracy-equivalence concern as in Table III applies here: without a statement about the numerical error of the emulated DGEMM relative to FP64, the comparison is misleading for readers who might infer that the emulated math is a drop-in replacement for native FP64 matrix multiply.
minor comments (5)
- [Figure 3 caption] The caption contains a typo: 'Fontier' should be 'Frontier'.
- [Throughout text and references] The name 'Volta' appears as 'V olta' in multiple places (e.g., Section IV-A, VI-B, and reference [50]); the spacing should be removed.
- [Table III] The variable 's' is used without definition. Define it as the number of integer slices used in the emulation scheme (presumably following [38]) and explain why s=7 was chosen.
- [Section V-A] The sentence 'as seen in Figure 3' for the precision/efficiency trade-off is confusing, because Figure 3 plots the HPL-MxP-to-HPL Rmax ratio over time, which is not directly a precision-versus-efficiency trade-off. Please re-reference or rephrase.
- [Table I / Section XI-A] Bytes/FLOP is an informative figure of merit, but the table should clarify whether the ratios are computed from the listed peak TFLOP/s values or from some measured values, and state the source or formula for the memory bandwidth numbers.
Circularity Check
No circular derivation: the paper's claims are contextual observations supported by benchmarks and vendor measurements, with only non-load-bearing self-citations.
full rationale
No load-bearing step reduces to the paper's own inputs. The central assertions—AI-driven reduced precision, mixed-precision gains, emulation benefits, and hardware/software co-evolution—are presented as review-level interpretations of external benchmarks (HPL, HPCG, Green500, TOP500, HPL-MxP) and of directly reported vendor measurements (Tables I-III, Figures 2-4). Table III reports a measured 2.0-2.3x HPL speedup for INT8 emulation versus native FP64; this is a benchmark ratio, not a quantity derived from an assumed accuracy equivalence, so it is not circular by construction. The absence of residual or validation data for the emulated run is a legitimate correctness/validation concern, but not a circularity. Self-authored citations ([5], [9], [10], [31]-[33]) appear as benchmark definitions and prior algorithm descriptions; the paper's trend conclusions do not rest on those citations as unverified premises. Therefore no specific circular step can be identified; the low score only notes the presence of self-citations and vendor-sourced measurements.
Assumptions & free parameters
assumptions (3)
- domain assumption Benchmark metrics (TOP500/HPL, Green500, HPCG, HPL-MxP) are meaningful proxies for real-world supercomputing performance.
- domain assumption The emulated HPL computation using INT8 slices (s=7) achieves the same numerical accuracy as native FP64.
- domain assumption Vendor-supplied performance and power measurements (Tables I to III) are accurate and representative.
Cite this review
Pith. "Pith review of Hardware Trends Impacting Floating-Point Computations In Scientific Applications." pith.science (2026). https://pith.science/paper/6GT5Z7QQ
@misc{pith2026241112090,
author = {Pith},
title = {Pith review of: Hardware Trends Impacting Floating-Point Computations In Scientific Applications},
year = {2026},
howpublished = {\url{https://pith.science/paper/6GT5Z7QQ}},
note = {Machine review of arXiv:2411.12090}
}
read the original abstract
The evolution of floating-point computation has been shaped by algorithmic advancements, architectural innovations, and the increasing computational demands of modern technologies, such as artificial intelligence (AI) and high-performance computing (HPC). This paper examines the historical progression of floating-point computation in scientific applications and contextualizes recent trends driven by AI, particularly the adoption of reduced-precision floating-point types. The challenges posed by these trends, including the trade-offs between performance, efficiency, and precision, are discussed, as are innovations in mixed-precision computing and emulation algorithms that offer solutions to these challenges. This paper also explores architectural shifts, including the role of specialized and general-purpose hardware, and how these trends will influence future advancements in scientific computing, energy efficiency, and system design.
Figures
Forward citations
Cited by 3 Pith papers
-
Ascend to Science: Exploration of AI Chips for Scientific Computing
AI-oriented Ascend NPUs can run scientific workloads with FP32-like accuracy and competitive throughput when algorithms are reformulated and data movement is explicitly orchestrated.
-
CHAMB-GA: A Containerized HPC Scalable Microservice-Based Framework for Genetic Algorithms
CHAMB-GA provides a microservice architecture with containers and a message broker to decouple genetic operations from fitness evaluations, enabling consistent scaling from small machines to over 3500 CPU cores on clo...
-
Mixed-precision numerics in scientific applications: survey and perspectives
A survey of mixed-precision numerical methods across CFD, climate, chemistry, and genomics, reporting speedups up to 8x on benchmarks and recommending co-design to unlock them.
Reference graph
Works this paper leans on
-
[38]
Performance enhancement of the ozaki scheme on integer matrix multiplication unit
Y . Uchino, K. Ozaki, and T. Imamura, “Performance enhancement of the ozaki scheme on integer matrix multiplication unit.” arXiv:2409.13313 [cs.DC], Sept. 2024
arXiv 2024
-
[1]
Deep learning,
Y . LeCun, Y . Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, p. 436, 2015
2015
-
[2]
A study of bfloat16 for deep learning training
D. Kalamkar, D. Mudigere, N. Mellempudi, D. Das, K. Banerjee, S. Avancha, D. T. V ooturi, N. Jammalamadaka, J. Huang, H. Yuen, J. Yang, J. Park, A. Heinecke, E. Georganas, S. Srinivasan, A. Kundu, M. Smelyanskiy, B. Kaul, and P. Dubey, “A study of bfloat16 for deep learning training.” arXiv:1905.12322 [cs.LG], May 2019
arXiv 1905
-
[3]
P. Micikevicius, D. Stosic, N. Burgess, M. Cornea, P. Dubey, R. Grisen- thwaite, S. Ha, A. Heinecke, P. Judd, J. Kamalu, N. Mellempudi, S. Oberman, M. Shoeybi, M. Siu, and H. Wu, “Fp8 formats for deep learning.” arXiv:2209.05433 [cs.LG], Sept. 2022. 8
arXiv 2022
-
[4]
R. C. Murphy, K. Pingali, J. D. Feo, and D. A. Bader, “Introducing the graph 500,” Cray User Group (CUG) , 2010
work page 2010
-
[5]
The linpack benchmark: Past, present, and future,
J. J. Dongarra, “The linpack benchmark: Past, present, and future,” Concurrency and Computation: Practice and Experience , vol. 15, no. 9, pp. 803–820, 2003
work page 2003
-
[6]
J. Dongarra and P. Luszczek, “Top500,” in Encyclopedia of Parallel Computing (D. Padua, ed.), pp. 2055–2057, Boston, MA: Springer US, 2011
work page 2011
-
[7]
E. Strohmaier, J. Dongarra, H. Simon, and M. Meuer, “Top 500. the list..” https://top500.org, June 2024
work page 2024
Show all 70 references
-
[8]
The green500 list: Encouraging sustainable supercomputing,
W. Feng and K. Cameron, “The green500 list: Encouraging sustainable supercomputing,” Computer, vol. 40, no. 12, pp. 50–55, 2007
2007
-
[9]
A new benchmark for ranking high performance computing systems,
J. J. Dongarra, M. A. Heroux, and P. Luszczek, “A new benchmark for ranking high performance computing systems,” Tech. Rep. UT-EECS- 13-736, University of Tennessee, 2013
2013
-
[10]
Hpl- ai mixed-precision benchmark: The next frontier of supercomputing,
J. J. Dongarra, P. Luszczek, S. Tomov, and M. A. Heroux, “Hpl- ai mixed-precision benchmark: The next frontier of supercomputing,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis , 2021
2021
-
[11]
The state of the transistor in 3 charts
IEEE Spectrum, “The state of the transistor in 3 charts.” https://spectr um.ieee.org/transistor-density. [Accessed 16-11-2024]
2024
-
[12]
The IBM 701 Speedcoding system,
J. W. Backus, “The IBM 701 Speedcoding system,” Journal of the ACM, vol. 1, no. 1, pp. 4–6, 1954
1954
-
[13]
The intel 8087 numeric data processor,
J. Palmer, “The intel 8087 numeric data processor,” in Proceedings of the 7th Annual Symposium on Computer Architecture, La Baule, France, May 6-8, 1980 (J. Lenfant, B. R. Borgerson, D. E. Atkins, K. B. Irani, D. Kinniment, and H. Aiso, eds.), pp. 174–181, ACM, 1980
1980
-
[14]
The MC68881 floating-point copro- cessor,
C. Huntsman and D. Cawthron, “The MC68881 floating-point copro- cessor,” IEEE Micro, vol. 3, pp. 44–54, Nov./Dec. 1983
1983
-
[15]
Developing the WTL3170/3171 Sparc floating-point copro- cessors,
M. Birman, A. Samuels, G. Chu, T. Chuk, L. Hu, J. McLeod, and J. Barnes, “Developing the WTL3170/3171 Sparc floating-point copro- cessors,” IEEE Micro, vol. 10, pp. 55–64, Jan./Feb. 1990
1990
-
[16]
Heinrich, MIPS R4000 user’s manual
J. Heinrich, MIPS R4000 user’s manual. USA: Prentice-Hall, Inc., 1993
1993
-
[17]
W. A. Triebel, The 80386, 80486, and Pentium Microprocessors: Hard- ware, Software, and Interfacing. Simon & Schuster Trade, 1st ed., 1997
1997
-
[18]
The 68040 processor. i. design and implementation,
R. Edenfield, M. Gallup, W. Ledbetter, R. McGarity, E. Quintana, and R. Reininger, “The 68040 processor. i. design and implementation,” IEEE Micro, vol. 10, no. 1, pp. 66–78, 1990
1990
-
[19]
Design considerations for the powerpc 601 microprocessor,
M. T. Vaden, L. J. Merkel, C. R. Moore, T. M. Potter, and R. J. Reese, “Design considerations for the powerpc 601 microprocessor,” IBM Journal of Research and Development , vol. 38, no. 5, pp. 605– 620, 1994
1994
-
[20]
Evolution of the graphics processing unit (gpu),
W. J. Dally, S. W. Keckler, and D. B. Kirk, “Evolution of the graphics processing unit (gpu),” IEEE Micro, vol. 41, no. 6, pp. 42–51, 2021
2021
-
[21]
Brook for gpus: stream computing on graphics hardware,
I. Buck, T. Foley, D. Horn, J. Sugerman, K. Fatahalian, M. Houston, and P. Hanrahan, “Brook for gpus: stream computing on graphics hardware,” ACM Trans. Graph., vol. 23, p. 777–786, Aug. 2004
2004
-
[22]
Scalable parallel programming with cuda.,
J. Nickolls, I. Buck, M. Garland, and K. Skadron, “Scalable parallel programming with cuda.,” in SIGGRAPH Classes , pp. 16:1–16:14, ACM, 2008
2008
-
[23]
Cuda: Scalable parallel programming for high-performance scientific computing,
D. Luebke, “Cuda: Scalable parallel programming for high-performance scientific computing,” in 2008 5th IEEE International Symposium on Biomedical Imaging: From Nano to Macro , pp. 836–838, 2008
2008
-
[24]
Accelerating molecular dynamics simulations using graphics processing units with cuda,
W. Liu, B. Schmidt, G. V oss, and W. M ¨uller-Wittig, “Accelerating molecular dynamics simulations using graphics processing units with cuda,” Computer physics communications, vol. 179, no. 9, pp. 634–641, 2008
2008
-
[25]
Gpu computing,
J. D. Owens, M. Houston, D. Luebke, S. Green, J. E. Stone, and J. C. Phillips, “Gpu computing,” Proceedings of the IEEE , vol. 96, no. 5, pp. 879–899, 2012
2012
-
[26]
Large-scale deep unsupervised learning using graphics processors,
R. Raina, A. Madhavan, and A. Y . Ng, “Large-scale deep unsupervised learning using graphics processors,” in Proceedings of the 26th annual international conference on machine learning, pp. 873–880, ACM, 2009
2009
-
[27]
Imagenet classifica- tion with deep convolutional neural networks,
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classifica- tion with deep convolutional neural networks,” in Advances in Neural Information Processing Systems (F. Pereira, C. Burges, L. Bottou, and K. Weinberger, eds.), vol. 25, Curran Associates, Inc., 2012
2012
-
[28]
OCP 8-bit Floating Point Specification (OFP8),
P. Micikevicius, S. Oberman, P. Dubey, M. Cornea, A. Rodriguez, I. Bratt, R. Grisenthwaite, N. Jouppi, C. Chou, A. Huffman, M. Schulte, R. Wittig, D. Jani, and S. Deng, “OCP 8-bit Floating Point Specification (OFP8),” Open Compute Project , 2023
2023
-
[29]
Microscal- ing data formats for deep learning
B. D. Rouhani, R. Zhao, A. More, M. Hall, A. Khodamoradi, S. Deng, D. Choudhary, M. Cornea, E. Dellinger, K. Denolf, S. Dusan, V . Elango, M. Golub, A. Heinecke, P. James-Roxby, D. Jani, G. Kolhe, M. Lang- hammer, A. Li, L. Melnick, M. Mesmakhosroshahi, A. Rodriguez, M. Schult...
-
[30]
Nvidia v100 gpu architecture
NVIDIA Corporation, “Nvidia v100 gpu architecture.” https://images.n vidia.com/content/volta-architecture/pdf/volta-architecture-whitepaper. pdf, 2017. White paper
2017
-
[31]
Accelerating scientific computations with mixed precision algorithms,
M. Baboulin, A. Buttari, J. Dongarra, J. Kurzak, J. Langou, J. Lan- gou, P. Luszczek, and S. Tomov, “Accelerating scientific computations with mixed precision algorithms,” Computer Physics Communications , vol. 180, no. 12, pp. 2526–2533, 2009
2009
-
[32]
A survey of numerical linear algebra methods utilizing mixed-precision arithmetic,
A. Abdelfattah, H. Anzt, E. G. Boman, E. Carson, T. Cojean, J. Don- garra, A. Fox, M. Gates, N. J. Higham, X. S. Li, J. Loe, P. Luszczek, S. Pranesh, S. Rajamanickam, T. Ribizel, B. F. Smith, K. Swirydowicz, S. Thomas, S. Tomov, Y . M. Tsai, and U. M. Yang, “A survey of numeri...
2021
-
[33]
The design of fast and energy-efficient linear solvers: On the potential of half-precision arithmetic and iterative refinement techniques,
A. Haidar, A. Abdelfattah, M. Zounon, P. Wu, S. Pranesh, S. Tomov, and J. Dongarra, “The design of fast and energy-efficient linear solvers: On the potential of half-precision arithmetic and iterative refinement techniques,” in International Conference on Computational Science...
2018
-
[34]
Deep learning with limited numerical precision,
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan, “Deep learning with limited numerical precision,” in Proceedings of the 32nd International Conference on Machine Learning (F. Bach and D. Blei, eds.), vol. 37 of Proceedings of Machine Learning Research , (Lille, Franc...
2015
-
[35]
Solving lattice qcd systems of equations using mixed precision solvers on gpus,
M. Clark, R. Babich, K. Barros, R. Brower, and C. Rebbi, “Solving lattice qcd systems of equations using mixed precision solvers on gpus,” Computer Physics Communications , vol. 181, no. 9, pp. 1517–1528, 2010
2010
-
[36]
Stochastic rounding: implementation, error analysis and applications,
M. Croci, M. Fasi, N. J. Higham, T. Mary, and M. Mikaitis, “Stochastic rounding: implementation, error analysis and applications,” Royal Soci- ety Open Science , vol. 9, no. 3, 2022
2022
-
[37]
Recovering single precision accuracy from tensor cores while surpassing the FP32 theoretical peak performance,
H. Ootomo and R. Yokota, “Recovering single precision accuracy from tensor cores while surpassing the FP32 theoretical peak performance,” Int. J. High Performance Computing Applications , vol. 36, p. 475–491, June 2022
2022
-
[39]
Leveraging the bfloat16 artificial intelligence datatype for higher-precision computations
G. Henry, P. T. P. Tang, and A. Heinecke, “Leveraging the bfloat16 artificial intelligence datatype for higher-precision computations.” arXiv:1904.06376 [cs.MS], Apr. 2019
1904 arXiv
-
[40]
Simulation intelligence: Towards a new generation of scientific methods
A. Lavin, D. Krakauer, H. Zenil, J. Gottschlich, T. Mattson, J. Brehmer, A. Anandkumar, S. Choudry, K. Rocki, A. G. Baydin, C. Prunkl, B. Paige, O. Isayev, E. Peterson, P. L. McMahon, J. Macke, K. Cranmer, J. Zhang, H. Wainwright, A. Hanuka, M. Veloso, S. Assefa, S. Zheng, and...
2021 arXiv
-
[41]
Functionality and performance of nvlink with ibm power9 processors,
IBM POWER9 NPU team, “Functionality and performance of nvlink with ibm power9 processors,” IBM J. Res. Dev. , vol. 62, p. 9:1–9:10, July 2018
2018
-
[42]
Nvidia gh200 grace hopper superchip architec- ture
NVIDIA Corporation, “Nvidia gh200 grace hopper superchip architec- ture.” https://resources.nvidia.com/en-us-grace-cpu/nvidia-grace-hopper, 2023
2023
-
[43]
Porting hpc applications to amd instinct TM mi300a using unified memory and openmp
S. Tandon, L. Grinberg, G.-T. Bercea, C. Bertolli, M. Olesen, S. Bn `a, and N. Malaya, “Porting hpc applications to amd instinct TM mi300a using unified memory and openmp.” arXiv.org:2405.00436 [cs.DC], May 2024
2024 arXiv
-
[44]
AMD CDNA Architecture
AMD Corporation, “AMD CDNA Architecture.” https://www.amd.com/ en/technologies/cdna.html. [Accessed 5-12-2024]
2024
-
[45]
Apple unveils m3, m3 pro, and m3 max, the most advanced chips for a personal computer
“Apple unveils m3, m3 pro, and m3 max, the most advanced chips for a personal computer.” https://www.apple.com/newsroom/2023/10/apple -unveils-m3-m3-pro-and-m3-max-the-most-advanced-chips-for-a-per sonal-computer/. [Accessed 14-11-2024]
2023
-
[46]
First impressions of the nvidia grace cpu superchip and nvidia grace hopper superchip for scientific workloads,
N. A. Simakov, M. D. Jones, T. R. Furlani, E. Siegmann, and R. J. Harrison, “First impressions of the nvidia grace cpu superchip and nvidia grace hopper superchip for scientific workloads,” in Proceedings of the International Conference on High Performance Computing in Asia- P...
2024
-
[47]
A survey of cpu-gpu heterogeneous computing techniques,
S. Mittal and J. S. Vetter, “A survey of cpu-gpu heterogeneous computing techniques,” ACM Comput. Surv., vol. 47, July 2015
2015
-
[48]
Programming model for a heterogeneous x86 platform,
B. Saha, X. Zhou, H. Chen, Y . Gao, S. Yan, M. Rajagopalan, J. Fang, P. Zhang, R. Ronen, and A. Mendelson, “Programming model for a heterogeneous x86 platform,” in Proceedings of the 30th ACM SIGPLAN Conference on Programming Language Design and Implementation , PLDI ’09, (New...
2009
-
[49]
Designing a unified program- ming model for heterogeneous machines,
M. Garland, M. Kudlur, and Y . Zheng, “Designing a unified program- ming model for heterogeneous machines,” in SC ’12: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis , pp. 1–11, 2012
2012
-
[50]
Dis- secting the nvidia volta gpu architecture via microbenchmarking
Z. Jia, M. Maggioni, B. Staiger, and D. P. Scarpazza, “Dis- secting the nvidia volta gpu architecture via microbenchmarking.” arXiv:1804.06826 [cs.DC], Apr. 2018
2018 arXiv
-
[51]
Tensor cores
NVIDIA Corporation, “Tensor cores.” https://www.nvidia.com/en-us/da ta-center/tensor-cores/, 2024
2024
-
[52]
CUDA PTX ISA
NVIDIA Corporation, “CUDA PTX ISA.” https://docs.nvidia.com/cuda /pdf/ptx isa 8.5.pdf, May 2024
2024
-
[53]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , NIPS’17, (Red Hook, NY , USA), p. 6000–6010, Curran...
2017
-
[54]
Mixed precision training,
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, et al. , “Mixed precision training,” International Conference on Learning Representa- tions (ICLR), 2018
2018
-
[55]
Simulating low precision floating-point arithmetic,
N. J. Higham and S. Pranesh, “Simulating low precision floating-point arithmetic,” SIAM Journal on Scientific Computing , vol. 41, no. 5, pp. C585–C602, 2019
2019
-
[56]
Improving weather forecast skill through reduced-precision data assimilation,
S. Hatfield, A. Subramanian, T. Palmer, and P. D ¨uben, “Improving weather forecast skill through reduced-precision data assimilation,” Monthly Weather Review, vol. 146, no. 1, pp. 49 – 62, 2018
2018
-
[57]
Pytorch: An imperative style, high- performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. K ¨opf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-...
1912 arXiv
-
[58]
Tensorflow: a system for large-scale machine learning,
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V . Vasudevan, P. Warden, M. Wicke, Y . Yu, and X. Zheng, “Tensorflow: a system for large-sca...
2016
-
[59]
JAX: composable transformations of Python+NumPy pro- grams
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclau- rin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang, “JAX: composable transformations of Python+NumPy pro- grams.” http://github.com/jax-ml/jax, 2018
2018
-
[60]
In-datacenter performance analysis of a tensor processing unit,
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P.-l. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V . Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D. ...
2017
-
[61]
Gpt-4 technical report
OpenAI, “Gpt-4 technical report.” https://arxiv.org/abs/2303.08774, 2024
2024 arXiv
-
[62]
Dynamic voltage and frequency scaling: the laws of diminishing returns,
E. Le Sueur and G. Heiser, “Dynamic voltage and frequency scaling: the laws of diminishing returns,” in Proceedings of the 2010 International Conference on Power Aware Computing and Systems , HotPower’10, (USA), p. 1–8, USENIX Association, 2010
2010
-
[63]
Intel® pentium® m processor power estimation, budgeting, optimization, and validation.,
D. Genossar and N. Shamir, “Intel® pentium® m processor power estimation, budgeting, optimization, and validation.,” Intel Technology Journal, vol. 7, no. 2, 2003
2003
-
[64]
Evaluating and modeling power consumption of multi-core processors,
R. Basmadjian and H. de Meer, “Evaluating and modeling power consumption of multi-core processors,” in Proceedings of the 3rd International Conference on Future Energy Systems: Where Energy, Computing and Communication Meet , e-Energy ’12, (New York, NY , USA), Association for...
2012
-
[65]
Beating floating point at its own game: Posit arithmetic,
Gustafson and Yonemoto, “Beating floating point at its own game: Posit arithmetic,” Supercomput. Front. Innov.: Int. J. , vol. 4, p. 71–86, June 2017
2017
-
[66]
The spinnaker 2 processing element architec- ture for hybrid digital neuromorphic computing
S. H ¨oppner, Y . Yan, B. V ogginger, C. Liu, F. Kelber, A. Dixius, S. Scholze, J. Partzsch, M. Stolba, F. Neum ¨arker, G. Ellguth, S. Hart- mann, S. Schiefer, T. Hocker, D. Walter, G. Liu, M. Mikaitis, J. Garside, S. Furber, and C. Mayr, “The spinnaker 2 processing element ar...
2022 arXiv
-
[67]
Analog and digital, continuous and discrete,
C. J. Maley, “Analog and digital, continuous and discrete,” Philosophical Studies, vol. 155, pp. 117–131, 2011
2011
-
[68]
Hbm (high bandwidth memory) dram technology and architecture,
H. Jun, J. Cho, K. Lee, H.-Y . Son, K. Kim, H. Jin, and K. Kim, “Hbm (high bandwidth memory) dram technology and architecture,” in 2017 IEEE International Memory Workshop (IMW) , pp. 1–4, 2017
2017
-
[69]
Flashattention: Fast and memory-efficient exact attention with io-awareness
T. Dao, D. Y . Fu, S. Ermon, A. Rudra, and C. R ´e, “Flashattention: Fast and memory-efficient exact attention with io-awareness.” https: //arxiv.org/abs/2205.14135, 2022
2022 arXiv
-
[70]
Flashattention-2: Faster attention with better parallelism and work partitioning
T. Dao, “Flashattention-2: Faster attention with better parallelism and work partitioning.” https://arxiv.org/abs/2307.08691, 2023. 10
2023 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.