Pith. sign in

REVIEW 5 minor 81 references

How to keep pushing ML accelerator performance? Know your rooflines!

T0 review · 0 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read An enhanced two-line roofline model, one for throughput and one for energy efficiency, can unify how ML accelerator designers reason about compute, memory, and data movement, and guide which optimization to pursue.

desk verdict A competent, well-organized survey that repackages the throughput and energy rooflines into a two-line framework for ML accelerators; useful for designers, but not a new research result. read the letter →

arxiv 2505.16346 v2 pith:D72YGMIR submitted 2025-05-22 cs.AR

classification cs.AR
keywords MLacceleratorsrooflinemodelenergyefficiencyarithmeticintensitydatareusequantizationsparsityin-memorycomputing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that a pair of rooflines, one for throughput and one for energy efficiency, gives a unifying picture of what limits ML accelerator performance. The throughput roofline is extended to multiple memory levels, each with its own arithmetic intensity, while the energy roofline is a curved bound separating regimes where memory energy or compute energy dominates. Under this view, every major accelerator technique, including parallelism, data reuse, quantization, sparsity, and near/in-memory compute, shows up either as raising a roofline, shifting workloads to higher arithmetic intensity, or improving utilization toward the roofline. The practical payoff is that a designer can look at a workload's position on the two rooflines and decide whether to invest in more compute, more memory bandwidth or energy efficiency, or better mapping and reuse.

What carries the argument

The load-bearing object is the per-memory-level arithmetic intensity $AI_{L_i}=N_{op}/N_{L_i}$, the number of operations per byte fetched from memory level $L_i$. It carries the argument because both rooflines are functions of this quantity: the throughput roofline of Eq. (4) takes a minimum over the per-level bandwidth products and the compute peak, while the energy roofline of Eq. (5) sums each level's energy per access divided by the per-level intensity. Data reuse, quantization, sparsity, and memory-proximity techniques all enter the model as changes in these intensities or in the roofline parameters, which is what makes the framework unifying.

What would settle it

A concrete test: on a real accelerator, run a sparse tensor kernel and measure its per-level memory traffic, peak bandwidths, and peak MAC rate, then compute the predicted operating points from Eqs. (4) and (5). If the measured throughput and energy efficiency lie far below both rooflines and the gap cannot be explained by the utilization factors described in Section II-B, the claim that the two rooflines bound and explain accelerator efficiency for that workload is refuted. A softer check is to take two architectures with identical rooflines but different dataflow flexibility and show that workload-level efficiency differs in a way the rooflines cannot express.

Watch

Extended reading notes

Core claim

The paper's central claim is that attainable throughput is $P_{TP}=f_{\mathrm{clk}}\min(AI_{L_n}B_{L_n},\dots,AI_{L_1}B_{L_1},A_{op})$ and attainable energy efficiency is $P_E=1/(E_{op}+\sum_i E_{L_i}/AI_{L_i})$, where $AI_{L_i}=N_{op}/N_{L_i}$ is the arithmetic intensity toward memory level $L_i$. Plotted together, these two expressions are claimed to explain where each execution regime falls and what action improves it. A distinctive consequence is that the throughput roofline has a sharp knee while the energy roofline is curved, and the two knees can sit at different arithmetic intensities, so an accelerator can be compute-bound for throughput and memory-bound for energy at the same operating point. The paper then reads the major efficiency techniques through this lens, treating sparsity as an intensity-lowering, utilization-changing effect and in-memory computing as a way to raise the compute roofline while eliminating the first-level memory diagonal, at the cost of new storage-compute coupling and utilization losses.

Load-bearing premise

The load-bearing premise is that a workload is well summarized by its arithmetic intensity at each memory level and that compute and memory latency overlap perfectly, so the minimum in the throughput equation captures latency; sparse and irregular workloads violate this, and the paper itself notes that they fall below the roofline and need hardware-specific utilization factors.

Editorial extensions

If this is right

  • A designer can classify any proposed optimization as raising the compute roofline, raising a memory roofline, moving the operating point to higher arithmetic intensity, or improving utilization, and can choose the category that addresses the actual bottleneck.
  • Because the throughput and energy knees depend on different parameters, improving peak TOPS and improving TOPS/W can require different changes to the same architecture, and a system can be compute-bound for one while memory-bound for the other.
  • Sparsity, especially unstructured sparsity, can lower effective arithmetic intensity and push a workload away from the roofline even when it reduces total operations; the paper concludes that roofline position alone is not enough to judge sparse hardware, and end-to-end energy and latency must be considered.
  • Near- and in-memory computing raise the compute roofline and reduce data-movement energy, but they couple storage and compute in ways that cause spatial and temporal utilization losses, so the full benefit depends on adding a second memory level and carefully mapping workloads.
  • Larger parallel compute arrays raise the roofline but shift the knee to higher arithmetic intensity, making the system more memory-bound unless data reuse keeps the workload's intensity high.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper does not formalize is a utilization-aware design rule: choose compute parallelism and memory bandwidth so that both knees sit near the arithmetic intensity of the target workload mix; the paper's area-allocation discussion points that way but stops short of a closed-form rule.
  • The framework suggests a reporting standard for published accelerators: give the throughput and energy rooflines together with the achieved utilization at the workload point, since identical rooflines can hide large efficiency differences; this is measurable and would make survey comparisons fairer.
  • For sparse workloads, defining a sparsity-aware arithmetic intensity that counts only non-zero operations and effective bytes with indices might restore the roofline's predictive power; the paper notes the gap but does not propose such a metric.
  • Treating the inter-chip network in multi-chip LLM systems as an additional memory level with its own bandwidth and energy per byte would extend the framework to scale-out architectures, which the paper mentions only as future outlook.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. This manuscript proposes an enhanced roofline modeling framework for ML accelerators, comprising a throughput roofline (Eq. 4) and an energy-efficiency roofline (Eq. 5) that account for per-level arithmetic intensities across a multi-level memory hierarchy. After deriving the two roofline equations from an additive energy model (Eq. 1) and a max-latency model (Eq. 2), the paper surveys five hardware techniques—parallelism, spatial/temporal data reuse, quantization, sparsity, and near/in-memory computing—and illustrates how each affects the roofline position and the workload operating point. It closes with trade-offs between parallelism and utilization, programmability and specialization, and an outlook on emerging technologies.

Significance. The paper's value lies in its organized synthesis and explicit two-roofline formulation. The equations are correct under the stated additive-energy and max-latency assumptions, and the paper is unusually candid about the model's limits, particularly in Section III-D and Section IV-B, where it acknowledges that sparse/irregular workloads fall below the roofline due to utilization losses and that roofline analysis alone is not sufficient for end-to-end speedup or energy decisions. No data fitting or fabricated results are involved; the illustrative parameters are clearly labeled as assumptions. These strengths make the paper a useful reference for practitioners and students, though it does not present new experimental data. The central claim is appropriately scoped as a framework for understanding rather than for prediction, so the absence of an integrated utilization factor is a limitation but not a fatal one.

minor comments (5)
  1. [Section III-F1, Eq. (6)] Equation (6) is self-referential as written: D_y appears on both sides, and the expression is dimensionally inconsistent with the surrounding prose. The text in Section III-F2 correctly states that the dynamic range in bits is the sum of input bits, weight bits, and log2(P_R). Please correct Eq. (6) to a bits-based form such as B_y = B_x + B_w + log2(P_R), or provide the equivalent linear-scale expression.
  2. [Abstract and title] The abstract and title contain the typo "rooline" instead of "roofline"; this should be corrected throughout.
  3. [Section I and Section V] The introduction states "nearly 100 × performance improvement every 24 months, maintained over the past 8 years" (which would imply an enormous cumulative factor), while the conclusion states "roughly 1000 × increase in throughput and energy efficiency." These two statements are inconsistent; please reconcile them or clarify the time scales being referenced.
  4. [Figure 3 caption] The caption's expression "AI=AIL3/16=AIL2=AIL1 ∗ 16" is ambiguous and potentially misleading. Please define the assumed relationships among the per-level arithmetic intensities explicitly, for instance by stating that each successive level differs by a factor of 16.
  5. [Section III-A] There is a typo in the first paragraph: "paralelization" should be "parallelization." Similar minor typographical errors appear elsewhere, such as "rooline" in the abstract and "the Samsung's" in Section III-A; a careful proofreading pass is recommended.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the two-roofline equations are algebraic restatements of the paper's own latency and energy definitions, and the central framing rests on external roofline literature.

full rationale

The paper's central equations are derived directly from its own stated definitions. Equation (4) is obtained by substituting the arithmetic-intensity definition AILi = Nop/NLi into the latency expression Ltask = (1/fclk)*max(NLn/BLn,...,NL1/BL1,Nop/Aop) and writing PTP = Nop/Ltask; Equation (5) is obtained by substituting the same definition into the energy expression Etask = Nop*Eop + sum(NLi*ELi) and writing PE = Nop/Etask. These are algebraic identities given the definitions, not empirical predictions, and no parameter is fitted to data and then renamed as a prediction. The paper makes no claim to predict measured accelerator performance from first principles; it presents an organizing framework whose validity rests on the external roofline model of Williams et al. [3] and the energy roofline of Choi et al. [9]. The self-citations that are present, such as [55] for the in-memory-compute dynamic-range trade-off, are explicitly attributed as prior work ('Simplifying the more detailed analysis in [55]') and are used as surveyed examples rather than as load-bearing justification for the paper's roofline framework. The paper even concedes in Section III-D that 'Roofline analysis alone, albeit useful, is not sufficient' for sparse workloads, which is a limitation of scope rather than a circular step. No load-bearing step in the derivation chain reduces by construction to its own inputs, so the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The framework introduces no fitted parameters; the numbers in Figure 3 are illustrative and clearly labeled as assumptions from a 22nm process. The load-bearing assumptions are the standard roofline model, the additive energy model, the monotonic arithmetic-intensity hierarchy, and the perfect-overlap latency model. The IMC trade-off analysis is adopted from the authors' own [55], with provenance stated.

assumptions (5)
  • standard math The classic roofline model of Williams et al. [3] and its energy variant by Choi et al. [9] are valid starting points.
    Section II.B builds Eq. (4) and (5) directly on these prior models, which are cited. The paper adds a multi-level memory hierarchy but does not question the underlying assumptions.
  • domain assumption Energy and latency of ML accelerators are dominated by MAC operations and data movement, each with constant per-byte/per-op costs (Eq. 1).
    Section II.A. This additive energy model ignores leakage, clock gating, and non-linear scaling, but is a common first-order model in the cited literature.
  • domain assumption Arithmetic intensity varies monotonically across memory levels (AIL3 > AIL2 > AIL1) for typical ML workloads.
    Section II.B, citing [8]. This drives the multi-level roofline. It does not hold for all layers, as the paper notes for batch-1 fully connected layers.
  • domain assumption Compute and memory transfers overlap such that latency is the max of the individual times (Eq. 4).
    Section II.B. If overlap is not perfect, the sharp roof becomes a rounded curve, which the paper acknowledges in the text and attributes to [10].
  • domain assumption The dynamic-range versus parallelism trade-off in IMC is as described in [55] by Verma et al.
    Section III-F.1 says 'Simplifying the more detailed analysis in [55]', so it adopts a prior result from one of the authors as the basis for the IMC trade-off discussion.

how reviews work

0 comments
Cite this review

Pith. "Pith review of How to keep pushing ML accelerator performance? Know your rooflines!." pith.science (2026). https://pith.science/paper/D72YGMIR

@misc{pith2026250516346,
  author       = {Pith},
  title        = {Pith review of: How to keep pushing ML accelerator performance? Know your rooflines!},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/D72YGMIR}},
  note         = {Machine review of arXiv:2505.16346}
}
read the original abstract

The rapidly growing importance of Machine Learning (ML) applications, coupled with their ever-increasing model size and inference energy footprint, has created a strong need for specialized ML hardware architectures. Numerous ML accelerators have been explored and implemented, primarily to increase task-level throughput per unit area and reduce task-level energy consumption. This paper surveys key trends toward these objectives for more efficient ML accelerators and provides a unifying framework to understand how compute and memory technologies/architectures interact to enhance system-level efficiency and performance. To achieve this, the paper introduces an enhanced version of the roofline model and applies it to ML accelerators as an effective tool for understanding where various execution regimes fall within roofline bounds and how to maximize performance and efficiency under the rooline. Key concepts are illustrated with examples from state-of-the-art designs, with a view towards open research opportunities to further advance accelerator performance.

Figures

Figures reproduced from arXiv: 2505.16346 by the authors.

Figure 1
Figure 1. Evolution of ML model size, GPU performance and Moore’s law [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Typical architecture of an ML accelerator. Note that the amount of [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Example roofline models for performance and energy effi [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (15 more)
Figure 4
Figure 4. Figure 4: Overview of various architectural and mapping techniques to maximize performance by raising the rooflines and more closely approaching the rooflines. [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: shows the performance of subsequent generations of Google’s TPU. As can be seen, different techniques, such as a wider the DRAM bandwidth, increased parallelism and a higher clock frequency are used to raise the roofline(1). Secondly, hardware support for quantization …
Figure 6
Figure 6. Figure 6: It is, however, important to realize that the selected paralel￾lization strategy strongly impacts the efficiency with which the parallel resources are used. This can be quantified in terms of the utilization of the processing elements. As detailed in [PITH_FULL_IMAGE:…
Figure 7
Figure 7. Figure 7: Spatial data reuse concepts B. Exploiting spatial and temporal data reuse The operands of the MAC operations in neural networks have a potential data reuse factor which can easily go up to 10,000’s. For example, the weights in a Conv layer or the operands of a large Ma…
Figure 10
Figure 10. Figure 10: RISC-V mixed-precision SIMD dot-product unit [PITH_FULL_IMAGE:figures/full_fig_p007_10.png]
Figure 11
Figure 11. Figure 11: (a) Block quantization and dot-product of block-quantized vector [PITH_FULL_IMAGE:figures/full_fig_p008_11.png]
Figure 12
Figure 12. Figure 12: Neureka mixed-precision accelerator architecture design [PITH_FULL_IMAGE:figures/full_fig_p008_12.png]
Figure 13
Figure 13. Figure 13: Nested loop for execution of a 3×3 convolutional layer in Neureka. Note that shift-and-add loop is executed sequentially bit-by-bit for the 8bit weights can be seen as one additional dimension in the nested for-loop notation discussed in section III: this is shown in …
Figure 16
Figure 16. Figure 16: Comparison of traditional, near-, and in-memory architectures [PITH_FULL_IMAGE:figures/full_fig_p010_16.png]
Figure 17
Figure 17. Figure 17: Basics of in-memory computing for MVM operations. [PITH_FULL_IMAGE:figures/full_fig_p011_17.png]
Figure 18
Figure 18. Figure 18: Fundamental dynamic-range trade-off of in-memory computing [55]. [PITH_FULL_IMAGE:figures/full_fig_p012_18.png]
Figure 19
Figure 19. Figure 19: Generalized column computation in in-memory computing. [PITH_FULL_IMAGE:figures/full_fig_p012_19.png]
Figure 21
Figure 21. Figure 21: Switched-capacitor in-memory computing for high-SNR computa [PITH_FULL_IMAGE:figures/full_fig_p013_21.png]
Figure 20
Figure 20. Figure 20: Analysis of digital in-memory computing versus standard digital [PITH_FULL_IMAGE:figures/full_fig_p013_20.png]
Figure 22
Figure 22. Figure 22: Roofline optimizations to different workload arithmetic intensities. [PITH_FULL_IMAGE:figures/full_fig_p014_22.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

81 extracted references · 63 canonical work pages

  1. [55]

    In-memory computing: Advances and prospects,

    N. Verma, H. Jia, H. Valavi, Y . Tang, M. Ozatay, L.-Y . Chen, B. Zhang, and P. Deaville, “In-memory computing: Advances and prospects,”IEEE Solid-State Circuits Magazine , vol. 11, no. 3, pp. 43–55, 2019

  2. [1]

    Visualizing size of large language models,

    G. Anil, “Visualizing size of large language models,” 2023, accessed: 2024-10-10. [Online]. Available: https://medium.com/@georgeanil/ visualizing-size-of-large-language-models-ec576caa5557

  3. [2]

    Trends in deep learning hardware,

    B. Dally, “Trends in deep learning hardware,” 2024, talk, presented by NVIDIA. PREPRINT OF ARTICLE PUBLISHED IN JOURNAL OF SOLID STATE CIRCUITS 17

  4. [3]

    Roofline: an insightful visual performance model for multicore architectures,

    S. Williams, A. Waterman, and D. Patterson, “Roofline: an insightful visual performance model for multicore architectures,” Communications of the ACM , vol. 52, no. 4, pp. 65–76, 2009

  5. [4]

    Eyeriss: A apatial architecture for energy-efficient dataflow for convolutional neural networks,

    Y .-H. Chen, J. Emer, and V . Sze, “Eyeriss: A apatial architecture for energy-efficient dataflow for convolutional neural networks,” in 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), 2016, pp. 367–379

  6. [5]

    Roofline performance analysis of dnn architectures on cpu and gpu systems,

    H. Prashanth and M. Rao, “Roofline performance analysis of dnn architectures on cpu and gpu systems,” in 2024 25th International Symposium on Quality Electronic Design (ISQED) . IEEE, 2024, pp. 1–8

  7. [6]

    S. W. Williams, Book: The roofline model . University of California, 2010

  8. [8]

    Understanding reuse, performance, and hardware cost of dnn dataflow: A data-centric approach,

    H. Kwon, P. Chatarasi, M. Pellauer, A. Parashar, V . Sarkar, and T. Kr- ishna, “Understanding reuse, performance, and hardware cost of dnn dataflow: A data-centric approach,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture , 2019, pp. 754–768

Show all 81 references
  1. [9]

    A roofline model of energy,

    J. W. Choi, D. Bedard, R. Fowler, and R. Vuduc, “A roofline model of energy,” in 2013 IEEE 27th International Symposium on Parallel and Distributed Processing. IEEE, 2013, pp. 661–672

  2. [10]

    Symphony: Orchestrating sparse and dense tensors with hierarchical heterogeneous processing,

    M. Pellauer, J. Clemons, V . Balaji, N. Crago, A. Jaleel, D. Lee, M. O’Connor, A. Parashar, S. Treichler, P.-A. Tsai et al. , “Symphony: Orchestrating sparse and dense tensors with hierarchical heterogeneous processing,” ACM Transactions on Computer Systems , vol. 41, no. 1-4,...

  3. [11]

    Lots of questions on Google’s “Trillium

    T. P. Morgan, “Lots of questions on Google’s “Trillium” TPU v6, a few answers,” Oct 2024. [Online]. Available: https://www.nextplatform.com/2024/06/10/ lots-of-questions-on-googles-trillium-tpu-v6-a-few-answers/

  4. [12]

    Envi- sion: A 0.26-to-10tops/w subword-parallel dynamic-voltage-accuracy- frequency-scalable convolutional neural network processor in 28nm fdsoi,

    B. Moons, R. Uytterhoeven, W. Dehaene, and M. Verhelst, “Envi- sion: A 0.26-to-10tops/w subword-parallel dynamic-voltage-accuracy- frequency-scalable convolutional neural network processor in 28nm fdsoi,” in 2017 IEEE International Solid-State Circuits Conference (ISSCC). IEEE...

  5. [13]

    9.5 a 6k-mac feature-map-sparsity-aware neural processing unit in 5nm flagship mobile soc,

    J.-S. Park, J.-W. Jang, H. Lee, D. Lee, S. Lee, H. Jung, S. Lee, S. Kwon, K. Jeong, J.-H. Song et al. , “9.5 a 6k-mac feature-map-sparsity-aware neural processing unit in 5nm flagship mobile soc,” in 2021 IEEE International Solid-State Circuits Conference (ISSCC) , vol. 64. IE...

  6. [14]

    Compute solution for tesla’s full self-driving computer,

    E. Talpes, D. D. Sarma, G. Venkataramanan, P. Bannon, B. McGee, B. Floering, A. Jalote, C. Hsiong, S. Arora, A. Gorti et al., “Compute solution for tesla’s full self-driving computer,” IEEE Micro , vol. 40, no. 2, pp. 25–35, 2020

  7. [15]

    7.2 a 12nm programmable convolution-efficient neural- processing-unit chip achieving 825tops,

    Y . Jiao, L. Han, R. Jin, Y .-J. Su, C. Ho, L. Yin, Y . Li, L. Chen, Z. Chen, L. Liu et al. , “7.2 a 12nm programmable convolution-efficient neural- processing-unit chip achieving 825tops,” in 2020 IEEE International Solid-State Circuits Conference-(ISSCC) . IEEE, 2020, pp. 136–140

  8. [16]

    Groq rocks neural networks,

    L. Gwennap, “Groq rocks neural networks,” Microprocessor Report, Tech. Rep., jan, 2020

  9. [17]

    9.1 a 7nm 4-core ai chip with 25.6tflops hybrid fp8 training, 102.4tops int4 inference and workload-aware throttling,

    A. Agrawal, S. K. Lee, J. Silberman, M. Ziegler, M. Kang, S. Venkatara- mani, N. Cao, B. Fleischer, M. Guillorn, M. Cohen, S. Mueller, J. Oh, M. Lutz, J. Jung, S. Koswatta, C. Zhou, V . Zalani, J. Bonanno, R. Casat- uta, C.-Y . Chen, J. Choi, H. Haynie, A. Herbert, R. Jain, M....

  10. [18]

    16.7 a 40-310tops/w sram-based all-digital up to 4b in-memory computing multi-tiled nn accelerator in fd-soi 18nm for deep-learning edge applications,

    G. Desoli, N. Chawla, T. Boesch, M. Avodhyawasi, H. Rawat, H. Chawla, V . Abhijith, P. Zambotti, A. Sharma, C. Cappetta, M. Rossi, A. De Vita, and F. Girardi, “16.7 a 40-310tops/w sram-based all-digital up to 4b in-memory computing multi-tiled nn accelerator in fd-soi 18nm for...

  11. [19]

    Charm: Composing heterogeneous accelerators for matrix multiply on versal acap architecture,

    J. Zhuang, J. Lau, H. Ye, Z. Yang, Y . Du, J. Lo, K. Denolf, S. Neuendorffer, A. Jones, J. Hu, D. Chen, J. Cong, and P. Zhou, “Charm: Composing heterogeneous accelerators for matrix multiply on versal acap architecture,” in Proceedings of the 2023 ACM/SIGDA International Sympo...

  12. [20]

    Davinci: A scalable architecture for neural network computing,

    H. Liao, J. Tu, J. Xia, and X. Zhou, “Davinci: A scalable architecture for neural network computing,” in 2019 IEEE Hot Chips 31 Symposium (HCS). IEEE Computer Society, 2019, pp. 1–44

  13. [21]

    Nvidia tensor core programmability, performance & precision,

    S. Markidis, S. W. Der Chien, E. Laure, I. B. Peng, and J. S. Vetter, “Nvidia tensor core programmability, performance & precision,” in 2018 IEEE international parallel and distributed processing symposium workshops (IPDPSW). IEEE, 2018, pp. 522–531

  14. [22]

    A charge domain sram compute-in-memory macro with c-2c ladder- based 8-bit mac unit in 22-nm finfet process for edge inference,

    H. Wang, R. Liu, R. Dorrance, D. Dasalukunte, D. Lake, and B. Carlton, “A charge domain sram compute-in-memory macro with c-2c ladder- based 8-bit mac unit in 22-nm finfet process for edge inference,” IEEE Journal of Solid-State Circuits , vol. 58, no. 4, pp. 1037–1050, 2023

  15. [23]

    A 22 nm, 1540 top/s/w, 12.1 top/s/mm 2 in-memory analog matrix-vector-multiplier for dnn acceleration,

    I. A. Papistas, S. Cosemans, B. Rooseleer, J. Doevenspeck, M.-H. Na, A. Mallik, P. Debacker, and D. Verkest, “A 22 nm, 1540 top/s/w, 12.1 top/s/mm 2 in-memory analog matrix-vector-multiplier for dnn acceleration,” in 2021 IEEE Custom Integrated Circuits Conference (CICC). IEEE...

  16. [24]

    A 64-tile 2.4- mb in-memory-computing cnn accelerator employing charge-domain compute,

    H. Valavi, P. J. Ramadge, E. Nestler, and N. Verma, “A 64-tile 2.4- mb in-memory-computing cnn accelerator employing charge-domain compute,” IEEE Journal of Solid-State Circuits, vol. 54, no. 6, pp. 1789– 1799, 2019

  17. [25]

    Compute Solution for Tesla’s Full Self-Driving Computer,

    E. Talpes, D. D. Sarma, G. Venkataramanan, P. Bannon, B. McGee, B. Floering, A. Jalote, C. Hsiong, S. Arora, A. Gorti, and G. S. Sachdev, “Compute Solution for Tesla’s Full Self-Driving Computer,” IEEE Micro, vol. 40, no. 2, pp. 25–35, 2020

  18. [26]

    Hardware for deep learning,

    B. Dally, “Hardware for deep learning,” in IEEE Hot Chips Symposium (HCS), vol. 35. IEEE, 2023, pp. 1–58

  19. [27]

    Lincoln ai computing survey (laics) update,

    A. Reuther, P. Michaleas, M. Jones, V . Gadepally, S. Samsi, and J. Kepner, “Lincoln ai computing survey (laics) update,” in 2023 IEEE High Performance Extreme Computing Conference (HPEC) . IEEE, 2023, pp. 1–7

  20. [28]

    Neural network accelerator comparison

    K. Guo, W. Li, K. Zhong, Z. Zhu, S. Zeng, T. Xie, S. Han, Y . Xie, P. Debacker, M. Verhelst, and Y . Wang, “Neural network accelerator comparison.” [Online]. Available: https://nicsefc.ee.tsinghua. edu.cn/project.html

  21. [29]

    Llm inference unveiled: Survey and roofline model insights,

    Z. Yuan, Y . Shang, Y . Zhou, Z. Dong, Z. Zhou, C. Xue, B. Wu, Z. Li, Q. Gu, Y . J. Lee et al. , “Llm inference unveiled: Survey and roofline model insights,” arXiv preprint arXiv:2402.16363 , 2024

  22. [30]

    Minifloats on risc-v cores: Isa extensions with mixed- precision short dot products,

    L. Bertaccini, G. Paulin, M. Cavalcante, T. Fischer, S. Mach, and L. Benini, “Minifloats on risc-v cores: Isa extensions with mixed- precision short dot products,” IEEE Transactions on Emerging Topics in Computing, 2024

  23. [31]

    Cutie: Beyond petaop/s/w ternary dnn inference acceleration with better-than-binary energy efficiency,

    M. Scherer, G. Rutishauser, L. Cavigelli, and L. Benini, “Cutie: Beyond petaop/s/w ternary dnn inference acceleration with better-than-binary energy efficiency,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , vol. 41, no. 4, pp. 1020–1033, 2021

  24. [32]

    Binareye: An always-on energy-accuracy-scalable binary cnn processor with all memory on chip in 28nm cmos,

    B. Moons, D. Bankman, L. Yang, B. Murmann, and M. Verhelst, “Binareye: An always-on energy-accuracy-scalable binary cnn processor with all memory on chip in 28nm cmos,” in 2018 IEEE Custom Integrated Circuits Conference (CICC) . IEEE, 2018, pp. 1–4

  25. [33]

    Bitnet: Scaling 1-bit transformers for large language models,

    H. Wang, S. Ma, L. Dong, S. Huang, H. Wang, L. Ma, F. Yang, R. Wang, Y . Wu, and F. Wei, “Bitnet: Scaling 1-bit transformers for large language models,” arXiv preprint arXiv:2310.11453 , 2023

  26. [34]

    A 3 tops/w risc-v parallel cluster for inference of fine-grain mixed-precision quantized neural networks,

    A. Nadalini, G. Rutishauser, A. Burrello, N. Bruschi, A. Garofalo, L. Benini, F. Conti, and D. Rossi, “A 3 tops/w risc-v parallel cluster for inference of fine-grain mixed-precision quantized neural networks,” in 2023 IEEE Computer Society Annual Symposium on VLSI (ISVLSI) . I...

  27. [35]

    Marsellus: A heterogeneous risc-v ai-iot end-node soc with 2–8 b dnn acceleration and 30%-boost adaptive body biasing,

    F. Conti, G. Paulin, A. Garofalo, D. Rossi, A. Di Mauro, G. Rutishauser, G. Ottavi, M. Eggiman, H. Okuhara, and L. Benini, “Marsellus: A heterogeneous risc-v ai-iot end-node soc with 2–8 b dnn acceleration and 30%-boost adaptive body biasing,” IEEE Journal of Solid-State Circu...

  28. [36]

    Microscaling data formats for deep learning,

    B. D. Rouhani, R. Zhao, A. More, M. Hall, A. Khodamoradi, S. Deng, D. Choudhary, M. Cornea, E. Dellinger, K. Denolf et al., “Microscaling data formats for deep learning,” arXiv preprint arXiv:2310.10537, 2023

  29. [37]

    Nvidia blackwell platform: Advancing generative ai and accelerated computing,

    A. Tirumala and R. Wong, “Nvidia blackwell platform: Advancing generative ai and accelerated computing,” in 2024 IEEE Hot Chips 36 Symposium (HCS), 2024, pp. 1–33

  30. [38]

    Siracusa: A 16 nm heterogenous risc-v soc for extended reality with at-mram neural engine,

    A. S. Prasad, M. Scherer, F. Conti, D. Rossi, A. Di Mauro, M. Eggimann, J. T. G ´omez, Z. Li, S. S. Sarwar, Z. Wang et al. , “Siracusa: A 16 nm heterogenous risc-v soc for extended reality with at-mram neural engine,” IEEE Journal of Solid-State Circuits , 2024

  31. [39]

    Onyx: A 12nm 756 gops/w coarse-grained reconfigurable array for accelerating dense and sparse applications,

    K. Koul, M. Strange, J. Melchert, A. Carsello, Y . Mei, O. Hsu, T. Kong, P.-H. Chen, H. Ke, K. Zhang et al. , “Onyx: A 12nm 756 gops/w coarse-grained reconfigurable array for accelerating dense and sparse applications,” in 2024 IEEE Symposium on VLSI Technology and Circuits (V...

  32. [40]

    Learning n: m fine-grained structured sparse neural networks from scratch,

    A. Zhou, Y . Ma, J. Zhu, J. Liu, Z. Zhang, K. Yuan, W. Sun, and H. Li, “Learning n: m fine-grained structured sparse neural networks from scratch,” arXiv preprint arXiv:2102.04010 , 2021

  33. [41]

    3.2 the a100 datacenter gpu and ampere architecture,

    J. Choquette, E. Lee, R. Krashinsky, V . Balan, and B. Khailany, “3.2 the a100 datacenter gpu and ampere architecture,” in 2021 IEEE International Solid-State Circuits Conference (ISSCC) , vol. 64. IEEE, 2021, pp. 48–50

  34. [42]

    Venom: A vectorized n: M format for unleashing the power of sparse tensor cores,

    R. L. Castro, A. Ivanov, D. Andrade, T. Ben-Nun, B. B. Fraguela, and T. Hoefler, “Venom: A vectorized n: M format for unleashing the power of sparse tensor cores,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis ,...

  35. [43]

    Occamy: A 432-core dual-chiplet dual-hbm2e 768-dp-gflop/s risc-v system for 8- to-64-bit dense and sparse computing in 12-nm finfet,

    P. Scheffler, T. Benz, V . Potocnik, T. Fischer, L. Colagrande, N. Wistoff, Y . Zhang, L. Bertaccini, G. Ottavi, M. Eggimann, M. Cavalcante, G. Paulin, F. K. G ¨urkaynak, D. Rossi, and L. Benini, “Occamy: A 432-core dual-chiplet dual-hbm2e 768-dp-gflop/s risc-v system for 8- t...

  36. [44]

    Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing,

    G. Heo, S. Lee, J. Cho, H. Choi, S. Lee, H. Ham, G. Kim, D. Mahajan, and J. Park, “Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating...

  37. [45]

    Inclusive-pim: Hardware-software co-design for broad acceleration on commercial pim architectures,

    J. Alsop, S. Aga, M. Ibrahim, M. Islam, A. Mccrabb, and N. Jayasena, “Inclusive-pim: Hardware-software co-design for broad acceleration on commercial pim architectures,” 2024. [Online]. Available: https://arxiv.org/abs/2309.07984

  38. [46]

    In-memory computation of a machine-learning classifier in a standard 6t sram array,

    J. Zhang, Z. Wang, and N. Verma, “In-memory computation of a machine-learning classifier in a standard 6t sram array,” IEEE Journal of Solid-State Circuits , vol. 52, no. 4, pp. 915–924, 2017

  39. [47]

    An energy-efficient memory-based high-throughput vlsi architecture for convolutional networks,

    M. Kang, S. K. Gonugondla, M.-S. Keel, and N. R. Shanbhag, “An energy-efficient memory-based high-throughput vlsi architecture for convolutional networks,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 1037– 1041

  40. [48]

    Fast, energy-efficient, robust, and reproducible mixed-signal neuromorphic classifier based on embedded nor flash memory technology,

    X. Guo, F. M. Bayat, M. Bavandpour, M. Klachko, M. R. Mahmoodi, M. Prezioso, K. K. Likharev, and D. B. Strukov, “Fast, energy-efficient, robust, and reproducible mixed-signal neuromorphic classifier based on embedded nor flash memory technology,” in 2017 IEEE International Ele...

  41. [49]

    Analog in-memory subthreshold deep neural network accelerator,

    L. Fick, D. Blaauw, D. Sylvester, S. Skrzyniarz, M. Parikh, and D. Fick, “Analog in-memory subthreshold deep neural network accelerator,” in 2017 IEEE Custom Integrated Circuits Conference (CICC) , 2017, pp. 1–4

  42. [50]

    A 5-nm 254-tops/w 221-tops/mm2 fully-digital computing-in-memory macro supporting wide-range dynamic-voltage- frequency scaling and simultaneous mac and write operations,

    H. Fujiwara, H. Mori, W.-C. Zhao, M.-C. Chuang, R. Naous, C.-K. Chuang, T. Hashizume, D. Sun, C.-F. Lee, K. Akarvardar, S. Adham, T.- L. Chou, M. E. Sinangil, Y . Wang, Y .-D. Chih, Y .-H. Chen, H.-J. Liao, and T.-Y . J. Chang, “A 5-nm 254-tops/w 221-tops/mm2 fully-digital com...

  43. [51]

    16.4 an 89tops/w and 16.3tops/mm2 all-digital sram-based full-precision compute-in memory macro in 22nm for machine-learning edge applications,

    Y .-D. Chih, P.-H. Lee, H. Fujiwara, Y .-C. Shih, C.-F. Lee, R. Naous, Y .-L. Chen, C.-P. Lo, C.-H. Lu, H. Mori, W.-C. Zhao, D. Sun, M. E. Sinangil, Y .-H. Chen, T.-L. Chou, K. Akarvardar, H.-J. Liao, Y . Wang, M.-F. Chang, and T.-Y . J. Chang, “16.4 an 89tops/w and 16.3tops/m...

  44. [52]

    A maximally row- parallel mram in-memory-computing macro addressing readout circuit sensitivity and area,

    P. Deaville, B. Zhang, L.-Y . Chen, and N. Verma, “A maximally row- parallel mram in-memory-computing macro addressing readout circuit sensitivity and area,” in ESSCIRC 2021 - IEEE 47th European Solid State Circuits Conference (ESSCIRC) , 2021, pp. 75–78

  45. [53]

    A programmable heterogeneous microprocessor based on bit-scalable in-memory comput- ing,

    H. Jia, H. Valavi, Y . Tang, J. Zhang, and N. Verma, “A programmable heterogeneous microprocessor based on bit-scalable in-memory comput- ing,” IEEE Journal of Solid-State Circuits, vol. 55, no. 9, pp. 2609–2621, 2020

  46. [54]

    A crossbar array of magnetoresistive memory devices for in-memory computing,

    S. Jung, H. Lee, S. Myung, H. Kim, S. K. Yoon, S.-W. Kwon, Y . Ju, M. Kim, W. Yi, S. Han, B. Kwon, B. Seo, K. Lee, G.-H. Koh, K. Lee, Y . Song, C. Choi, D. Ham, and S. J. Kim, “A crossbar array of magnetoresistive memory devices for in-memory computing,” Nature, vol. 601, no. ...

  47. [56]

    14.2 a compute sram with bit-serial integer/floating-point operations for programmable in-memory vector acceleration,

    J. Wang, X. Wang, C. Eckert, A. Subramaniyan, R. Das, D. Blaauw, and D. Sylvester, “14.2 a compute sram with bit-serial integer/floating-point operations for programmable in-memory vector acceleration,” in 2019 IEEE International Solid-State Circuits Conference - (ISSCC), 2019...

  48. [57]

    A 40nm 64kb 26.56tops/w 2.37mb/mm2rram binary/compute-in-memory macro with 4.23x im- provement in density and > 75% use of sensing dynamic range,

    S. D. Spetalnick, M. Chang, B. Crafton, W.-S. Khwa, Y .-D. Chih, M.-F. Chang, and A. Raychowdhury, “A 40nm 64kb 26.56tops/w 2.37mb/mm2rram binary/compute-in-memory macro with 4.23x im- provement in density and > 75% use of sensing dynamic range,” in2022 IEEE International Soli...

  49. [58]

    Funda- mental limits on energy-delay-accuracy of in-memory architectures in inference applications,

    S. K. Gonugondla, C. Sakr, H. Dbouk, and N. R. Shanbhag, “Funda- mental limits on energy-delay-accuracy of in-memory architectures in inference applications,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 41, no. 10, pp. 3188–3201, 2022

  50. [59]

    11.3 metis aipu: A 12nm 15tops/w 209.6tops soc for cost- and energy-efficient inference at the edge,

    P. A. Hager, B. Moons, S. Cosemans, I. A. Papistas, B. Rooseleer, J. V . Loon, R. Uytterhoeven, F. Zaruba, S. Koumousi, M. Stanisavljevic, S. Mach, S. Mutsaards, R. K. Aljameh, G. H. Khov, B. Machiels, C. Olar, A. Psarras, S. Geursen, J. Vermeeren, Y . Lu, A. Maringanti, D. Am...

  51. [60]

    Benchmarking in-memory computing architectures,

    N. R. Shanbhag and S. K. Roy, “Benchmarking in-memory computing architectures,” IEEE Open Journal of the Solid-State Circuits Society , vol. 2, pp. 288–300, 2022

  52. [61]

    A 22nm 128-kb mram row/column-parallel in-memory computing macro with memory- resistance boosting and multi-column adc readout,

    P. Deaville, B. Zhang, and N. Verma, “A 22nm 128-kb mram row/column-parallel in-memory computing macro with memory- resistance boosting and multi-column adc readout,” in 2022 IEEE Symposium on VLSI Technology and Circuits (VLSI Technology and Circuits), 2022, pp. 268–269

  53. [62]

    A 64-core mixed-signal in-memory compute chip based on phase-change memory for deep neural network inference,

    M. Le Gallo, R. Khaddam-Aljameh, M. Stanisavljevic, A. Vasilopoulos, B. Kersting, M. Dazzi, G. Karunaratne, M. Br ¨andli, A. Singh, S. M. M ¨uller, J. B ¨uchel, X. Timoneda, V . Joshi, M. J. Rasch, U. Egger, A. Garofalo, A. Petropoulos, T. Antonakopoulos, K. Brew, S. Choi, I. ...

  54. [63]

    An n40 256k×44 embedded rram macro with sl-precharge sa and low-voltage current limiter to improve read and write performance,

    C.-C. Chou, Z.-J. Lin, P.-L. Tseng, C.-F. Li, C.-Y . Chang, W.-C. Chen, Y .-D. Chih, and T.-Y . J. Chang, “An n40 256k×44 embedded rram macro with sl-precharge sa and low-voltage current limiter to improve read and write performance,” in 2018 IEEE International Solid-State Cir...

  55. [64]

    Cmos- embedded stt-mram arrays in 2x nm nodes for gp-mcu applications,

    D. Shum, D. Houssameddine, S. T. Woo, Y . S. You, J. Wong, K. W. Wong, C. C. Wang, K. H. Lee, K. Yamane, V . B. Naik, C. S. Seet, T. Tahmasebi, C. Hai, H. W. Yang, N. Thiyagarajah, R. Chao, J. W. Ting, N. L. Chung, T. Ling, T. H. Chan, S. Y . Siah, R. Nair, S. Deshpande, R. Wh...

  56. [65]

    A switched-capacitor sram in-memory computing macro with high-precision, high-efficiency differential archi- tecture,

    J. Lee, B. Zhang, and N. Verma, “A switched-capacitor sram in-memory computing macro with high-precision, high-efficiency differential archi- tecture,” in 2024 European Conference on Solid-State Circuits , 2024

  57. [66]

    Scalable and Programmable Neural Network Inference Accelerator Based on In-Memory Computing,

    H. Jia, M. Ozatay, Y . Tang, H. Valavi, R. Pathak, J. Lee, and N. Verma, “Scalable and Programmable Neural Network Inference Accelerator Based on In-Memory Computing,” IEEE Journal of Solid-State Circuits, vol. 57, no. 1, pp. 198–211, 2022

  58. [67]

    Interstellar: Using Halide’s Scheduling Language to Analyze DNN Accelerators,

    X. Yang, M. Gao, Q. Liu, J. Setter, J. Pu, A. Nayak, S. Bell, K. Cao, H. Ha, P. Raina, C. Kozyrakis, and M. Horowitz, “Interstellar: Using Halide’s Scheduling Language to Analyze DNN Accelerators,” in Proceedings of the Twenty-Fifth International Conference on Architectural Su...

  59. [68]

    MAESTRO: A Data-Centric Approach to Understand Reuse, Performance, and Hardware Cost of DNN Mappings,

    H. Kwon, P. Chatarasi, V . Sarkar, T. Krishna, M. Pellauer, and A. Parashar, “MAESTRO: A Data-Centric Approach to Understand Reuse, Performance, and Hardware Cost of DNN Mappings,” IEEE Micro, vol. 40, no. 3, pp. 20–29, 2020

  60. [69]

    Timeloop: A Systematic Approach to DNN Accelerator Evaluation,

    A. Parashar, P. Raina, Y . S. Shao, Y .-H. Chen, V . A. Ying, A. Mukkara, R. Venkatesan, B. Khailany, S. W. Keckler, and J. Emer, “Timeloop: A Systematic Approach to DNN Accelerator Evaluation,” in 2019 IEEE International Symposium on Performance Analysis of Systems and Softwa...

  61. [70]

    ZigZag: Enlarging Joint Architecture-Mapping Design Space Exploration for DNN Accelerators,

    L. Mei, P. Houshmand, V . Jain, S. Giraldo, and M. Verhelst, “ZigZag: Enlarging Joint Architecture-Mapping Design Space Exploration for DNN Accelerators,” IEEE Transactions on Computers , vol. 70, no. 8, pp. 1160–1174, 2021

  62. [71]

    CoSA: Scheduling by constrained op- timization for spatial accelerators,

    Q. Huang, A. Kalaiah, M. Kang, J. Demmel, G. Dinh, J. Wawrzynek, T. Norell, and Y . S. Shao, “CoSA: Scheduling by constrained op- timization for spatial accelerators,” in 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) , 2021, pp. 554–566

  63. [72]

    Mind Mappings: Enabling Efficient Algorithm-Accelerator Mapping Space Search,

    K. Hegde, P.-A. Tsai, S. Huang, V . Chandra, A. Parashar, and C. W. Fletcher, “Mind Mappings: Enabling Efficient Algorithm-Accelerator Mapping Space Search,” in Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operatin...

  64. [73]

    GAMMA: Automating the HW Mapping of DNN Models on Accelerators via Genetic Algorithm,

    S.-C. Kao and T. Krishna, “GAMMA: Automating the HW Mapping of DNN Models on Accelerators via Genetic Algorithm,” in Proceedings of the 39th International Conference on Computer-Aided Design , ser. ICCAD ’20. New York, NY , USA: Association for Computing Machinery, 2020. [Onli...

  65. [74]

    Stream: Design space exploration of layer-fused dnns on hetero- geneous dataflow accelerators,

    A. Symons, L. Mei, S. Colleman, P. Houshmand, S. Karl, and M. Ver- helst, “Stream: Design space exploration of layer-fused dnns on hetero- geneous dataflow accelerators,” IEEE Transactions on Computers, 2024

  66. [75]

    The groq software-defined scale-out tensor streaming multiprocessor : From chips-to-systems architectural overview,

    D. Abts, J. Kim, G. Kimmell, M. Boyd, K. Kang, S. Parmar, A. Ling, A. Bitar, I. Ahmed, and J. Ross, “The groq software-defined scale-out tensor streaming multiprocessor : From chips-to-systems architectural overview,” in 2022 IEEE Hot Chips 34 Symposium (HCS) , 2022, pp. 1–69

  67. [76]

    Application specific instruction processor based implementation of a gnss receiver on an fpga,

    K. G ¨otz and T. Noll, “Application specific instruction processor based implementation of a gnss receiver on an fpga,” in Design & Test in Europe Conference, 2006, pp. 58–63

  68. [77]

    How flexible is your com- puting system?

    S. Huang, L. Waeijen, and H. Corporaal, “How flexible is your com- puting system?” ACM Transactions on Embedded Computing Systems (TECS), vol. 21, no. 4, pp. 1–41, 2022

  69. [78]

    Tandem processor: Grappling with emerging operators in neural networks,

    S. Ghodrati, S. Kinzer, H. Xu, R. Mahapatra, Y . Kim, B. H. Ahn, D. K. Wang, L. Karthikeyan, A. Yazdanbakhsh, J. Park et al., “Tandem processor: Grappling with emerging operators in neural networks,” in Proceedings of the 29th ACM International Conference on Architectural Supp...

  70. [79]

    Mec: memory-efficient convolution for deep neural network,

    M. M. Cho and D. Brand, “Mec: memory-efficient convolution for deep neural network,” in ICML-34: Proceedings of the International Conference on Machine Learning - Volume 70 . JMLR.org, 2017, p. 815–824

  71. [80]

    A formalism of dnn accelerator flexibility,

    S.-C. Kao, H. Kwon, M. Pellauer, A. Parashar, and T. Krishna, “A formalism of dnn accelerator flexibility,” Proceedings of the ACM on Measurement and Analysis of Computing Systems , vol. 6, no. 2, pp. 1–23, 2022

  72. [81]

    Mlir: Scaling compiler infrastructure for domain specific computation,

    C. Lattner, M. Amini, U. Bondhugula, A. Cohen, A. Davis, J. Pien- aar, R. Riddle, T. Shpeisman, N. Vasilache, and O. Zinenko, “Mlir: Scaling compiler infrastructure for domain specific computation,” in 2021 IEEE/ACM International Symposium on Code Generation and Optimization (...

  73. [82]

    The hardware lottery,

    S. Hooker, “The hardware lottery,” Communications of the ACM, vol. 64, no. 12, pp. 58–65, 2021. Marian Verhelst Marian Verhelst is a professor at the MICAS labs of KU Leuven and a research director at imec. Her research focuses on embedded machine learning, hardware accelerato...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.