Pith. sign in

REVIEW 3 major objections 6 minor 234 references

Memory-Centric Computing: Recent Advances in Processing-in-DRAM

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Processing-in-DRAM turns DRAM into a compute substrate, with unmodified chips performing Boolean logic at 94–99% success.

desk verdict A clear, well-illustrated survey of the authors' own Processing-in-DRAM work, but the COTS 'high reliability' claim outruns the data: 95% per-operation success is undependable without an error model. read the letter →

arxiv 2412.19275 v1 pith:HL5XMXY6 submitted 2024-12-26 cs.AR cs.DC

classification cs.ARcs.DC
keywords processing-in-DRAMmemory-centriccomputingbulkbitwiseoperationsDRAMtimingviolationsimultaneousrowactivationMIMDRAMSectoredin-memorycomputation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that DRAM, the working memory of nearly every computer, can be turned into a place where computation happens rather than merely data storage, and that doing so removes the expensive movement of data between processor and memory. It reviews recent advances in Processing-in-DRAM in three parts: hardware/software support (MIMDRAM) that lets different memory arrays execute different bit-serial instructions at fine grain; experimental evidence that unmodified commercial DRAM chips can perform NOT, NAND, NOR, AND, OR, and multi-row copy by deliberately violating timing parameters, with per-operation success rates above 94%; and a DRAM design (Sectored DRAM) that activates only part of a row to save energy. If these results hold, memory-centric designs could substantially reduce data movement overheads in memory-bound workloads, strengthening the case for treating memory as a combined storage-and-computation substrate.

What carries the argument

The load-bearing mechanism is charge sharing on DRAM bitlines during simultaneous row activation. In an Ambit-style triple-row activation, three cells connected to one bitline share charge and a sense amplifier resolves the majority, yielding AND and OR; in the COTS experiments, violating tRP lets a second activate interrupt precharge so the same charge-sharing principle performs NOT, NAND/NOR, and Multi-RowCopy on unmodified chips. MIMDRAM's additional machinery is mat-level isolation: mat isolation transistors, row decoder latches, and mat selectors let different mats in a subarray run different instructions, with inter-mat and intra-mat interconnects moving data between mats. Sectored DRAM's machinery is wordline segmentation with sector latches and transistors to activate a chosen subset of mats.

What would settle it

Run the NOT, AND/NAND, and Multi-RowCopy command sequences from Section V on a large sample of commercial DRAM chips, comparing every result against a known-good pattern across extended temperature, voltage, and chip-to-chip variation; if the measured success rate falls below what the paper claims or below a level usable without correction, the claim of reliable unmodified-chip computation is falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that Processing-in-DRAM has moved from proposal to demonstrated capability. First, MIMDRAM shows a hardware/software co-designed system in which each DRAM mat within a subarray can independently execute different bit-serial SIMD operations; evaluating twelve real applications and 495 multi-programmed mixes, it reports 13.2x/0.22x/173x the performance of CPU/GPU/SIMDRAM baselines, and 15.6x the SIMD utilization of SIMDRAM. Second, experiments on 224 and 120 commercial off-the-shelf DRAM chips show that by issuing back-to-back ACT/PRE commands with tRP below the manufacturer specification (e.g., <3ns), a memory controller can make the chips perform functionally complete Boolean operations: NOT, NAND, NOR, AND, OR, and one-row-to-31-row copy, with average success rates of 94.94% to 99.98%. Third, Sectored DRAM splits wordlines so a single activate touches only a subset of mats, cutting DRAM energy by up to 33% and improving performance by up to 36% for data-intensive workloads.

Load-bearing premise

The central claim of dependable computation on unmodified DRAM chips rests on the assumption that deliberately violating the manufacturer's timing parameters produces correct results reliably enough for real workloads; the paper reports 94.94% to 99.98% success rates and does not propose an error-detection or correction mechanism for the failed operations.

Editorial extensions

If this is right

  • If MIMDRAM's results stand, bulk bitwise operations in DRAM can be programmed at the granularity of vectorizing compilers, making in-memory computation usable for real workloads without hand-tuning.
  • If COTS DRAM chips truly perform NOT/NAND/NOR at above 94% success under timing violation, then a memory controller with no DRAM modification can implement functionally complete Boolean logic on existing hardware.
  • The demonstrated Multi-RowCopy, one row copied concurrently into up to 31 rows, implies that bulk data copy and initialization can be accelerated inside DRAM, and may also offer a defense against cold-boot attacks.
  • Sectored DRAM's fine-grained activation implies that the energy cost of over-fetching can be cut by roughly a fifth to a third in data-intensive workloads, which directly improves the efficiency of Processing-in-DRAM systems that rely on row activation.
  • Together, the three advances imply that the memory controller and DRAM interface should be redesigned to expose row-level and mat-level computation commands rather than treating DRAM solely as a byte-addressable store.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not specify an error-detection or correction mechanism for COTS DRAM computation; one implicit consequence is that practical use would need to pair these operations with ECC or re-execution, which would reduce the net gains.
  • The COTS success rates are averages across tested chips; the paper leaves open whether a production system could tolerate the worst-case chips, so the strongest defensible version of the claim is that the capability exists, not that it is already dependable at scale.
  • A natural extension is to combine the two approaches: use MIMDRAM-style mat-level control with COTS-style timing-violation operations to build a hybrid that keeps the programmability of the former and the zero-modification cost of the latter.
  • If the fine-grained activation of Sectored DRAM is combined with in-memory compute, the energy savings from avoiding over-fetch would compound the savings from avoiding data movement, a combination the paper does not quantify.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This invited overview paper summarizes recent Processing-in-DRAM (PiD) advances from the authors' group. It describes MIMDRAM, a hardware/software co-designed Processing-using-DRAM system that enables fine-grained, multiple-instruction multiple-data execution within DRAM mats; experimental results on commercial off-the-shelf (COTS) DRAM chips showing NOT, AND/NAND/OR/NOR, many-row activation, Multi-RowCopy, and true random number generation; and Sectored DRAM, a fine-grained DRAM architecture. The paper reports large performance and energy gains for MIMDRAM (13.2x/0.22x/173x performance versus CPU/GPU/SIMDRAM), high success rates (>94%) for COTS DRAM bulk-bitwise operations, and up to 33% DRAM energy savings for Sectored DRAM.

Significance. If the claims hold, the paper provides useful evidence that Processing-in-DRAM can reduce data-movement overheads: MIMDRAM's compiler/runtime support addresses programmability, the COTS experiments demonstrate functional completeness on unmodified chips, and Sectored DRAM targets access-granularity inefficiencies. The experimental breadth is a strength: the COTS results are based on hundreds of real chips, the MIMDRAM evaluation uses gem5 and CACTI against external baselines, and the paper explicitly identifies interconnect throughput as a limiter. However, the paper is a condensed summary of prior peer-reviewed works rather than a self-contained technical contribution, and the practical-usability claim for COTS DRAM computation is not fully supported by the reported success-rate statistics.

major comments (3)
  1. [Section V, Figs. 8-11] The paper states that COTS DRAM chips perform NOT, AND/NAND/OR/NOR, and Multi-RowCopy operations "at high success rates (>94%)" and describes these as "functionally-complete bulk-bitwise Boolean operations." The reported success rates are per-operation, not per-bit: for example, the average success rate for 16-input AND/NAND is 94.94%. A computation composed of k such operations has expected success roughly 0.95^k, so a 10-gate expression would be wrong in more than 40% of runs without correction. The paper does not report per-bit error rates, error correlation across rows/data patterns/temperature/voltage, whether errors are deterministic and profileable or random, or any detection/correction mechanism. The evidence supports a feasibility demonstration, but the claim that unmodified COTS DRAM can serve as a dependable substrate for real workloads is not supported. Please add an explicit error model, or restrict the claim to "feasibility" and discuss error mitigation (e.g., verification, recomputation, ECC).
  2. [Section IV, Fig. 6] The headline MIMDRAM results are compared against CPU, GPU, and SIMDRAM baselines, but the evaluation methodology for the CPU and GPU baselines is not described. The text only says that MIMDRAM and SIMDRAM are implemented in gem5 and evaluated with CACTI, and that 12 memory-bound applications were selected from five benchmark suites. It is unclear whether the CPU and GPU results come from real hardware, full-system simulation, or vendor simulators, and whether the 12 applications are identical across all four target systems. Without this information, the 13.2x/0.22x/173x performance and 0.0017x/0.00007x/0.004x energy ratios cannot be independently interpreted. Please specify the baseline configurations, simulation methodology, and whether the reported ratios are end-to-end or kernel-only.
  3. [Section VI] The Sectored DRAM section makes strong quantitative claims: "reduces the DRAM energy consumption of data intensive workloads by up to 33% (20% on average), while improving performance by up to 36% (17% on average)" and "system-wide energy savings of up to 23%." However, the section provides no evaluation setup: no workload list, no baseline DRAM configuration, no simulator description, and no sensitivity analysis. The reader cannot verify these numbers from this paper. Please add a concise methodology summary, or state explicitly that the results are taken from [222] and summarize the setup used there.
minor comments (6)
  1. [Section V, Fig. 11 caption] The caption says the results were "measured in 224, 224, and 120 COTS DRAM chips" without spelling out which operation corresponds to which count; please clarify the mapping (e.g., NOT: 224, AND/NAND/OR/NOR: 224, Multi-RowCopy: 120).
  2. [Section V, Fig. 8 and surrounding text] The phrase "with violated manufacturer-recommended tRP timing, e.g., <3ns" is awkward; it should read "with tRP set below the manufacturer-recommended value (e.g., <3ns)."
  3. [Section IV, Fig. 6] The notation "13.2×/0.22×/173× the performance" is ambiguous because the baseline order is only given later; spell out "relative to CPU/GPU/SIMDRAM" in the caption or text.
  4. [Section V, Fig. 12] The figure describes a TRNG command sequence with ACT R0 and ACT R3 but says this "activates four rows simultaneously"; please explain how four rows are activated, since the text otherwise implies a single-row ACT activates only one row.
  5. [References] The reference list contains several typos, e.g., [5] "V olos" should be "Volos," [64] "opthers" should be "others," and [138] "Boroum" should be "Boroumand." A careful proofread is needed.
  6. [Section III, first paragraph] The sentence "First, enabling widespread use of PiM on a wide variety of important workloads requires PiM systems to i) be easy to program and seamlessly compile workloads into" is grammatically incomplete; please revise.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a review that reports experimental and simulation results from cited prior works, with no fitted parameter or definitional reduction.

full rationale

The paper is an invited review summarizing the authors' prior peer-reviewed work on Processing-in-DRAM. Its central quantitative claims are measurements or simulations taken from the cited original papers: MIMDRAM speedups are obtained from gem5 and CACTI evaluations described in [219]; COTS DRAM operation success rates are experimental characterizations reported in [220,221,223]; Sectored DRAM energy/performance numbers are simulation results from [222]. None of these numbers is derived from a parameter fitted in this paper, and none is defined in terms of the conclusion it supports. The paper does not invoke a uniqueness theorem, does not smuggle an ansatz through a citation, and does not rename an empirical pattern as a new derivation. The only notable concern is that the COTS DRAM >94% success-rate claim lacks a per-bit error model and error-correction mechanism, which makes the practical reliability conclusion under-supported; however, that is a correctness/completeness issue, not circularity. Heavy self-citation is present, but per the review rules self-citation is not circularity when, as here, the cited works are independent experimental and simulation studies with external baselines (CPU, GPU, SIMDRAM) and the claims do not reduce to the survey's own definitions.

Assumptions & free parameters 0 free parameters · 3 assumptions · 2 invented entities

This review introduces no new natural entities or fitted parameters. The hardware structures listed are proposed designs evaluated in simulation; they have no independent fabricated evidence. The central assumptions are that timing-violation DRAM operations are reliable and that the simulation models reflect real hardware.

assumptions (3)
  • domain assumption DRAM charge-sharing and sense-amplifier behavior can be reliably exploited by violating timing parameters
    The COTS results (Section V) rely on the assumption that many-row activation leads to deterministic majority/AND/OR behavior across chips.
  • domain assumption Simulation models (gem5, CACTI) accurately capture MIMDRAM and Sectored DRAM behavior
    MIMDRAM and Sectored DRAM evaluations (Sections IV and VI) use simulation; their speedup and energy numbers depend on model fidelity.
  • domain assumption The 12 selected applications are representative of data-intensive workloads
    The authors selected 12 applications from 117 based on memory-bound and auto-vectorizable criteria, which may favor the evaluated PiD mechanisms.
invented entities (2)
  • MIMDRAM hardware modifications (mat isolation transistor, row decoder latch, mat selector, inter-mat and intra-mat interconnects)
    purpose: Enable fine-grained processing-using-DRAM with multiple-instruction multiple-data execution
    These structures are proposed and evaluated only through gem5/CACTI simulation in the cited work; no fabricated chip is demonstrated.
  • Sectored DRAM sector latches, sector transistors, and additional local wordline drivers
    purpose: Enable fine-grained DRAM row activation and data transfer to reduce wasted energy
    The design is simulated, not fabricated; independent silicon evidence is absent.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Memory-Centric Computing: Recent Advances in Processing-in-DRAM." pith.science (2026). https://pith.science/paper/HL5XMXY6

@misc{pith2026241219275,
  author       = {Pith},
  title        = {Pith review of: Memory-Centric Computing: Recent Advances in Processing-in-DRAM},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HL5XMXY6}},
  note         = {Machine review of arXiv:2412.19275}
}
read the original abstract

Memory-centric computing aims to enable computation capability in and near all places where data is generated and stored. As such, it can greatly reduce the large negative performance and energy impact of data access and data movement, by 1) fundamentally avoiding data movement, 2) reducing data access latency & energy, and 3) exploiting large parallelism of memory arrays. Many recent studies show that memory-centric computing can largely improve system performance & energy efficiency. Major industrial vendors and startup companies have recently introduced memory chips with sophisticated computation capabilities. Going forward, both hardware and software stack should be revisited and designed carefully to take advantage of memory-centric computing. This work describes several major recent advances in memory-centric computing, specifically in Processing-in-DRAM, a paradigm where the operational characteristics of a DRAM chip are exploited and enhanced to perform computation on data stored in DRAM. Specifically, we describe 1) new techniques that slightly modify DRAM chips to enable both enhanced computation capability and easier programmability, 2) new experimental studies that demonstrate the functionally-complete bulk-bitwise computational capability of real commercial off-the-shelf DRAM chips, without any modifications to the DRAM chip or the interface, and 3) new DRAM designs that improve access granularity & efficiency, unleashing the true potential of Processing-in-DRAM.

Figures

Figures reproduced from arXiv: 2412.19275 by the authors.

Figure 1
Figure 1. An example of performing the MAJority-of-three operation (i.e., MAJ3(A,B,C)) (a) and the NOT operation (i.e., dst=NOT(src)) in Ambit [122]. In (a), we focus on DRAM cell and sense amplifier operations ( 0 ). Initially, cells A, B, C, and bitline have voltage levels of GND, VDD, VDD, and VDD/2, respectively ( 1 ). We first perform a triple-row activation (TRA) to simultaneously activate cells A, B, and C ( 2 ). When … view at source ↗
Figure 5
Figure 5. An example of a PuD vector reduction, i.e., out+=(A[i]+B[i]), in MIMDRAM [219]. For illustration purposes, we assume that DRAM has only two mats, and the 10 1-bit data elements of the input arrays A and B are evenly distributed across the two DRAM mats. MIMDRAM executes a vector reduction in three steps. In the first step, MIMDRAM executes a PuD addition operation over the data in the two DRAM mats ( 1 ), storing th… view at source ↗
Figure 6
Figure 6. CPU-normalized performance (a), energy (b), and energy efficiency (performance/Watt) (c) results for processor-centric (i.e., Intel Skylake CPU [225] and NVIDIA A100 GPU [226]) and memory-centric (i.e., SIMDRAM [185] and MIMDRAM [219]) architectures executing 12 real-world applications. We implement MIMDRAM and SIMDRAM using gem5 and evaluate their energy consumption using CACTI. We analyze 117 applications from SPE… view at source ↗
Figures from the paper (7 more)
Figure 7
Figure 7. Figure 7 [PITH_FULL_IMAGE:figures/full_fig_p003_7.png]
Figure 9
Figure 9. Figure 9: Command sequence for performing the two-input AND and NAND operations (i.e., AND(X, Y) and NAND(X, Y)) in COTS DRAM chips and the state of cells during each related step. The memory controller issues each command (shown in orange boxes below the time axis) at the corre…
Figure 10
Figure 10. Figure 10: Command sequence for performing the Multi-RowCopy operation (i.e., copying src row to N other dst rows simultaneously) in COTS DRAM chips and the state of cells during each related step. The memory controller issues each command (shown in orange boxes below the time a…
Figure 8
Figure 8. Figure 8: Command sequence for performing the NOT operation (dst = NOT(src)) in COTS DRAM chips and the state of cells during each related step. The memory controller issues each command (shown in orange boxes below the time axis) at the corresponding tick mark on the time axis …
Figure 11
Figure 11. Figure 11: Success rates of the NOT operation with varying numbers of destination rows (a) AND, NAND, OR, and NOR operations with varying numbers of input operands (b) the Multi-RowCopy operation with varying numbers of destination rows (c), as measured in 224, 224, and 120 COTS…
Figure 12
Figure 12. Figure 12: Command sequence for true random number generation in COTS DRAM chips and the state of cells during each related step. The memory controller issues each command (shown in orange boxes below the time axis) at the corresponding tick mark and asserted signals are highlig…
Figure 14
Figure 14. Figure 14: A DRAM module (a), a DRAM chip with multiple banks (b), baseline DRAM bank organization with multiple subarrays (c), and a Sectored DRAM subarray (d) [222]. The global row decoder ( 1 ) enables a global wordline ( 2 ) based on the higher-order bits of a DRAM row addre…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

234 extracted references · 75 canonical work pages

  1. [222]

    Sectored DRAM: An Energy-Efficient High-Throughput and Practical Fine-Grained DRAM Architecture,

    A. Olgun, F. Bostanci, G. F. Oliveira, Y . C. Tugrul, R. Bera, A. G. Yaglikci, H. Hassan, O. Ergin, and O. Mutlu, “Sectored DRAM: An Energy-Efficient High-Throughput and Practical Fine-Grained DRAM Architecture,” TACO, 2024

  2. [1]

    Memory Scaling: A Systems Architecture Perspective,

    O. Mutlu, “Memory Scaling: A Systems Architecture Perspective,” in IMW, 2013

  3. [2]

    Research Problems and Opportunities in Memory Systems,

    O. Mutlu and L. Subramanian, “Research Problems and Opportunities in Memory Systems,” SUPERFRI, 2014

  4. [3]

    The Tail at Scale,

    J. Dean and L. A. Barroso, “The Tail at Scale,” CACM, 2013

  5. [4]

    Profiling a Warehouse-Scale Computer,

    S. Kanev, J. P. Darago, K. Hazelwood, P. Ranganathan, T. Moseley, G.-Y . Wei, and D. Brooks, “Profiling a Warehouse-Scale Computer,” in ISCA, 2015

  6. [5]

    Clearing the Clouds: A Study of Emerging Scale-Out Workloads on Modern Hardware,

    M. Ferdman, A. Adileh, O. Kocberber, S. V olos, M. Alisafaee, D. Jevdjic, C. Kaynak, A. D. Popescu, A. Ailamaki, and B. Falsafi, “Clearing the Clouds: A Study of Emerging Scale-Out Workloads on Modern Hardware,” in ASPLOS, 2012

  7. [6]

    BigDataBench: A Big Data Benchmark Suite from Internet Services,

    L. Wang, J. Zhan, C. Luo, Y . Zhu, Q. Yang, Y . He, W. Gao, Z. Jia, Y . Shi, S. Zhang et al., “BigDataBench: A Big Data Benchmark Suite from Internet Services,” in HPCA, 2014

  8. [7]

    Enabling Practical Processing in and near Memory for Data-Intensive Computing,

    O. Mutlu, S. Ghose, J. Gómez-Luna, and R. Ausavarungnirun, “Enabling Practical Processing in and near Memory for Data-Intensive Computing,” in DAC, 2019

Show all 234 references
  1. [8]

    Process- ing Data Where It Makes Sense: Enabling In-Memory Computation,

    O. Mutlu, S. Ghose, J. Gómez-Luna, and R. Ausavarungnirun, “Process- ing Data Where It Makes Sense: Enabling In-Memory Computation,” MicPro, 2019

  2. [9]

    Intelligent Architectures for Intelligent Machines,

    O. Mutlu, “Intelligent Architectures for Intelligent Machines,” in VLSI- DAT, 2020

  3. [10]

    Processing-in-Memory: A Workload-Driven Perspective,

    S. Ghose, A. Boroumand, J. S. Kim, J. Gómez-Luna, and O. Mutlu, “Processing-in-Memory: A Workload-Driven Perspective,” IBM JRD, 2019

  4. [11]

    A Modern Primer on Processing in Memory,

    O. Mutlu, S. Ghose, J. Gómez-Luna, and R. Ausavarungnirun, “A Modern Primer on Processing in Memory,” in Emerging Computing: From Devices to Systems — Looking Beyond Moore and Von Neumann . Springer, 2022. 5

  5. [12]

    DAMOV: A New Methodology and Benchmark Suite for Evaluating Data Movement Bottlenecks,

    G. F. Oliveira, J. Gómez-Luna, L. Orosa, S. Ghose, N. Vijaykumar, I. Fernandez, M. Sadrosadati, and O. Mutlu, “DAMOV: A New Methodology and Benchmark Suite for Evaluating Data Movement Bottlenecks,” IEEE Access, 2021

  6. [13]

    Google Workloads for Consumer Devices: Mitigating Data Movement Bottlenecks,

    A. Boroumand, S. Ghose, Y . Kim, R. Ausavarungnirun, E. Shiu, R. Thakur, D. Kim, A. Kuusela, A. Knies, P. Ranganathan et al. , “Google Workloads for Consumer Devices: Mitigating Data Movement Bottlenecks,” in ASPLOS, 2018

  7. [14]

    Google Neural Network Models for Edge Devices: Analyzing and Mitigating Machine Learning Inference Bottlenecks,

    A. Boroumand, S. Ghose, B. Akin, R. Narayanaswami, G. F. Oliveira, X. Ma, E. Shiu, and O. Mutlu, “Google Neural Network Models for Edge Devices: Analyzing and Mitigating Machine Learning Inference Bottlenecks,” in PACT, 2021

  8. [15]

    Reducing Data Movement Energy via Online Data Clustering and Encoding,

    S. Wang and E. Ipek, “Reducing Data Movement Energy via Online Data Clustering and Encoding,” in MICRO, 2016

  9. [16]

    Quantifying the Energy Cost of Data Movement for Emerging Smart Phone Workloads on Mobile Platforms,

    D. Pandiyan and C.-J. Wu, “Quantifying the Energy Cost of Data Movement for Emerging Smart Phone Workloads on Mobile Platforms,” in IISWC, 2014

  10. [17]

    EDEN: Enabling Energy-Efficient, High-Performance Deep Neural Network Inference Using Approximate DRAM,

    S. Koppula, L. Orosa, A. G. Ya ˘glıkçı, R. Azizi, T. Shahroodi, K. Kanellopoulos, and O. Mutlu, “EDEN: Enabling Energy-Efficient, High-Performance Deep Neural Network Inference Using Approximate DRAM,” in MICRO, 2019

  11. [18]

    Co-Architecting Controllers and DRAM to Enhance DRAM Process Scaling,

    U. Kang, H.-S. Yu, C. Park, H. Zheng, J. Halbert, K. Bains, S. Jang, and J. S. Choi, “Co-Architecting Controllers and DRAM to Enhance DRAM Process Scaling,” in The Memory Forum, 2014

  12. [19]

    Reflections on the Memory Wall,

    S. A. McKee, “Reflections on the Memory Wall,” in CF, 2004

  13. [20]

    The Memory Gap and the Future of High Performance Memories,

    M. V . Wilkes, “The Memory Gap and the Future of High Performance Memories,” CAN, 2001

  14. [21]

    A Case for Exploiting Subarray-Level Parallelism (SALP) in DRAM,

    Y . Kim, V . Seshadri, D. Lee, J. Liu, and O. Mutlu, “A Case for Exploiting Subarray-Level Parallelism (SALP) in DRAM,” in ISCA, 2012

  15. [22]

    Hitting the Memory Wall: Implications of the Obvious,

    W. A. Wulf and S. A. McKee, “Hitting the Memory Wall: Implications of the Obvious,” CAN, 1995

  16. [23]

    Demystifying Complex Workload–DRAM Interactions: An Experimental Study,

    S. Ghose, T. Li, N. Hajinazar, D. S. Cali, and O. Mutlu, “Demystifying Complex Workload–DRAM Interactions: An Experimental Study,” in SIGMETRICS, 2020

  17. [24]

    A Scalable Processing- in-Memory Accelerator for Parallel Graph Processing,

    J. Ahn, S. Hong, S. Yoo, O. Mutlu, and K. Choi, “A Scalable Processing- in-Memory Accelerator for Parallel Graph Processing,” in ISCA, 2015

  18. [25]

    PIM-Enabled Instructions: A Low-Overhead, Locality-Aware Processing-in-Memory Architecture,

    J. Ahn, S. Yoo, O. Mutlu, and K. Choi, “PIM-Enabled Instructions: A Low-Overhead, Locality-Aware Processing-in-Memory Architecture,” in ISCA, 2015

  19. [26]

    Transparent Offloading and Mapping (TOM) Enabling Programmer-Transparent Near-Data Processing in GPU Systems,

    K. Hsieh, E. Ebrahimi, G. Kim, N. Chatterjee, M. O’Connor, N. Vijayku- mar, O. Mutlu, and S. W. Keckler, “Transparent Offloading and Mapping (TOM) Enabling Programmer-Transparent Near-Data Processing in GPU Systems,” in ISCA, 2016

  20. [27]

    FIGARO: Improving System Performance via Fine-Grained In-DRAM Data Relocation and Caching,

    Y . Wang, L. Orosa, X. Peng, Y . Guo, S. Ghose, M. Patel, J. S. Kim, J. G. Luna, M. Sadrosadati, N. M. Ghiasi et al., “FIGARO: Improving System Performance via Fine-Grained In-DRAM Data Relocation and Caching,” in MICRO, 2020

  21. [28]

    It’s the Memory, Stupid!

    R. Sites, “It’s the Memory, Stupid!” MPR, 1996

  22. [29]

    Language Models are Few-Shot Learners,

    T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert- V oss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B...

  23. [30]

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in NAACL, 2019

  24. [31]

    Accelerating Neural Network Inference with Processing-in-DRAM: From the Edge to the Cloud,

    G. F. Oliveira, J. Gómez-Luna, S. Ghose, A. Boroumand, and O. Mutlu, “Accelerating Neural Network Inference with Processing-in-DRAM: From the Edge to the Cloud,” IEEE Micro, 2022

  25. [32]

    NEUPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing,

    G. Heo, S. Lee, J. Cho, H. Choi, S. Lee, H. Ham, G. Kim, D. Mahajan, and J. Park, “NEUPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing,” in ASPLOS, 2024

  26. [33]

    TransPIM: A Memory-Based Acceleration via Software-Hardware Co-Design for Transformer,

    M. Zhou, W. Xu, J. Kang, and T. Rosing, “TransPIM: A Memory-Based Acceleration via Software-Hardware Co-Design for Transformer,” in HPCA, 2022

  27. [34]

    AttAcc! Unleashing the Power of PIM for Batched Transformer- based Generative Model Inference,

    J. Park, J. Choi, K. Kyung, M. J. Kim, Y . Kwon, N. S. Kim, and J. H. Ahn, “AttAcc! Unleashing the Power of PIM for Batched Transformer- based Generative Model Inference,” in ASPLOS, 2024

  28. [35]

    IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System,

    M. Seo, X. T. Nguyen, S. J. Hwang, Y . Kwon, G. Kim, C. Park, I. Kim, J. Park, J. Kim, W. Shin et al., “IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System,” in ASPLOS, 2024

  29. [36]

    PIM-Opt: Demystifying Distributed Optimization Algorithms on a Real-World Processing-In-Memory System,

    S. Rhyner, H. Luo, J. Gómez-Luna, M. Sadrosadati, J. Jiang, A. Olgun, H. Gupta, C. Zhang, and O. Mutlu, “PIM-Opt: Demystifying Distributed Optimization Algorithms on a Real-World Processing-In-Memory System,” in PACT, 2024

  30. [37]

    Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching,

    S. Yun, K. Kyung, J. Cho, J. Choi, J. Kim, B. Kim, S. Lee, K. Sohn, and J. H. Ahn, “Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching,” in MICRO, 2024

  31. [38]

    Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System,

    H. Jang, J. Song, J. Jung, J. Park, Y . Kim, and J. Lee, “Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System,” in HPCA, 2024

  32. [39]

    Accelerating Genome Analysis: A Primer on an Ongoing Journey,

    M. Alser, Z. Bingöl, D. S. Cali, J. Kim, S. Ghose, C. Alkan, and O. Mutlu, “Accelerating Genome Analysis: A Primer on an Ongoing Journey,” IEEE Micro, 2020

  33. [40]

    FPGA-Based Near-Memory Acceleration of Modern Data-Intensive Applications,

    G. Singh, M. Alser, D. S. Cali, D. Diamantopoulos, J. Gómez-Luna, H. Corporaal, and O. Mutlu, “FPGA-Based Near-Memory Acceleration of Modern Data-Intensive Applications,” IEEE Micro, 2021

  34. [41]

    From Molecules to Genomic Varia- tions: Accelerating Genome Analysis via Intelligent Algorithms and Architectures,

    M. Alser, J. Lindegger, C. Firtina, N. Almadhoun, H. Mao, G. Singh, J. Gomez-Luna, and O. Mutlu, “From Molecules to Genomic Varia- tions: Accelerating Genome Analysis via Intelligent Algorithms and Architectures,” CSBJ, 2022

  35. [42]

    GRIM-Filter: Fast Seed Loca- tion Filtering in DNA Read Mapping Using Processing-in-Memory Technologies,

    J. S. Kim, D. S. Cali, H. Xin, D. Lee, S. Ghose, M. Alser, H. Hassan, O. Ergin, C. Alkan, and O. Mutlu, “GRIM-Filter: Fast Seed Loca- tion Filtering in DNA Read Mapping Using Processing-in-Memory Technologies,” BMC Genomics, 2018

  36. [43]

    GenStore: A High- Performance and Energy-Efficient In-Storage Computing System for Genome Sequence Analysis,

    N. M. Ghiasi, J. Park, H. Mustafa, J. Kim, A. Olgun, A. Gollwitzer, D. S. Cali, C. Firtina, H. Mao, N. A. Alserr et al., “GenStore: A High- Performance and Energy-Efficient In-Storage Computing System for Genome Sequence Analysis,” in ASPLOS, 2022

  37. [44]

    MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Anal- ysis with In-Storage Processing,

    N. M. Ghiasi, M. Sadrosadati, H. Mustafa, A. Gollwitzer, C. Firtina, J. Eudine, H. Mao, J. Lindegger, M. B. Cavlak, M. Alser et al., “MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Anal- ysis with In-Storage Processing,” in ISCA, 2024

  38. [45]

    GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis,

    D. S. Cali, G. S. Kalsi, Z. Bingöl, C. Firtina, L. Subramanian, J. S. Kim, R. Ausavarungnirun, M. Alser, J. Gomez-Luna, A. Boroumand et al., “GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis,” in MICRO, 2020

  39. [46]

    SeGraM: A Universal Hardware Accelerator for Genomic Sequence-to-Graph and Sequence-to-Sequence Mapping,

    D. S. Cali, K. Kanellopoulos, J. Lindegger, Z. Bingöl, G. S. Kalsi, Z. Zuo, C. Firtina, M. B. Cavlak, J. Kim, N. M. Ghiasi et al., “SeGraM: A Universal Hardware Accelerator for Genomic Sequence-to-Graph and Sequence-to-Sequence Mapping,” in ISCA, 2022

  40. [47]

    Nanopore Sequencing Technology and Tools for Genome Assembly: Computational Analysis of the Current State, Bottlenecks and Future Directions,

    D. Senol Cali, J. S. Kim, S. Ghose, C. Alkan, and O. Mutlu, “Nanopore Sequencing Technology and Tools for Genome Assembly: Computational Analysis of the Current State, Bottlenecks and Future Directions,” Briefings in Bioinformatics , 2018

  41. [48]

    NDA: Near-DRAM Acceleration Architecture Leveraging Commodity DRAM Devices and Standard Memory Modules,

    A. Farmahini-Farahani, J. H. Ahn, K. Morrow, and N. S. Kim, “NDA: Near-DRAM Acceleration Architecture Leveraging Commodity DRAM Devices and Standard Memory Modules,” in HPCA, 2015

  42. [49]

    JAFAR: Near-Data Processing for Databases,

    O. O. Babarinsa and S. Idreos, “JAFAR: Near-Data Processing for Databases,” in SIGMOD, 2015

  43. [50]

    The True Processing in Memory Accelerator,

    F. Devaux, “The True Processing in Memory Accelerator,” in Hot Chips, 2019

  44. [51]

    Benchmarking Memory-Centric Computing Systems: Analysis of Real Processing-in-Memory Hardware,

    J. Gómez-Luna, I. El Hajj, I. Fernandez, C. Giannoula, G. F. Oliveira, and O. Mutlu, “Benchmarking Memory-Centric Computing Systems: Analysis of Real Processing-in-Memory Hardware,” in CUT, 2021

  45. [52]

    Benchmarking a New Paradigm: Experimental Analysis and Characterization of a Real Processing-in-Memory System,

    J. Gómez-Luna, I. El Hajj, I. Fernandez, C. Giannoula, G. F. Oliveira, and O. Mutlu, “Benchmarking a New Paradigm: Experimental Analysis and Characterization of a Real Processing-in-Memory System,” IEEE Access, 2022

  46. [53]

    SynCron: Efficient Synchronization Support for Near-Data-Processing Architectures,

    C. Giannoula, N. Vijaykumar, N. Papadopoulou, V . Karakostas, I. Fer- nandez, J. Gómez-Luna, L. Orosa, N. Koziris, G. Goumas, and O. Mutlu, “SynCron: Efficient Synchronization Support for Near-Data-Processing Architectures,” in HPCA, 2021

  47. [54]

    NERO: A Near High-Bandwidth Memory Stencil Accelerator for Weather Prediction Modeling,

    G. Singh, D. Diamantopoulos, C. Hagleitner, J. Gomez-Luna, S. Stuijk, O. Mutlu, and H. Corporaal, “NERO: A Near High-Bandwidth Memory Stencil Accelerator for Weather Prediction Modeling,” in FPL, 2020

  48. [55]

    A 1ynm 1.25V 8Gb, 16Gb/s/pin GDDR6- Based Accelerator-in-Memory Supporting 1TFLOPS MAC Operation and Various Activation Functions for Deep-Learning Applications,

    S. Lee, K. Kim, S. Oh, J. Park, G. Hong, D. Ka, K. Hwang, J. Park, K. Kang, J. Kim, J. Jeon, N. Kim, Y . Kwon, K. Vladimir, W. Shin, J. Won, M. Lee, H. Joo et al., “A 1ynm 1.25V 8Gb, 16Gb/s/pin GDDR6- Based Accelerator-in-Memory Supporting 1TFLOPS MAC Operation and Various Act...

  49. [56]

    Near-Memory Processing in Action: Accelerating Personalized Recommendation with AxDIMM,

    L. Ke, X. Zhang, J. So, J.-G. Lee, S.-H. Kang, S. Lee, S. Han, Y . Cho, J. H. Kim, Y . Kwon et al. , “Near-Memory Processing in Action: Accelerating Personalized Recommendation with AxDIMM,” IEEE Micro, 2021

  50. [57]

    SparseP: Towards Efficient Sparse Matrix Vector Multipli- cation on Real Processing-in-Memory Architectures,

    C. Giannoula, I. Fernandez, J. G. Luna, N. Koziris, G. Goumas, and O. Mutlu, “SparseP: Towards Efficient Sparse Matrix Vector Multipli- cation on Real Processing-in-Memory Architectures,” in SIGMETRICS, 2022

  51. [58]

    McDRAM: Low Latency and Energy-Efficient Matrix Computations in DRAM,

    H. Shin, D. Kim, E. Park, S. Park, Y . Park, and S. Yoo, “McDRAM: Low Latency and Energy-Efficient Matrix Computations in DRAM,” TCADICS, 2018

  52. [59]

    McDRAM v2: In-Dynamic Random Access Memory Systolic Array Accelerator to Address the Large Model Problem in Deep Neural Networks on the Edge,

    S. Cho, H. Choi, E. Park, H. Shin, and S. Yoo, “McDRAM v2: In-Dynamic Random Access Memory Systolic Array Accelerator to Address the Large Model Problem in Deep Neural Networks on the Edge,” IEEE Access, 2020

  53. [60]

    Casper: Accelerating Stencil Computation using Near-Cache Processing,

    A. Denzler, R. Bera, N. Hajinazar, G. Singh, G. F. Oliveira, J. Gómez- Luna, and O. Mutlu, “Casper: Accelerating Stencil Computation using Near-Cache Processing,” IEEE Access, 2023

  54. [61]

    Chameleon: Versatile and Practical Near-DRAM Acceleration Archi- tecture for Large Memory Systems,

    H. Asghari-Moghaddam, Y . H. Son, J. H. Ahn, and N. S. Kim, “Chameleon: Versatile and Practical Near-DRAM Acceleration Archi- tecture for Large Memory Systems,” in MICRO, 2016

  55. [62]

    A Case for Intelligent RAM,

    D. Patterson, T. Anderson, N. Cardwell et al., “A Case for Intelligent RAM,” IEEE Micro, 1997

  56. [63]

    Computational RAM: Implementing Processors in Memory,

    D. G. Elliott, M. Stumm, W. M. Snelgrove et al., “Computational RAM: Implementing Processors in Memory,” D&T, 1999. 6

  57. [64]

    Saving Memory Movements Through Vector Processing in the DRAM,

    M. A. Z. Alves, P. C. Santos, F. B. Moreira, and opthers, “Saving Memory Movements Through Vector Processing in the DRAM,” in CASES, 2015

  58. [65]

    Beyond the Wall: Near-Data Processing for Databases,

    S. L. Xi, O. Babarinsa, M. Athanassoulis, and S. Idreos, “Beyond the Wall: Near-Data Processing for Databases,” in DaMoN, 2015

  59. [66]

    ABC-DIMM: Alleviating the Bottleneck of Communication in DIMM-Based Near-Memory Processing with Inter-DIMM Broadcast,

    W. Sun, Z. Li, S. Yin, S. Wei, and L. Liu, “ABC-DIMM: Alleviating the Bottleneck of Communication in DIMM-Based Near-Memory Processing with Inter-DIMM Broadcast,” in ISCA, 2021

  60. [67]

    GraphSSD: Graph Semantics Aware SSD,

    K. K. Matam, G. Koo, H. Zha, H.-W. Tseng, and M. Annavaram, “GraphSSD: Graph Semantics Aware SSD,” in ISCA, 2019

  61. [68]

    Processing in Memory: The Terasys Massively Parallel PIM Array,

    M. Gokhale, B. Holmes, and K. Iobst, “Processing in Memory: The Terasys Massively Parallel PIM Array,” Computer, 1995

  62. [69]

    Mapping Irregular Applications to DIV A, a PIM-Based Data-Intensive Architecture,

    M. Hall, P. Kogge, J. Koller, P. Diniz, J. Chame, J. Draper, J. LaCoss, J. Granacki, J. Brockman, A. Srivastava et al. , “Mapping Irregular Applications to DIV A, a PIM-Based Data-Intensive Architecture,” in SC, 1999

  63. [70]

    Opportunities and Challenges of Performing Vector Operations Inside the DRAM,

    M. A. Z. Alves, P. C. Santos, M. Diener, and L. Carro, “Opportunities and Challenges of Performing Vector Operations Inside the DRAM,” in MEMSYS, 2015

  64. [71]

    Livia: Data-Centric Computing Throughout the Memory Hierarchy,

    E. Lockerman, A. Feldmann, M. Bakhshalipour, A. Stanescu, S. Gupta, D. Sanchez, and N. Beckmann, “Livia: Data-Centric Computing Throughout the Memory Hierarchy,” in ASPLOS, 2020

  65. [72]

    GraphPIM: Enabling Instruction-Level PIM Offloading in Graph Computing Frameworks,

    L. Nai, R. Hadidi, J. Sim, H. Kim, P. Kumar, and H. Kim, “GraphPIM: Enabling Instruction-Level PIM Offloading in Graph Computing Frameworks,” in HPCA, 2017

  66. [73]

    LazyPIM: An Efficient Cache Coherence Mechanism for Processing-in-Memory,

    A. Boroumand, S. Ghose, B. Lucia, K. Hsieh, K. Malladi, H. Zheng, and O. Mutlu, “LazyPIM: An Efficient Cache Coherence Mechanism for Processing-in-Memory,” CAL, 2017

  67. [74]

    TOP-PIM: Throughput-Oriented Programmable Processing in Memory,

    D. Zhang, N. Jayasena, A. Lyashevsky, J. L. Greathouse, L. Xu, and M. Ignatowski, “TOP-PIM: Throughput-Oriented Programmable Processing in Memory,” in HPDC, 2014

  68. [75]

    HRL: Efficient and Flexible Reconfigurable Logic for Near-Data Processing,

    M. Gao and C. Kozyrakis, “HRL: Efficient and Flexible Reconfigurable Logic for Near-Data Processing,” in HPCA, 2016

  69. [76]

    The Mondrian Data Engine,

    M. Drumond, A. Daglis, N. Mirzadeh, D. Ustiugov, J. Picorel, B. Falsafi, B. Grot, and D. Pnevmatikatos, “The Mondrian Data Engine,” in ISCA, 2017

  70. [77]

    Operand Size Reconfiguration for Big Data Processing in Memory,

    P. C. Santos, G. F. Oliveira, D. G. Tomé, M. A. Z. Alves, E. C. Almeida, and L. Carro, “Operand Size Reconfiguration for Big Data Processing in Memory,” in DATE, 2017

  71. [78]

    NIM: An HMC-Based Machine for Neuron Computation,

    G. F. Oliveira, P. C. Santos, M. A. Alves, and L. Carro, “NIM: An HMC-Based Machine for Neuron Computation,” in ARC, 2017

  72. [79]

    TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory,

    M. Gao, J. Pu, X. Yang, M. Horowitz, and C. Kozyrakis, “TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory,” in ASPLOS, 2017

  73. [80]

    Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory,

    D. Kim, J. Kung, S. Chai, S. Yalamanchili, and S. Mukhopadhyay, “Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory,” in ISCA, 2016

  74. [81]

    Leveraging 3D Technologies for Hardware Security: Opportunities and Challenges,

    P. Gu, S. Li, D. Stow, R. Barnes, L. Liu, Y . Xie, and E. Kursun, “Leveraging 3D Technologies for Hardware Security: Opportunities and Challenges,” in GLSVLSI, 2016

  75. [82]

    CoNDA: Efficient Cache Coherence Support for Near-Data Accelerators,

    A. Boroumand, S. Ghose, M. Patel, H. Hassan, B. Lucia, R. Ausavarung- nirun, K. Hsieh, N. Hajinazar, K. T. Malladi, H. Zheng et al., “CoNDA: Efficient Cache Coherence Support for Near-Data Accelerators,” in ISCA, 2019

  76. [83]

    NDC: Analyzing the Impact of 3D-Stacked Memory+Logic Devices on MapReduce Workloads,

    S. H. Pugsley, J. Jestes, H. Zhang, R. Balasubramonian et al., “NDC: Analyzing the Impact of 3D-Stacked Memory+Logic Devices on MapReduce Workloads,” in ISPASS, 2014

  77. [84]

    Scheduling Techniques for GPU Architectures with Processing-in-Memory Capabilities,

    A. Pattnaik, X. Tang, A. Jog, O. Kayiran, A. K. Mishra, M. T. Kandemir, O. Mutlu, and C. R. Das, “Scheduling Techniques for GPU Architectures with Processing-in-Memory Capabilities,” in PACT, 2016

  78. [85]

    Data Reorganization in Memory Using 3D-Stacked DRAM,

    B. Akin, F. Franchetti, and J. C. Hoe, “Data Reorganization in Memory Using 3D-Stacked DRAM,” in ISCA, 2015

  79. [86]

    Accelerating Pointer Chasing in 3D-Stacked Memory: Challenges, Mechanisms, Evaluation,

    K. Hsieh, S. Khan, N. Vijaykumar, K. K. Chang, A. Boroumand, S. Ghose, and O. Mutlu, “Accelerating Pointer Chasing in 3D-Stacked Memory: Challenges, Mechanisms, Evaluation,” in ICCD, 2016

  80. [87]

    BSSync: Processing Near Memory for Machine Learning Workloads with Bounded Staleness Consistency Models,

    J. H. Lee, J. Sim, and H. Kim, “BSSync: Processing Near Memory for Machine Learning Workloads with Bounded Staleness Consistency Models,” in PACT, 2015

  81. [88]

    Polynesia: Enabling High-Performance and Energy-Efficient Hybrid Transaction- al/Analytical Databases with Hardware/Software Co-Design,

    A. Boroumand, S. Ghose, G. F. Oliveira, and O. Mutlu, “Polynesia: Enabling High-Performance and Energy-Efficient Hybrid Transaction- al/Analytical Databases with Hardware/Software Co-Design,” in ICDE, 2022

  82. [89]

    Practical Mechanisms for Reducing Processor-Memory Data Movement in Modern Workloads,

    A. Boroumand, “Practical Mechanisms for Reducing Processor-Memory Data Movement in Modern Workloads,” Ph.D. dissertation, Carnegie Mellon University, 2020

  83. [90]

    SISA: Set-Centric Instruction Set Architecture for Graph Mining on Processing-in-Memory Systems,

    M. Besta, R. Kanakagiri, G. Kwasniewski, R. Ausavarungnirun, J. Beránek, K. Kanellopoulos, K. Janda, Z. V onarburg-Shmaria, L. Gian- inazzi, I. Stefan et al., “SISA: Set-Centric Instruction Set Architecture for Graph Mining on Processing-in-Memory Systems,” in MICRO, 2021

  84. [91]

    NATSA: A Near-Data Processing Accelerator for Time Series Analysis,

    I. Fernandez, R. Quislant, E. Gutiérrez, O. Plata, C. Giannoula, M. Alser, J. Gómez-Luna, and O. Mutlu, “NATSA: A Near-Data Processing Accelerator for Time Series Analysis,” in ICCD, 2020

  85. [92]

    NAPEL: Near-Memory Computing Application Performance Prediction via Ensemble Learning,

    G. Singh, G. , G. F. Oliveira, S. Corda, S. Stuijk, O. Mutlu, and H. Cor- poraal, “NAPEL: Near-Memory Computing Application Performance Prediction via Ensemble Learning,” in DAC, 2019

  86. [93]

    A 20nm 6GB Function- in-Memory DRAM, Based on HBM2 with a 1.2 TFLOPS Programmable Computing Unit using Bank-Level Parallelism, for Machine Learning Applications,

    Y .-C. Kwon, S. H. Lee, J. Lee, S.-H. Kwon, J. M. Ryu, J.-P. Son, O. Seongil, H.-S. Yu, H. Lee, S. Y . Kim et al., “A 20nm 6GB Function- in-Memory DRAM, Based on HBM2 with a 1.2 TFLOPS Programmable Computing Unit using Bank-Level Parallelism, for Machine Learning Applications,...

  87. [94]

    Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology: Industrial Product,

    S. Lee, S.-h. Kang, J. Lee, H. Kim, E. Lee, S. Seo, H. Yoon, S. Lee, K. Lim, H. Shin et al., “Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology: Industrial Product,” in ISCA, 2021

  88. [95]

    184QPS/W 64Mb/ mm2 3D Logic- to-DRAM Hybrid Bonding with Process-Near-Memory Engine for Recommendation System,

    D. Niu, S. Li, Y . Wang, W. Han, Z. Zhang, Y . Guan, T. Guan, F. Sun, F. Xue, L. Duan et al. , “184QPS/W 64Mb/ mm2 3D Logic- to-DRAM Hybrid Bonding with Process-Near-Memory Engine for Recommendation System,” in ISSCC, 2022

  89. [96]

    Acceler- ating Sparse Matrix-Matrix Multiplication with 3D-Stacked Logic-in- Memory Hardware,

    Q. Zhu, T. Graf, H. E. Sumbul, L. Pileggi, and F. Franchetti, “Acceler- ating Sparse Matrix-Matrix Multiplication with 3D-Stacked Logic-in- Memory Hardware,” in HPEC, 2013

  90. [97]

    Logic-Base Interconnect Design for Near Memory Computing in the Smart Memory Cube,

    E. Azarkhish, C. Pfister, D. Rossi, I. Loi, and L. Benini, “Logic-Base Interconnect Design for Near Memory Computing in the Smart Memory Cube,” IEEE VLSI, 2016

  91. [98]

    Neurostream: Scalable and Energy Efficient Deep Learning with Smart Memory Cubes,

    E. Azarkhish, D. Rossi, I. Loi, and L. Benini, “Neurostream: Scalable and Energy Efficient Deep Learning with Smart Memory Cubes,” TPDS, 2018

  92. [99]

    3D-Stacked Memory-Side Acceleration: Accelerator and System Design,

    Q. Guo, N. Alachiotis, B. Akin, F. Sadi, G. Xu, T. M. Low, L. Pileggi, J. C. Hoe, and F. Franchetti, “3D-Stacked Memory-Side Acceleration: Accelerator and System Design,” in WoNDP, 2014

  93. [100]

    HAMLeT: Hardware Accelerated Memory Layout Transform within 3D-Stacked DRAM,

    B. Akın, J. C. Hoe, and F. Franchetti, “HAMLeT: Hardware Accelerated Memory Layout Transform within 3D-Stacked DRAM,” in HPEC, 2014

  94. [101]

    A Heterogeneous PIM Hardware-Software Co-Design for Energy-Efficient Graph Processing,

    Y . Huang, L. Zheng, P. Yao, J. Zhao, X. Liao, H. Jin, and J. Xue, “A Heterogeneous PIM Hardware-Software Co-Design for Energy-Efficient Graph Processing,” in IPDPS, 2020

  95. [102]

    GraphH: A Processing-in-Memory Architecture for Large- Scale Graph Processing,

    G. Dai, T. Huang, Y . Chi, J. Zhao, G. Sun, Y . Liu, Y . Wang, Y . Xie, and H. Yang, “GraphH: A Processing-in-Memory Architecture for Large- Scale Graph Processing,” TCAD, 2018

  96. [103]

    Processing-in- Memory for Energy-Efficient Neural Network Training: A Heteroge- neous Approach,

    J. Liu, H. Zhao, M. A. Ogleari, D. Li, and J. Zhao, “Processing-in- Memory for Energy-Efficient Neural Network Training: A Heteroge- neous Approach,” in MICRO, 2018

  97. [104]

    Adaptive Scheduling for Systems with Asymmetric Memory Hierarchies,

    P.-A. Tsai, C. Chen, and D. Sanchez, “Adaptive Scheduling for Systems with Asymmetric Memory Hierarchies,” in MICRO, 2018

  98. [105]

    iPIM: Programmable In-Memory Image Processing Accelerator using Near-Bank Architecture,

    P. Gu, X. Xie, Y . Ding, G. Chen, W. Zhang, D. Niu, and Y . Xie, “iPIM: Programmable In-Memory Image Processing Accelerator using Near-Bank Architecture,” in ISCA, 2020

  99. [106]

    DRAMA: An Architecture for Accelerated Processing Near Memory,

    A. Farmahini-Farahani, J. H. Ahn, K. Compton, and N. S. Kim, “DRAMA: An Architecture for Accelerated Processing Near Memory,” CAL, 2014

  100. [107]

    Near-DRAM Acceleration with Single-ISA Heterogeneous Processing in Standard Memory Modules,

    H. Asghari-Moghaddam, A. Farmahini-Farahani, K. Morrow et al. , “Near-DRAM Acceleration with Single-ISA Heterogeneous Processing in Standard Memory Modules,” IEEE Micro, 2016

  101. [108]

    Active-Routing: Compute on the Way for Near-Data Processing,

    J. Huang, R. R. Puli, P. Majumder, S. Kim, R. Boyapati, K. H. Yum, and E. J. Kim, “Active-Routing: Compute on the Way for Near-Data Processing,” in HPCA, 2019

  102. [109]

    Lightweight SIMT Core Designs for Intelligent 3D Stacked DRAM,

    C. D. Kersey, H. Kim, and S. Yalamanchili, “Lightweight SIMT Core Designs for Intelligent 3D Stacked DRAM,” in MEMSYS, 2017

  103. [110]

    PIMS: A Lightweight Processing-in-Memory Accelerator for Stencil Computations,

    J. Li, X. Wang, A. Tumeo, B. Williams, J. D. Leidel, and Y . Chen, “PIMS: A Lightweight Processing-in-Memory Accelerator for Stencil Computations,” in MEMSYS, 2019

  104. [111]

    GraphQ: Scalable PIM-Based Graph Processing,

    Y . Zhuo, C. Wang, M. Zhang, R. Wang, D. Niu, Y . Wang, and X. Qian, “GraphQ: Scalable PIM-Based Graph Processing,” in MICRO, 2019

  105. [112]

    GraphP: Reducing Communication for PIM-Based Graph Processing with Efficient Data Partition,

    M. Zhang, Y . Zhuo, C. Wang, M. Gao, Y . Wu, K. Chen, C. Kozyrakis, and X. Qian, “GraphP: Reducing Communication for PIM-Based Graph Processing with Efficient Data Partition,” in HPCA, 2018

  106. [113]

    Triple Engine Processor (TEP): A Heterogeneous Near-Memory Processor for Diverse Kernel Operations,

    H. Lim and G. Park, “Triple Engine Processor (TEP): A Heterogeneous Near-Memory Processor for Diverse Kernel Operations,” TACO, 2017

  107. [114]

    A Case for Near Memory Computation Inside the Smart Memory Cube,

    E. Azarkhish, D. Rossi, I. Loi, and L. Benini, “A Case for Near Memory Computation Inside the Smart Memory Cube,” in EMS, 2016

  108. [115]

    Large Vector Extensions Inside the HMC,

    M. A. Z. Alves, M. Diener, P. C. Santos, and L. Carro, “Large Vector Extensions Inside the HMC,” in DATE, 2016

  109. [116]

    Charon: Specialized Near-Memory Processing Architecture for Clearing Dead Objects in Memory,

    J. Jang, J. Heo, Y . Lee, J. Won, S. Kim, S. J. Jung, H. Jang, T. J. Ham, and J. W. Lee, “Charon: Specialized Near-Memory Processing Architecture for Clearing Dead Objects in Memory,” in MICRO, 2019

  110. [117]

    Active Memory Cube: A Processing-in-Memory Architecture for Exascale Systems,

    R. Nair, S. F. Antao, C. Bertolli, P. Bose et al., “Active Memory Cube: A Processing-in-Memory Architecture for Exascale Systems,” IBM JRD, 2015

  111. [118]

    CAIRO: A Compiler-Assisted Technique for Enabling Instruction-Level Offloading of Processing-in- Memory,

    R. Hadidi, L. Nai, H. Kim, and H. Kim, “CAIRO: A Compiler-Assisted Technique for Enabling Instruction-Level Offloading of Processing-in- Memory,” TACO, 2017

  112. [119]

    Processing in 3D Memories to Speed Up Operations on Complex Data Structures,

    P. C. Santos, G. F. Oliveira, J. P. Lima, M. A. Alves, L. Carro, and A. C. Beck, “Processing in 3D Memories to Speed Up Operations on Complex Data Structures,” in DATE, 2018

  113. [120]

    PRIME: A Novel Processing-in-Memory Architecture for Neural Network Computation in ReRAM-Based Main Memory,

    P. Chi, S. Li, C. Xu, T. Zhang, J. Zhao, Y . Liu, Y . Wang, and Y . Xie, “PRIME: A Novel Processing-in-Memory Architecture for Neural Network Computation in ReRAM-Based Main Memory,” in ISCA, 2016

  114. [121]

    ISAAC: A Convo- lutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars,

    A. Shafiee, A. Nag, N. Muralimanohar, R. Balasubramonian, J. P. Strachan, M. Hu, R. S. Williams, and V . Srikumar, “ISAAC: A Convo- lutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars,” in ISCA, 2016. 7

  115. [122]

    Ambit: In-Memory Accelerator for Bulk Bitwise Operations Using Commodity DRAM Technology,

    V . Seshadri, D. Lee, T. Mullins, H. Hassan, A. Boroumand, J. Kim, M. A. Kozuch, O. Mutlu, P. B. Gibbons, and T. C. Mowry, “Ambit: In-Memory Accelerator for Bulk Bitwise Operations Using Commodity DRAM Technology,” in MICRO, 2017

  116. [123]

    In-DRAM Bulk Bitwise Execution Engine,

    V . Seshadri and O. Mutlu, “In-DRAM Bulk Bitwise Execution Engine,” arXiv:1905.09822, 2019

  117. [124]

    DRISA: A DRAM-Based Reconfigurable In-Situ Accelerator,

    S. Li, D. Niu, K. T. Malladi, H. Zheng, B. Brennan, and Y . Xie, “DRISA: A DRAM-Based Reconfigurable In-Situ Accelerator,” in MICRO, 2017

  118. [125]

    RowClone: Fast and Energy-Efficient In-DRAM Bulk Data Copy and Initialization,

    V . Seshadri, Y . Kim, C. Fallin, D. Lee, R. Ausavarungnirun, G. Pekhi- menko, Y . Luo, O. Mutlu, P. B. Gibbons, M. A. Kozuch et al. , “RowClone: Fast and Energy-Efficient In-DRAM Bulk Data Copy and Initialization,” in MICRO, 2013

  119. [126]

    The Processing Using Memory Paradigm: In-DRAM Bulk Copy, Initialization, Bitwise AND and OR,

    V . Seshadri and O. Mutlu, “The Processing Using Memory Paradigm: In-DRAM Bulk Copy, Initialization, Bitwise AND and OR,” arXiv:1610.09603, 2016

  120. [127]

    DrAcc: A DRAM Based Accelerator for Accurate CNN Inference,

    Q. Deng, L. Jiang, Y . Zhang, M. Zhang, and J. Yang, “DrAcc: A DRAM Based Accelerator for Accurate CNN Inference,” in DAC, 2018

  121. [128]

    ELP2IM: Efficient and Low Power Bitwise Operation Processing in DRAM,

    X. Xin, Y . Zhang, and J. Yang, “ELP2IM: Efficient and Low Power Bitwise Operation Processing in DRAM,” in HPCA, 2020

  122. [129]

    GraphR: Accelerating Graph Processing Using ReRAM,

    L. Song, Y . Zhuo, X. Qian, H. Li, and Y . Chen, “GraphR: Accelerating Graph Processing Using ReRAM,” in HPCA, 2018

  123. [130]

    PipeLayer: A Pipelined ReRAM- Based Accelerator for Deep Learning,

    L. Song, X. Qian, H. Li, and Y . Chen, “PipeLayer: A Pipelined ReRAM- Based Accelerator for Deep Learning,” in HPCA, 2017

  124. [131]

    ComputeDRAM: In- Memory Compute Using Off-the-Shelf DRAMs,

    F. Gao, G. Tziantzioulis, and D. Wentzlaff, “ComputeDRAM: In- Memory Compute Using Off-the-Shelf DRAMs,” in MICRO, 2019

  125. [132]

    Neural Cache: Bit-Serial In-Cache Acceleration of Deep Neural Networks,

    C. Eckert, X. Wang, J. Wang, A. Subramaniyan, R. Iyer, D. Sylvester, D. Blaauw, and R. Das, “Neural Cache: Bit-Serial In-Cache Acceleration of Deep Neural Networks,” in ISCA, 2018

  126. [133]

    Compute Caches,

    S. Aga, S. Jeloka, A. Subramaniyan, S. Narayanasamy, D. Blaauw, and R. Das, “Compute Caches,” in HPCA, 2017

  127. [134]

    Duality Cache for Data Parallel Acceleration,

    D. Fujiki, S. Mahlke, and R. Das, “Duality Cache for Data Parallel Acceleration,” in ISCA, 2019

  128. [135]

    Buddy-RAM: Improving the Performance and Efficiency of Bulk Bitwise Operations Using DRAM,

    V . Seshadri, D. Lee, T. Mullins, H. Hassan, A. Boroumand, J. Kim, M. A. Kozuch, O. Mutlu, P. B. Gibbons, and T. C. Mowry, “Buddy-RAM: Improving the Performance and Efficiency of Bulk Bitwise Operations Using DRAM,” arXiv:1611.09988, 2016

  129. [136]

    Simple Operations in Memory to Reduce Data Movement,

    V . Seshadri and O. Mutlu, “Simple Operations in Memory to Reduce Data Movement,” in Advances in Computers, Volume 106 , 2017

  130. [137]

    RowClone: Accelerating Data Movement and Initialization Using DRAM,

    V . Seshadri, Y . Kim, C. Fallin, D. Lee, R. Ausavarungnirun, G. Pekhi- menko, Y . Luo, O. Mutlu, P. B. Gibbons, M. A. Kozuch et al. , “RowClone: Accelerating Data Movement and Initialization Using DRAM,” arXiv:1805.03502, 2018

  131. [138]

    Fast Bulk Bitwise AND and OR in DRAM,

    V . Seshadri, K. Hsieh, A. Boroum, D. Lee, M. A. Kozuch, O. Mutlu, P. B. Gibbons, and T. C. Mowry, “Fast Bulk Bitwise AND and OR in DRAM,” CAL, 2015

  132. [139]

    Pinatubo: A Processing-in-Memory Architecture for Bulk Bitwise Operations in Emerging Non-V olatile Memories,

    S. Li, C. Xu, Q. Zou, J. Zhao, Y . Lu, and Y . Xie, “Pinatubo: A Processing-in-Memory Architecture for Bulk Bitwise Operations in Emerging Non-V olatile Memories,” in DAC, 2016

  133. [140]

    pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables,

    J. D. Ferreira, G. Falcao, J. Gómez-Luna, M. Alser, L. Orosa, M. Sadrosadati, J. S. Kim, G. F. Oliveira, T. Shahroodi, A. Nori et al., “pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables,” in MICRO, 2022

  134. [141]

    FloatPIM: In-Memory Acceleration of Deep Neural Network Training with High Precision,

    M. Imani, S. Gupta, Y . Kim, and T. Rosing, “FloatPIM: In-Memory Acceleration of Deep Neural Network Training with High Precision,” in ISCA, 2019

  135. [142]

    Sparse BD-Net: A Multiplication-Less DNN with Sparse Binarized Depth-Wise Separable Convolution,

    Z. He, L. Yang, S. Angizi, A. S. Rakin, and D. Fan, “Sparse BD-Net: A Multiplication-Less DNN with Sparse Binarized Depth-Wise Separable Convolution,” JETC, 2020

  136. [143]

    Flash-Cosmos: In-Flash Bulk Bitwise Operations Using Inherent Computation Capability of NAND Flash Memory,

    J. Park, R. Azizi, G. F. Oliveira, M. Sadrosadati, R. Nadig, D. Novo, J. Gómez-Luna, M. Kim, and O. Mutlu, “Flash-Cosmos: In-Flash Bulk Bitwise Operations Using Inherent Computation Capability of NAND Flash Memory,” in MICRO, 2022

  137. [144]

    Adapting the RACER Architecture to Integrate Improved In-ReRAM Logic Primitives,

    M. S. Truong, L. Shen, A. Glass, A. Hoffmann, L. R. Carley, J. A. Bain, and S. Ghose, “Adapting the RACER Architecture to Integrate Improved In-ReRAM Logic Primitives,” JETCAS, 2022

  138. [145]

    RACER: Bit-Pipelined Processing Using Resistive Memory,

    M. S. Truong, E. Chen, D. Su, L. Shen, A. Glass, L. R. Carley, J. A. Bain, and S. Ghose, “RACER: Bit-Pipelined Processing Using Resistive Memory,” in MICRO, 2021

  139. [146]

    QUAC-TRNG: High-Throughput True Random Number Generation Using Quadruple Row Activation in Commodity DRAMs,

    A. Olgun, M. Patel, A. G. Ya ˘glıkçı, H. Luo, J. S. Kim, F. N. Bostancı, N. Vijaykumar, O. Ergin, and O. Mutlu, “QUAC-TRNG: High-Throughput True Random Number Generation Using Quadruple Row Activation in Commodity DRAMs,” in ISCA, 2021

  140. [147]

    D-RaNGe: Using Commodity DRAM Devices to Generate True Random Numbers With Low Latency and High Throughput,

    J. S. Kim, M. Patel, H. Hassan, L. Orosa, and O. Mutlu, “D-RaNGe: Using Commodity DRAM Devices to Generate True Random Numbers With Low Latency and High Throughput,” in HPCA, 2019

  141. [148]

    The DRAM Latency PUF: Quickly Evaluating Physical Unclonable Functions by Exploiting the Latency-Reliability Tradeoff in Modern Commodity DRAM Devices,

    J. S. Kim, M. Patel, H. Hassan, and O. Mutlu, “The DRAM Latency PUF: Quickly Evaluating Physical Unclonable Functions by Exploiting the Latency-Reliability Tradeoff in Modern Commodity DRAM Devices,” in HPCA, 2018

  142. [149]

    DR-STRaNGe: End-to-End System Design for DRAM-Based True Random Number Generators,

    F. N. Bostancı, A. Olgun, L. Orosa, A. G. Ya˘glıkçı, J. S. Kim, H. Hassan, O. Ergin, and O. Mutlu, “DR-STRaNGe: End-to-End System Design for DRAM-Based True Random Number Generators,” in HPCA, 2022

  143. [150]

    PiDRAM: A Holistic End-to-End FPGA- Based Framework for Processing-in-DRAM,

    A. Olgun, J. G. Luna, K. Kanellopoulos, B. Salami, H. Hassan, O. Ergin, and O. Mutlu, “PiDRAM: A Holistic End-to-End FPGA- Based Framework for Processing-in-DRAM,” TACO, 2022

  144. [151]

    In-Memory Low-Cost Bit-Serial Addition Using Commodity DRAM Technology,

    M. F. Ali, A. Jaiswal, and K. Roy, “In-Memory Low-Cost Bit-Serial Addition Using Commodity DRAM Technology,” in TCAS-I, 2019

  145. [152]

    GraphiDe: A Graph Processing Accelerator Leveraging In-DRAM-Computing,

    S. Angizi and D. Fan, “GraphiDe: A Graph Processing Accelerator Leveraging In-DRAM-Computing,” in GLSVLSI, 2019

  146. [153]

    SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator,

    S. Li, A. O. Glova, X. Hu, P. Gu, D. Niu, K. T. Malladi, H. Zheng, B. Brennan, and Y . Xie, “SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator,” in MICRO, 2018

  147. [154]

    Parallel Automata Processor,

    A. Subramaniyan and R. Das, “Parallel Automata Processor,” in ISCA, 2017

  148. [155]

    Hyper-AP: Enhancing Associative Processing Through A Full-Stack Optimization,

    Y . Zha and J. Li, “Hyper-AP: Enhancing Associative Processing Through A Full-Stack Optimization,” in ISCA, 2020

  149. [156]

    In-Memory Data Parallel Processor,

    D. Fujiki, S. Mahlke, and R. Das, “In-Memory Data Parallel Processor,” in ASPLOS, 2018

  150. [157]

    CODIC: A Low-Cost Substrate for Enabling Custom In- DRAM Functionalities and Optimizations,

    L. Orosa, Y . Wang, M. Sadrosadati, J. Kim, M. Patel, I. Puddu, H. Luo, K. Razavi, J. Gómez-Luna, H. Hassan, N. M. Ghiasi, S. Ghose, and O. Mutlu, “CODIC: A Low-Cost Substrate for Enabling Custom In- DRAM Functionalities and Optimizations,” in ISCA, 2021

  151. [158]

    Ultra Low Power Associative Computing with Spin Neurons and Resistive Crossbar Memory,

    M. Sharad, D. Fan, and K. Roy, “Ultra Low Power Associative Computing with Spin Neurons and Resistive Crossbar Memory,” in DAC, 2013

  152. [159]

    NoM: Network-on-Memory for Inter- Bank Data Transfer in Highly-Banked Memories,

    S. H. S. Rezaei, M. Modarressi, R. Ausavarungnirun, M. Sadrosadati, O. Mutlu, and M. Daneshtalab, “NoM: Network-on-Memory for Inter- Bank Data Transfer in Highly-Banked Memories,” CAL, 2020

  153. [160]

    ParaBit: Processing Parallel Bitwise Operations in NAND Flash Memory Based SSDs,

    C. Gao, X. Xin, Y . Lu, Y . Zhang, J. Yang, and J. Shu, “ParaBit: Processing Parallel Bitwise Operations in NAND Flash Memory Based SSDs,” in MICRO, 2021

  154. [161]

    An In-Flash Binary Neural Network Accelerator with SLC NAND Flash Array,

    W. H. Choi, P.-F. Chiu, W. Ma, G. Hemink, T. T. Hoang, M. Lueker- Boden, and Z. Bandic, “An In-Flash Binary Neural Network Accelerator with SLC NAND Flash Array,” in ISCAS, 2020

  155. [162]

    A Novel Convolution Computing Paradigm Based on NOR Flash Array with High Computing Speed and Energy Efficiency,

    R. Han, P. Huang, Y . Xiang, C. Liu, Z. Dong, Z. Su, Y . Liu, L. Liu, X. Liu, and J. Kang, “A Novel Convolution Computing Paradigm Based on NOR Flash Array with High Computing Speed and Energy Efficiency,” TCAS-I, 2019

  156. [163]

    High-Performance Mixed-Signal Neurocomputing with Nanoscale Floating-Gate Memory Cell Arrays,

    F. Merrikh-Bayat, X. Guo, M. Klachko, M. Prezioso, K. K. Likharev, and D. B. Strukov, “High-Performance Mixed-Signal Neurocomputing with Nanoscale Floating-Gate Memory Cell Arrays,” TNNLS, 2017

  157. [164]

    Three- Dimensional NAND Flash for Vector–Matrix Multiplication,

    P. Wang, F. Xu, B. Wang, B. Gao, H. Wu, H. Qian, and S. Yu, “Three- Dimensional NAND Flash for Vector–Matrix Multiplication,” TVLSI, 2018

  158. [165]

    Lue, P.-K

    H.-T. Lue, P.-K. Hsu, M.-L. Wei, T.-H. Yeh, P.-Y . Du, W.-C. Chen, K.-C. Wang, and C.-Y . Lu, “Optimal Design Methods to Transform 3D NAND Flash into a High-Density, High-Bandwidth and Low-Power Nonvolatile Computing in Memory (nvCIM) Accelerator for Deep-Learning Neural Netwo...

  159. [166]

    Behemoth: A Flash-Centric Training Accelerator for Extreme-Scale DNNs,

    S. Kim, Y . Jin, G. Sohn, J. Bae, T. J. Ham, and J. W. Lee, “Behemoth: A Flash-Centric Training Accelerator for Extreme-Scale DNNs,” in FAST, 2021

  160. [167]

    MemCore: Computing-in-Flash Design for Deep Neural Network Acceleration,

    S. Wang, “MemCore: Computing-in-Flash Design for Deep Neural Network Acceleration,” in EDTM, 2022

  161. [168]

    Flash Memory Array for Efficient Implementation of Deep Neural Networks,

    R. Han, Y . Xiang, P. Huang, Y . Shan, X. Liu, and J. Kang, “Flash Memory Array for Efficient Implementation of Deep Neural Networks,” Adv. Intell. Syst. , 2021

  162. [169]

    S-FLASH: A NAND Flash-Based Deep Neural Network Accelerator Exploiting Bit-Level Sparsity,

    M. Kang, H. Kim, H. Shin, J. Sim, K. Kim, and L.-S. Kim, “S-FLASH: A NAND Flash-Based Deep Neural Network Accelerator Exploiting Bit-Level Sparsity,” TC, 2021

  163. [170]

    Neuromorphic Computing Using NAND Flash Memory Architecture with Pulse Width Modulation Scheme,

    S.-T. Lee and J.-H. Lee, “Neuromorphic Computing Using NAND Flash Memory Architecture with Pulse Width Modulation Scheme,” Front. Neurosci., 2020

  164. [171]

    3D-FPIM: An Extreme Energy-Efficient DNN Acceleration System Using 3D NAND Flash-Based In-Situ PIM Unit,

    H. Lee, M. Kim, D. Min, J. Kim, J. Back, H. Yoo, J.-H. Lee, and J. Kim, “3D-FPIM: An Extreme Energy-Efficient DNN Acceleration System Using 3D NAND Flash-Based In-Situ PIM Unit,” in MICRO, 2022

  165. [172]

    A Dual-Split 6T SRAM-Based Computing-in-Memory Unit-Macro with Fully Parallel Product-Sum Operation for Binarized DNN Edge Processors,

    X. Si, W.-S. Khwa, J.-J. Chen, J.-F. Li, X. Sun, R. Liu, S. Yu, H. Yamauchi, Q. Li, and M.-F. Chang, “A Dual-Split 6T SRAM-Based Computing-in-Memory Unit-Macro with Fully Parallel Product-Sum Operation for Binarized DNN Edge Processors,” TCAS-I, 2019

  166. [173]

    BLADE: An In-Cache Computing Architecture for Edge Devices,

    W. A. Simon, Y . M. Qureshi, M. Rios, A. Levisse, M. Zapater, and D. Atienza, “BLADE: An In-Cache Computing Architecture for Edge Devices,” TC, 2020

  167. [174]

    GenCache: Leveraging In-Cache Operators for Efficient Sequence Alignment,

    A. Nag, C. Ramachandra, R. Balasubramonian, R. Stutsman, E. Gi- acomin, H. Kambalasubramanyam, and P.-E. Gaillardon, “GenCache: Leveraging In-Cache Operators for Efficient Sequence Alignment,” in MICRO, 2019

  168. [175]

    Bit Prudent In- Cache Acceleration of Deep Convolutional Neural Networks,

    X. Wang, J. Yu, C. Augustine, R. Iyer, and R. Das, “Bit Prudent In- Cache Acceleration of Deep Convolutional Neural Networks,” in HPCA, 2019

  169. [176]

    Towards a Reconfigurable Bit-Serial/Bit-Parallel Vector Accelerator Using In-Situ Processing-in-SRAM,

    K. Al-Hawaj, O. Afuye, S. Agwa, A. Apsel, and C. Batten, “Towards a Reconfigurable Bit-Serial/Bit-Parallel Vector Accelerator Using In-Situ Processing-in-SRAM,” in ISCAS, 2020

  170. [177]

    An Energy-Efficient VLSI Architecture for Pattern Recognition via Deep Embedding of Computation in SRAM,

    M. Kang, M.-S. Keel, N. R. Shanbhag, S. Eilert, and K. Curewitz, “An Energy-Efficient VLSI Architecture for Pattern Recognition via Deep Embedding of Computation in SRAM,” in ICASSP, 2014

  171. [178]

    Colonnade: A Reconfig- urable SRAM-Based Digital Bit-Serial Compute-in-Memory Macro for Processing Neural Networks,

    H. Kim, T. Yoo, T. T.-H. Kim, and B. Kim, “Colonnade: A Reconfig- urable SRAM-Based Digital Bit-Serial Compute-in-Memory Macro for Processing Neural Networks,” JSSC, 2021

  172. [179]

    C3SRAM: An In-Memory- Computing SRAM Macro Based on Robust Capacitive Coupling Computing Mechanism,

    Z. Jiang, S. Yin, J.-S. Seo, and M. Seok, “C3SRAM: An In-Memory- Computing SRAM Macro Based on Robust Capacitive Coupling Computing Mechanism,” JSSC, 2020. 8

  173. [180]

    A 28 nm Configurable Memory (TCAM/BCAM/SRAM) Using Push-Rule 6T Bit Cell Enabling Logic-in-Memory,

    S. Jeloka, N. B. Akesh, D. Sylvester, and D. Blaauw, “A 28 nm Configurable Memory (TCAM/BCAM/SRAM) Using Push-Rule 6T Bit Cell Enabling Logic-in-Memory,” JSSC, 2016

  174. [181]

    Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory Fusion,

    Z. Wang, C. Liu, A. Arora, L. John, and T. Nowatzki, “Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory Fusion,” in ASPLOS, 2023

  175. [182]

    Energy-Efficient and High Throughput Sparse Distributed Memory Architecture,

    M. Kang, E. P. Kim, M.-s. Keel, and N. R. Shanbhag, “Energy-Efficient and High Throughput Sparse Distributed Memory Architecture,” in ISCAS, 2015

  176. [183]

    DUAL: Acceleration of Clustering Algorithms Using Digital-Based Processing in-Memory,

    M. Imani, S. Pampana, S. Gupta, M. Zhou, Y . Kim, and T. Rosing, “DUAL: Acceleration of Clustering Algorithms Using Digital-Based Processing in-Memory,” in MICRO, 2020

  177. [184]

    Low-Cost Inter-Linked Subarrays (LISA): Enabling Fast Inter-Subarray Data Movement in DRAM,

    K. K. Chang, P. J. Nair, D. Lee, S. Ghose, M. K. Qureshi, and O. Mutlu, “Low-Cost Inter-Linked Subarrays (LISA): Enabling Fast Inter-Subarray Data Movement in DRAM,” in HPCA, 2016

  178. [185]

    SIMDRAM: A Framework for Bit-Serial SIMD Processing Using DRAM,

    N. Hajinazar, G. F. Oliveira, S. Gregorio, J. D. Ferreira, N. M. Ghiasi, M. Patel, M. Alser, S. Ghose, J. Gómez-Luna, and O. Mutlu, “SIMDRAM: A Framework for Bit-Serial SIMD Processing Using DRAM,” in ASPLOS, 2021

  179. [186]

    LAcc: Exploiting Lookup Table-Based Fast and Accurate Vector Multiplication in DRAM-Based CNN Accelerator,

    Q. Deng, Y . Zhang, M. Zhang, and J. Yang, “LAcc: Exploiting Lookup Table-Based Fast and Accurate Vector Multiplication in DRAM-Based CNN Accelerator,” in DAC, 2019

  180. [187]

    Look-Up-Table Based Processing- in-Memory Architecture with Programmable Precision-Scaling for Deep Learning Applications,

    P. R. Sutradhar, S. Bavikadi, M. Connolly, S. Prajapati, M. A. Indovina, S. M. P. Dinakarrao, and A. Ganguly, “Look-Up-Table Based Processing- in-Memory Architecture with Programmable Precision-Scaling for Deep Learning Applications,” TPDS, 2021

  181. [188]

    pPIM: A Programmable Processor-in- Memory Architecture with Precision-Scaling for Deep Learning,

    P. R. Sutradhar, M. Connolly, S. Bavikadi, S. M. P. Dinakarrao, M. A. Indovina, and A. Ganguly, “pPIM: A Programmable Processor-in- Memory Architecture with Precision-Scaling for Deep Learning,” CAL, 2020

  182. [189]

    Fulcrum: A Simplified Control and Access Mechanism Toward Flexible and Practical In-Situ Accelerators,

    M. Lenjani, P. Gonzalez, E. Sadredini, S. Li, Y . Xie, A. Akel, S. Eilert, M. R. Stan, and K. Skadron, “Fulcrum: A Simplified Control and Access Mechanism Toward Flexible and Practical In-Situ Accelerators,” in HPCA, 2020

  183. [190]

    CHOPPER: A Compiler Infras- tructure for Programmable Bit-Serial SIMD Processing Using Memory In DRAM,

    X. Peng, Y . Wang, and M.-C. Yang, “CHOPPER: A Compiler Infras- tructure for Programmable Bit-Serial SIMD Processing Using Memory In DRAM,” in HPCA, 2023

  184. [191]

    DaPPA: A Data-Parallel Framework for Processing-in-Memory Archi- tectures,

    G. F. Oliveira, A. Kohli, D. Novo, J. Gómez-Luna, and O. Mutlu, “DaPPA: A Data-Parallel Framework for Processing-in-Memory Archi- tectures,” arXiv:2310.10168, 2023

  185. [192]

    Methodologies, Workloads, and Tools for Processing-in-Memory: Enabling the Adoption of Data-Centric Architectures,

    G. F. Oliveira, J. Gómez-Luna, S. Ghose, and O. Mutlu, “Methodologies, Workloads, and Tools for Processing-in-Memory: Enabling the Adoption of Data-Centric Architectures,” in ISVLSI, 2022

  186. [193]

    Heterogeneous Data-Centric Architectures for Modern Data-Intensive Applications: Case Studies in Machine Learning and Databases,

    G. F. Oliveira, A. Boroumand, S. Ghose, J. Gómez-Luna, and O. Mutlu, “Heterogeneous Data-Centric Architectures for Modern Data-Intensive Applications: Case Studies in Machine Learning and Databases,” in ISVLSI, 2022

  187. [194]

    Swordfish: A Framework for Evaluating Deep Neural Network-Based Basecalling Using Computation- In-Memory with Non-Ideal Memristors,

    T. Shahroodi, G. Singh, M. Zahedi, H. Mao, J. Lindegger, C. Firtina, S. Wong, O. Mutlu, and S. Hamdioui, “Swordfish: A Framework for Evaluating Deep Neural Network-Based Basecalling Using Computation- In-Memory with Non-Ideal Memristors,” in MICRO, 2023

  188. [195]

    SimplePIM: A Software Framework for Productive and Efficient Processing-In- Memory,

    J. Chen, J. Gómez-Luna, I. E. Hajj, Y . Guo, and O. Mutlu, “SimplePIM: A Software Framework for Productive and Efficient Processing-In- Memory,” in PACT, 2023

  189. [196]

    Evaluating Homomorphic Operations on a Real-World Processing-In- Memory System,

    H. Gupta, M. Kabra, J. Gómez-Luna, K. Kanellopoulos, and O. Mutlu, “Evaluating Homomorphic Operations on a Real-World Processing-In- Memory System,” in IISWC, 2023

  190. [197]

    Evaluating Machine LearningWork- loads on Memory-Centric Computing Systems,

    J. Gómez-Luna, Y . Guo, S. Brocard, J. Legriel, R. Cimadomo, G. F. Oliveira, G. Singh, and O. Mutlu, “Evaluating Machine LearningWork- loads on Memory-Centric Computing Systems,” in ISPASS, 2023

  191. [198]

    TransPimLib: Efficient Transcendental Functions for Processing-in-Memory Systems,

    M. Item, J. Gómez-Luna, Y . Guo, G. F. Oliveira, M. Sadrosadati, and O. Mutlu, “TransPimLib: Efficient Transcendental Functions for Processing-in-Memory Systems,” in ISPASS, 2023

  192. [199]

    A Framework for High-Throughput Sequence Alignment Using Real Processing-In-Memory Systems,

    S. Diab, A. Nassereldine, M. Alser, J. Gómez Luna, O. Mutlu, and I. El Hajj, “A Framework for High-Throughput Sequence Alignment Using Real Processing-In-Memory Systems,” Bioinformatics, 2023

  193. [200]

    GenPIP: In-Memory Acceleration of Genome Analysis via Tight Integration of Basecalling and Read Mapping,

    H. Mao, M. Alser, M. Sadrosadati, C. Firtina, A. Baranwal, D. S. Cali, A. Manglik, N. A. Alserr, and O. Mutlu, “GenPIP: In-Memory Acceleration of Genome Analysis via Tight Integration of Basecalling and Read Mapping,” in MICRO, 2022

  194. [201]

    Accelerating Weather Prediction Using Near-Memory Reconfigurable Fabric,

    G. Singh, D. Diamantopoulos, J. Gómez-Luna, C. Hagleitner, S. Stuijk, H. Corporaal, and O. Mutlu, “Accelerating Weather Prediction Using Near-Memory Reconfigurable Fabric,” TRETS, 2022

  195. [202]

    Cellular logic-in-memory arrays,

    W. H. Kautz, “Cellular logic-in-memory arrays,” IEEE ToC, 1969

  196. [203]

    A Logic-in-Memory Computer,

    H. S. Stone, “A Logic-in-Memory Computer,” IEEE ToC, 1970

  197. [204]

    HMC Specification Rev. 2.0,

    HMC Consortium, “HMC Specification Rev. 2.0,” www.hybridmemory cube.org/

  198. [205]

    A 1.2V 8Gb 8-Channel 128GB/s High-Bandwidth Memory (HBM) Stacked DRAM with Effective Microbump I/O Test Methods Using 29nm Process and TSV,

    D. U. Lee, K. W. Kim, K. W. Kim, H. Kim, J. Y . Kim, Y . J. Park, J. H. Kim, D. S. Kim, H. B. Park, J. W. Shin et al. , “A 1.2V 8Gb 8-Channel 128GB/s High-Bandwidth Memory (HBM) Stacked DRAM with Effective Microbump I/O Test Methods Using 29nm Process and TSV,” in ISSCC, 2014

  199. [206]

    Simultane- ous Multi-Layer Access: Improving 3D-Stacked Memory Bandwidth at Low Cost,

    D. Lee, S. Ghose, G. Pekhimenko, S. Khan, and O. Mutlu, “Simultane- ous Multi-Layer Access: Improving 3D-Stacked Memory Bandwidth at Low Cost,” TACO, 2016

  200. [207]

    JEDEC, JESD23-5D: High Bandwidth Memory (HBM) DRAM Standard, 2021

  201. [208]

    JEDEC, JESD23-8A: High Bandwidth Memory (HBM3) DRAM Stan- dard, 2021

  202. [209]

    Present and Future, Challenges of High Bandwith Memory (HBM),

    K. Kim and M.-j. Park, “Present and Future, Challenges of High Bandwith Memory (HBM),” in IMW, 2024

  203. [210]

    MATSA: An MRAM-Based Energy-Efficient Accelerator for Time Series Analysis,

    I. Fernandez, C. Giannoula, A. Manglik, R. Quislant, N. M. Ghiasi, J. Gómez-Luna, E. Gutierrez, O. Plata, and O. Mutlu, “MATSA: An MRAM-Based Energy-Efficient Accelerator for Time Series Analysis,” IEEE Access, 2024

  204. [211]

    Design Space Exploration for PIM Architectures in 3D-Stacked Memories,

    J. P. C. de Lima, P. C. Santos, M. A. Alves, A. Beck, and L. Carro, “Design Space Exploration for PIM Architectures in 3D-Stacked Memories,” in CF, 2018

  205. [212]

    BlueDBM: An appliance for big data analytics,

    S.-W. Jun, M. Liu, S. Lee, J. Hicks, J. Ankcorn, M. King, S. Xu, and Arvind, “BlueDBM: An appliance for big data analytics,” ISCA, 2015

  206. [213]

    Field-Effect Transistor Memory,

    R. H. Dennard, “Field-Effect Transistor Memory,” U.S. Patent 3,387,286, 1968

  207. [214]

    SwiftRL: Towards Efficient Reinforcement Learning on Real Processing-In- Memory Systems,

    K. Gogineni, S. S. Dayapule, J. Gómez-Luna, K. Gogineni, P. Wei, T. Lan, M. Sadrosadati, O. Mutlu, and G. Venkataramani, “SwiftRL: Towards Efficient Reinforcement Learning on Real Processing-In- Memory Systems,” in ISPASS, 2024

  208. [215]

    SecNDP: Secure Near-Data Processing with Untrusted Memory,

    W. Xiong, L. Ke, D. Jankov, M. Kounavis, X. Wang, E. Northup, J. A. Yang, B. Acun, C.-J. Wu, P. T. P. Tang et al., “SecNDP: Secure Near-Data Processing with Untrusted Memory,” in HPCA, 2022

  209. [216]

    Invisimem: Smart Memory Defenses For Memory Bus Side Channel,

    S. Aga and S. Narayanasamy, “Invisimem: Smart Memory Defenses For Memory Bus Side Channel,” ISCA, 2017

  210. [217]

    FracDRAM: Fractional Values in Off-the-Shelf DRAM,

    F. Gao, G. Tziantzioulis, and D. Wentzlaff, “FracDRAM: Fractional Values in Off-the-Shelf DRAM,” in MICRO, 2022

  211. [218]

    DRAM Bender: An Extensible and Versatile FPGA-based Infrastructure to Easily Test State-of-the-art DRAM Chips,

    A. Olgun, H. Hassan, A. G. Ya ˘glıkçı, Y . C. Tu˘grul, L. Orosa, H. Luo, M. Patel, O. Ergin, and O. Mutlu, “DRAM Bender: An Extensible and Versatile FPGA-based Infrastructure to Easily Test State-of-the-art DRAM Chips,” TCAD, 2023

  212. [219]

    MIMDRAM: An End-to-End Processing- Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data Comput- ing,

    G. F. Oliveira, A. Olgun, A. G. G. Yaglikçi, N. Bostanci, J. Gómez-Luna, S. Ghose, and O. Mutlu, “MIMDRAM: An End-to-End Processing- Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data Comput- ing,” in HPCA, 2024

  213. [220]

    Functionally-Complete Boolean Logic in Real DRAM Chips: Experi- mental Characterization and Analysis,

    I. E. Yuksel, Y . C. Tugrul, A. Olgun, F. N. Bostanci, A. G. Yaglikci, G. F. de Oliveira, H. Luo, J. G. Luna, M. Sadrosadati, and O. Mutlu, “Functionally-Complete Boolean Logic in Real DRAM Chips: Experi- mental Characterization and Analysis,” in HPCA, 2024

  214. [221]

    Simultaneous Many-Row Activation in Off-the-Shelf DRAM Chips: Experimental Characterization and Analysis,

    I. E. Yuksel, Y . C. Tugrul, F. N. Bostanci, G. F. de Oliveira, A. G. Yaglikci, A. Olgun, M. Soysal, H. Luo, J. G. Luna, M. Sadrosadati, and O. Mutlu, “Simultaneous Many-Row Activation in Off-the-Shelf DRAM Chips: Experimental Characterization and Analysis,” in DSN, 2024

  215. [223]

    PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips,

    I. E. Yuksel, Y . C. Tugrul, F. N. Bostanci, A. G. Yaglikci, A. Olgun, G. F. Oliveira, M. Soysal, H. Luo, J. G. Luna, M. Sadrosadati, and O. Mutlu, “PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips,” arXiv:2312.02...

  216. [224]

    Very High-Speed Computing Systems,

    M. J. Flynn, “Very High-Speed Computing Systems,” Proc. IEEE, 1966

  217. [225]

    6th Generation Intel Core Processor Family Datasheet,

    Intel Corp., “6th Generation Intel Core Processor Family Datasheet,” http://www.intel.com/content/www/us/en/processors/core/

  218. [226]

    NVIDIA A100 Tensor Core GPU Architecture,

    NVIDIA, “NVIDIA A100 Tensor Core GPU Architecture,” https://t.ly /rMUgA, 2020

  219. [227]

    SoftMC: A Flexible and Practical Open-Source Infrastructure for Enabling Experimental DRAM Studies,

    H. Hassan, N. Vijaykumar, S. Khan, S. Ghose, K. Chang, G. Pekhi- menko, D. Lee, O. Ergin, and O. Mutlu, “SoftMC: A Flexible and Practical Open-Source Infrastructure for Enabling Experimental DRAM Studies,” in HPCA, 2017

  220. [228]

    A Low-Cost Reduced-Latency DRAM Architecture with Dynamic Reconfiguration of Row Decoder,

    F. Bai, S. Wang, X. Jia, Y . Guo, B. Yu, H. Wang, C. Lai, Q. Ren, and H. Sun, “A Low-Cost Reduced-Latency DRAM Architecture with Dynamic Reconfiguration of Row Decoder,” TVLSI, 2022

  221. [229]

    N. H. Weste and D. Harris, CMOS VLSI Design: A Circuits and Systems Perspective. Pearson Education India, 2015

  222. [230]

    High-Performance Low-Power Selective Precharge Schemes for Address Decoders,

    M. A. Turi and J. G. Delgado-Frias, “High-Performance Low-Power Selective Precharge Schemes for Address Decoders,” TCAS-II, 2008

  223. [231]

    180-2: Secure Hash Standard (SHS),

    National Institute of Standards and Technology (NIST), “180-2: Secure Hash Standard (SHS),” Federal Information Processing Standards Publications, 2012

  224. [232]

    A Statistical Test Suite for Random and Pseudorandom Number Generators for Cryptographic Applications,

    L. Bassham, A. Rukhin, J. Soto, J. Nechvatal, M. Smid, S. Leigh, M. Levenson, M. Vangel, N. Heckert, and D. Banks, “A Statistical Test Suite for Random and Pseudorandom Number Generators for Cryptographic Applications,” NIST SP, 2010

  225. [233]

    JEDEC, JESD79-4C: DDR4 SDRAM , 2017

  226. [234]

    Memory-Centric Computing: Recent Advances in Processing-in-DRAM (Invited),

    O. Mutlu, G. F. Oliveira, A. Olgun, and I. E. Yuksel, “Memory-Centric Computing: Recent Advances in Processing-in-DRAM (Invited),” in IEDM, 2024. 9

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.