REVIEW 3 major objections 6 minor 234 references
Memory-Centric Computing: Recent Advances in Processing-in-DRAM
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Processing-in-DRAM turns DRAM into a compute substrate, with unmodified chips performing Boolean logic at 94–99% success.
desk verdict A clear, well-illustrated survey of the authors' own Processing-in-DRAM work, but the COTS 'high reliability' claim outruns the data: 95% per-operation success is undependable without an error model. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is charge sharing on DRAM bitlines during simultaneous row activation. In an Ambit-style triple-row activation, three cells connected to one bitline share charge and a sense amplifier resolves the majority, yielding AND and OR; in the COTS experiments, violating tRP lets a second activate interrupt precharge so the same charge-sharing principle performs NOT, NAND/NOR, and Multi-RowCopy on unmodified chips. MIMDRAM's additional machinery is mat-level isolation: mat isolation transistors, row decoder latches, and mat selectors let different mats in a subarray run different instructions, with inter-mat and intra-mat interconnects moving data between mats. Sectored DRAM's machinery is wordline segmentation with sector latches and transistors to activate a chosen subset of mats.
What would settle it
Run the NOT, AND/NAND, and Multi-RowCopy command sequences from Section V on a large sample of commercial DRAM chips, comparing every result against a known-good pattern across extended temperature, voltage, and chip-to-chip variation; if the measured success rate falls below what the paper claims or below a level usable without correction, the claim of reliable unmodified-chip computation is falsified.
Extended reading notes
Core claim
The paper's central claim is that Processing-in-DRAM has moved from proposal to demonstrated capability. First, MIMDRAM shows a hardware/software co-designed system in which each DRAM mat within a subarray can independently execute different bit-serial SIMD operations; evaluating twelve real applications and 495 multi-programmed mixes, it reports 13.2x/0.22x/173x the performance of CPU/GPU/SIMDRAM baselines, and 15.6x the SIMD utilization of SIMDRAM. Second, experiments on 224 and 120 commercial off-the-shelf DRAM chips show that by issuing back-to-back ACT/PRE commands with tRP below the manufacturer specification (e.g., <3ns), a memory controller can make the chips perform functionally complete Boolean operations: NOT, NAND, NOR, AND, OR, and one-row-to-31-row copy, with average success rates of 94.94% to 99.98%. Third, Sectored DRAM splits wordlines so a single activate touches only a subset of mats, cutting DRAM energy by up to 33% and improving performance by up to 36% for data-intensive workloads.
Load-bearing premise
The central claim of dependable computation on unmodified DRAM chips rests on the assumption that deliberately violating the manufacturer's timing parameters produces correct results reliably enough for real workloads; the paper reports 94.94% to 99.98% success rates and does not propose an error-detection or correction mechanism for the failed operations.
Editorial extensions
If this is right
- If MIMDRAM's results stand, bulk bitwise operations in DRAM can be programmed at the granularity of vectorizing compilers, making in-memory computation usable for real workloads without hand-tuning.
- If COTS DRAM chips truly perform NOT/NAND/NOR at above 94% success under timing violation, then a memory controller with no DRAM modification can implement functionally complete Boolean logic on existing hardware.
- The demonstrated Multi-RowCopy, one row copied concurrently into up to 31 rows, implies that bulk data copy and initialization can be accelerated inside DRAM, and may also offer a defense against cold-boot attacks.
- Sectored DRAM's fine-grained activation implies that the energy cost of over-fetching can be cut by roughly a fifth to a third in data-intensive workloads, which directly improves the efficiency of Processing-in-DRAM systems that rely on row activation.
- Together, the three advances imply that the memory controller and DRAM interface should be redesigned to expose row-level and mat-level computation commands rather than treating DRAM solely as a byte-addressable store.
Reading between the lines
- The paper does not specify an error-detection or correction mechanism for COTS DRAM computation; one implicit consequence is that practical use would need to pair these operations with ECC or re-execution, which would reduce the net gains.
- The COTS success rates are averages across tested chips; the paper leaves open whether a production system could tolerate the worst-case chips, so the strongest defensible version of the claim is that the capability exists, not that it is already dependable at scale.
- A natural extension is to combine the two approaches: use MIMDRAM-style mat-level control with COTS-style timing-violation operations to build a hybrid that keeps the programmability of the former and the zero-modification cost of the latter.
- If the fine-grained activation of Sectored DRAM is combined with in-memory compute, the energy savings from avoiding over-fetch would compound the savings from avoiding data movement, a combination the paper does not quantify.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This invited overview paper summarizes recent Processing-in-DRAM (PiD) advances from the authors' group. It describes MIMDRAM, a hardware/software co-designed Processing-using-DRAM system that enables fine-grained, multiple-instruction multiple-data execution within DRAM mats; experimental results on commercial off-the-shelf (COTS) DRAM chips showing NOT, AND/NAND/OR/NOR, many-row activation, Multi-RowCopy, and true random number generation; and Sectored DRAM, a fine-grained DRAM architecture. The paper reports large performance and energy gains for MIMDRAM (13.2x/0.22x/173x performance versus CPU/GPU/SIMDRAM), high success rates (>94%) for COTS DRAM bulk-bitwise operations, and up to 33% DRAM energy savings for Sectored DRAM.
Significance. If the claims hold, the paper provides useful evidence that Processing-in-DRAM can reduce data-movement overheads: MIMDRAM's compiler/runtime support addresses programmability, the COTS experiments demonstrate functional completeness on unmodified chips, and Sectored DRAM targets access-granularity inefficiencies. The experimental breadth is a strength: the COTS results are based on hundreds of real chips, the MIMDRAM evaluation uses gem5 and CACTI against external baselines, and the paper explicitly identifies interconnect throughput as a limiter. However, the paper is a condensed summary of prior peer-reviewed works rather than a self-contained technical contribution, and the practical-usability claim for COTS DRAM computation is not fully supported by the reported success-rate statistics.
major comments (3)
- [Section V, Figs. 8-11] The paper states that COTS DRAM chips perform NOT, AND/NAND/OR/NOR, and Multi-RowCopy operations "at high success rates (>94%)" and describes these as "functionally-complete bulk-bitwise Boolean operations." The reported success rates are per-operation, not per-bit: for example, the average success rate for 16-input AND/NAND is 94.94%. A computation composed of k such operations has expected success roughly 0.95^k, so a 10-gate expression would be wrong in more than 40% of runs without correction. The paper does not report per-bit error rates, error correlation across rows/data patterns/temperature/voltage, whether errors are deterministic and profileable or random, or any detection/correction mechanism. The evidence supports a feasibility demonstration, but the claim that unmodified COTS DRAM can serve as a dependable substrate for real workloads is not supported. Please add an explicit error model, or restrict the claim to "feasibility" and discuss error mitigation (e.g., verification, recomputation, ECC).
- [Section IV, Fig. 6] The headline MIMDRAM results are compared against CPU, GPU, and SIMDRAM baselines, but the evaluation methodology for the CPU and GPU baselines is not described. The text only says that MIMDRAM and SIMDRAM are implemented in gem5 and evaluated with CACTI, and that 12 memory-bound applications were selected from five benchmark suites. It is unclear whether the CPU and GPU results come from real hardware, full-system simulation, or vendor simulators, and whether the 12 applications are identical across all four target systems. Without this information, the 13.2x/0.22x/173x performance and 0.0017x/0.00007x/0.004x energy ratios cannot be independently interpreted. Please specify the baseline configurations, simulation methodology, and whether the reported ratios are end-to-end or kernel-only.
- [Section VI] The Sectored DRAM section makes strong quantitative claims: "reduces the DRAM energy consumption of data intensive workloads by up to 33% (20% on average), while improving performance by up to 36% (17% on average)" and "system-wide energy savings of up to 23%." However, the section provides no evaluation setup: no workload list, no baseline DRAM configuration, no simulator description, and no sensitivity analysis. The reader cannot verify these numbers from this paper. Please add a concise methodology summary, or state explicitly that the results are taken from [222] and summarize the setup used there.
minor comments (6)
- [Section V, Fig. 11 caption] The caption says the results were "measured in 224, 224, and 120 COTS DRAM chips" without spelling out which operation corresponds to which count; please clarify the mapping (e.g., NOT: 224, AND/NAND/OR/NOR: 224, Multi-RowCopy: 120).
- [Section V, Fig. 8 and surrounding text] The phrase "with violated manufacturer-recommended tRP timing, e.g., <3ns" is awkward; it should read "with tRP set below the manufacturer-recommended value (e.g., <3ns)."
- [Section IV, Fig. 6] The notation "13.2×/0.22×/173× the performance" is ambiguous because the baseline order is only given later; spell out "relative to CPU/GPU/SIMDRAM" in the caption or text.
- [Section V, Fig. 12] The figure describes a TRNG command sequence with ACT R0 and ACT R3 but says this "activates four rows simultaneously"; please explain how four rows are activated, since the text otherwise implies a single-row ACT activates only one row.
- [References] The reference list contains several typos, e.g., [5] "V olos" should be "Volos," [64] "opthers" should be "others," and [138] "Boroum" should be "Boroumand." A careful proofread is needed.
- [Section III, first paragraph] The sentence "First, enabling widespread use of PiM on a wide variety of important workloads requires PiM systems to i) be easy to program and seamlessly compile workloads into" is grammatically incomplete; please revise.
Circularity Check
No significant circularity: the paper is a review that reports experimental and simulation results from cited prior works, with no fitted parameter or definitional reduction.
full rationale
The paper is an invited review summarizing the authors' prior peer-reviewed work on Processing-in-DRAM. Its central quantitative claims are measurements or simulations taken from the cited original papers: MIMDRAM speedups are obtained from gem5 and CACTI evaluations described in [219]; COTS DRAM operation success rates are experimental characterizations reported in [220,221,223]; Sectored DRAM energy/performance numbers are simulation results from [222]. None of these numbers is derived from a parameter fitted in this paper, and none is defined in terms of the conclusion it supports. The paper does not invoke a uniqueness theorem, does not smuggle an ansatz through a citation, and does not rename an empirical pattern as a new derivation. The only notable concern is that the COTS DRAM >94% success-rate claim lacks a per-bit error model and error-correction mechanism, which makes the practical reliability conclusion under-supported; however, that is a correctness/completeness issue, not circularity. Heavy self-citation is present, but per the review rules self-citation is not circularity when, as here, the cited works are independent experimental and simulation studies with external baselines (CPU, GPU, SIMDRAM) and the claims do not reduce to the survey's own definitions.
Assumptions & free parameters
assumptions (3)
- domain assumption DRAM charge-sharing and sense-amplifier behavior can be reliably exploited by violating timing parameters
- domain assumption Simulation models (gem5, CACTI) accurately capture MIMDRAM and Sectored DRAM behavior
- domain assumption The 12 selected applications are representative of data-intensive workloads
invented entities (2)
-
MIMDRAM hardware modifications (mat isolation transistor, row decoder latch, mat selector, inter-mat and intra-mat interconnects)
-
Sectored DRAM sector latches, sector transistors, and additional local wordline drivers
Cite this review
Pith. "Pith review of Memory-Centric Computing: Recent Advances in Processing-in-DRAM." pith.science (2026). https://pith.science/paper/HL5XMXY6
@misc{pith2026241219275,
author = {Pith},
title = {Pith review of: Memory-Centric Computing: Recent Advances in Processing-in-DRAM},
year = {2026},
howpublished = {\url{https://pith.science/paper/HL5XMXY6}},
note = {Machine review of arXiv:2412.19275}
}
read the original abstract
Memory-centric computing aims to enable computation capability in and near all places where data is generated and stored. As such, it can greatly reduce the large negative performance and energy impact of data access and data movement, by 1) fundamentally avoiding data movement, 2) reducing data access latency & energy, and 3) exploiting large parallelism of memory arrays. Many recent studies show that memory-centric computing can largely improve system performance & energy efficiency. Major industrial vendors and startup companies have recently introduced memory chips with sophisticated computation capabilities. Going forward, both hardware and software stack should be revisited and designed carefully to take advantage of memory-centric computing. This work describes several major recent advances in memory-centric computing, specifically in Processing-in-DRAM, a paradigm where the operational characteristics of a DRAM chip are exploited and enhanced to perform computation on data stored in DRAM. Specifically, we describe 1) new techniques that slightly modify DRAM chips to enable both enhanced computation capability and easier programmability, 2) new experimental studies that demonstrate the functionally-complete bulk-bitwise computational capability of real commercial off-the-shelf DRAM chips, without any modifications to the DRAM chip or the interface, and 3) new DRAM designs that improve access granularity & efficiency, unleashing the true potential of Processing-in-DRAM.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[222]
Sectored DRAM: An Energy-Efficient High-Throughput and Practical Fine-Grained DRAM Architecture,
A. Olgun, F. Bostanci, G. F. Oliveira, Y . C. Tugrul, R. Bera, A. G. Yaglikci, H. Hassan, O. Ergin, and O. Mutlu, “Sectored DRAM: An Energy-Efficient High-Throughput and Practical Fine-Grained DRAM Architecture,” TACO, 2024
work page 2024
-
[1]
Memory Scaling: A Systems Architecture Perspective,
O. Mutlu, “Memory Scaling: A Systems Architecture Perspective,” in IMW, 2013
2013
-
[2]
Research Problems and Opportunities in Memory Systems,
O. Mutlu and L. Subramanian, “Research Problems and Opportunities in Memory Systems,” SUPERFRI, 2014
2014
-
[3]
The Tail at Scale,
J. Dean and L. A. Barroso, “The Tail at Scale,” CACM, 2013
2013
-
[4]
Profiling a Warehouse-Scale Computer,
S. Kanev, J. P. Darago, K. Hazelwood, P. Ranganathan, T. Moseley, G.-Y . Wei, and D. Brooks, “Profiling a Warehouse-Scale Computer,” in ISCA, 2015
2015
-
[5]
Clearing the Clouds: A Study of Emerging Scale-Out Workloads on Modern Hardware,
M. Ferdman, A. Adileh, O. Kocberber, S. V olos, M. Alisafaee, D. Jevdjic, C. Kaynak, A. D. Popescu, A. Ailamaki, and B. Falsafi, “Clearing the Clouds: A Study of Emerging Scale-Out Workloads on Modern Hardware,” in ASPLOS, 2012
2012
-
[6]
BigDataBench: A Big Data Benchmark Suite from Internet Services,
L. Wang, J. Zhan, C. Luo, Y . Zhu, Q. Yang, Y . He, W. Gao, Z. Jia, Y . Shi, S. Zhang et al., “BigDataBench: A Big Data Benchmark Suite from Internet Services,” in HPCA, 2014
2014
-
[7]
Enabling Practical Processing in and near Memory for Data-Intensive Computing,
O. Mutlu, S. Ghose, J. Gómez-Luna, and R. Ausavarungnirun, “Enabling Practical Processing in and near Memory for Data-Intensive Computing,” in DAC, 2019
2019
Show all 234 references
-
[8]
Process- ing Data Where It Makes Sense: Enabling In-Memory Computation,
O. Mutlu, S. Ghose, J. Gómez-Luna, and R. Ausavarungnirun, “Process- ing Data Where It Makes Sense: Enabling In-Memory Computation,” MicPro, 2019
2019
-
[9]
Intelligent Architectures for Intelligent Machines,
O. Mutlu, “Intelligent Architectures for Intelligent Machines,” in VLSI- DAT, 2020
2020
-
[10]
Processing-in-Memory: A Workload-Driven Perspective,
S. Ghose, A. Boroumand, J. S. Kim, J. Gómez-Luna, and O. Mutlu, “Processing-in-Memory: A Workload-Driven Perspective,” IBM JRD, 2019
2019
-
[11]
A Modern Primer on Processing in Memory,
O. Mutlu, S. Ghose, J. Gómez-Luna, and R. Ausavarungnirun, “A Modern Primer on Processing in Memory,” in Emerging Computing: From Devices to Systems — Looking Beyond Moore and Von Neumann . Springer, 2022. 5
2022
-
[12]
DAMOV: A New Methodology and Benchmark Suite for Evaluating Data Movement Bottlenecks,
G. F. Oliveira, J. Gómez-Luna, L. Orosa, S. Ghose, N. Vijaykumar, I. Fernandez, M. Sadrosadati, and O. Mutlu, “DAMOV: A New Methodology and Benchmark Suite for Evaluating Data Movement Bottlenecks,” IEEE Access, 2021
2021
-
[13]
Google Workloads for Consumer Devices: Mitigating Data Movement Bottlenecks,
A. Boroumand, S. Ghose, Y . Kim, R. Ausavarungnirun, E. Shiu, R. Thakur, D. Kim, A. Kuusela, A. Knies, P. Ranganathan et al. , “Google Workloads for Consumer Devices: Mitigating Data Movement Bottlenecks,” in ASPLOS, 2018
2018
-
[14]
Google Neural Network Models for Edge Devices: Analyzing and Mitigating Machine Learning Inference Bottlenecks,
A. Boroumand, S. Ghose, B. Akin, R. Narayanaswami, G. F. Oliveira, X. Ma, E. Shiu, and O. Mutlu, “Google Neural Network Models for Edge Devices: Analyzing and Mitigating Machine Learning Inference Bottlenecks,” in PACT, 2021
2021
-
[15]
Reducing Data Movement Energy via Online Data Clustering and Encoding,
S. Wang and E. Ipek, “Reducing Data Movement Energy via Online Data Clustering and Encoding,” in MICRO, 2016
2016
-
[16]
Quantifying the Energy Cost of Data Movement for Emerging Smart Phone Workloads on Mobile Platforms,
D. Pandiyan and C.-J. Wu, “Quantifying the Energy Cost of Data Movement for Emerging Smart Phone Workloads on Mobile Platforms,” in IISWC, 2014
2014
-
[17]
EDEN: Enabling Energy-Efficient, High-Performance Deep Neural Network Inference Using Approximate DRAM,
S. Koppula, L. Orosa, A. G. Ya ˘glıkçı, R. Azizi, T. Shahroodi, K. Kanellopoulos, and O. Mutlu, “EDEN: Enabling Energy-Efficient, High-Performance Deep Neural Network Inference Using Approximate DRAM,” in MICRO, 2019
2019
-
[18]
Co-Architecting Controllers and DRAM to Enhance DRAM Process Scaling,
U. Kang, H.-S. Yu, C. Park, H. Zheng, J. Halbert, K. Bains, S. Jang, and J. S. Choi, “Co-Architecting Controllers and DRAM to Enhance DRAM Process Scaling,” in The Memory Forum, 2014
2014
-
[19]
Reflections on the Memory Wall,
S. A. McKee, “Reflections on the Memory Wall,” in CF, 2004
2004
-
[20]
The Memory Gap and the Future of High Performance Memories,
M. V . Wilkes, “The Memory Gap and the Future of High Performance Memories,” CAN, 2001
2001
-
[21]
A Case for Exploiting Subarray-Level Parallelism (SALP) in DRAM,
Y . Kim, V . Seshadri, D. Lee, J. Liu, and O. Mutlu, “A Case for Exploiting Subarray-Level Parallelism (SALP) in DRAM,” in ISCA, 2012
2012
-
[22]
Hitting the Memory Wall: Implications of the Obvious,
W. A. Wulf and S. A. McKee, “Hitting the Memory Wall: Implications of the Obvious,” CAN, 1995
1995
-
[23]
Demystifying Complex Workload–DRAM Interactions: An Experimental Study,
S. Ghose, T. Li, N. Hajinazar, D. S. Cali, and O. Mutlu, “Demystifying Complex Workload–DRAM Interactions: An Experimental Study,” in SIGMETRICS, 2020
2020
-
[24]
A Scalable Processing- in-Memory Accelerator for Parallel Graph Processing,
J. Ahn, S. Hong, S. Yoo, O. Mutlu, and K. Choi, “A Scalable Processing- in-Memory Accelerator for Parallel Graph Processing,” in ISCA, 2015
2015
-
[25]
PIM-Enabled Instructions: A Low-Overhead, Locality-Aware Processing-in-Memory Architecture,
J. Ahn, S. Yoo, O. Mutlu, and K. Choi, “PIM-Enabled Instructions: A Low-Overhead, Locality-Aware Processing-in-Memory Architecture,” in ISCA, 2015
2015
-
[26]
Transparent Offloading and Mapping (TOM) Enabling Programmer-Transparent Near-Data Processing in GPU Systems,
K. Hsieh, E. Ebrahimi, G. Kim, N. Chatterjee, M. O’Connor, N. Vijayku- mar, O. Mutlu, and S. W. Keckler, “Transparent Offloading and Mapping (TOM) Enabling Programmer-Transparent Near-Data Processing in GPU Systems,” in ISCA, 2016
2016
-
[27]
FIGARO: Improving System Performance via Fine-Grained In-DRAM Data Relocation and Caching,
Y . Wang, L. Orosa, X. Peng, Y . Guo, S. Ghose, M. Patel, J. S. Kim, J. G. Luna, M. Sadrosadati, N. M. Ghiasi et al., “FIGARO: Improving System Performance via Fine-Grained In-DRAM Data Relocation and Caching,” in MICRO, 2020
2020
-
[28]
It’s the Memory, Stupid!
R. Sites, “It’s the Memory, Stupid!” MPR, 1996
1996
-
[29]
Language Models are Few-Shot Learners,
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert- V oss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B...
2020
-
[30]
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in NAACL, 2019
2019
-
[31]
Accelerating Neural Network Inference with Processing-in-DRAM: From the Edge to the Cloud,
G. F. Oliveira, J. Gómez-Luna, S. Ghose, A. Boroumand, and O. Mutlu, “Accelerating Neural Network Inference with Processing-in-DRAM: From the Edge to the Cloud,” IEEE Micro, 2022
2022
-
[32]
NEUPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing,
G. Heo, S. Lee, J. Cho, H. Choi, S. Lee, H. Ham, G. Kim, D. Mahajan, and J. Park, “NEUPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing,” in ASPLOS, 2024
2024
-
[33]
TransPIM: A Memory-Based Acceleration via Software-Hardware Co-Design for Transformer,
M. Zhou, W. Xu, J. Kang, and T. Rosing, “TransPIM: A Memory-Based Acceleration via Software-Hardware Co-Design for Transformer,” in HPCA, 2022
2022
-
[34]
AttAcc! Unleashing the Power of PIM for Batched Transformer- based Generative Model Inference,
J. Park, J. Choi, K. Kyung, M. J. Kim, Y . Kwon, N. S. Kim, and J. H. Ahn, “AttAcc! Unleashing the Power of PIM for Batched Transformer- based Generative Model Inference,” in ASPLOS, 2024
2024
-
[35]
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System,
M. Seo, X. T. Nguyen, S. J. Hwang, Y . Kwon, G. Kim, C. Park, I. Kim, J. Park, J. Kim, W. Shin et al., “IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System,” in ASPLOS, 2024
2024
-
[36]
PIM-Opt: Demystifying Distributed Optimization Algorithms on a Real-World Processing-In-Memory System,
S. Rhyner, H. Luo, J. Gómez-Luna, M. Sadrosadati, J. Jiang, A. Olgun, H. Gupta, C. Zhang, and O. Mutlu, “PIM-Opt: Demystifying Distributed Optimization Algorithms on a Real-World Processing-In-Memory System,” in PACT, 2024
2024
-
[37]
Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching,
S. Yun, K. Kyung, J. Cho, J. Choi, J. Kim, B. Kim, S. Lee, K. Sohn, and J. H. Ahn, “Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching,” in MICRO, 2024
2024
-
[38]
Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System,
H. Jang, J. Song, J. Jung, J. Park, Y . Kim, and J. Lee, “Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System,” in HPCA, 2024
2024
-
[39]
Accelerating Genome Analysis: A Primer on an Ongoing Journey,
M. Alser, Z. Bingöl, D. S. Cali, J. Kim, S. Ghose, C. Alkan, and O. Mutlu, “Accelerating Genome Analysis: A Primer on an Ongoing Journey,” IEEE Micro, 2020
2020
-
[40]
FPGA-Based Near-Memory Acceleration of Modern Data-Intensive Applications,
G. Singh, M. Alser, D. S. Cali, D. Diamantopoulos, J. Gómez-Luna, H. Corporaal, and O. Mutlu, “FPGA-Based Near-Memory Acceleration of Modern Data-Intensive Applications,” IEEE Micro, 2021
2021
-
[41]
From Molecules to Genomic Varia- tions: Accelerating Genome Analysis via Intelligent Algorithms and Architectures,
M. Alser, J. Lindegger, C. Firtina, N. Almadhoun, H. Mao, G. Singh, J. Gomez-Luna, and O. Mutlu, “From Molecules to Genomic Varia- tions: Accelerating Genome Analysis via Intelligent Algorithms and Architectures,” CSBJ, 2022
2022
-
[42]
GRIM-Filter: Fast Seed Loca- tion Filtering in DNA Read Mapping Using Processing-in-Memory Technologies,
J. S. Kim, D. S. Cali, H. Xin, D. Lee, S. Ghose, M. Alser, H. Hassan, O. Ergin, C. Alkan, and O. Mutlu, “GRIM-Filter: Fast Seed Loca- tion Filtering in DNA Read Mapping Using Processing-in-Memory Technologies,” BMC Genomics, 2018
2018
-
[43]
GenStore: A High- Performance and Energy-Efficient In-Storage Computing System for Genome Sequence Analysis,
N. M. Ghiasi, J. Park, H. Mustafa, J. Kim, A. Olgun, A. Gollwitzer, D. S. Cali, C. Firtina, H. Mao, N. A. Alserr et al., “GenStore: A High- Performance and Energy-Efficient In-Storage Computing System for Genome Sequence Analysis,” in ASPLOS, 2022
2022
-
[44]
MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Anal- ysis with In-Storage Processing,
N. M. Ghiasi, M. Sadrosadati, H. Mustafa, A. Gollwitzer, C. Firtina, J. Eudine, H. Mao, J. Lindegger, M. B. Cavlak, M. Alser et al., “MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Anal- ysis with In-Storage Processing,” in ISCA, 2024
2024
-
[45]
GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis,
D. S. Cali, G. S. Kalsi, Z. Bingöl, C. Firtina, L. Subramanian, J. S. Kim, R. Ausavarungnirun, M. Alser, J. Gomez-Luna, A. Boroumand et al., “GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis,” in MICRO, 2020
2020
-
[46]
SeGraM: A Universal Hardware Accelerator for Genomic Sequence-to-Graph and Sequence-to-Sequence Mapping,
D. S. Cali, K. Kanellopoulos, J. Lindegger, Z. Bingöl, G. S. Kalsi, Z. Zuo, C. Firtina, M. B. Cavlak, J. Kim, N. M. Ghiasi et al., “SeGraM: A Universal Hardware Accelerator for Genomic Sequence-to-Graph and Sequence-to-Sequence Mapping,” in ISCA, 2022
2022
-
[47]
Nanopore Sequencing Technology and Tools for Genome Assembly: Computational Analysis of the Current State, Bottlenecks and Future Directions,
D. Senol Cali, J. S. Kim, S. Ghose, C. Alkan, and O. Mutlu, “Nanopore Sequencing Technology and Tools for Genome Assembly: Computational Analysis of the Current State, Bottlenecks and Future Directions,” Briefings in Bioinformatics , 2018
2018
-
[48]
NDA: Near-DRAM Acceleration Architecture Leveraging Commodity DRAM Devices and Standard Memory Modules,
A. Farmahini-Farahani, J. H. Ahn, K. Morrow, and N. S. Kim, “NDA: Near-DRAM Acceleration Architecture Leveraging Commodity DRAM Devices and Standard Memory Modules,” in HPCA, 2015
2015
-
[49]
JAFAR: Near-Data Processing for Databases,
O. O. Babarinsa and S. Idreos, “JAFAR: Near-Data Processing for Databases,” in SIGMOD, 2015
2015
-
[50]
The True Processing in Memory Accelerator,
F. Devaux, “The True Processing in Memory Accelerator,” in Hot Chips, 2019
2019
-
[51]
Benchmarking Memory-Centric Computing Systems: Analysis of Real Processing-in-Memory Hardware,
J. Gómez-Luna, I. El Hajj, I. Fernandez, C. Giannoula, G. F. Oliveira, and O. Mutlu, “Benchmarking Memory-Centric Computing Systems: Analysis of Real Processing-in-Memory Hardware,” in CUT, 2021
2021
-
[52]
Benchmarking a New Paradigm: Experimental Analysis and Characterization of a Real Processing-in-Memory System,
J. Gómez-Luna, I. El Hajj, I. Fernandez, C. Giannoula, G. F. Oliveira, and O. Mutlu, “Benchmarking a New Paradigm: Experimental Analysis and Characterization of a Real Processing-in-Memory System,” IEEE Access, 2022
2022
-
[53]
SynCron: Efficient Synchronization Support for Near-Data-Processing Architectures,
C. Giannoula, N. Vijaykumar, N. Papadopoulou, V . Karakostas, I. Fer- nandez, J. Gómez-Luna, L. Orosa, N. Koziris, G. Goumas, and O. Mutlu, “SynCron: Efficient Synchronization Support for Near-Data-Processing Architectures,” in HPCA, 2021
2021
-
[54]
NERO: A Near High-Bandwidth Memory Stencil Accelerator for Weather Prediction Modeling,
G. Singh, D. Diamantopoulos, C. Hagleitner, J. Gomez-Luna, S. Stuijk, O. Mutlu, and H. Corporaal, “NERO: A Near High-Bandwidth Memory Stencil Accelerator for Weather Prediction Modeling,” in FPL, 2020
2020
-
[55]
A 1ynm 1.25V 8Gb, 16Gb/s/pin GDDR6- Based Accelerator-in-Memory Supporting 1TFLOPS MAC Operation and Various Activation Functions for Deep-Learning Applications,
S. Lee, K. Kim, S. Oh, J. Park, G. Hong, D. Ka, K. Hwang, J. Park, K. Kang, J. Kim, J. Jeon, N. Kim, Y . Kwon, K. Vladimir, W. Shin, J. Won, M. Lee, H. Joo et al., “A 1ynm 1.25V 8Gb, 16Gb/s/pin GDDR6- Based Accelerator-in-Memory Supporting 1TFLOPS MAC Operation and Various Act...
2022
-
[56]
Near-Memory Processing in Action: Accelerating Personalized Recommendation with AxDIMM,
L. Ke, X. Zhang, J. So, J.-G. Lee, S.-H. Kang, S. Lee, S. Han, Y . Cho, J. H. Kim, Y . Kwon et al. , “Near-Memory Processing in Action: Accelerating Personalized Recommendation with AxDIMM,” IEEE Micro, 2021
2021
-
[57]
SparseP: Towards Efficient Sparse Matrix Vector Multipli- cation on Real Processing-in-Memory Architectures,
C. Giannoula, I. Fernandez, J. G. Luna, N. Koziris, G. Goumas, and O. Mutlu, “SparseP: Towards Efficient Sparse Matrix Vector Multipli- cation on Real Processing-in-Memory Architectures,” in SIGMETRICS, 2022
2022
-
[58]
McDRAM: Low Latency and Energy-Efficient Matrix Computations in DRAM,
H. Shin, D. Kim, E. Park, S. Park, Y . Park, and S. Yoo, “McDRAM: Low Latency and Energy-Efficient Matrix Computations in DRAM,” TCADICS, 2018
2018
-
[59]
McDRAM v2: In-Dynamic Random Access Memory Systolic Array Accelerator to Address the Large Model Problem in Deep Neural Networks on the Edge,
S. Cho, H. Choi, E. Park, H. Shin, and S. Yoo, “McDRAM v2: In-Dynamic Random Access Memory Systolic Array Accelerator to Address the Large Model Problem in Deep Neural Networks on the Edge,” IEEE Access, 2020
2020
-
[60]
Casper: Accelerating Stencil Computation using Near-Cache Processing,
A. Denzler, R. Bera, N. Hajinazar, G. Singh, G. F. Oliveira, J. Gómez- Luna, and O. Mutlu, “Casper: Accelerating Stencil Computation using Near-Cache Processing,” IEEE Access, 2023
2023
-
[61]
Chameleon: Versatile and Practical Near-DRAM Acceleration Archi- tecture for Large Memory Systems,
H. Asghari-Moghaddam, Y . H. Son, J. H. Ahn, and N. S. Kim, “Chameleon: Versatile and Practical Near-DRAM Acceleration Archi- tecture for Large Memory Systems,” in MICRO, 2016
2016
-
[62]
A Case for Intelligent RAM,
D. Patterson, T. Anderson, N. Cardwell et al., “A Case for Intelligent RAM,” IEEE Micro, 1997
1997
-
[63]
Computational RAM: Implementing Processors in Memory,
D. G. Elliott, M. Stumm, W. M. Snelgrove et al., “Computational RAM: Implementing Processors in Memory,” D&T, 1999. 6
1999
-
[64]
Saving Memory Movements Through Vector Processing in the DRAM,
M. A. Z. Alves, P. C. Santos, F. B. Moreira, and opthers, “Saving Memory Movements Through Vector Processing in the DRAM,” in CASES, 2015
2015
-
[65]
Beyond the Wall: Near-Data Processing for Databases,
S. L. Xi, O. Babarinsa, M. Athanassoulis, and S. Idreos, “Beyond the Wall: Near-Data Processing for Databases,” in DaMoN, 2015
2015
-
[66]
ABC-DIMM: Alleviating the Bottleneck of Communication in DIMM-Based Near-Memory Processing with Inter-DIMM Broadcast,
W. Sun, Z. Li, S. Yin, S. Wei, and L. Liu, “ABC-DIMM: Alleviating the Bottleneck of Communication in DIMM-Based Near-Memory Processing with Inter-DIMM Broadcast,” in ISCA, 2021
2021
-
[67]
GraphSSD: Graph Semantics Aware SSD,
K. K. Matam, G. Koo, H. Zha, H.-W. Tseng, and M. Annavaram, “GraphSSD: Graph Semantics Aware SSD,” in ISCA, 2019
2019
-
[68]
Processing in Memory: The Terasys Massively Parallel PIM Array,
M. Gokhale, B. Holmes, and K. Iobst, “Processing in Memory: The Terasys Massively Parallel PIM Array,” Computer, 1995
1995
-
[69]
Mapping Irregular Applications to DIV A, a PIM-Based Data-Intensive Architecture,
M. Hall, P. Kogge, J. Koller, P. Diniz, J. Chame, J. Draper, J. LaCoss, J. Granacki, J. Brockman, A. Srivastava et al. , “Mapping Irregular Applications to DIV A, a PIM-Based Data-Intensive Architecture,” in SC, 1999
1999
-
[70]
Opportunities and Challenges of Performing Vector Operations Inside the DRAM,
M. A. Z. Alves, P. C. Santos, M. Diener, and L. Carro, “Opportunities and Challenges of Performing Vector Operations Inside the DRAM,” in MEMSYS, 2015
2015
-
[71]
Livia: Data-Centric Computing Throughout the Memory Hierarchy,
E. Lockerman, A. Feldmann, M. Bakhshalipour, A. Stanescu, S. Gupta, D. Sanchez, and N. Beckmann, “Livia: Data-Centric Computing Throughout the Memory Hierarchy,” in ASPLOS, 2020
2020
-
[72]
GraphPIM: Enabling Instruction-Level PIM Offloading in Graph Computing Frameworks,
L. Nai, R. Hadidi, J. Sim, H. Kim, P. Kumar, and H. Kim, “GraphPIM: Enabling Instruction-Level PIM Offloading in Graph Computing Frameworks,” in HPCA, 2017
2017
-
[73]
LazyPIM: An Efficient Cache Coherence Mechanism for Processing-in-Memory,
A. Boroumand, S. Ghose, B. Lucia, K. Hsieh, K. Malladi, H. Zheng, and O. Mutlu, “LazyPIM: An Efficient Cache Coherence Mechanism for Processing-in-Memory,” CAL, 2017
2017
-
[74]
TOP-PIM: Throughput-Oriented Programmable Processing in Memory,
D. Zhang, N. Jayasena, A. Lyashevsky, J. L. Greathouse, L. Xu, and M. Ignatowski, “TOP-PIM: Throughput-Oriented Programmable Processing in Memory,” in HPDC, 2014
2014
-
[75]
HRL: Efficient and Flexible Reconfigurable Logic for Near-Data Processing,
M. Gao and C. Kozyrakis, “HRL: Efficient and Flexible Reconfigurable Logic for Near-Data Processing,” in HPCA, 2016
2016
-
[76]
The Mondrian Data Engine,
M. Drumond, A. Daglis, N. Mirzadeh, D. Ustiugov, J. Picorel, B. Falsafi, B. Grot, and D. Pnevmatikatos, “The Mondrian Data Engine,” in ISCA, 2017
2017
-
[77]
Operand Size Reconfiguration for Big Data Processing in Memory,
P. C. Santos, G. F. Oliveira, D. G. Tomé, M. A. Z. Alves, E. C. Almeida, and L. Carro, “Operand Size Reconfiguration for Big Data Processing in Memory,” in DATE, 2017
2017
-
[78]
NIM: An HMC-Based Machine for Neuron Computation,
G. F. Oliveira, P. C. Santos, M. A. Alves, and L. Carro, “NIM: An HMC-Based Machine for Neuron Computation,” in ARC, 2017
2017
-
[79]
TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory,
M. Gao, J. Pu, X. Yang, M. Horowitz, and C. Kozyrakis, “TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory,” in ASPLOS, 2017
2017
-
[80]
Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory,
D. Kim, J. Kung, S. Chai, S. Yalamanchili, and S. Mukhopadhyay, “Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory,” in ISCA, 2016
2016
-
[81]
Leveraging 3D Technologies for Hardware Security: Opportunities and Challenges,
P. Gu, S. Li, D. Stow, R. Barnes, L. Liu, Y . Xie, and E. Kursun, “Leveraging 3D Technologies for Hardware Security: Opportunities and Challenges,” in GLSVLSI, 2016
2016
-
[82]
CoNDA: Efficient Cache Coherence Support for Near-Data Accelerators,
A. Boroumand, S. Ghose, M. Patel, H. Hassan, B. Lucia, R. Ausavarung- nirun, K. Hsieh, N. Hajinazar, K. T. Malladi, H. Zheng et al., “CoNDA: Efficient Cache Coherence Support for Near-Data Accelerators,” in ISCA, 2019
2019
-
[83]
NDC: Analyzing the Impact of 3D-Stacked Memory+Logic Devices on MapReduce Workloads,
S. H. Pugsley, J. Jestes, H. Zhang, R. Balasubramonian et al., “NDC: Analyzing the Impact of 3D-Stacked Memory+Logic Devices on MapReduce Workloads,” in ISPASS, 2014
2014
-
[84]
Scheduling Techniques for GPU Architectures with Processing-in-Memory Capabilities,
A. Pattnaik, X. Tang, A. Jog, O. Kayiran, A. K. Mishra, M. T. Kandemir, O. Mutlu, and C. R. Das, “Scheduling Techniques for GPU Architectures with Processing-in-Memory Capabilities,” in PACT, 2016
2016
-
[85]
Data Reorganization in Memory Using 3D-Stacked DRAM,
B. Akin, F. Franchetti, and J. C. Hoe, “Data Reorganization in Memory Using 3D-Stacked DRAM,” in ISCA, 2015
2015
-
[86]
Accelerating Pointer Chasing in 3D-Stacked Memory: Challenges, Mechanisms, Evaluation,
K. Hsieh, S. Khan, N. Vijaykumar, K. K. Chang, A. Boroumand, S. Ghose, and O. Mutlu, “Accelerating Pointer Chasing in 3D-Stacked Memory: Challenges, Mechanisms, Evaluation,” in ICCD, 2016
2016
-
[87]
BSSync: Processing Near Memory for Machine Learning Workloads with Bounded Staleness Consistency Models,
J. H. Lee, J. Sim, and H. Kim, “BSSync: Processing Near Memory for Machine Learning Workloads with Bounded Staleness Consistency Models,” in PACT, 2015
2015
-
[88]
Polynesia: Enabling High-Performance and Energy-Efficient Hybrid Transaction- al/Analytical Databases with Hardware/Software Co-Design,
A. Boroumand, S. Ghose, G. F. Oliveira, and O. Mutlu, “Polynesia: Enabling High-Performance and Energy-Efficient Hybrid Transaction- al/Analytical Databases with Hardware/Software Co-Design,” in ICDE, 2022
2022
-
[89]
Practical Mechanisms for Reducing Processor-Memory Data Movement in Modern Workloads,
A. Boroumand, “Practical Mechanisms for Reducing Processor-Memory Data Movement in Modern Workloads,” Ph.D. dissertation, Carnegie Mellon University, 2020
2020
-
[90]
SISA: Set-Centric Instruction Set Architecture for Graph Mining on Processing-in-Memory Systems,
M. Besta, R. Kanakagiri, G. Kwasniewski, R. Ausavarungnirun, J. Beránek, K. Kanellopoulos, K. Janda, Z. V onarburg-Shmaria, L. Gian- inazzi, I. Stefan et al., “SISA: Set-Centric Instruction Set Architecture for Graph Mining on Processing-in-Memory Systems,” in MICRO, 2021
2021
-
[91]
NATSA: A Near-Data Processing Accelerator for Time Series Analysis,
I. Fernandez, R. Quislant, E. Gutiérrez, O. Plata, C. Giannoula, M. Alser, J. Gómez-Luna, and O. Mutlu, “NATSA: A Near-Data Processing Accelerator for Time Series Analysis,” in ICCD, 2020
2020
-
[92]
NAPEL: Near-Memory Computing Application Performance Prediction via Ensemble Learning,
G. Singh, G. , G. F. Oliveira, S. Corda, S. Stuijk, O. Mutlu, and H. Cor- poraal, “NAPEL: Near-Memory Computing Application Performance Prediction via Ensemble Learning,” in DAC, 2019
2019
-
[93]
A 20nm 6GB Function- in-Memory DRAM, Based on HBM2 with a 1.2 TFLOPS Programmable Computing Unit using Bank-Level Parallelism, for Machine Learning Applications,
Y .-C. Kwon, S. H. Lee, J. Lee, S.-H. Kwon, J. M. Ryu, J.-P. Son, O. Seongil, H.-S. Yu, H. Lee, S. Y . Kim et al., “A 20nm 6GB Function- in-Memory DRAM, Based on HBM2 with a 1.2 TFLOPS Programmable Computing Unit using Bank-Level Parallelism, for Machine Learning Applications,...
2021
-
[94]
Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology: Industrial Product,
S. Lee, S.-h. Kang, J. Lee, H. Kim, E. Lee, S. Seo, H. Yoon, S. Lee, K. Lim, H. Shin et al., “Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology: Industrial Product,” in ISCA, 2021
2021
-
[95]
184QPS/W 64Mb/ mm2 3D Logic- to-DRAM Hybrid Bonding with Process-Near-Memory Engine for Recommendation System,
D. Niu, S. Li, Y . Wang, W. Han, Z. Zhang, Y . Guan, T. Guan, F. Sun, F. Xue, L. Duan et al. , “184QPS/W 64Mb/ mm2 3D Logic- to-DRAM Hybrid Bonding with Process-Near-Memory Engine for Recommendation System,” in ISSCC, 2022
2022
-
[96]
Acceler- ating Sparse Matrix-Matrix Multiplication with 3D-Stacked Logic-in- Memory Hardware,
Q. Zhu, T. Graf, H. E. Sumbul, L. Pileggi, and F. Franchetti, “Acceler- ating Sparse Matrix-Matrix Multiplication with 3D-Stacked Logic-in- Memory Hardware,” in HPEC, 2013
2013
-
[97]
Logic-Base Interconnect Design for Near Memory Computing in the Smart Memory Cube,
E. Azarkhish, C. Pfister, D. Rossi, I. Loi, and L. Benini, “Logic-Base Interconnect Design for Near Memory Computing in the Smart Memory Cube,” IEEE VLSI, 2016
2016
-
[98]
Neurostream: Scalable and Energy Efficient Deep Learning with Smart Memory Cubes,
E. Azarkhish, D. Rossi, I. Loi, and L. Benini, “Neurostream: Scalable and Energy Efficient Deep Learning with Smart Memory Cubes,” TPDS, 2018
2018
-
[99]
3D-Stacked Memory-Side Acceleration: Accelerator and System Design,
Q. Guo, N. Alachiotis, B. Akin, F. Sadi, G. Xu, T. M. Low, L. Pileggi, J. C. Hoe, and F. Franchetti, “3D-Stacked Memory-Side Acceleration: Accelerator and System Design,” in WoNDP, 2014
2014
-
[100]
HAMLeT: Hardware Accelerated Memory Layout Transform within 3D-Stacked DRAM,
B. Akın, J. C. Hoe, and F. Franchetti, “HAMLeT: Hardware Accelerated Memory Layout Transform within 3D-Stacked DRAM,” in HPEC, 2014
2014
-
[101]
A Heterogeneous PIM Hardware-Software Co-Design for Energy-Efficient Graph Processing,
Y . Huang, L. Zheng, P. Yao, J. Zhao, X. Liao, H. Jin, and J. Xue, “A Heterogeneous PIM Hardware-Software Co-Design for Energy-Efficient Graph Processing,” in IPDPS, 2020
2020
-
[102]
GraphH: A Processing-in-Memory Architecture for Large- Scale Graph Processing,
G. Dai, T. Huang, Y . Chi, J. Zhao, G. Sun, Y . Liu, Y . Wang, Y . Xie, and H. Yang, “GraphH: A Processing-in-Memory Architecture for Large- Scale Graph Processing,” TCAD, 2018
2018
-
[103]
Processing-in- Memory for Energy-Efficient Neural Network Training: A Heteroge- neous Approach,
J. Liu, H. Zhao, M. A. Ogleari, D. Li, and J. Zhao, “Processing-in- Memory for Energy-Efficient Neural Network Training: A Heteroge- neous Approach,” in MICRO, 2018
2018
-
[104]
Adaptive Scheduling for Systems with Asymmetric Memory Hierarchies,
P.-A. Tsai, C. Chen, and D. Sanchez, “Adaptive Scheduling for Systems with Asymmetric Memory Hierarchies,” in MICRO, 2018
2018
-
[105]
iPIM: Programmable In-Memory Image Processing Accelerator using Near-Bank Architecture,
P. Gu, X. Xie, Y . Ding, G. Chen, W. Zhang, D. Niu, and Y . Xie, “iPIM: Programmable In-Memory Image Processing Accelerator using Near-Bank Architecture,” in ISCA, 2020
2020
-
[106]
DRAMA: An Architecture for Accelerated Processing Near Memory,
A. Farmahini-Farahani, J. H. Ahn, K. Compton, and N. S. Kim, “DRAMA: An Architecture for Accelerated Processing Near Memory,” CAL, 2014
2014
-
[107]
Near-DRAM Acceleration with Single-ISA Heterogeneous Processing in Standard Memory Modules,
H. Asghari-Moghaddam, A. Farmahini-Farahani, K. Morrow et al. , “Near-DRAM Acceleration with Single-ISA Heterogeneous Processing in Standard Memory Modules,” IEEE Micro, 2016
2016
-
[108]
Active-Routing: Compute on the Way for Near-Data Processing,
J. Huang, R. R. Puli, P. Majumder, S. Kim, R. Boyapati, K. H. Yum, and E. J. Kim, “Active-Routing: Compute on the Way for Near-Data Processing,” in HPCA, 2019
2019
-
[109]
Lightweight SIMT Core Designs for Intelligent 3D Stacked DRAM,
C. D. Kersey, H. Kim, and S. Yalamanchili, “Lightweight SIMT Core Designs for Intelligent 3D Stacked DRAM,” in MEMSYS, 2017
2017
-
[110]
PIMS: A Lightweight Processing-in-Memory Accelerator for Stencil Computations,
J. Li, X. Wang, A. Tumeo, B. Williams, J. D. Leidel, and Y . Chen, “PIMS: A Lightweight Processing-in-Memory Accelerator for Stencil Computations,” in MEMSYS, 2019
2019
-
[111]
GraphQ: Scalable PIM-Based Graph Processing,
Y . Zhuo, C. Wang, M. Zhang, R. Wang, D. Niu, Y . Wang, and X. Qian, “GraphQ: Scalable PIM-Based Graph Processing,” in MICRO, 2019
2019
-
[112]
GraphP: Reducing Communication for PIM-Based Graph Processing with Efficient Data Partition,
M. Zhang, Y . Zhuo, C. Wang, M. Gao, Y . Wu, K. Chen, C. Kozyrakis, and X. Qian, “GraphP: Reducing Communication for PIM-Based Graph Processing with Efficient Data Partition,” in HPCA, 2018
2018
-
[113]
Triple Engine Processor (TEP): A Heterogeneous Near-Memory Processor for Diverse Kernel Operations,
H. Lim and G. Park, “Triple Engine Processor (TEP): A Heterogeneous Near-Memory Processor for Diverse Kernel Operations,” TACO, 2017
2017
-
[114]
A Case for Near Memory Computation Inside the Smart Memory Cube,
E. Azarkhish, D. Rossi, I. Loi, and L. Benini, “A Case for Near Memory Computation Inside the Smart Memory Cube,” in EMS, 2016
2016
-
[115]
Large Vector Extensions Inside the HMC,
M. A. Z. Alves, M. Diener, P. C. Santos, and L. Carro, “Large Vector Extensions Inside the HMC,” in DATE, 2016
2016
-
[116]
Charon: Specialized Near-Memory Processing Architecture for Clearing Dead Objects in Memory,
J. Jang, J. Heo, Y . Lee, J. Won, S. Kim, S. J. Jung, H. Jang, T. J. Ham, and J. W. Lee, “Charon: Specialized Near-Memory Processing Architecture for Clearing Dead Objects in Memory,” in MICRO, 2019
2019
-
[117]
Active Memory Cube: A Processing-in-Memory Architecture for Exascale Systems,
R. Nair, S. F. Antao, C. Bertolli, P. Bose et al., “Active Memory Cube: A Processing-in-Memory Architecture for Exascale Systems,” IBM JRD, 2015
2015
-
[118]
CAIRO: A Compiler-Assisted Technique for Enabling Instruction-Level Offloading of Processing-in- Memory,
R. Hadidi, L. Nai, H. Kim, and H. Kim, “CAIRO: A Compiler-Assisted Technique for Enabling Instruction-Level Offloading of Processing-in- Memory,” TACO, 2017
2017
-
[119]
Processing in 3D Memories to Speed Up Operations on Complex Data Structures,
P. C. Santos, G. F. Oliveira, J. P. Lima, M. A. Alves, L. Carro, and A. C. Beck, “Processing in 3D Memories to Speed Up Operations on Complex Data Structures,” in DATE, 2018
2018
-
[120]
PRIME: A Novel Processing-in-Memory Architecture for Neural Network Computation in ReRAM-Based Main Memory,
P. Chi, S. Li, C. Xu, T. Zhang, J. Zhao, Y . Liu, Y . Wang, and Y . Xie, “PRIME: A Novel Processing-in-Memory Architecture for Neural Network Computation in ReRAM-Based Main Memory,” in ISCA, 2016
2016
-
[121]
ISAAC: A Convo- lutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars,
A. Shafiee, A. Nag, N. Muralimanohar, R. Balasubramonian, J. P. Strachan, M. Hu, R. S. Williams, and V . Srikumar, “ISAAC: A Convo- lutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars,” in ISCA, 2016. 7
2016
-
[122]
Ambit: In-Memory Accelerator for Bulk Bitwise Operations Using Commodity DRAM Technology,
V . Seshadri, D. Lee, T. Mullins, H. Hassan, A. Boroumand, J. Kim, M. A. Kozuch, O. Mutlu, P. B. Gibbons, and T. C. Mowry, “Ambit: In-Memory Accelerator for Bulk Bitwise Operations Using Commodity DRAM Technology,” in MICRO, 2017
2017
-
[123]
In-DRAM Bulk Bitwise Execution Engine,
V . Seshadri and O. Mutlu, “In-DRAM Bulk Bitwise Execution Engine,” arXiv:1905.09822, 2019
1905 arXiv
-
[124]
DRISA: A DRAM-Based Reconfigurable In-Situ Accelerator,
S. Li, D. Niu, K. T. Malladi, H. Zheng, B. Brennan, and Y . Xie, “DRISA: A DRAM-Based Reconfigurable In-Situ Accelerator,” in MICRO, 2017
2017
-
[125]
RowClone: Fast and Energy-Efficient In-DRAM Bulk Data Copy and Initialization,
V . Seshadri, Y . Kim, C. Fallin, D. Lee, R. Ausavarungnirun, G. Pekhi- menko, Y . Luo, O. Mutlu, P. B. Gibbons, M. A. Kozuch et al. , “RowClone: Fast and Energy-Efficient In-DRAM Bulk Data Copy and Initialization,” in MICRO, 2013
2013
-
[126]
The Processing Using Memory Paradigm: In-DRAM Bulk Copy, Initialization, Bitwise AND and OR,
V . Seshadri and O. Mutlu, “The Processing Using Memory Paradigm: In-DRAM Bulk Copy, Initialization, Bitwise AND and OR,” arXiv:1610.09603, 2016
2016 arXiv
-
[127]
DrAcc: A DRAM Based Accelerator for Accurate CNN Inference,
Q. Deng, L. Jiang, Y . Zhang, M. Zhang, and J. Yang, “DrAcc: A DRAM Based Accelerator for Accurate CNN Inference,” in DAC, 2018
2018
-
[128]
ELP2IM: Efficient and Low Power Bitwise Operation Processing in DRAM,
X. Xin, Y . Zhang, and J. Yang, “ELP2IM: Efficient and Low Power Bitwise Operation Processing in DRAM,” in HPCA, 2020
2020
-
[129]
GraphR: Accelerating Graph Processing Using ReRAM,
L. Song, Y . Zhuo, X. Qian, H. Li, and Y . Chen, “GraphR: Accelerating Graph Processing Using ReRAM,” in HPCA, 2018
2018
-
[130]
PipeLayer: A Pipelined ReRAM- Based Accelerator for Deep Learning,
L. Song, X. Qian, H. Li, and Y . Chen, “PipeLayer: A Pipelined ReRAM- Based Accelerator for Deep Learning,” in HPCA, 2017
2017
-
[131]
ComputeDRAM: In- Memory Compute Using Off-the-Shelf DRAMs,
F. Gao, G. Tziantzioulis, and D. Wentzlaff, “ComputeDRAM: In- Memory Compute Using Off-the-Shelf DRAMs,” in MICRO, 2019
2019
-
[132]
Neural Cache: Bit-Serial In-Cache Acceleration of Deep Neural Networks,
C. Eckert, X. Wang, J. Wang, A. Subramaniyan, R. Iyer, D. Sylvester, D. Blaauw, and R. Das, “Neural Cache: Bit-Serial In-Cache Acceleration of Deep Neural Networks,” in ISCA, 2018
2018
-
[133]
Compute Caches,
S. Aga, S. Jeloka, A. Subramaniyan, S. Narayanasamy, D. Blaauw, and R. Das, “Compute Caches,” in HPCA, 2017
2017
-
[134]
Duality Cache for Data Parallel Acceleration,
D. Fujiki, S. Mahlke, and R. Das, “Duality Cache for Data Parallel Acceleration,” in ISCA, 2019
2019
-
[135]
Buddy-RAM: Improving the Performance and Efficiency of Bulk Bitwise Operations Using DRAM,
V . Seshadri, D. Lee, T. Mullins, H. Hassan, A. Boroumand, J. Kim, M. A. Kozuch, O. Mutlu, P. B. Gibbons, and T. C. Mowry, “Buddy-RAM: Improving the Performance and Efficiency of Bulk Bitwise Operations Using DRAM,” arXiv:1611.09988, 2016
2016 arXiv
-
[136]
Simple Operations in Memory to Reduce Data Movement,
V . Seshadri and O. Mutlu, “Simple Operations in Memory to Reduce Data Movement,” in Advances in Computers, Volume 106 , 2017
2017
-
[137]
RowClone: Accelerating Data Movement and Initialization Using DRAM,
V . Seshadri, Y . Kim, C. Fallin, D. Lee, R. Ausavarungnirun, G. Pekhi- menko, Y . Luo, O. Mutlu, P. B. Gibbons, M. A. Kozuch et al. , “RowClone: Accelerating Data Movement and Initialization Using DRAM,” arXiv:1805.03502, 2018
2018 arXiv
-
[138]
Fast Bulk Bitwise AND and OR in DRAM,
V . Seshadri, K. Hsieh, A. Boroum, D. Lee, M. A. Kozuch, O. Mutlu, P. B. Gibbons, and T. C. Mowry, “Fast Bulk Bitwise AND and OR in DRAM,” CAL, 2015
2015
-
[139]
Pinatubo: A Processing-in-Memory Architecture for Bulk Bitwise Operations in Emerging Non-V olatile Memories,
S. Li, C. Xu, Q. Zou, J. Zhao, Y . Lu, and Y . Xie, “Pinatubo: A Processing-in-Memory Architecture for Bulk Bitwise Operations in Emerging Non-V olatile Memories,” in DAC, 2016
2016
-
[140]
pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables,
J. D. Ferreira, G. Falcao, J. Gómez-Luna, M. Alser, L. Orosa, M. Sadrosadati, J. S. Kim, G. F. Oliveira, T. Shahroodi, A. Nori et al., “pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables,” in MICRO, 2022
2022
-
[141]
FloatPIM: In-Memory Acceleration of Deep Neural Network Training with High Precision,
M. Imani, S. Gupta, Y . Kim, and T. Rosing, “FloatPIM: In-Memory Acceleration of Deep Neural Network Training with High Precision,” in ISCA, 2019
2019
-
[142]
Sparse BD-Net: A Multiplication-Less DNN with Sparse Binarized Depth-Wise Separable Convolution,
Z. He, L. Yang, S. Angizi, A. S. Rakin, and D. Fan, “Sparse BD-Net: A Multiplication-Less DNN with Sparse Binarized Depth-Wise Separable Convolution,” JETC, 2020
2020
-
[143]
Flash-Cosmos: In-Flash Bulk Bitwise Operations Using Inherent Computation Capability of NAND Flash Memory,
J. Park, R. Azizi, G. F. Oliveira, M. Sadrosadati, R. Nadig, D. Novo, J. Gómez-Luna, M. Kim, and O. Mutlu, “Flash-Cosmos: In-Flash Bulk Bitwise Operations Using Inherent Computation Capability of NAND Flash Memory,” in MICRO, 2022
2022
-
[144]
Adapting the RACER Architecture to Integrate Improved In-ReRAM Logic Primitives,
M. S. Truong, L. Shen, A. Glass, A. Hoffmann, L. R. Carley, J. A. Bain, and S. Ghose, “Adapting the RACER Architecture to Integrate Improved In-ReRAM Logic Primitives,” JETCAS, 2022
2022
-
[145]
RACER: Bit-Pipelined Processing Using Resistive Memory,
M. S. Truong, E. Chen, D. Su, L. Shen, A. Glass, L. R. Carley, J. A. Bain, and S. Ghose, “RACER: Bit-Pipelined Processing Using Resistive Memory,” in MICRO, 2021
2021
-
[146]
QUAC-TRNG: High-Throughput True Random Number Generation Using Quadruple Row Activation in Commodity DRAMs,
A. Olgun, M. Patel, A. G. Ya ˘glıkçı, H. Luo, J. S. Kim, F. N. Bostancı, N. Vijaykumar, O. Ergin, and O. Mutlu, “QUAC-TRNG: High-Throughput True Random Number Generation Using Quadruple Row Activation in Commodity DRAMs,” in ISCA, 2021
2021
-
[147]
D-RaNGe: Using Commodity DRAM Devices to Generate True Random Numbers With Low Latency and High Throughput,
J. S. Kim, M. Patel, H. Hassan, L. Orosa, and O. Mutlu, “D-RaNGe: Using Commodity DRAM Devices to Generate True Random Numbers With Low Latency and High Throughput,” in HPCA, 2019
2019
-
[148]
The DRAM Latency PUF: Quickly Evaluating Physical Unclonable Functions by Exploiting the Latency-Reliability Tradeoff in Modern Commodity DRAM Devices,
J. S. Kim, M. Patel, H. Hassan, and O. Mutlu, “The DRAM Latency PUF: Quickly Evaluating Physical Unclonable Functions by Exploiting the Latency-Reliability Tradeoff in Modern Commodity DRAM Devices,” in HPCA, 2018
2018
-
[149]
DR-STRaNGe: End-to-End System Design for DRAM-Based True Random Number Generators,
F. N. Bostancı, A. Olgun, L. Orosa, A. G. Ya˘glıkçı, J. S. Kim, H. Hassan, O. Ergin, and O. Mutlu, “DR-STRaNGe: End-to-End System Design for DRAM-Based True Random Number Generators,” in HPCA, 2022
2022
-
[150]
PiDRAM: A Holistic End-to-End FPGA- Based Framework for Processing-in-DRAM,
A. Olgun, J. G. Luna, K. Kanellopoulos, B. Salami, H. Hassan, O. Ergin, and O. Mutlu, “PiDRAM: A Holistic End-to-End FPGA- Based Framework for Processing-in-DRAM,” TACO, 2022
2022
-
[151]
In-Memory Low-Cost Bit-Serial Addition Using Commodity DRAM Technology,
M. F. Ali, A. Jaiswal, and K. Roy, “In-Memory Low-Cost Bit-Serial Addition Using Commodity DRAM Technology,” in TCAS-I, 2019
2019
-
[152]
GraphiDe: A Graph Processing Accelerator Leveraging In-DRAM-Computing,
S. Angizi and D. Fan, “GraphiDe: A Graph Processing Accelerator Leveraging In-DRAM-Computing,” in GLSVLSI, 2019
2019
-
[153]
SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator,
S. Li, A. O. Glova, X. Hu, P. Gu, D. Niu, K. T. Malladi, H. Zheng, B. Brennan, and Y . Xie, “SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator,” in MICRO, 2018
2018
-
[154]
Parallel Automata Processor,
A. Subramaniyan and R. Das, “Parallel Automata Processor,” in ISCA, 2017
2017
-
[155]
Hyper-AP: Enhancing Associative Processing Through A Full-Stack Optimization,
Y . Zha and J. Li, “Hyper-AP: Enhancing Associative Processing Through A Full-Stack Optimization,” in ISCA, 2020
2020
-
[156]
In-Memory Data Parallel Processor,
D. Fujiki, S. Mahlke, and R. Das, “In-Memory Data Parallel Processor,” in ASPLOS, 2018
2018
-
[157]
CODIC: A Low-Cost Substrate for Enabling Custom In- DRAM Functionalities and Optimizations,
L. Orosa, Y . Wang, M. Sadrosadati, J. Kim, M. Patel, I. Puddu, H. Luo, K. Razavi, J. Gómez-Luna, H. Hassan, N. M. Ghiasi, S. Ghose, and O. Mutlu, “CODIC: A Low-Cost Substrate for Enabling Custom In- DRAM Functionalities and Optimizations,” in ISCA, 2021
2021
-
[158]
Ultra Low Power Associative Computing with Spin Neurons and Resistive Crossbar Memory,
M. Sharad, D. Fan, and K. Roy, “Ultra Low Power Associative Computing with Spin Neurons and Resistive Crossbar Memory,” in DAC, 2013
2013
-
[159]
NoM: Network-on-Memory for Inter- Bank Data Transfer in Highly-Banked Memories,
S. H. S. Rezaei, M. Modarressi, R. Ausavarungnirun, M. Sadrosadati, O. Mutlu, and M. Daneshtalab, “NoM: Network-on-Memory for Inter- Bank Data Transfer in Highly-Banked Memories,” CAL, 2020
2020
-
[160]
ParaBit: Processing Parallel Bitwise Operations in NAND Flash Memory Based SSDs,
C. Gao, X. Xin, Y . Lu, Y . Zhang, J. Yang, and J. Shu, “ParaBit: Processing Parallel Bitwise Operations in NAND Flash Memory Based SSDs,” in MICRO, 2021
2021
-
[161]
An In-Flash Binary Neural Network Accelerator with SLC NAND Flash Array,
W. H. Choi, P.-F. Chiu, W. Ma, G. Hemink, T. T. Hoang, M. Lueker- Boden, and Z. Bandic, “An In-Flash Binary Neural Network Accelerator with SLC NAND Flash Array,” in ISCAS, 2020
2020
-
[162]
A Novel Convolution Computing Paradigm Based on NOR Flash Array with High Computing Speed and Energy Efficiency,
R. Han, P. Huang, Y . Xiang, C. Liu, Z. Dong, Z. Su, Y . Liu, L. Liu, X. Liu, and J. Kang, “A Novel Convolution Computing Paradigm Based on NOR Flash Array with High Computing Speed and Energy Efficiency,” TCAS-I, 2019
2019
-
[163]
High-Performance Mixed-Signal Neurocomputing with Nanoscale Floating-Gate Memory Cell Arrays,
F. Merrikh-Bayat, X. Guo, M. Klachko, M. Prezioso, K. K. Likharev, and D. B. Strukov, “High-Performance Mixed-Signal Neurocomputing with Nanoscale Floating-Gate Memory Cell Arrays,” TNNLS, 2017
2017
-
[164]
Three- Dimensional NAND Flash for Vector–Matrix Multiplication,
P. Wang, F. Xu, B. Wang, B. Gao, H. Wu, H. Qian, and S. Yu, “Three- Dimensional NAND Flash for Vector–Matrix Multiplication,” TVLSI, 2018
2018
-
[165]
Lue, P.-K
H.-T. Lue, P.-K. Hsu, M.-L. Wei, T.-H. Yeh, P.-Y . Du, W.-C. Chen, K.-C. Wang, and C.-Y . Lu, “Optimal Design Methods to Transform 3D NAND Flash into a High-Density, High-Bandwidth and Low-Power Nonvolatile Computing in Memory (nvCIM) Accelerator for Deep-Learning Neural Netwo...
2019
-
[166]
Behemoth: A Flash-Centric Training Accelerator for Extreme-Scale DNNs,
S. Kim, Y . Jin, G. Sohn, J. Bae, T. J. Ham, and J. W. Lee, “Behemoth: A Flash-Centric Training Accelerator for Extreme-Scale DNNs,” in FAST, 2021
2021
-
[167]
MemCore: Computing-in-Flash Design for Deep Neural Network Acceleration,
S. Wang, “MemCore: Computing-in-Flash Design for Deep Neural Network Acceleration,” in EDTM, 2022
2022
-
[168]
Flash Memory Array for Efficient Implementation of Deep Neural Networks,
R. Han, Y . Xiang, P. Huang, Y . Shan, X. Liu, and J. Kang, “Flash Memory Array for Efficient Implementation of Deep Neural Networks,” Adv. Intell. Syst. , 2021
2021
-
[169]
S-FLASH: A NAND Flash-Based Deep Neural Network Accelerator Exploiting Bit-Level Sparsity,
M. Kang, H. Kim, H. Shin, J. Sim, K. Kim, and L.-S. Kim, “S-FLASH: A NAND Flash-Based Deep Neural Network Accelerator Exploiting Bit-Level Sparsity,” TC, 2021
2021
-
[170]
Neuromorphic Computing Using NAND Flash Memory Architecture with Pulse Width Modulation Scheme,
S.-T. Lee and J.-H. Lee, “Neuromorphic Computing Using NAND Flash Memory Architecture with Pulse Width Modulation Scheme,” Front. Neurosci., 2020
2020
-
[171]
3D-FPIM: An Extreme Energy-Efficient DNN Acceleration System Using 3D NAND Flash-Based In-Situ PIM Unit,
H. Lee, M. Kim, D. Min, J. Kim, J. Back, H. Yoo, J.-H. Lee, and J. Kim, “3D-FPIM: An Extreme Energy-Efficient DNN Acceleration System Using 3D NAND Flash-Based In-Situ PIM Unit,” in MICRO, 2022
2022
-
[172]
A Dual-Split 6T SRAM-Based Computing-in-Memory Unit-Macro with Fully Parallel Product-Sum Operation for Binarized DNN Edge Processors,
X. Si, W.-S. Khwa, J.-J. Chen, J.-F. Li, X. Sun, R. Liu, S. Yu, H. Yamauchi, Q. Li, and M.-F. Chang, “A Dual-Split 6T SRAM-Based Computing-in-Memory Unit-Macro with Fully Parallel Product-Sum Operation for Binarized DNN Edge Processors,” TCAS-I, 2019
2019
-
[173]
BLADE: An In-Cache Computing Architecture for Edge Devices,
W. A. Simon, Y . M. Qureshi, M. Rios, A. Levisse, M. Zapater, and D. Atienza, “BLADE: An In-Cache Computing Architecture for Edge Devices,” TC, 2020
2020
-
[174]
GenCache: Leveraging In-Cache Operators for Efficient Sequence Alignment,
A. Nag, C. Ramachandra, R. Balasubramonian, R. Stutsman, E. Gi- acomin, H. Kambalasubramanyam, and P.-E. Gaillardon, “GenCache: Leveraging In-Cache Operators for Efficient Sequence Alignment,” in MICRO, 2019
2019
-
[175]
Bit Prudent In- Cache Acceleration of Deep Convolutional Neural Networks,
X. Wang, J. Yu, C. Augustine, R. Iyer, and R. Das, “Bit Prudent In- Cache Acceleration of Deep Convolutional Neural Networks,” in HPCA, 2019
2019
-
[176]
Towards a Reconfigurable Bit-Serial/Bit-Parallel Vector Accelerator Using In-Situ Processing-in-SRAM,
K. Al-Hawaj, O. Afuye, S. Agwa, A. Apsel, and C. Batten, “Towards a Reconfigurable Bit-Serial/Bit-Parallel Vector Accelerator Using In-Situ Processing-in-SRAM,” in ISCAS, 2020
2020
-
[177]
An Energy-Efficient VLSI Architecture for Pattern Recognition via Deep Embedding of Computation in SRAM,
M. Kang, M.-S. Keel, N. R. Shanbhag, S. Eilert, and K. Curewitz, “An Energy-Efficient VLSI Architecture for Pattern Recognition via Deep Embedding of Computation in SRAM,” in ICASSP, 2014
2014
-
[178]
Colonnade: A Reconfig- urable SRAM-Based Digital Bit-Serial Compute-in-Memory Macro for Processing Neural Networks,
H. Kim, T. Yoo, T. T.-H. Kim, and B. Kim, “Colonnade: A Reconfig- urable SRAM-Based Digital Bit-Serial Compute-in-Memory Macro for Processing Neural Networks,” JSSC, 2021
2021
-
[179]
C3SRAM: An In-Memory- Computing SRAM Macro Based on Robust Capacitive Coupling Computing Mechanism,
Z. Jiang, S. Yin, J.-S. Seo, and M. Seok, “C3SRAM: An In-Memory- Computing SRAM Macro Based on Robust Capacitive Coupling Computing Mechanism,” JSSC, 2020. 8
2020
-
[180]
A 28 nm Configurable Memory (TCAM/BCAM/SRAM) Using Push-Rule 6T Bit Cell Enabling Logic-in-Memory,
S. Jeloka, N. B. Akesh, D. Sylvester, and D. Blaauw, “A 28 nm Configurable Memory (TCAM/BCAM/SRAM) Using Push-Rule 6T Bit Cell Enabling Logic-in-Memory,” JSSC, 2016
2016
-
[181]
Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory Fusion,
Z. Wang, C. Liu, A. Arora, L. John, and T. Nowatzki, “Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory Fusion,” in ASPLOS, 2023
2023
-
[182]
Energy-Efficient and High Throughput Sparse Distributed Memory Architecture,
M. Kang, E. P. Kim, M.-s. Keel, and N. R. Shanbhag, “Energy-Efficient and High Throughput Sparse Distributed Memory Architecture,” in ISCAS, 2015
2015
-
[183]
DUAL: Acceleration of Clustering Algorithms Using Digital-Based Processing in-Memory,
M. Imani, S. Pampana, S. Gupta, M. Zhou, Y . Kim, and T. Rosing, “DUAL: Acceleration of Clustering Algorithms Using Digital-Based Processing in-Memory,” in MICRO, 2020
2020
-
[184]
Low-Cost Inter-Linked Subarrays (LISA): Enabling Fast Inter-Subarray Data Movement in DRAM,
K. K. Chang, P. J. Nair, D. Lee, S. Ghose, M. K. Qureshi, and O. Mutlu, “Low-Cost Inter-Linked Subarrays (LISA): Enabling Fast Inter-Subarray Data Movement in DRAM,” in HPCA, 2016
2016
-
[185]
SIMDRAM: A Framework for Bit-Serial SIMD Processing Using DRAM,
N. Hajinazar, G. F. Oliveira, S. Gregorio, J. D. Ferreira, N. M. Ghiasi, M. Patel, M. Alser, S. Ghose, J. Gómez-Luna, and O. Mutlu, “SIMDRAM: A Framework for Bit-Serial SIMD Processing Using DRAM,” in ASPLOS, 2021
2021
-
[186]
LAcc: Exploiting Lookup Table-Based Fast and Accurate Vector Multiplication in DRAM-Based CNN Accelerator,
Q. Deng, Y . Zhang, M. Zhang, and J. Yang, “LAcc: Exploiting Lookup Table-Based Fast and Accurate Vector Multiplication in DRAM-Based CNN Accelerator,” in DAC, 2019
2019
-
[187]
Look-Up-Table Based Processing- in-Memory Architecture with Programmable Precision-Scaling for Deep Learning Applications,
P. R. Sutradhar, S. Bavikadi, M. Connolly, S. Prajapati, M. A. Indovina, S. M. P. Dinakarrao, and A. Ganguly, “Look-Up-Table Based Processing- in-Memory Architecture with Programmable Precision-Scaling for Deep Learning Applications,” TPDS, 2021
2021
-
[188]
pPIM: A Programmable Processor-in- Memory Architecture with Precision-Scaling for Deep Learning,
P. R. Sutradhar, M. Connolly, S. Bavikadi, S. M. P. Dinakarrao, M. A. Indovina, and A. Ganguly, “pPIM: A Programmable Processor-in- Memory Architecture with Precision-Scaling for Deep Learning,” CAL, 2020
2020
-
[189]
Fulcrum: A Simplified Control and Access Mechanism Toward Flexible and Practical In-Situ Accelerators,
M. Lenjani, P. Gonzalez, E. Sadredini, S. Li, Y . Xie, A. Akel, S. Eilert, M. R. Stan, and K. Skadron, “Fulcrum: A Simplified Control and Access Mechanism Toward Flexible and Practical In-Situ Accelerators,” in HPCA, 2020
2020
-
[190]
CHOPPER: A Compiler Infras- tructure for Programmable Bit-Serial SIMD Processing Using Memory In DRAM,
X. Peng, Y . Wang, and M.-C. Yang, “CHOPPER: A Compiler Infras- tructure for Programmable Bit-Serial SIMD Processing Using Memory In DRAM,” in HPCA, 2023
2023
-
[191]
DaPPA: A Data-Parallel Framework for Processing-in-Memory Archi- tectures,
G. F. Oliveira, A. Kohli, D. Novo, J. Gómez-Luna, and O. Mutlu, “DaPPA: A Data-Parallel Framework for Processing-in-Memory Archi- tectures,” arXiv:2310.10168, 2023
2023 arXiv
-
[192]
Methodologies, Workloads, and Tools for Processing-in-Memory: Enabling the Adoption of Data-Centric Architectures,
G. F. Oliveira, J. Gómez-Luna, S. Ghose, and O. Mutlu, “Methodologies, Workloads, and Tools for Processing-in-Memory: Enabling the Adoption of Data-Centric Architectures,” in ISVLSI, 2022
2022
-
[193]
Heterogeneous Data-Centric Architectures for Modern Data-Intensive Applications: Case Studies in Machine Learning and Databases,
G. F. Oliveira, A. Boroumand, S. Ghose, J. Gómez-Luna, and O. Mutlu, “Heterogeneous Data-Centric Architectures for Modern Data-Intensive Applications: Case Studies in Machine Learning and Databases,” in ISVLSI, 2022
2022
-
[194]
Swordfish: A Framework for Evaluating Deep Neural Network-Based Basecalling Using Computation- In-Memory with Non-Ideal Memristors,
T. Shahroodi, G. Singh, M. Zahedi, H. Mao, J. Lindegger, C. Firtina, S. Wong, O. Mutlu, and S. Hamdioui, “Swordfish: A Framework for Evaluating Deep Neural Network-Based Basecalling Using Computation- In-Memory with Non-Ideal Memristors,” in MICRO, 2023
2023
-
[195]
SimplePIM: A Software Framework for Productive and Efficient Processing-In- Memory,
J. Chen, J. Gómez-Luna, I. E. Hajj, Y . Guo, and O. Mutlu, “SimplePIM: A Software Framework for Productive and Efficient Processing-In- Memory,” in PACT, 2023
2023
-
[196]
Evaluating Homomorphic Operations on a Real-World Processing-In- Memory System,
H. Gupta, M. Kabra, J. Gómez-Luna, K. Kanellopoulos, and O. Mutlu, “Evaluating Homomorphic Operations on a Real-World Processing-In- Memory System,” in IISWC, 2023
2023
-
[197]
Evaluating Machine LearningWork- loads on Memory-Centric Computing Systems,
J. Gómez-Luna, Y . Guo, S. Brocard, J. Legriel, R. Cimadomo, G. F. Oliveira, G. Singh, and O. Mutlu, “Evaluating Machine LearningWork- loads on Memory-Centric Computing Systems,” in ISPASS, 2023
2023
-
[198]
TransPimLib: Efficient Transcendental Functions for Processing-in-Memory Systems,
M. Item, J. Gómez-Luna, Y . Guo, G. F. Oliveira, M. Sadrosadati, and O. Mutlu, “TransPimLib: Efficient Transcendental Functions for Processing-in-Memory Systems,” in ISPASS, 2023
2023
-
[199]
A Framework for High-Throughput Sequence Alignment Using Real Processing-In-Memory Systems,
S. Diab, A. Nassereldine, M. Alser, J. Gómez Luna, O. Mutlu, and I. El Hajj, “A Framework for High-Throughput Sequence Alignment Using Real Processing-In-Memory Systems,” Bioinformatics, 2023
2023
-
[200]
GenPIP: In-Memory Acceleration of Genome Analysis via Tight Integration of Basecalling and Read Mapping,
H. Mao, M. Alser, M. Sadrosadati, C. Firtina, A. Baranwal, D. S. Cali, A. Manglik, N. A. Alserr, and O. Mutlu, “GenPIP: In-Memory Acceleration of Genome Analysis via Tight Integration of Basecalling and Read Mapping,” in MICRO, 2022
2022
-
[201]
Accelerating Weather Prediction Using Near-Memory Reconfigurable Fabric,
G. Singh, D. Diamantopoulos, J. Gómez-Luna, C. Hagleitner, S. Stuijk, H. Corporaal, and O. Mutlu, “Accelerating Weather Prediction Using Near-Memory Reconfigurable Fabric,” TRETS, 2022
2022
-
[202]
Cellular logic-in-memory arrays,
W. H. Kautz, “Cellular logic-in-memory arrays,” IEEE ToC, 1969
1969
-
[203]
A Logic-in-Memory Computer,
H. S. Stone, “A Logic-in-Memory Computer,” IEEE ToC, 1970
1970
-
[204]
HMC Specification Rev. 2.0,
HMC Consortium, “HMC Specification Rev. 2.0,” www.hybridmemory cube.org/
-
[205]
A 1.2V 8Gb 8-Channel 128GB/s High-Bandwidth Memory (HBM) Stacked DRAM with Effective Microbump I/O Test Methods Using 29nm Process and TSV,
D. U. Lee, K. W. Kim, K. W. Kim, H. Kim, J. Y . Kim, Y . J. Park, J. H. Kim, D. S. Kim, H. B. Park, J. W. Shin et al. , “A 1.2V 8Gb 8-Channel 128GB/s High-Bandwidth Memory (HBM) Stacked DRAM with Effective Microbump I/O Test Methods Using 29nm Process and TSV,” in ISSCC, 2014
2014
-
[206]
Simultane- ous Multi-Layer Access: Improving 3D-Stacked Memory Bandwidth at Low Cost,
D. Lee, S. Ghose, G. Pekhimenko, S. Khan, and O. Mutlu, “Simultane- ous Multi-Layer Access: Improving 3D-Stacked Memory Bandwidth at Low Cost,” TACO, 2016
2016
-
[207]
JEDEC, JESD23-5D: High Bandwidth Memory (HBM) DRAM Standard, 2021
2021
-
[208]
JEDEC, JESD23-8A: High Bandwidth Memory (HBM3) DRAM Stan- dard, 2021
2021
-
[209]
Present and Future, Challenges of High Bandwith Memory (HBM),
K. Kim and M.-j. Park, “Present and Future, Challenges of High Bandwith Memory (HBM),” in IMW, 2024
2024
-
[210]
MATSA: An MRAM-Based Energy-Efficient Accelerator for Time Series Analysis,
I. Fernandez, C. Giannoula, A. Manglik, R. Quislant, N. M. Ghiasi, J. Gómez-Luna, E. Gutierrez, O. Plata, and O. Mutlu, “MATSA: An MRAM-Based Energy-Efficient Accelerator for Time Series Analysis,” IEEE Access, 2024
2024
-
[211]
Design Space Exploration for PIM Architectures in 3D-Stacked Memories,
J. P. C. de Lima, P. C. Santos, M. A. Alves, A. Beck, and L. Carro, “Design Space Exploration for PIM Architectures in 3D-Stacked Memories,” in CF, 2018
2018
-
[212]
BlueDBM: An appliance for big data analytics,
S.-W. Jun, M. Liu, S. Lee, J. Hicks, J. Ankcorn, M. King, S. Xu, and Arvind, “BlueDBM: An appliance for big data analytics,” ISCA, 2015
2015
-
[213]
Field-Effect Transistor Memory,
R. H. Dennard, “Field-Effect Transistor Memory,” U.S. Patent 3,387,286, 1968
1968
-
[214]
SwiftRL: Towards Efficient Reinforcement Learning on Real Processing-In- Memory Systems,
K. Gogineni, S. S. Dayapule, J. Gómez-Luna, K. Gogineni, P. Wei, T. Lan, M. Sadrosadati, O. Mutlu, and G. Venkataramani, “SwiftRL: Towards Efficient Reinforcement Learning on Real Processing-In- Memory Systems,” in ISPASS, 2024
2024
-
[215]
SecNDP: Secure Near-Data Processing with Untrusted Memory,
W. Xiong, L. Ke, D. Jankov, M. Kounavis, X. Wang, E. Northup, J. A. Yang, B. Acun, C.-J. Wu, P. T. P. Tang et al., “SecNDP: Secure Near-Data Processing with Untrusted Memory,” in HPCA, 2022
2022
-
[216]
Invisimem: Smart Memory Defenses For Memory Bus Side Channel,
S. Aga and S. Narayanasamy, “Invisimem: Smart Memory Defenses For Memory Bus Side Channel,” ISCA, 2017
2017
-
[217]
FracDRAM: Fractional Values in Off-the-Shelf DRAM,
F. Gao, G. Tziantzioulis, and D. Wentzlaff, “FracDRAM: Fractional Values in Off-the-Shelf DRAM,” in MICRO, 2022
2022
-
[218]
DRAM Bender: An Extensible and Versatile FPGA-based Infrastructure to Easily Test State-of-the-art DRAM Chips,
A. Olgun, H. Hassan, A. G. Ya ˘glıkçı, Y . C. Tu˘grul, L. Orosa, H. Luo, M. Patel, O. Ergin, and O. Mutlu, “DRAM Bender: An Extensible and Versatile FPGA-based Infrastructure to Easily Test State-of-the-art DRAM Chips,” TCAD, 2023
2023
-
[219]
MIMDRAM: An End-to-End Processing- Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data Comput- ing,
G. F. Oliveira, A. Olgun, A. G. G. Yaglikçi, N. Bostanci, J. Gómez-Luna, S. Ghose, and O. Mutlu, “MIMDRAM: An End-to-End Processing- Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data Comput- ing,” in HPCA, 2024
2024
-
[220]
Functionally-Complete Boolean Logic in Real DRAM Chips: Experi- mental Characterization and Analysis,
I. E. Yuksel, Y . C. Tugrul, A. Olgun, F. N. Bostanci, A. G. Yaglikci, G. F. de Oliveira, H. Luo, J. G. Luna, M. Sadrosadati, and O. Mutlu, “Functionally-Complete Boolean Logic in Real DRAM Chips: Experi- mental Characterization and Analysis,” in HPCA, 2024
2024
-
[221]
Simultaneous Many-Row Activation in Off-the-Shelf DRAM Chips: Experimental Characterization and Analysis,
I. E. Yuksel, Y . C. Tugrul, F. N. Bostanci, G. F. de Oliveira, A. G. Yaglikci, A. Olgun, M. Soysal, H. Luo, J. G. Luna, M. Sadrosadati, and O. Mutlu, “Simultaneous Many-Row Activation in Off-the-Shelf DRAM Chips: Experimental Characterization and Analysis,” in DSN, 2024
2024
-
[223]
PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips,
I. E. Yuksel, Y . C. Tugrul, F. N. Bostanci, A. G. Yaglikci, A. Olgun, G. F. Oliveira, M. Soysal, H. Luo, J. G. Luna, M. Sadrosadati, and O. Mutlu, “PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips,” arXiv:2312.02...
2023 arXiv
-
[224]
Very High-Speed Computing Systems,
M. J. Flynn, “Very High-Speed Computing Systems,” Proc. IEEE, 1966
1966
-
[225]
6th Generation Intel Core Processor Family Datasheet,
Intel Corp., “6th Generation Intel Core Processor Family Datasheet,” http://www.intel.com/content/www/us/en/processors/core/
-
[226]
NVIDIA A100 Tensor Core GPU Architecture,
NVIDIA, “NVIDIA A100 Tensor Core GPU Architecture,” https://t.ly /rMUgA, 2020
2020
-
[227]
SoftMC: A Flexible and Practical Open-Source Infrastructure for Enabling Experimental DRAM Studies,
H. Hassan, N. Vijaykumar, S. Khan, S. Ghose, K. Chang, G. Pekhi- menko, D. Lee, O. Ergin, and O. Mutlu, “SoftMC: A Flexible and Practical Open-Source Infrastructure for Enabling Experimental DRAM Studies,” in HPCA, 2017
2017
-
[228]
A Low-Cost Reduced-Latency DRAM Architecture with Dynamic Reconfiguration of Row Decoder,
F. Bai, S. Wang, X. Jia, Y . Guo, B. Yu, H. Wang, C. Lai, Q. Ren, and H. Sun, “A Low-Cost Reduced-Latency DRAM Architecture with Dynamic Reconfiguration of Row Decoder,” TVLSI, 2022
2022
-
[229]
N. H. Weste and D. Harris, CMOS VLSI Design: A Circuits and Systems Perspective. Pearson Education India, 2015
2015
-
[230]
High-Performance Low-Power Selective Precharge Schemes for Address Decoders,
M. A. Turi and J. G. Delgado-Frias, “High-Performance Low-Power Selective Precharge Schemes for Address Decoders,” TCAS-II, 2008
2008
-
[231]
180-2: Secure Hash Standard (SHS),
National Institute of Standards and Technology (NIST), “180-2: Secure Hash Standard (SHS),” Federal Information Processing Standards Publications, 2012
2012
-
[232]
A Statistical Test Suite for Random and Pseudorandom Number Generators for Cryptographic Applications,
L. Bassham, A. Rukhin, J. Soto, J. Nechvatal, M. Smid, S. Leigh, M. Levenson, M. Vangel, N. Heckert, and D. Banks, “A Statistical Test Suite for Random and Pseudorandom Number Generators for Cryptographic Applications,” NIST SP, 2010
2010
-
[233]
JEDEC, JESD79-4C: DDR4 SDRAM , 2017
2017
-
[234]
Memory-Centric Computing: Recent Advances in Processing-in-DRAM (Invited),
O. Mutlu, G. F. Oliveira, A. Olgun, and I. E. Yuksel, “Memory-Centric Computing: Recent Advances in Processing-in-DRAM (Invited),” in IEDM, 2024. 9
2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.