REVIEW 4 major objections 4 minor 102 references
EasyDRAM: An FPGA-based Infrastructure for Fast and Accurate End-to-End Evaluation of Emerging DRAM Techniques
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read EasyDRAM claims that a C++ software memory controller plus time scaling lets FPGA platforms evaluate DRAM techniques on real DRAM chips with average execution-time error below 0.1% against a 1 GHz reference, and that this changes in-DRAM…
desk verdict Open-source FPGA infrastructure with a genuinely useful time-scaling idea, but the headline accuracy claim needs much stronger OoO validation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is time scaling, implemented as per-domain emulation counters plus clock-gating. Each domain (processors, memory controller) has a counter tracking its emulated clock cycles; when a memory request is outstanding, the processor is stopped and the memory controller enters critical mode, so a software controller that takes hundreds of FPGA cycles per decision cannot let the processor execute extra instructions or observe responses early. Responses are tagged with a processor-cycle value before release, and the command sequence prepared in software is handed to a hardware command executor that issues DRAM commands at nanosecond timing, so the slow software part and the fast real DRAM chip are decoupled.
What would settle it
Run the benchmark workloads to completion on a real system with the same processor and cache setup, and compare the measured end-to-end times against EasyDRAM's time-scaled results; errors well above the reported 0.1% average, or divergence that grows with workload, would show the time-scaling premise fails. A second check is to validate against an independently designed 1 GHz memory controller instead of the one EasyDRAM provides.
Extended reading notes
Core claim
The paper's central claim is that accurate end-to-end DRAM technique evaluation does not require a hardware description language or a fast FPGA processor. A software memory controller running on a simple programmable core can make scheduling decisions and prepare DRAM command sequences, while time scaling counters and processor clock-gating keep the emulated processor from running ahead of or behind the memory system. The result is that a low-clock FPGA processor emulates a high-clock modern processor's timing behavior on real DDR4 chips; the paper's evidence is average execution-time and memory-latency error below 0.1% against its own RTL reference, a main-memory latency profile resembling a real embedded system, and end-to-end case studies where time scaling reduces RowClone speedups from 306.7x to 15.0x and yields a 2.75% average speedup from tRCD reduction.
Load-bearing premise
The whole accuracy claim depends on whether pretending the FPGA processor is running at a higher clock speed, by pausing it and counting cycles, produces the same memory behavior as a real modern processor would; the paper's main accuracy check compares EasyDRAM to its own reference design, not to an independent real processor.
Editorial extensions
If this is right
- DRAM techniques can be prototyped and modified in C++ (roughly 325 lines for the two case studies) with no RTL changes, which lowers the bar for memory-system researchers who are not hardware designers.
- FPGA emulators that ignore the processor-DRAM clock-frequency gap can overstate DRAM technique benefits by about 20x; time-scaled results show in-DRAM bulk copy still gives 15.0x average speedup for copying and 1.8x for initialization, mainly for large arrays.
- Real-chip characterization changes conclusions that pure simulation reaches: because some DRAM rows cannot be cloned reliably and require CPU fallback, software simulation overestimates RowClone initialization speedups.
- On the evaluated workloads, tRCD reduction improves performance by 2.75% on average (9.76% maximum), and the paper expects larger gains on more memory-intensive workloads.
- Execution speed is 5.9x faster on average (20.3x maximum) than a cycle-level software simulator, allowing workloads to be run to completion rather than truncated.
Reading between the lines
- Editorial inference: time scaling should generalize beyond DRAM evaluation to any FPGA prototype where a slow soft processor interacts with a fast peripheral, although the paper only demonstrates it for memory latency and execution time.
- Editorial inference: the execution-time accuracy claim currently rests on comparison with an RTL reference that shares EasyDRAM's memory-controller scheduling logic, so an independent real-processor execution-time comparison is the natural next test.
- Editorial inference: because the real-system latency profile was matched after tuning and with a smaller L2 cache than the modeled board, some of the resemblance may come from compensating parameter choices; sweeping cache size and processor parameters would show how sensitive the match is.
- Editorial inference: a direct application would be to evaluate other DRAM techniques such as refresh reduction or RowHammer mitigations under time scaling and compare the speedups against measurements on a real system, which would test whether the 20x correction factor holds across techniques.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents EasyDRAM, an FPGA-based framework for end-to-end evaluation of DRAM techniques on real DRAM chips. The two key contributions are (1) a programmable memory controller written in C++ that replaces hardwired RTL memory controllers, lowering the barrier to implementing DRAM techniques, and (2) a 'time scaling' mechanism that decouples processor and DRAM clock domains so that a slow FPGA processor can emulate a faster modern processor's timing. The authors validate time scaling against a 1 GHz RTL reference (reporting <0.1% average execution-time error), compare a memory-latency profile against an NVIDIA Jetson Nano, and demonstrate two case studies: RowClone and tRCD reduction. They report that without time scaling RowClone appears 306.7x faster than a CPU baseline, while with time scaling it is 15.0x, and interpret this as roughly 20x improved accuracy over prior FPGA platforms. The manuscript is open-source and reproducible, with code released on GitHub.
Significance. If the accuracy claims hold, EasyDRAM would be a genuinely useful infrastructure: it combines real-DRAM operation, a high-level C++ API for memory-controller design, and fast full-workload evaluation. The paper is also valuable as a concrete demonstration that FPGA-based emulators that ignore processor-DRAM clock-frequency disparity can overstate DRAM technique benefits by an order of magnitude. The RowClone case study is a compelling cautionary result. However, the significance is conditional on the time-scaling mechanism faithfully modeling a modern out-of-order CPU's memory behavior. The current validation does not establish that: the RTL reference shares the same memory-controller logic, and the only real-system comparison is a latency profile, not an execution-time comparison. These gaps are load-bearing because the paper's central claim is 'accurate' end-to-end evaluation.
major comments (4)
- [§4.3–4.4, Figs. 5 and 6] The time-scaling mechanism as described clock-gates the entire processor as soon as a main-memory request is placed in the hardware buffer, and un-gates it only after the software memory controller advances the time-scaling counters. This collapses the out-of-order core's memory-level parallelism into a blocking, in-order model. A real OoO core (BOOM, and the modeled Cortex A57) issues multiple outstanding misses and continues executing independent instructions while a miss is in flight, so immediate full-core gating overestimates memory stall time. The <0.1% validation in §6 does not settle this because the PolyBench workloads used there have very low LLC miss rates (the paper itself reports only 2.2 misses per kilo-cycle in §8.3), where MLP effects are small. I therefore do not see evidence that time scaling faithfully reproduces OoO memory behavior, which is the load-bearing premise for the 'accurate' claim. A concrete test would be to run MLP-sensitive workloads (e.g., multiple independent pointer-chasing chains or many-stream memory access) and compare against an RTL reference that does not gate the processor on the first outstanding miss.
- [§6, Time Scaling Validation] The RTL reference is not an independent ground truth. The manuscript states that the reference 'implements EasyDRAM's memory controller in hardware' and takes 'the same scheduling decisions as EasyDRAM's memory controller.' This validates the internal consistency of the time-scaling cycle counting, but it does not validate that the modeled processor-memory interaction matches a real modern system. The two compared systems share the same processor model, the same memory-controller scheduling policy, and presumably the same cache hierarchy; the only difference is the 100-MHz-with-time-scaling vs. 1-GHz-without-time-scaling configuration. This is a self-referential check, not an independent validation of realistic OoO processor behavior.
- [§6, Fig. 8 and Abstract] The only real-system validation is a memory-latency profile (cycles per load instruction) obtained after tuning BOOM parameters to match the Jetson Nano, and the manuscript acknowledges that the L2 cache size differs (512 KiB vs. 2 MiB). A latency profile after parameter tuning does not validate full-workload execution time, nor does it exercise memory-level parallelism, scheduling interactions, or the effects of time scaling on end-to-end performance. Consequently, the abstract's claim that EasyDRAM yields results 'by ≈20× for execution time' more accurately than prior platforms is not directly supported: the 306.7x vs. 15.0x comparison in §7.2 shows a large difference between two EasyDRAM configurations, but no independent ground truth establishes that 15.0x is the correct answer.
- [§7.2, Footnote 5 and Fig. 13] The Ramulator 2.0 comparisons are presented as supporting evidence for accuracy, but the paper itself concedes in footnote 5 that the Ramulator processor model (a simple out-of-order core and last-level cache) 'significantly differs' from EasyDRAM's real processor system. This means the Ramulator values cannot serve as a ground-truth reference for EasyDRAM's accuracy; they are just another point of comparison. The conclusions in §7.2 that time scaling is 'more accurate' rely on assuming Ramulator is closer to reality than the no-time-scaling configuration, which the paper does not independently establish. This should be stated explicitly in the main text, and the accuracy conclusions should be tempered accordingly.
minor comments (4)
- [§6] The manuscript reports only aggregate '<0.1% average' and '<1% maximum' errors; a table or plot with per-benchmark execution times and memory-latency errors for all 29 workloads would let readers assess outliers and strengthen reproducibility.
- [§6, Fig. 8] The L2 cache size mismatch (512 KiB vs. 2 MiB) is mentioned but its effect on the latency-profile comparison is not discussed; the steep latency rise in the real system at 2 MiB vs. EasyDRAM at 512 KiB may reflect this difference and should be analyzed or at least acknowledged as a confound.
- [§8.1] The strong/weak tRCD threshold of 9.0 ns is presented without justification or sensitivity analysis; since this threshold determines the fraction of strong rows and the resulting performance gains, reporting a range of thresholds would make the results more robust.
- [Abstract and §7.2] The phrase 'more accurate results (e.g., by ≈20× for execution time)' overstates what is measured; it would be more precise to say that time scaling changes the estimated RowClone speedup from 306.7x to 15.0x compared to a no-time-scaling configuration, and to avoid claiming a factor-of-20 accuracy improvement without an external ground truth.
Circularity Check
The real-system latency-profile validation of time scaling is partly self-definitional: the main-memory plateau is forced by the emulated-frequency conversion, though the RTL consistency check and the real-DRAM case studies give the paper independent content.
-
self definitional
[Section 4.3 (Time Scaling, Figure 5) and Section 6 (Real System, Figure 8)]
"SMC obtains the time spent for the ACT command from DRAM Bender 4, calculates the cycles the processor domain should execute at the emulated system's clock frequency, and advances the memory controller cycle counter by the number of cycles spent for the ACT command 5. ... In contrast, EasyDRAM - Time Scaling (i.e., a system emulation that mimics a board with Cortex A57) has a similar memory latency profile compared to a real system."
Under time scaling, the processor-cycle cost of a main-memory access is defined as the real DRAM command time multiplied by the user-selected emulated processor frequency. On the real Jetson Nano, the same real DRAM time expressed in Cortex A57 cycles is the same product with the A57 clock frequency. Hence, once EasyDRAM is configured to emulate the A57 frequency, the main-memory plateau in Figure 8 must approximately coincide with the real system's plateau; that part of the profile resemblance is the definition of the counter update, not an independent confirmation that the modeled processor faithfully captures memory behavior.
full rationale
The paper's central innovation is time scaling, and its two validation strands are not equally circular. The RTL validation (Section 6) compares a 100 MHz time-scaled EasyDRAM against a 1 GHz RTL reference that implements the same memory-controller scheduling decisions; this is an internal consistency check between two different implementations, not a reduction to the same equation, so I do not score it as circular. The real-system validation, however, relies on Figure 8, whose main-memory latency plateau is largely determined by the time-scaling definition: real DRAM time multiplied by the chosen emulated frequency equals the real system's cycles-per-access once the frequency is set to the A57 clock. That makes the plateau resemblance self-definitional for the main-memory component. The case studies are genuine end-to-end measurements on real DRAM chips, and the Ramulator 2.0 comparisons and RTL consistency check provide independent content, so the paper does not collapse entirely into its own assumptions. The '≈20x more accurate' conclusion is a sensitivity result between two EasyDRAM configurations whose 'accuracy' label inherits the time-scaling validation weakness, but it is not itself a fitted input renamed as a prediction. There is no load-bearing self-citation chain or imported uniqueness theorem. Overall, the circularity burden is moderate: one central validation component reduces by construction, but the platform's execution-time consistency check and real-hardware case studies keep the central claim partially independent.
Assumptions & free parameters
free parameters (3)
- Strong/weak tRCD classification threshold (9.0 ns) =
9.0 ns
- BOOM core and memory-hierarchy configuration parameters =
Not fully specified
- Time scaling validation clock frequencies (100 MHz emulating 1 GHz) =
100 MHz to 1 GHz
assumptions (3)
- domain assumption DRAM Bender reliably issues DRAM commands to real DDR4 chips at specified nanosecond timings
- domain assumption The 1 GHz RTL reference system used for time-scaling validation is a correct model of a real processor's memory behavior
- domain assumption The Jetson Nano SoC with Cortex A57 is representative of a modern computing system for DRAM evaluation
invented entities (1)
-
Time scaling mechanism
independent evidence
Cite this review
Pith. "Pith review of EasyDRAM: An FPGA-based Infrastructure for Fast and Accurate End-to-End Evaluation of Emerging DRAM Techniques." pith.science (2026). https://pith.science/paper/CB44JEUX
@misc{pith2026250610441,
author = {Pith},
title = {Pith review of: EasyDRAM: An FPGA-based Infrastructure for Fast and Accurate End-to-End Evaluation of Emerging DRAM Techniques},
year = {2026},
howpublished = {\url{https://pith.science/paper/CB44JEUX}},
note = {Machine review of arXiv:2506.10441}
}
read the original abstract
DRAM is a critical component of modern computing systems. Recent works propose numerous techniques (that we call DRAM techniques) to enhance DRAM-based computing systems' throughput, reliability, and computing capabilities (e.g., in-DRAM bulk data copy). Evaluating the system-wide benefits of DRAM techniques is challenging as they often require modifications across multiple layers of the computing stack. Prior works propose FPGA-based platforms for rapid end-to-end evaluation of DRAM techniques on real DRAM chips. Unfortunately, existing platforms fall short in two major aspects: (1) they require deep expertise in hardware description languages, limiting accessibility; and (2) they are not designed to accurately model modern computing systems. We introduce EasyDRAM, an FPGA-based framework for rapid and accurate end-to-end evaluation of DRAM techniques on real DRAM chips. EasyDRAM overcomes the main drawbacks of prior FPGA-based platforms with two key ideas. First, EasyDRAM removes the need for hardware description language expertise by enabling developers to implement DRAM techniques using a high-level language (C++). At runtime, EasyDRAM executes the software-defined memory system design in a programmable memory controller. Second, EasyDRAM tackles a fundamental challenge in accurately modeling modern systems: real processors typically operate at higher clock frequencies than DRAM, a disparity that is difficult to replicate on FPGA platforms. EasyDRAM addresses this challenge by decoupling the processor-DRAM interface and advancing the system state using a novel technique we call time scaling, which faithfully captures the timing behavior of the modeled system. We believe and hope that EasyDRAM will enable innovative ideas in memory system design to rapidly come to fruition. To aid future research EasyDRAM implementation is open sourced at https://github.com/CMU-SAFARI/EasyDRAM.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
Memory Scaling: A Systems Architecture Perspective,
O. Mutlu, “Memory Scaling: A Systems Architecture Perspective, ” inIMW, 2013
2013
-
[2]
A Modern Primer On Processing In Memory,
O. Mutlu, S. Ghose, J. Gómez-Luna, and R. Ausavarungnirun, “A Modern Primer On Processing In Memory, ” inEmerging Computing: From Devices To Systems - Looking Beyond Moore And Von Neumann, 2021
2021
-
[3]
Processing-in- memory: A Workload-driven Perspective,
S. Ghose, A. Boroumand, J. S. Kim, J. Gómez-Luna, and O. Mutlu, “Processing-in- memory: A Workload-driven Perspective, ”IBM JRD, 2019
2019
-
[4]
A Modern Primer on Processing in Memory,
O. Mutlu, S. Ghose, J. Gómez-Luna, R. Ausavarungnirun, M. Sadrosadati, and G. F. Oliveira, “A Modern Primer on Processing in Memory, ” arXiv:2012.03112 [cs.AR], 2025
arXiv 2012
-
[5]
Solar-DRAM: Reducing DRAM Access Latency by Exploiting the Variation in Local Bitlines,
J. Kim, M. Patel, H. Hassan, and O. Mutlu, “Solar-DRAM: Reducing DRAM Access Latency by Exploiting the Variation in Local Bitlines, ” inICCD, 2018
2018
-
[6]
Reducing DRAM Latency Via Charge-level-aware Look-ahead Partial Restoration,
Y. Wang, A. Tavakkol, L. Orosa, S. Ghose, N. M. Ghiasi, M. Patel, J. S. Kim, H. Hassan, M. Sadrosadati, and O. Mutlu, “Reducing DRAM Latency Via Charge-level-aware Look-ahead Partial Restoration, ” inMICRO, 2018
2018
-
[7]
HiRA: Hidden Row Activation for Reducing Refresh Latency of Off-the- shelf DRAM Chips,
A. G. Yağlikci, A. Olgun, M. Patel, H. Luo, H. Hassan, L. Orosa, O. Ergin, and O. Mutlu, “HiRA: Hidden Row Activation for Reducing Refresh Latency of Off-the- shelf DRAM Chips, ” inMICRO, 2022
2022
-
[8]
CLR-DRAM: A Low-cost DRAM Architecture Enabling Dynamic Capacity-latency Trade-off,
H. Luo, T. Shahroodi, H. Hassan, M. Patel, A. G. Yaglikci, L. Orosa, J. Park, and O. Mutlu, “CLR-DRAM: A Low-cost DRAM Architecture Enabling Dynamic Capacity-latency Trade-off, ” inISCA, 2020
2020
Show all 102 references
-
[9]
VRL-DRAM: Improving DRAM Performance via Variable Refresh Latency,
A. Das, H. Hassan, and O. Mutlu, “VRL-DRAM: Improving DRAM Performance via Variable Refresh Latency, ” inDAC, 2018
2018
-
[10]
ChargeCache: Reducing DRAM Latency by Exploiting Row Access Locality,
H. Hassan, G. Pekhimenko, N. Vijaykumar, V. Seshadri, D. Lee, O. Ergin, and O. Mutlu, “ChargeCache: Reducing DRAM Latency by Exploiting Row Access Locality, ” inHPCA, 2016
2016
-
[11]
Tiered-latency DRAM: A Low Latency and Low Cost DRAM Architecture,
D. Lee, Y. Kim, V. Seshadri, J. Liu, L. Subramanian, and O. Mutlu, “Tiered-latency DRAM: A Low Latency and Low Cost DRAM Architecture, ” inHPCA, 2013
2013
-
[12]
Adaptive-latency DRAM: Optimizing DRAM Timing for the Common-case,
D. Lee, Y. Kim, G. Pekhimenko, S. Khan, V. Seshadri, K. Chang, and O. Mutlu, “Adaptive-latency DRAM: Optimizing DRAM Timing for the Common-case, ” in HPCA, 2015
2015
-
[13]
Understanding Latency Variation in Modern DRAM Chips: Experimental Characterization, Analysis, and Optimization,
K. K. Chang, A. Kashyap, H. Hassan, S. Ghose, K. Hsieh, D. Lee, T. Li, G. Pekhimenko, S. Khan, and O. Mutlu, “Understanding Latency Variation in Modern DRAM Chips: Experimental Characterization, Analysis, and Optimization, ” inSIGMETRICS, 2016
2016
-
[14]
Understanding and Improving the Latency of DRAM-based Memory Systems,
K. K. Chang, “Understanding and Improving the Latency of DRAM-based Memory Systems, ” Ph.D. dissertation, Carnegie Mellon Univ., 2017
2017
-
[15]
Multiple Clone Row DRAM: A Low Latency and Area Optimized DRAM,
J. Choi, W. Shin, J. Jang, J. Suh, Y. Kwon, Y. Moon, and L.-S. Kim, “Multiple Clone Row DRAM: A Low Latency and Area Optimized DRAM, ” inISCA, 2015
2015
-
[16]
Low-cost Inter-linked Subarrays (LISA): Enabling Fast Inter-subarray Data Movement in DRAM,
K. K. Chang, P. J. Nair, D. Lee, S. Ghose, M. K. Qureshi, and O. Mutlu, “Low-cost Inter-linked Subarrays (LISA): Enabling Fast Inter-subarray Data Movement in DRAM, ” inHPCA, 2016
2016
-
[17]
Reducing Memory Access Latency with Asymmetric DRAM Bank Organizations,
Y. H. Son, O. Seongil, Y. Ro, J. W. Lee, and J. H. Ahn, “Reducing Memory Access Latency with Asymmetric DRAM Bank Organizations, ” inISCA, 2013
2013
-
[18]
Simultaneous Many- row Activation in Off-the-shelf DRAM Chips: Experimental Characterization and Analysis,
İ. E. Yüksel, Y. C. Tuğrul, F. N. Bostancı, G. F. Oliveira, A. G. Yağlıkçı, A. Olgun, M. Soysal, H. Luo, J. Gómez-Luna, M. Sadrosadatiet al., “Simultaneous Many- row Activation in Off-the-shelf DRAM Chips: Experimental Characterization and Analysis, ” inDSN, 2024
2024
-
[19]
Functionally-complete Boolean Logic in Real DRAM Chips: Experimental Characterization and Analysis,
İ. E. Yüksel, Y. C. Tuğrul, A. Olgun, F. N. Bostancı, A. G. Yağlıkçı, G. F. Oliveira, H. Luo, J. Gómez-Luna, M. Sadrosadati, and O. Mutlu, “Functionally-complete Boolean Logic in Real DRAM Chips: Experimental Characterization and Analysis, ” inHPCA, 2024
2024
-
[20]
QUAC-TRNG: High-throughput True Random Number 13 Generation using Quadruple Row Activation in Commodity DRAM Chips,
A. Olgun, M. Patel, A. G. Yağlıkçı, H. Luo, J. S. Kim, F. Nisa Bostancı, N. Vijaykumar, O. Ergin, and O. Mutlu, “QUAC-TRNG: High-throughput True Random Number 13 Generation using Quadruple Row Activation in Commodity DRAM Chips, ” inISCA, 2021
2021
-
[21]
CROW: A Low-cost Substrate for Improving DRAM Performance, Energy Efficiency, and Reliability,
H. Hassan, M. Patel, J. S. Kim, A. G. Yaglikci, N. Vijaykumar, N. M. Ghiasi, S. Ghose, and O. Mutlu, “CROW: A Low-cost Substrate for Improving DRAM Performance, Energy Efficiency, and Reliability, ” inISCA, 2019
2019
-
[22]
Using Run-Time Reverse-Engineering to Optimize DRAM Refresh,
D. M. Mathew, E. F. Zulian, M. Jung, K. Kraft, C. Weis, B. Jacob, and N. Wehn, “Using Run-Time Reverse-Engineering to Optimize DRAM Refresh, ” inMEMSYS, 2017
2017
-
[23]
EDEN: Enabling Energy-efficient, High-performance Deep Neural Network Inference Using Approximate DRAM,
S. Koppula, L. Orosa, A. G. Yağlıkçı, R. Azizi, T. Shahroodi, K. Kanellopoulos, and O. Mutlu, “EDEN: Enabling Energy-efficient, High-performance Deep Neural Network Inference Using Approximate DRAM, ” inMICRO, 2019
2019
-
[24]
Orosa, S
L. Orosa, S. Koppula, K. Kanellopoulos, A. G. Yağlıkçı, and O. Mutlu,Using Approxi- mate DRAM For Enabling Energy-efficient, High-performance Deep Neural Network Inference. Springer, 10 2023, pp. 275–314
2023
-
[25]
Exploiting Expendable Process-margins In Drams For Run-time Performance Optimization,
K. Chandrasekar, S. Goossens, C. Weis, M. Koedam, B. Akesson, N. Wehn, and K. Goossens, “Exploiting Expendable Process-margins In Drams For Run-time Performance Optimization, ” inDATE, 2014
2014
-
[26]
CDAR-DRAM: Enabling Run- time DRAM Performance And Energy Optimization Via In-situ Charge Detection And Adaptive Data Restoration,
Y. Qin, C. Lin, W. He, Y. Sun, Z. Mao, and M. Seok, “CDAR-DRAM: Enabling Run- time DRAM Performance And Energy Optimization Via In-situ Charge Detection And Adaptive Data Restoration, ”IEEE Transactions On Computer-aided Design Of Integrated Circuits And Systems, 2023
2023
-
[27]
Restore Truncation For Performance Improvement In Future DRAM Systems,
X. Zhang, Y. Zhang, B. R. Childers, and J. Yang, “Restore Truncation For Performance Improvement In Future DRAM Systems, ” inHPCA, 2016
2016
-
[28]
NUAT: A Non-uniform Access Time Memory Controller,
W. Shin, J. Yang, J. Choi, and L.-S. Kim, “NUAT: A Non-uniform Access Time Memory Controller, ” inHPCA, 2014
2014
-
[29]
Dram-latency Optimization Inspired By Relationship Between Row-access Time And Refresh Timing,
W. Shin, J. Choi, J. Jang, J. Suh, Y. Moon, Y. Kwon, and L.-S. Kim, “Dram-latency Optimization Inspired By Relationship Between Row-access Time And Refresh Timing, ”IEEE TC, 2015
2015
-
[30]
Prelatpuf: Exploiting DRAM Latency Variations For Generating Robust Device Signatures,
B. M. S. Bahar Talukder, B. Ray, D. Forte, and M. T. Rahman, “Prelatpuf: Exploiting DRAM Latency Variations For Generating Robust Device Signatures, ”IEEE Access, 2019
2019
-
[31]
D-RaNGe: Using Commodity DRAM Devices to Generate True Random Numbers with Low Latency and High Throughput,
J. S. Kim, M. Patel, H. Hassan, L. Orosa, and O. Mutlu, “D-RaNGe: Using Commodity DRAM Devices to Generate True Random Numbers with Low Latency and High Throughput, ” inHPCA, 2019
2019
-
[32]
FracDRAM: Fractional Values in Off- the-shelf DRAM,
F. Gao, G. Tziantzioulis, and D. Wentzlaff, “FracDRAM: Fractional Values in Off- the-shelf DRAM, ” inMICRO, 2022
2022
-
[33]
CODIC: A Low-cost Substrate for Enabling Custom In-DRAM Functionalities and Optimiza- tions,
L. Orosa, Y. Wang, M. Sadrosadati, J. S. Kim, M. Patel, I. Puddu, H. Luo, K. Razavi, J. Gómez-Luna, H. Hassan, N. Mansouri-Ghiasi, S. Ghose, and O. Mutlu, “CODIC: A Low-cost Substrate for Enabling Custom In-DRAM Functionalities and Optimiza- tions, ” inISCA, 2021
2021
-
[34]
DR-STRaNGe: End-to-end System Design For Dram-based True Random Number Generators,
F. N. Bostanci, A. Olgun, L. Orosa, A. G. Yaglikci, J. S. Kim, H. Hassan, O. Ergin, and O. Mutlu, “DR-STRaNGe: End-to-end System Design For Dram-based True Random Number Generators, ” inHPCA, 2022
2022
-
[35]
ComputeDRAM: In-memory Compute using Off-the-shelf DRAMs,
F. Gao, G. Tziantzioulis, and D. Wentzlaff, “ComputeDRAM: In-memory Compute using Off-the-shelf DRAMs, ” inMICRO, 2019
2019
-
[36]
The DRAM Latency PUF: Quickly Evaluating Physical Unclonable Functions by Exploiting The Latency-Reliability Tradeoff in Modern Commodity DRAM Devices,
J. S. Kim, M. Patel, H. Hassan, and O. Mutlu, “The DRAM Latency PUF: Quickly Evaluating Physical Unclonable Functions by Exploiting The Latency-Reliability Tradeoff in Modern Commodity DRAM Devices, ” inHPCA, 2018
2018
-
[37]
JEDEC,JESD79-4C: DDR4 SDRAM Standard, 2020
2020
-
[38]
JEDEC,JESD79-5: DDR5 SDRAM Standard, 2020
2020
-
[39]
RowClone: Fast and Energy-efficient In-DRAM Bulk Data Copy and Initialization,
V. Seshadri, Y. Kim, C. Fallin, D. Lee, R. Ausavarungnirun, G. Pekhimenko, Y. Luo, O. Mutlu, P. B. Gibbons, M. A. Kozuchet al., “RowClone: Fast and Energy-efficient In-DRAM Bulk Data Copy and Initialization, ” inMICRO, 2013
2013
-
[40]
PiDRAM: A Holistic End-to-end FPGA-based Framework for Processing-in-DRAM,
A. Olgun, J. G. Luna, K. Kanellopoulos, B. Salami, H. Hassan, O. Ergin, and O. Mutlu, “PiDRAM: A Holistic End-to-end FPGA-based Framework for Processing-in-DRAM, ” TACO, 2022
2022
-
[41]
DRAM Bender: An Extensible and Versatile FPGA-based Infrastructure to Easily Test State-of-the-art DRAM Chips,
A. Olgun, H. Hassan, A. G. Yağlıkcı, Y. C. Tuğrul, L. Orosa, H. Luo, M. Patel, O. Ergin, and O. Mutlu, “DRAM Bender: An Extensible and Versatile FPGA-based Infrastructure to Easily Test State-of-the-art DRAM Chips, ”TCAD, 2023
2023
-
[42]
JEDEC,JESD209-4B: Low Power Double Data Rate 4 (LPDDR4) Standard, 2017
2017
-
[43]
JEDEC,JESD209-5A: LPDDR5 SDRAM Standard, 2020
2020
-
[44]
The gem5 Simulator,
N. Binkert, B. Beckmann, G. Black, S. K. Reinhardt, A. Saidi, A. Basu, J. Hestness, D. R. Hower, T. Krishna, S. Sardashtiet al., “The gem5 Simulator, ”ACM SIGARCH Computer Architecture News, 2011
2011
-
[45]
Ramulator: A Fast And Extensible DRAM Simulator,
Y. Kim, W. Yang, and O. Mutlu, “Ramulator: A Fast And Extensible DRAM Simulator, ” CAL, 2016
2016
-
[46]
Sim2PIM: A Complete Simulation Framework for Processing-in-memory,
B. E. Forlin, P. C. Santos, A. E. Becker, M. A. Alves, and L. Carro, “Sim2PIM: A Complete Simulation Framework for Processing-in-memory, ”Journal Of Systems Architecture, 2022
2022
-
[47]
Gem5-GPU: A Hetero- geneous CPU-GPU Simulator,
J. Power, J. Hestness, M. S. Orr, M. D. Hill, and D. A. Wood, “Gem5-GPU: A Hetero- geneous CPU-GPU Simulator, ”IEEE CAL, 2014
2014
-
[48]
Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator,
H. Luo, Y. C. Tuğrul, F. N. Bostancı, A. Olgun, A. G. Yağlıkçı, , and O. Mutlu, “Ramulator 2.0: A Modern, Modular, and Extensible DRAM Simulator, ”IEEE CAL, 2024
2024
-
[49]
Flipping Bits in Memory without Accessing Them: An Experimental Study of DRAM Disturbance Errors,
Y. Kim, R. Daly, J. Kim, C. Fallin, J. H. Lee, D. Lee, C. Wilkerson, K. Lai, and O. Mutlu, “Flipping Bits in Memory without Accessing Them: An Experimental Study of DRAM Disturbance Errors, ” inISCA, 2014
2014
-
[50]
A Deeper Look Into RowHammer’s Sensitivities: Experimental Analysis of Real DRAM Chips and Implications on Future Attacks and Defenses,
L. Orosa, A. G. Yağlıkçı, H. Luo, A. Olgun, J. Park, H. Hassan, M. Patel, J. S. Kim, and O. Mutlu, “A Deeper Look Into RowHammer’s Sensitivities: Experimental Analysis of Real DRAM Chips and Implications on Future Attacks and Defenses, ” inMICRO, 2021
2021
-
[51]
Spatial Variation-aware Read Disturbance Defenses: Experimental Analysis of Real DRAM Chips and Implications on Future Solutions,
A. G. Yağlıkçı, Y. C. Tuğrul, G. F. Oliveira, İ. E. Yüksel, A. Olgun, H. Luo, and O. Mutlu, “Spatial Variation-aware Read Disturbance Defenses: Experimental Analysis of Real DRAM Chips and Implications on Future Solutions, ” inHPCA, 2024
2024
-
[52]
Understanding Rowhammer Under Reduced Wordline Voltage: An Experimental Study Using Real DRAM Devices,
A. G. Yağlıkcı, H. Luo, G. F. De Oliviera, A. Olgun, M. Patel, J. Park, H. Hassan, J. S. Kim, L. Orosa, and O. Mutlu, “Understanding Rowhammer Under Reduced Wordline Voltage: An Experimental Study Using Real DRAM Devices, ” inDSN, 2022
2022
-
[53]
Strober: Fast and Accurate Sample-based Energy Simulation for Arbitrary RTL,
D. Kim, A. Izraelevitz, C. Celio, H. Kim, B. Zimmer, Y. Lee, J. Bachrach, and K. Asanović, “Strober: Fast and Accurate Sample-based Energy Simulation for Arbitrary RTL, ”ACM SIGARCH Computer Architecture News, vol. 44, no. 3, pp. 128–139, 2016
2016
-
[54]
Chipyard: Integrated Design, Simulation, and Implementation Framework for Custom SoCs,
A. Amid, D. Biancolin, A. Gonzalez, D. Grubb, S. Karandikar, H. Liew, A. Magyar, H. Mao, A. Ou, N. Pembertonet al., “Chipyard: Integrated Design, Simulation, and Implementation Framework for Custom SoCs, ”IEEE MICRO, 2020
2020
-
[55]
SoftMC: A Flexible and Practical Open-source Infrastruc- ture for Enabling Experimental DRAM Studies,
H. Hassan, N. Vijaykumar, S. Khan, S. Ghose, K. Chang, G. Pekhimenko, D. Lee, O. Ergin, and O. Mutlu, “SoftMC: A Flexible and Practical Open-source Infrastruc- ture for Enabling Experimental DRAM Studies, ” inHPCA, 2017
2017
-
[56]
FASED: FPGA-accelerated Simulation and Evaluation of DRAM,
D. Biancolin, S. Karandikar, D. Kim, J. Koenig, A. Waterman, J. Bachrach, and K. Asanovic, “FASED: FPGA-accelerated Simulation and Evaluation of DRAM, ” in ISFPGA, 2019
2019
-
[57]
FireSim: FPGA-accelerated Cycle-exact Scale-out System Simulation in the Public Cloud,
S. Karandikar, H. Mao, D. Kim, D. Biancolin, A. Amid, D. Lee, N. Pemberton, E. Amaro, C. Schmidt, A. Chopraet al., “FireSim: FPGA-accelerated Cycle-exact Scale-out System Simulation in the Public Cloud, ” inISCA, 2018
2018
-
[58]
PiMulator: a Fast and Flexible Processing-in-Memory Emulation Platform,
S. Mosanu, M. N. Sakib, T. Tracy, E. Cukurtas, A. Ahmed, P. Ivanov, S. Khan, K. Skadron, and M. Stan, “PiMulator: a Fast and Flexible Processing-in-Memory Emulation Platform, ” inDATE, 2022
2022
-
[59]
FreezeTime: Towards System Emulation through Architectural Virtualization,
S. Mosanu, J. Fixelle, M. N. Sakib, K. Skadron, and M. Stan, “FreezeTime: Towards System Emulation through Architectural Virtualization, ” inIPDPSW, 2023
2023
-
[60]
The Reach Profiler (REAPER): Enabling the Mitigation of DRAM Retention Failures via Profiling at Aggressive Conditions,
M. Patel, J. S. Kim, and O. Mutlu, “The Reach Profiler (REAPER): Enabling the Mitigation of DRAM Retention Failures via Profiling at Aggressive Conditions, ” in ISCA, 2017
2017
-
[61]
A Reconfigurable Real-time SDRAM Controller for Mixed Time-criticality Systems,
S. Goossens, J. Kuijsten, B. Akesson, and K. Goossens, “A Reconfigurable Real-time SDRAM Controller for Mixed Time-criticality Systems, ” inCODES+ISSS, 2013
2013
-
[62]
PARDIS: A Programmable Memory Controller For The DDRx Interfacing Standards,
M. N. Bojnordi and E. Ipek, “PARDIS: A Programmable Memory Controller For The DDRx Interfacing Standards, ” inISCA, 2012
2012
-
[63]
Polybench: The Polyhedral Benchmark Suite,
L.-N. Pouchet, “Polybench: The Polyhedral Benchmark Suite, ” https://www.cs. colostate.edu/~pouchet/software/polybench/
-
[64]
NVIDIA Jetson Nano,
NVIDIA, “NVIDIA Jetson Nano, ” https://developer.nvidia.com/embedded/ jetson-nano
-
[65]
Impulse: Building a Smarter Memory Controller,
J. Carter, W. Hsieh, L. Stoller, M. Swanson, L. Zhang, E. Brunvand, A. Davis, C.- C. Kuo, R. Kuramkote, M. Parkeret al., “Impulse: Building a Smarter Memory Controller, ” inHPCA, 1999
1999
-
[66]
Improving DRAM Performance by Parallelizing Refreshes with Accesses,
K. K.-W. Chang, D. Lee, Z. Chishti, A. R. Alameldeen, C. Wilkerson, Y. Kim, and O. Mutlu, “Improving DRAM Performance by Parallelizing Refreshes with Accesses, ” inHPCA, 2014
2014
-
[67]
Understanding Reduced-voltage Operation in Modern DRAM De- vices: Experimental Characterization, Analysis, and Mechanisms,
K. Changet al., “Understanding Reduced-voltage Operation in Modern DRAM De- vices: Experimental Characterization, Analysis, and Mechanisms, ” inSIGMETRICS, 2017
2017
-
[68]
TRRespass: Exploiting The Many Sides of Target Row Refresh,
P. Frigo, E. Vannacc, H. Hassan, V. Van Der Veen, O. Mutlu, C. Giuffrida, H. Bos, and K. Razavi, “TRRespass: Exploiting The Many Sides of Target Row Refresh, ” in S&P, 2020
2020
-
[69]
What Your DRAM Power Models Are Not Telling You: Lessons From A Detailed Experimental Study,
S. Ghose, A. G. Yaglikci, R. Gupta, D. Lee, K. Kudrolli, W. Liu, H. Hassan, K. Chang, N. Chatterjee, A. Agrawal, M. O’Connor, and O. Mutlu, “What Your DRAM Power Models Are Not Telling You: Lessons From A Detailed Experimental Study, ” in SIGMETRICS, 2018
2018
-
[70]
Self-optimizing Memory Con- trollers: A Reinforcement Learning Approach,
E. Ipek, O. Mutlu, J. F. Martinez, and R. Caruana, “Self-optimizing Memory Con- trollers: A Reinforcement Learning Approach, ” inISCA, 2008
2008
-
[71]
ABACuS: All-bank Activation Counters for Scalable and Low Overhead Rowhammer Mitigation,
A. Olgun, Y. C. Tugrul, N. Bostanci, I. E. Yuksel, H. Luo, S. Rhyner, A. G. Yaglikci, G. F. Oliveira, and O. Mutlu, “ABACuS: All-bank Activation Counters for Scalable and Low Overhead Rowhammer Mitigation, ” inUSENIX Security, 2024
2024
-
[72]
BreakHammer: Enhancing RowHammer Mitigations by Carefully Throttling Suspect Threads,
O. Canpolat, A. G. Yağlıkçı, A. Olgun, İ. E. Yüksel, Y. C. Tuğrul, K. Kanellopoulos, O. Ergin, and O. Mutlu, “BreakHammer: Enhancing RowHammer Mitigations by Carefully Throttling Suspect Threads, ”MICRO, 2024
2024
-
[73]
A Case for Exploiting Subarray-level Parallelism (SALP) in DRAM,
Y. Kim, V. Seshadri, D. Lee, J. Liu, O. Mutlu, Y. Kim, V. Seshadri, D. Lee, J. Liu, and O. Mutlu, “A Case for Exploiting Subarray-level Parallelism (SALP) in DRAM, ” in ISCA, 2012
2012
-
[74]
Controller for A Synchronous DRAM That Maximizes Throughput by Allowing Memory Requests And Commands to be Issued Out Of Order,
W. K. Zuravleff and T. Robinson, “Controller for A Synchronous DRAM That Maximizes Throughput by Allowing Memory Requests And Commands to be Issued Out Of Order, ” US Patent: 5,630,096, 1997
1997
-
[75]
Memory Access Scheduling,
S. Rixner, W. J. Dally, U. J. Kapasi, P. Mattson, and J. D. Owens, “Memory Access Scheduling, ” inISCA, 2000
2000
-
[76]
Parallelism-aware Batch Scheduling: Enhancing Both Performance And Fairness Of Shared DRAM Systems,
O. Mutlu and T. Moscibroda, “Parallelism-aware Batch Scheduling: Enhancing Both Performance And Fairness Of Shared DRAM Systems, ” inISCA, 2008
2008
-
[77]
Stall-time Fair Memory Access Scheduling For Chip Multiprocessors,
O. Mutlu and T. Moscibroda, “Stall-time Fair Memory Access Scheduling For Chip Multiprocessors, ” inMICRO, 2007
2007
-
[78]
The Blacklisting Memory Scheduler: Achieving High Performance And Fairness At Low Cost,
L. Subramanian, D. Lee, V. Seshadri, H. Rastogi, and O. Mutlu, “The Blacklisting Memory Scheduler: Achieving High Performance And Fairness At Low Cost, ” in ICCD, 2014
2014
-
[79]
BLISS: Balancing Performance, Fairness And Complexity In Memory Access Scheduling,
L. Subramanian, D. Lee, V. Seshadri, H. Rastogi, and O. Mutlu, “BLISS: Balancing Performance, Fairness And Complexity In Memory Access Scheduling, ”TPDS, 2016
2016
-
[80]
OpenHW Group – Website,
OpenHW Group, “OpenHW Group – Website, ” https://www.openhwgroup.org/
-
[81]
The Rocket Chip Generator,
K. Asanovic, R. Avizienis, J. Bachrach, S. Beamer, D. Biancolin, C. Celio, H. Cook, D. Dabbelt, J. Hauser, A. Izraelevitzet al., “The Rocket Chip Generator, ”EECS Department, University Of California, Berkeley, Tech. Rep. UCB/EECS-2016-17, 2016
2016
-
[82]
IBM z15: A 12-Core 5.2GHz Microprocessor,
C. Berry, B. Bell, A. Jatkowski, J. Surprise, J. Isakson, O. Geva, B. Deskin, M. Ci- chanowski, D. Hamid, C. Cavitt, G. Fredeman, A. Saporito, A. Mishra, A. Buyuk- tosunoglu, T. Webel, P. Lobo, P. Parashurama, R. Bertran, D. Chidambarrao, 14 D. Wolpert, and B. Bruen, “IBM z15:...
2020
-
[83]
IBM Telum: a 16-Core 5+ GHz DCM,
O. Geva, C. Berry, R. Sonnelitter, D. Wolpert, A. Collura, T. Strach, D. Phan, C. Lichtenau, A. Buyuktosunoglu, H. Harrer, J. Zitz, C. Marquart, D. Malone, T. Webel, A. Jatkowski, J. Isakson, D. Hamid, M. Cichanowski, M. Romain, F. Hasan, K. Williams, J. Surprise, C. Cavitt, a...
2022
-
[84]
DRAM Bender — Github Repository,
SAFARI Research Group, “DRAM Bender — Github Repository, ” https://github.com/ CMU-SAFARI/DRAM-Bender, 2022
2022
-
[85]
Chisel: Constructing Hardware in a Scala Embedded Language,
J. Bachrach, H. Vo, B. Richards, Y. Lee, A. Waterman, R. Avižienis, J. Wawrzynek, and K. Asanović, “Chisel: Constructing Hardware in a Scala Embedded Language, ” inDAC, 2012
2012
-
[86]
A Highly Productive Implementation of an Out-of-order Processor Gen- erator,
C. Celio, “A Highly Productive Implementation of an Out-of-order Processor Gen- erator, ” Ph.D. dissertation, UC Berkeley, 2018
2018
-
[87]
LMbench: Portable Tools for Performance Analysis
L. W. McVoy, C. Staelinet al., “LMbench: Portable Tools for Performance Analysis.” inUSENIX ATC, 1996
1996
-
[88]
Xilinx VCU 108 FPGA Board,
Xilinx Inc., “Xilinx VCU 108 FPGA Board, ” https://www.xilinx.com/products/ boards-and-kits/ek-u1-vcu108-g.html
-
[89]
Ramulator 2.0,
SAFARI Research Group, “Ramulator 2.0, ” https://github.com/CMU-SAFARI/ ramulator2
-
[90]
Design-induced Latency Variation in Modern DRAM Chips: Characterization, Analysis, and Latency Reduction Mechanisms,
D. Lee, S. Khan, L. Subramanian, S. Ghose, R. Ausavarungnirun, G. Pekhimenko, V. Seshadri, and O. Mutlu, “Design-induced Latency Variation in Modern DRAM Chips: Characterization, Analysis, and Latency Reduction Mechanisms, ” inSIG- METRICS, 2017
2017
-
[91]
DDR4 SDRAM EDY4016A - 256Mb x 16,
Micron, “DDR4 SDRAM EDY4016A - 256Mb x 16, ” https://mm.digikey.com/ Volume0/opasdata/d220001/medias/docus/1135/EDY4016A.pdf, 2014
2014
-
[92]
Space/time Trade-offs In Hash Coding With Allowable Errors,
B. H. Bloom, “Space/time Trade-offs In Hash Coding With Allowable Errors, ”Com- munications Of The ACM, 1970
1970
-
[93]
RAIDR: Retention-aware Intelligent DRAM Refresh,
J. Liuet al., “RAIDR: Retention-aware Intelligent DRAM Refresh, ” 2012
2012
-
[94]
MEG: A RISCV-based System Emu- lation Infrastructure for Near-data Processing Using FPGAs and High-bandwidth Memory,
J. Zhang, Y. Zha, N. Beckwith, B. Liu, and J. Li, “MEG: A RISCV-based System Emu- lation Infrastructure for Near-data Processing Using FPGAs and High-bandwidth Memory, ” inTRETS, 2020
2020
-
[95]
Fast Bulk Bitwise AND And OR In DRAM,
V. Seshadri, K. Hsieh, A. Boroumand, D. Lee, M. A. Kozuch†, O. Mutlu, P. B. Gibbons, and T. C. Mowry, “Fast Bulk Bitwise AND And OR In DRAM, ” inCAL, 2015
2015
-
[96]
Ambit: In-memory Accelerator For Bulk Bitwise Operations Using Commodity DRAM Technology,
V. Seshadri, D. Lee, T. Mullins, H. Hassan, A. Boroumand, J. Kim, M. A. Kozuch, O. Mutlu, P. B. Gibbons, and T. C. Mowry, “Ambit: In-memory Accelerator For Bulk Bitwise Operations Using Commodity DRAM Technology, ” inMICRO, 2017
2017
-
[97]
Softmc — Github Repository,
SAFARI Research Group, “Softmc — Github Repository, ” https://github.com/ CMU-SAFARI/softmc, 2017
2017
-
[98]
The DRAM Latency PUF: Quickly Evaluating Physical Unclonable Functions By Exploiting The Latency–reliability Tradeoff In Modern DRAM Devices,
J. Kim, M. Patel, H. Hassan, and O. Mutlu, “The DRAM Latency PUF: Quickly Evaluating Physical Unclonable Functions By Exploiting The Latency–reliability Tradeoff In Modern DRAM Devices, ” inHPCA, 2018
2018
-
[99]
Exploiting DRAM Latency Variations for Generating True Random Numbers,
B. M. S. Bahar Talukder, J. Kerns, B. Ray, T. Morris, and M. T. Rahman, “Exploiting DRAM Latency Variations for Generating True Random Numbers, ” inICCE, 2019
2019
-
[100]
AVATAR: A Variable- retention-time (VRT) Aware Refresh for DRAM Systems,
M. Qureshi, D.-H. Kim, S. Khan, P. Nair, and O. Mutlu, “AVATAR: A Variable- retention-time (VRT) Aware Refresh for DRAM Systems, ” inDSN, 2015
2015
-
[101]
RAIDR: Retention-aware Intelligent DRAM Refresh,
J. Liu, B. Jaiyen, R. Veras, and O. Mutlu, “RAIDR: Retention-aware Intelligent DRAM Refresh, ” inISCA, 2012
2012
-
[102]
Retention-aware Placement In DRAM (RAPID): Software Methods For Quasi-non-volatile DRAM,
R. K. Venkatesanet al., “Retention-aware Placement In DRAM (RAPID): Software Methods For Quasi-non-volatile DRAM, ” inHPCA, 2006. 15
2006
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.