REVIEW 4 cited by
Memory-Centric Computing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Memory-centric computing aims to enable computation capability in and near all places where data is generated and stored. As such, it can greatly reduce the large negative performance and energy impact of data access and data movement, by fundamentally avoiding data movement and reducing data access latency & energy. Many recent studies show that memory-centric computing can greatly improve system performance and energy efficiency. Major industrial vendors and startup companies have also recently introduced memory chips that have sophisticated computation capabilities. This talk describes promising ongoing research and development efforts in memory-centric computing. We classify such efforts into two major fundamental categories: 1) processing using memory, which exploits analog operational properties of memory structures to perform massively-parallel operations in memory, and 2) processing near memory, which integrates processing capability in memory controllers, the logic layer of 3D-stacked memory technologies, or memory chips to enable high-bandwidth and low-latency memory access to near-memory logic. We show both types of architectures (and their combination) can enable orders of magnitude improvements in performance and energy consumption of many important workloads, such as graph analytics, databases, machine learning, video processing, climate modeling, genome analysis. We discuss adoption challenges for the memory-centric computing paradigm and conclude with some research & development opportunities.
Forward citations
Cited by 4 Pith papers
-
Neural Weight Compression for Language Models
A single learned neural codec, trained once on real LLM weights, compresses Llama-scale models to 4-6 bits per weight with near-FP16 accuracy, beating hand-crafted quantization at those bitrates.
-
LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
LUT-DLA converts neural network layers into vector-quantized lookup tables, claiming sub-1-bit-equivalent inference with 1.4-7.0x power and 1.5-146.1x area efficiency gains over conventional accelerators.
-
New Tools, Programming Models, and System Support for Processing-in-Memory Architectures
A PhD dissertation contributing DAMOV (data-movement benchmark suite), MIMDRAM and Proteus (processing-using-DRAM designs), and DaPPA (near-memory programming framework), claiming large performance and energy gains fo...
-
Next-Gen Computing Systems with Compute Express Link: a Comprehensive Survey
A structured survey of CXL-based computing system research, organized by memory expansion, unified memory, and distributed memory pooling and sharing.
Discussion (0). Continue with ORCID to comment.