REVIEW 3 major objections 6 minor 2 cited by
Next-Gen Computing Systems with Compute Express Link: a Comprehensive Survey
T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This survey argues that CXL makes memory a coherent, first-class fabric resource and organizes CXL systems into memory expansion, unified memory, and distributed sharing.
desk verdict A useful, current CXL survey that is one revision away from being a dependable reference; the taxonomy is rougher than advertised. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery that carries the argument is the CXL protocol stack: CXL.io for device discovery and direct memory access (DMA), CXL.cache for devices caching host memory, and CXL.mem for hosts caching device-attached memory as host-managed device memory. These protocols compose into three device types, and the CXL Switch extends them from one machine to a pool of hosts and memory modules. That same stack generates the survey's taxonomy: Type-3 memory devices anchor the Memory Expansion category, Type-2 accelerator devices anchor Unified Memory, and switches plus CXL 3.0 multi-host sharing anchor distributed pooling and shared memory. The survey also leans on tiered-memory page placement and near-memory processing as the mechanisms that make expansion usable despite CXL's higher latency.
What would settle it
A concrete check: saturate a commercial CXL memory device with random reads and measure sustained latency; if loaded latency exceeds a few microseconds, or if a rack-scale path through CXL switches is slower than an equivalent remote direct memory access path, then the survey's claim that CXL provides near-DRAM, network-beating memory semantics for disaggregation fails.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that CXL changes what an interconnect is for: instead of shuttling data between disjoint memory domains, a CXL fabric makes memory itself a shared, addressable resource. In a single machine this takes the form of tiered memory systems that keep hot pages in fast DRAM and cold pages in slower CXL memory, near-memory processing that computes beside the data, and a hardware-coherent unified space that lets accelerators read and write host memory without a CPU relay. Across machines, CXL 2.0 memory pooling lets one host claim memory from a pool, while CXL 3.0 memory sharing lets several hosts share one memory module, turning shared CXL memory into a communication channel for remote procedure calls and synchronization. The paper also assembles real-hardware measurements placing CXL memory between DRAM and non-volatile memory in latency (roughly 170–250 nanoseconds) and argues that future systems should be engineered around this memory fabric rather than around individual nodes.
Load-bearing premise
The survey's organizing claim rests on the taxonomy announced in the introduction and used in Sections III and IV: CXL systems divide cleanly into Memory Expansion, Unified Memory, and distributed pooling or sharing, with near-memory processing counted under Memory Expansion; if those categories overlap, as near-memory processing's compute-offload nature suggests, the organization weakens.
Editorial extensions
If this is right
- If this picture is right, memory expansion is a practical pressure valve for the memory wall: applications that tolerate roughly 170–250 ns access can use terabyte-scale CXL memory, with tiered placement keeping hot pages in DRAM.
- If this picture is right, hardware cache coherence will let GPUs and other accelerators share memory directly, removing the CPU-mediated data copies that today dominate heterogeneous workloads.
- If this picture is right, memory pooling decouples memory capacity from individual servers, so data centers can allocate memory where demand actually is rather than where hardware was purchased.
- If this picture is right, CXL 3.0 shared memory becomes a low-latency communication substrate for remote procedure calls, distributed synchronization, and serverless state, bypassing the network stack inside a rack.
- If this picture is right, the design center of future systems shifts to memory-centric computing, in which the distance between data and compute is measured in CXL fabric hops and managed explicitly.
Reading between the lines
- Beyond the paper's classification, the line between Memory Expansion and near-memory processing is likely to blur: once CXL devices carry general-purpose compute cores, capacity expansion becomes a compute-placement problem as much as a capacity problem.
- The survey presents tiering and interleaving as competing policies, but the measurements it cites imply that workload-aware systems will combine both—tiering for latency-sensitive phases and interleaving for bandwidth-hungry phases—because DRAM latency degrades sharply near bandwidth saturation.
- An extension the survey does not run is a direct benchmark of CXL shared-memory RPC against remote direct memory access (RDMA) at identical intra-rack distances; its own numbers suggest CXL should win inside a rack and lose across racks, which would make CXL-over-Ethernet hybrids the natural scale-out path.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a survey of computing systems built around Compute Express Link (CXL). It introduces CXL protocols and device types, reviews evaluation and simulation platforms, proposes a taxonomy for single-machine systems (Memory Expansion and Unified Memory) and for distributed systems (Memory Pooling and Memory Sharing), discusses near-memory processing and heterogeneous accelerator integration, and outlines future research directions toward memory-centric computing. The survey compiles roughly 90 references from 2019 to 2024 and positions itself as a comprehensive review of recent CXL research from single-machine to distributed settings.
Significance. The manuscript is timely and covers a rapidly evolving area. Its main contribution is organizational: it collects and classifies a substantial body of recent CXL systems research. If the taxonomy were consistent, the survey would serve as a useful entry point for newcomers and as a reference for researchers. The paper contains no derivations or fitted models, so circularity is not a concern; its value depends on the accuracy and coherence of its literature organization. The current taxonomy, however, is applied inconsistently, and the unsupported publication-count claim in the Introduction weakens the survey's credibility. These issues are fixable, but they affect the central contribution and require revision.
major comments (3)
- [Abstract, Sections III-IV, Table I] The survey's central organizing claim is that single-machine CXL research partitions cleanly into Memory Expansion and Unified Memory based on interconnection type (processor-to-memory vs. processor-to-accelerator). This partition is not consistently applied. Section IV, 'Near-Memory Processing,' is listed under Memory Expansion in Table I and announced in the Section III opening paragraph, yet near-memory processing is a compute-offload mechanism and does not change the interconnect type that allegedly defines the taxonomy. Moreover, ReCXL [14] appears both in Section III.C (Application-Specific Optimizations) and in Section IV.A (Workload-Customized Near-Memory Process), and references [25] and [26] appear in multiple rows of Table I. The categories are therefore neither mutually exclusive nor based on a uniformly applied criterion. The taxonomy should be redefined—for example, by treating near-memory processing as an orthogonal dimension—or the paper should explicitly present the categories as overlapping research themes rather than as a partition.
- [Section I (Introduction)] The Introduction states that 'only 10 paper were published during 2019 to 2022, while 40 in 2023, 51 in 2024' with no citation, source, or counting methodology. This quantitative claim is used to justify the survey's timeliness and the need for a new survey. The authors should either provide the bibliographic corpus used for the count, state the inclusion criteria (e.g., which venues and search terms), or soften the claim to avoid an unverifiable statistic that undermines the paper's scholarly framing.
- [Section V.A] The subsection 'Memory Expansion with CPU Relay' is classified under Unified Memory, but the works it discusses (e.g., [17], [51]) describe using CXL to expand accelerator memory through a CPU relay. That is memory expansion for accelerators, not the unified-memory scenario defined in Section V's introduction, which emphasizes direct CXL access between coherent accelerators and processors. This placement blurs the boundary between the two headline categories and suggests that the taxonomy's criterion (interconnection type) is not consistently applied across all subsections. The authors should either reclassify these works or revise the category definitions to accommodate relay-based accelerator memory expansion.
minor comments (6)
- [Throughout] There are numerous typographical and grammatical errors, including 'Memory Expansion focus' (should be 'focuses'), 'Thet propose' in Section III.A, 'system‘s' with an unbalanced smart quote, and inconsistent British/American spelling ('synchronisation' vs. 'synchronization'). A careful proofreading pass is recommended.
- [Section II.B] The affiliation 'UIU' should be 'UIUC' (University of Illinois Urbana-Champaign) when referring to the authors of [24].
- [Reference [43]] Reference [43] lists the author as 'Y. T. ea tl.,' which is malformed; the author list should be completed or the entry replaced with a properly formatted citation.
- [Table I] Table I lists the same references, such as [25] and [26], in multiple rows (e.g., Latency/Bandwidth, Use Case: Regular Applications, Use Case: Large Language Models). While a single work can support multiple topics, the table should explicitly state that rows are not exclusive, or it will reinforce the taxonomy concerns raised about category disjointness.
- [Section VIII.A.1] The claim that memory interleaving makes 'the overall bandwidth can be the sum of CXL memory and DRAM in theory' should be qualified or cited, because the achievable aggregate bandwidth depends on the memory controller, CXL root port, and PCIe topology, and may not reach the theoretical sum.
- [Section IX] The 'Memory-Centric Computing in CXL Fabric' section presents a speculative research vision. It would be helpful to explicitly label this section as the authors' position or vision rather than as established results, and to distinguish it from the survey's descriptive contributions.
Circularity Check
No significant circularity: the survey is expository and makes no predictive or derived claims that reduce to its own inputs.
full rationale
This is a comprehensive survey paper, not a derivation or modeling paper. It contains no fitted parameters, no equations that transform inputs into outputs, and no prediction that could reduce by construction to its own assumptions. The central claim is a classification of existing CXL research into categories (Memory Expansion, Unified Memory, and distributed pooling/sharing), which is an organizational judgment rather than a derived result. The taxonomy could be criticized for consistency—for example, ReCXL [14] appears both under 'Application-Specific Optimizations' in Section III.C and under 'Workload-Customized Near-Memory Process' in Section IV.A, and near-memory processing is arguably a compute-offload mechanism rather than memory expansion—but such a critique concerns the quality or mutual exclusivity of a survey taxonomy, not circular reasoning. The paper does not define 'Memory Expansion' in terms of the conclusions of the surveyed works, nor does it invoke a self-citation chain to force its classification. Citations such as TPP [9] and Pond [18] provide external latency measurements and system results that are independent of this paper's claims. There are no self-citations by the authors that are load-bearing, and no uniqueness theorem or ansatz is imported from the authors' prior work. Because the survey is self-contained as an expository review and derives no quantitative result from its own categories, the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption CXL hardware measurements (latency, bandwidth, coherence behavior) reported in the cited evaluations are accurate and representative.
- domain assumption The surveyed papers are a representative sample of CXL research through 2024.
- domain assumption Future CXL hardware (CXL 3.0 switches, Type-2 devices) will become available as assumed in the vision sections.
Cite this review
Pith. "Pith review of Next-Gen Computing Systems with Compute Express Link: a Comprehensive Survey." pith.science (2026). https://pith.science/paper/PWK7ABSJ
@misc{pith2026241220249,
author = {Pith},
title = {Pith review of: Next-Gen Computing Systems with Compute Express Link: a Comprehensive Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/PWK7ABSJ}},
note = {Machine review of arXiv:2412.20249}
}
read the original abstract
Interconnection is crucial for computing systems. However, the current interconnection performance between processors and devices, such as memory devices and accelerators, significantly lags behind their computing performance, severely limiting the overall performance. To address this challenge, Intel proposes Compute Express Link (CXL), an open industry-standard interconnection. With memory semantics, CXL offers low-latency, scalable, and coherent interconnection between processors and devices. This paper introduces recent advances in CXL-based computing systems from single-machine to distributed. In single-machine systems, we classify existing research into two categories: Memory Expansion and Unified Memory. Memory Expansion focus on processors and memory, aims to address memory wall challenge. Unified memory focus on processors and accelerators, aims to enhance collaboration in heterogeneous computing systems. In distributed systems, we present how to build efficient disaggregation systems based on CXL infrastructure, enabling resource pooling and sharing. Finally, we discuss the future research and envision memory-centric computing with CXL.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 2 Pith papers
-
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving
KV-cache serving systems concentrate into five archetypes under a four-axis taxonomy, with ownership explaining residual distributed design variance and seven measurement gaps blocking next steps.
-
Bridging the Cognitive Gap: A Unified Memory Paradigm for 6G Agentic AI-RAN
Proposal to replace message-passing interfaces in AI-RAN with zero-copy CXL shared memory, organized as reflexive, contextual, and evolutionary cognitive loops.
Reference graph
Works this paper leans on
-
[2]
An introduction to the compute express link (cxl) interconnect,
D. Das Sharma, R. Blankenship, and D. Berger, “An introduction to the compute express link (cxl) interconnect,” ACM Comput. Surv., vol. 56, no. 11, Jul. 2024
2024
-
[14]
Enabling efficient large recommendation model training with near cxl memory processing,
H. Liu, L. Zheng, Y . Huang, J. Zhou, C. Liu, R. Wang, X. Liao, H. Jin, and J. Xue, “Enabling efficient large recommendation model training with near cxl memory processing,” in 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA) , 2024
2024
-
[25]
Exploring performance and cost optimization with asic-based cxl memory,
Y . Tang, P. Zhou, W. Zhang, H. Hu, Q. Yang, H. Xiang, T. Liu, J. Shan, R. Huang, C. Zhao et al., “Exploring performance and cost optimization with asic-based cxl memory,” inProceedings of the Nineteenth European Conference on Computer Systems , 2024, pp. 818–833
2024
-
[26]
Ex- ploring and evaluating real-world cxl: Use cases and system adoption,
J. Liu, X. Wang, J. Wu, S. Yang, J. Ren, B. Shankar, and D. Li, “Ex- ploring and evaluating real-world cxl: Use cases and system adoption,” arXiv preprint arXiv:2405.14209 , 2024
arXiv 2024
-
[17]
Gpu graph processing on cxl- based microsecond-latency external memory,
S. Sano, Y . Bando, K. Hiwada, H. Kajihara, T. Suzuki, Y . Nakanishi, D. Taki, A. Kaneko, and T. Shiozawa, “Gpu graph processing on cxl- based microsecond-latency external memory,” in Proceedings of the SC ’23 Workshops of The International Conference on High Performance Computing, Network, Storage, and Analysis, ser. SC-W ’23. New York, NY , USA: Associa...
work page 2023
-
[51]
Lmb: Augmenting pcie devices with cxl-linked memory buffer,
J. Wang, X. Zhang, C. Tang, X. Chen, and T. Lu, “Lmb: Augmenting pcie devices with cxl-linked memory buffer,” 2024
work page 2024
-
[1]
Ai and memory wall,
A. Gholami, Z. Yao, S. Kim, C. Hooper, M. W. Mahoney, and K. Keutzer, “Ai and memory wall,” IEEE Micro , vol. 44, no. 3, pp. 33–39, 2024
2024
-
[3]
Memory disaggregation: Advances and open challenges,
H. Al Maruf and M. Chowdhury, “Memory disaggregation: Advances and open challenges,” SIGOPS Oper. Syst. Rev., vol. 57, no. 1, p. 29–37, Jun. 2023. [Online]. Available: https://doi.org/10.1145/3606557.3606562
arXiv 2023
Show all 93 references
-
[4]
Distributed shared memory: A survey of issues and algorithms,
B. Nitzberg and V . Lo, “Distributed shared memory: A survey of issues and algorithms,” Computer, vol. 24, no. 8, p. 52–60, Aug. 1991. [Online]. Available: https://doi.org/10.1109/2.84877
1991 doi
-
[5]
Compute express link 3.0,
“Compute express link 3.0,” 2022. [Online]. Available: https://www.computeexpresslink.org/ files/ugd/0c1418 a8713008916044ae9604405d10a7773b.pdf
2022
-
[6]
Compute express link cxl 3.0 is the exciting building block for dis- aggregation
“Compute express link cxl 3.0 is the exciting building block for dis- aggregation.” 2022. [Online]. Available: https://www.servethehome.com/ compute-expresslink-cxl-3-0-is-the-exciting-building-block-for-disaggregation/
2022
-
[7]
Compute express link™: The breakthrough cpu-to-device intercon- nect
“Compute express link™: The breakthrough cpu-to-device intercon- nect.” 2022. [Online]. Available: https://www.computeexpresslink.org/ home
2022
-
[8]
Compute express link (cxl): Enabling heterogeneous data-centric computing with heterogeneous memory hierarchy,
D. D. Sharma, “Compute express link (cxl): Enabling heterogeneous data-centric computing with heterogeneous memory hierarchy,” IEEE Micro, vol. 43, no. 2, pp. 99–109, 2023
2023
-
[9]
Tpp: Transparent page placement for cxl-enabled tiered-memory,
H. Maruf, H. Wang, A. Dhanotia, J. Weiner, N. Agarwal, P. Bhat- tacharya, C. Petersen, M. Chowdhury, S. Kanaujia, and P. Chauhan, “Tpp: Transparent page placement for cxl-enabled tiered-memory,” ASPLOS, Jan 2023
2023
-
[10]
Tiered memory management: Access latency is the key!
M. Vuppalapati and R. Agarwal, “Tiered memory management: Access latency is the key!” in Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles, ser. SOSP ’24. New York, NY , USA: Association for Computing Machinery, 2024, p. 79–94
2024
-
[11]
CXL- ANNS: Software-Hardware collaborative memory disaggregation and computation for Billion-Scale approximate nearest neighbor search,
J. Jang, H. Choi, H. Bae, S. Lee, M. Kwon, and M. Jung, “CXL- ANNS: Software-Hardware collaborative memory disaggregation and computation for Billion-Scale approximate nearest neighbor search,” in 2023 USENIX Annual Technical Conference (USENIX ATC 23). Boston, MA: USENIX Asso...
2023
-
[12]
Bridging software-hardware for cxl memory disaggregation in billion-scale near- est neighbor search,
J. Jang, H. Choi, H. Bae, S. Lee, M. Kwon, and M. Jung, “Bridging software-hardware for cxl memory disaggregation in billion-scale near- est neighbor search,” ACM Trans. Storage, vol. 20, no. 2, Feb. 2024
2024
-
[13]
Neomem: Hardware/software co-design for cxl-native memory tiering,
Z. Zhou, Y . Chen, T. Zhang, Y . Wang, R. Shu, S. Xu, P. Cheng, L. Qu, Y . Xiong, J. Zhang, and G. Sun, “Neomem: Hardware/software co-design for cxl-native memory tiering,” in 2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2024
2024
-
[15]
Breaking barriers: Expanding gpu memory with sub-two digit nanosecond latency cxl controller,
D. Gouk, S. Kang, H. Bae, E. Ryu, S. Lee, D. Kim, J. Jang, and M. Jung, “Breaking barriers: Expanding gpu memory with sub-two digit nanosecond latency cxl controller,” in Proceedings of the 16th ACM Workshop on Hot Topics in Storage and File Systems , ser. HotStorage ’24. New ...
2024
-
[16]
Failure tolerant training with persistent memory disaggregation over cxl,
M. Kwon, J. Jang, H. Choi, S. Lee, and M. Jung, “Failure tolerant training with persistent memory disaggregation over cxl,” IEEE Micro, vol. 43, pp. 66–75, 2023
2023
-
[18]
Pond: Cxl-based memory pooling systems for cloud platforms,
H. Li, D. Berger, S. Novakovic, L. Hsu, D. Ernst, P. Zardoshti, M. Shah, S. Rajadnya, S. Lee, I. Agarwal, M. Hill, M. Fontoura, and R. Bian- chini, “Pond: Cxl-based memory pooling systems for cloud platforms,” ASPLOS, Jan 2023
2023
-
[19]
Direct access, high- performance memory disaggregation with directcxl,
D. Gouk, S. Lee, M. Kwon, and M. Jung, “Direct access, high- performance memory disaggregation with directcxl,” ATC, Jan 2022
2022
-
[20]
Ddc: A vision for a disaggregated datacenter,
M. Ewais and P. Chow, “Ddc: A vision for a disaggregated datacenter,” 2024
2024
-
[21]
Notnets: Accelerating microservices by bypassing the network,
P. Alvaro, M. Adiletta, A. Cockroft, F. Hady, R. Illikkal, E. Ramos, J. Tsai, and R. Soul´e, “Notnets: Accelerating microservices by bypassing the network,” 2024
2024
-
[22]
Partial failure resilient memory management system for (cxl-based) distributed shared memory,
M. Zhang, T. Ma, J. Hua, Z. Liu, K. Chen, N. Ding, F. Du, J. Jiang, T. Ma, and Y . Wu, “Partial failure resilient memory management system for (cxl-based) distributed shared memory,” in Proceedings of the 29th Symposium on Operating Systems Principles, ser. SOSP ’23. New York,...
2023
-
[23]
Redis ltd
“Redis ltd.” 2024. [Online]. Available: https://redis.io/
2024
-
[24]
Demystifying cxl memory with genuine cxl-ready systems and devices,
Y . Sun, Y . Yuan, Z. Yu, R. Kuper, C. Song, J. Huang, H. Ji, S. Agarwal, J. Lou, I. Jeong et al., “Demystifying cxl memory with genuine cxl-ready systems and devices,” in Proceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture , 2023, pp. 105–121
2023
-
[27]
Evaluating emerging cxl- enabled memory pooling for hpc systems,
J. Wahlgren, M. Gokhale, and I. B. Peng, “Evaluating emerging cxl- enabled memory pooling for hpc systems,” in 2022 IEEE/ACM Work- shop on Memory Centric High Performance Computing (MCHPC) . IEEE, 2022, pp. 11–20
2022
-
[28]
Zero-offload: Democratizing billion-scale model training,
J. Ren, S. Rajbhandari, R. Aminabadi, O. Ruwase, S. Yang, M. Zhang, L. Dong, and Y . He, “Zero-offload: Democratizing billion-scale model training,” Cornell University - arXiv,Cornell University - arXiv , Jan 2021
2021
-
[29]
High-throughput generative inference of large language model with a single gpu
Y . Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Zhang, I. Stoica, and C. R ´e, “High-throughput generative inference of large language model with a single gpu.”
-
[30]
Performance evaluation on cxl-enabled hybrid memory pool,
Q. Yang, R. Jin, B. Davis, D. Inupakutika, and M. Zhao, “Performance evaluation on cxl-enabled hybrid memory pool,” in 2022 IEEE Inter- national Conference on Networking, Architecture and Storage (NAS) , 2022, pp. 1–5
2022
-
[31]
Cxlmemsim: A pure software simulated cxl.mem for performance characterization,
Y . Yang, P. Safayenikoo, J. Ma, T. A. Khan, and A. Quinn, “Cxlmemsim: A pure software simulated cxl.mem for performance characterization,” 2023
2023
-
[32]
emucxl: an emulation framework for cxl- based disaggregated memory applications,
R. Gond and P. Kulkarni, “emucxl: an emulation framework for cxl- based disaggregated memory applications,” 2024
2024
-
[33]
A mess of memory system bench- marking, simulation and application profiling,
P. Esmaili-Dokht, F. Sgherzi, V . S. Girelli, I. Boixaderas, M. Carmin, A. Monemi, A. Armejach, E. Mercadal, G. Llort, P. Radojkovi ´c, M. Moreto, J. Gim ´enez, X. Martorell, E. Ayguad ´e, J. Labarta, E. Con- falonieri, R. Dubey, and J. Adlard, “A mess of memory system bench- ...
2024
-
[34]
Dracksim: Simulating cxl-enabled large-scale disaggre- gated memory systems,
A. Puri, K. Bellamkonda, K. Narreddy, J. Jose, V . Tamarapalli, and V . Narayanan, “Dracksim: Simulating cxl-enabled large-scale disaggre- gated memory systems,” in Proceedings of the 38th ACM SIGSIM Con- ference on Principles of Advanced Discrete Simulation , ser. SIGSIM- PAD...
2024
-
[35]
Nomad: Non-Exclusive memory tiering via transactional page migration,
L. Xiang, Z. Lin, W. Deng, H. Lu, J. Rao, Y . Yuan, and R. Wang, “Nomad: Non-Exclusive memory tiering via transactional page migration,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . Santa Clara, CA: USENIX Association, Jul. 2024, pp. 19–3...
2024
-
[36]
Smt: Software-defined memory tiering for heterogeneous computing systems with cxl memory expander,
K. Kim, H. Kim, J. So, W. Lee, J. Im, S. Park, J. Cho, and H. Song, “Smt: Software-defined memory tiering for heterogeneous computing systems with cxl memory expander,” IEEE Micro , vol. 43, no. 2, pp. 20–29, 2023
2023
-
[37]
Lightweight frequency- based tiering for cxl memory systems,
K. Song, J. Yang, S. Liu, and G. Pekhimenko, “Lightweight frequency- based tiering for cxl memory systems,” 2023. [Online]. Available: https://arxiv.org/abs/2312.04789
2023 arXiv
-
[38]
Exploring dram cache prefetching for pooled memory,
C. Tirumalasetty and N. Annapreddy, “Exploring dram cache prefetching for pooled memory,” 2024. [Online]. Available: https: //arxiv.org/abs/2406.14778
2024 arXiv
-
[39]
Overcoming the memory wall with CXL-Enabled SSDs,
S.-P. Yang, M. Kim, S. Nam, J. Park, J. yong Choi, E. H. Nam, E. Lee, S. Lee, and B. S. Kim, “Overcoming the memory wall with CXL-Enabled SSDs,” in 2023 USENIX Annual Technical Conference (USENIX ATC 23). Boston, MA: USENIX Association, Jul. 2023, pp. 601–617
2023
-
[40]
Cache in hand: Expander-driven cxl prefetcher for next generation cxl-ssd,
M. Kwon, S. Lee, and M. Jung, “Cache in hand: Expander-driven cxl prefetcher for next generation cxl-ssd,” in Proceedings of the 15th ACM Workshop on Hot Topics in Storage and File Systems , ser. HotStorage ’23. New York, NY , USA: Association for Computing Machinery, 2023, p. 24–30
2023
-
[41]
Hello bytes, bye blocks: Pcie storage meets compute express link for memory expansion (cxl-ssd),
M. Jung, “Hello bytes, bye blocks: Pcie storage meets compute express link for memory expansion (cxl-ssd),” in Proceedings of the 14th ACM Workshop on Hot Topics in Storage and File Systems , ser. HotStorage ’22. New York, NY , USA: Association for Computing Machinery, 2022, p. 45–51
2022
-
[42]
Exploiting cxl-based memory for distributed deep learning,
M. Arif, K. Assogba, M. M. Rafique, and S. Vazhkudai, “Exploiting cxl-based memory for distributed deep learning,” in Proceedings of the 51st International Conference on Parallel Processing , ser. ICPP ’22. New York, NY , USA: Association for Computing Machinery, 2023. [Online...
2023
-
[43]
Exploring cxl-based kv cache storage for llmserving,
Y . T. ea tl., “Exploring cxl-based kv cache storage for llmserving,” in Machine Learning for Systems Workshop at NeurIPS 2024 , 2024
2024
-
[44]
Clay: Cxl-based scalable ndp architecture accelerating embedding layers,
S. Yun, H. Nam, K. Kyung, J. Park, B. Kim, Y . Kwon, E. Lee, and J. H. Ahn, “Clay: Cxl-based scalable ndp architecture accelerating embedding layers,” in Proceedings of the 38th ACM International Conference on Supercomputing, ser. ICS ’24. New York, NY , USA: Association for C...
2024
-
[45]
Beacon: Scalable near-data-processing accelerators for genome analysis near memory pool with the cxl support,
W. Huangfu, K. T. Malladi, A. Chang, and Y . Xie, “Beacon: Scalable near-data-processing accelerators for genome analysis near memory pool with the cxl support,” in Proceedings of the 55th Annual IEEE/ACM International Symposium on Microarchitecture, ser. MICRO ’22. IEEE Press...
2023
-
[46]
An lpddr-based cxl-pnm platform for tco-efficient inference of transformer-based large language models,
S.-S. Park, K. Kim, J. So, J. Jung, J. Lee, K. Woo, N. Kim, Y . Lee, H. Kim, Y . Kwon, J. Kim, J. Lee, Y . Cho, Y . Tai, J. Cho, H. Song, J. H. Ahn, and N. S. Kim, “An lpddr-based cxl-pnm platform for tco-efficient inference of transformer-based large language models,” in 2024...
2024
-
[47]
Computational cxl-memory solution for accel- erating memory-intensive applications,
J. Sim, S. Ahn, T. Ahn, S. Lee, M. Rhee, J. Kim, K. Shin, D. Moon, E. Kim, and K. Park, “Computational cxl-memory solution for accel- erating memory-intensive applications,” IEEE Computer Architecture Letters, vol. 22, no. 1, pp. 5–8, 2023
2023
-
[48]
Polaris: Enhancing cxl-based memory expanders with memory-side prefetching,
Z. Zhou, S. Xu, Y . Chen, T. Zhang, R. Shu, L. Qu, P. Cheng, Y . Xiong, and G. Sun, “Polaris: Enhancing cxl-based memory expanders with memory-side prefetching,” in Advanced Parallel Processing Technolo- gies: 15th International Symposium, APPT 2023, Nanchang, China, August 4–...
2023
-
[49]
Low-overhead general-purpose near- data processing in cxl memory expanders,
H. Ham, J. Hong, G. Park, Y . Shin, O. Woo, W. Yang, J. Bae, E. Park, H. Sung, E. Lim, and G. Kim, “Low-overhead general-purpose near- data processing in cxl memory expanders,” in 2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2024
2024
-
[50]
Udon: A case for offloading to general purpose compute on cxl memory,
J. Hermes, J. Minor, M. Wu, A. Patil, and E. V . Hensbergen, “Udon: A case for offloading to general purpose compute on cxl memory,” ArXiv, vol. abs/2404.02868, 2024
2024 arXiv
-
[52]
Design and analysis of cxl performance models for tightly-coupled heterogeneous computing,
A. M. Cabrera, A. R. Young, and J. S. Vetter, “Design and analysis of cxl performance models for tightly-coupled heterogeneous computing,” in Proceedings of the 1st International Workshop on Extreme Heterogeneity Solutions, ser. ExHET ’22. New York, NY , USA: Association for C...
2022
-
[53]
Synergizing cxl with unified memory for scalable gpu memory expansion,
J. Lee and J. Kim, “Synergizing cxl with unified memory for scalable gpu memory expansion,” in 2024 International Conference on Electron- ics, Information, and Communication (ICEIC) , 2024, pp. 1–4
2024
-
[54]
Accelerating performance of gpu-based workloads using cxl,
M. Arif, A. Maurya, and M. M. Rafique, “Accelerating performance of gpu-based workloads using cxl,” Proceedings of the 13th Workshop on AI and Scientific Computing at Scale using Flexible Computing , 2023
2023
-
[55]
Salus: Efficient security support for cxl-expanded gpu memory,
R. Abdullah, H. Lee, H. Zhou, and A. Awad, “Salus: Efficient security support for cxl-expanded gpu memory,” 2024 IEEE International Sym- posium on High-Performance Computer Architecture (HPCA) , 2024
2024
-
[56]
Memory pooling with cxl,
D. Gouk, M. Kwon, H. Bae, S. Lee, and M. Jung, “Memory pooling with cxl,” IEEE Micro, vol. 43, no. 2, pp. 48–57, 2023
2023
-
[57]
Logical memory pools: Flexible and local disaggregated memory,
E. Amaro, S. Wang, A. Panda, and M. K. Aguilera, “Logical memory pools: Flexible and local disaggregated memory,” Proceedings of the 22nd ACM Workshop on Hot Topics in Networks , 2023. 15
2023
-
[58]
Fabric-centric computing,
M. Liu, “Fabric-centric computing,” in Proceedings of the 19th Work- shop on Hot Topics in Operating Systems, ser. HOTOS ’23. New York, NY , USA: Association for Computing Machinery, 2023, p. 118–126
2023
-
[59]
Working with disaggregated systems. what are the challenges and opportunities of rdma and cxl?
A. M. Geyer, D. Ritter, D. H. Lee, M. Ahn, J. Pietrzyk, A. Krause, D. Habich, and W. Lehner, “Working with disaggregated systems. what are the challenges and opportunities of rdma and cxl?” in Datenbanksys- teme f ¨ur Business, Technologie und Web , 2023
2023
-
[60]
Near to far: An evaluation of disaggregated memory for in-memory data processing,
A. Geyer, J. Pietrzyk, A. Krause, D. Habich, W. Lehner, C. F ¨arber, and T. Willhalm, “Near to far: An evaluation of disaggregated memory for in-memory data processing,” in Proceedings of the 1st Workshop on Disruptive Memory Systems , ser. DIMES ’23. New York, NY , USA: Assoc...
2023
-
[61]
Aurelia: Cxl fabric with tentacle,
S.-T. Wang and W. Wang, “Aurelia: Cxl fabric with tentacle,” Proceed- ings of the 4th Workshop on Resource Disaggregation and Serverless , 2023
2023
-
[62]
Cxl over ethernet: A novel fpga-based memory disaggregation design in data centers,
C. Wang, K. He, R. Fan, X. Wang, W. Wang, and Q. Hao, “Cxl over ethernet: A novel fpga-based memory disaggregation design in data centers,” in 2023 IEEE 31st Annual International Symposium on Field- Programmable Custom Computing Machines (FCCM), 2023, pp. 75–82
2023
-
[63]
Rcmp: Reconstructing rdma-based memory disaggregation via cxl,
Z. Wang, Y . Guo, K. Lu, J. Wan, D. Wang, T. Yao, and H. Wu, “Rcmp: Reconstructing rdma-based memory disaggregation via cxl,”ACM Trans. Archit. Code Optim. , vol. 21, no. 1, Jan. 2024
2024
-
[64]
Understanding routable PCIe performance for composable infrastructures,
W. Hou, J. Zhang, Z. Wang, and M. Liu, “Understanding routable PCIe performance for composable infrastructures,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24). Santa Clara, CA: USENIX Association, Apr. 2024, pp. 297–312
2024
-
[65]
A case against cxl memory pooling,
P. Levis, K. Lin, and A. Tai, “A case against cxl memory pooling,” in Proceedings of the 22nd ACM Workshop on Hot Topics in Networks , ser. HotNets ’23. New York, NY , USA: Association for Computing Machinery, 2023, p. 18–24
2023
-
[66]
Design tradeoffs in cxl-based memory pools for public cloud platforms,
D. S. Berger, D. Ernst, H. Li, P. Zardoshti, M. Shah, S. Rajadnya, S. Lee, L. Hsu, I. Agarwal, M. D. Hill, and R. Bianchini, “Design tradeoffs in cxl-based memory pools for public cloud platforms,” IEEE Micro , vol. 43, no. 2, p. 30–38, Mar. 2023
2023
-
[67]
Memory sharing with cxl: Hardware and software design approaches,
S. Jain, N. Yeleswarapu, H. A. Maruf, and R. Gupta, “Memory sharing with cxl: Hardware and software design approaches,” 2024
2024
-
[68]
Trenv: Transparently share serverless execution environments across different functions and nodes,
J. Huang, M. Zhang, T. Ma, Z. Liu, S. Lin, K. Chen, J. Jiang, X. Liao, Y . Shan, N. Zhang, M. Lu, T. Ma, H. Gong, and Y . Wu, “Trenv: Transparently share serverless execution environments across different functions and nodes,” in Proceedings of the ACM SIGOPS 30th Symposium on...
2024
-
[69]
HydraRPC: RPC in the CXL era,
T. Ma, Z. Liu, C. Wei, J. Huang, Y . Zhuo, H. Li, N. Zhang, Y . Guan, D. Niu, M. Zhang, and T. Ma, “HydraRPC: RPC in the CXL era,” in 2024 USENIX Annual Technical Conference (USENIX ATC 24) . Santa Clara, CA: USENIX Association, Jul. 2024, pp. 387–395
2024
-
[70]
Telepathic datacenters: Fast rpcs using shared cxl memory,
S. Mahar, E. Hajyjasini, S. Lee, Z. Zhang, M. Shen, and S. Swanson, “Telepathic datacenters: Fast rpcs using shared cxl memory,” 2024
2024
-
[71]
Cxl shared memory programming: Barely distributed and almost persistent,
Y . Xu, S. Mahar, Z. Liu, M. Shen, and S. Swanson, “Cxl shared memory programming: Barely distributed and almost persistent,” 2024
2024
-
[72]
Dissecting cxl memory performance at scale: Analysis, modeling, and optimization,
J. Liu, H. Hadian, H. Xu, D. S. Berger, and H. Li, “Dissecting cxl memory performance at scale: Analysis, modeling, and optimization,”
-
[73]
The hitchhiker’s guide to programming and optimizing cxl-based heterogeneous systems,
Z. Wang, S. Mahar, L. Li, J. Park, J. Kim, T. Michailidis, Y . Pan, T. Rosing, D. Tullsen, S. Swanson, K. C. Ryoo, S. Park, and J. Zhao, “The hitchhiker’s guide to programming and optimizing cxl-based heterogeneous systems,” 2024. [Online]. Available: https: //arxiv.org/abs/2411.02814
2024 arXiv
-
[74]
The landscape of parallel computing research: A view from berkeley,
K. Asanovi ´c, R. Bodik, B. C. Catanzaro, J. J. Gebis, P. Husbands, K. Keutzer, D. A. Patterson, W. L. Plishker, J. Shalf, S. W. Williams, and K. A. Yelick, “The landscape of parallel computing research: A view from berkeley,” Tech. Rep. UCB/EECS-2006-183, Dec 2006
2006
-
[75]
Benchmarking cloud serving systems with ycsb,
B. F. Cooper, A. Silberstein, E. Tam, R. Ramakrishnan, and R. Sears, “Benchmarking cloud serving systems with ycsb,” in Proceedings of the 1st ACM Symposium on Cloud Computing , ser. SoCC ’10. New York, NY , USA: Association for Computing Machinery, 2010, p. 143–154
-
[76]
Flexible i/o tester
“Flexible i/o tester.” 2024. [Online]. Available: https://github.com/ axboe/fio
2024
-
[77]
Merci: efficient embedding reduction on commodity hardware via sub-query memoization,
Y . Lee, S. H. Seo, H. Choi, H. U. Sul, S. Kim, J. W. Lee, and T. J. Ham, “Merci: efficient embedding reduction on commodity hardware via sub-query memoization,” in Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Oper...
2021
-
[78]
Managing memory tiers with CXL in virtualized environments,
Y . Zhong, D. S. Berger, C. Waldspurger, R. Wee, I. Agarwal, R. Agarwal, F. Hady, K. Kumar, M. D. Hill, M. Chowdhury, and A. Cidon, “Managing memory tiers with CXL in virtualized environments,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) ....
2024
-
[79]
Characterizing the dilemma of performance and index size in billion-scale vector search and breaking it with second-tier memory,
R. Cheng, Y . Peng, X. Wei, H. Xie, R. Chen, S. Shen, and H. Chen, “Characterizing the dilemma of performance and index size in billion-scale vector search and breaking it with second-tier memory,”
-
[80]
Efficient tensor offloading for large deep-learning model training based on compute express link,
D. Xu, Y . Feng, K. Shin, D. Kim, H. Jeon, and D. Li, “Efficient tensor offloading for large deep-learning model training based on compute express link,” in SC24: International Conference for High Performance Computing, Networking, Storage and Analysis , 2024, pp. 1–18
2024
-
[81]
Available: https://arxiv.org/abs/2405.03267
[Online]. Available: https://arxiv.org/abs/2405.03267
-
[82]
Cxl and the return of scale-up database engines,
A. Lerner and G. Alonso, “Cxl and the return of scale-up database engines,” Proc. VLDB Endow. , vol. 17, no. 10, p. 2568–2575, Aug
-
[83]
Enabling cxl memory expansion for in-memory database management systems,
M. Ahn, A. Chang, D. Lee, J. Gim, J. Kim, J. Jung, O. Rebholz, V . Pham, K. Malladi, and Y . S. Ki, “Enabling cxl memory expansion for in-memory database management systems,” in Proceedings of the 18th International Workshop on Data Management on New Hardware , ser. DaMoN ’22....
2022
-
[84]
Database kernels: Seamless integration of database systems and fast storage via cxl,
S. Lee, A. Lerner, P. Bonnet, and P. Cudr ´e-Mauroux, “Database kernels: Seamless integration of database systems and fast storage via cxl,” in Conference on Innovative Data Systems Research , 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:266732447
2024
-
[85]
Available: https://doi.org/10.14778/3675034.3675047
[Online]. Available: https://doi.org/10.14778/3675034.3675047
-
[86]
A three-tier buffer manager integrating cxl device memory for database systems,
N. Riekenbrauck, M. Weisgut, D. Lindner, and T. Rabl, “A three-tier buffer manager integrating cxl device memory for database systems,” in 2024 IEEE 40th International Conference on Data Engineering Workshops (ICDEW), 2024, pp. 395–401
2024
-
[87]
Understanding and optimizing serverless workloads in cxl-enabled tiered memory,
Y . Li and S. Yao, “Understanding and optimizing serverless workloads in cxl-enabled tiered memory,” 2023. [Online]. Available: https://arxiv.org/abs/2309.01736
2023 arXiv
-
[88]
Cxl-enabled enhanced memory functions,
D. Boles, D. Waddington, and D. A. Roberts, “Cxl-enabled enhanced memory functions,” IEEE Micro, vol. 43, no. 2, pp. 58–65, 2023
2023
-
[89]
¯Apta: Fault- tolerant object-granular cxl disaggregated memory for accelerating faas,
A. Patil, V . Nagarajan, N. Nikoleris, and N. Oswald, “ ¯Apta: Fault- tolerant object-granular cxl disaggregated memory for accelerating faas,” in 2023 53rd Annual IEEE/IFIP International Conference on Depend- able Systems and Networks (DSN) , 2023, pp. 201–215
2023
-
[90]
A survey on memory-centric computer architectures,
A. Gebregiorgis, H. A. Du Nguyen, J. Yu, R. Bishnoi, M. Taouil, F. Catthoor, and S. Hamdioui, “A survey on memory-centric computer architectures,” J. Emerg. Technol. Comput. Syst. , vol. 18, no. 4, Oct. 2022
2022
-
[91]
Memory protection keys
J. Corbet, “Memory protection keys.” 2015. [Online]. Available: https://lwn.net/Articles/643797/
2015
-
[92]
Memory-centric computing,
O. Mutlu, “Memory-centric computing,” 2023. [Online]. Available: https://arxiv.org/abs/2305.20000
2023 arXiv
-
[2024]
Available: https://arxiv.org/abs/2409.14317
[Online]. Available: https://arxiv.org/abs/2409.14317
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.