Pith. sign in

REVIEW 3 major objections 6 minor 2 cited by

Next-Gen Computing Systems with Compute Express Link: a Comprehensive Survey

T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This survey argues that CXL makes memory a coherent, first-class fabric resource and organizes CXL systems into memory expansion, unified memory, and distributed sharing.

desk verdict A useful, current CXL survey that is one revision away from being a dependable reference; the taxonomy is rougher than advertised. read the letter →

arxiv 2412.20249 v2 pith:PWK7ABSJ submitted 2024-12-28 cs.DC

classification cs.DC
keywords ComputeExpressLinkInterconnectionMemoryExpansionDisaggregationDistributedSharedUnifiedNear-MemoryProcessingTiered
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Computing performance has outrun interconnect performance, so processors wait on memory and device communication. This survey argues that Compute Express Link (CXL), an open industry-standard interconnect carrying memory semantics over PCIe, attacks that bottleneck by letting processors access device-attached memory with cache coherence and lower latency than plain PCIe. It organizes the field into two single-machine classes—Memory Expansion, which grows memory capacity and manages it as a tier below DRAM, and Unified Memory, which lets CPUs and accelerators share a coherent memory space—plus a distributed class where CXL switches pool and share memory across nodes. The survey's forward claim is that CXL pushes computing toward a memory-centric paradigm in which data movement, not raw compute, is the resource to orchestrate.

What carries the argument

The machinery that carries the argument is the CXL protocol stack: CXL.io for device discovery and direct memory access (DMA), CXL.cache for devices caching host memory, and CXL.mem for hosts caching device-attached memory as host-managed device memory. These protocols compose into three device types, and the CXL Switch extends them from one machine to a pool of hosts and memory modules. That same stack generates the survey's taxonomy: Type-3 memory devices anchor the Memory Expansion category, Type-2 accelerator devices anchor Unified Memory, and switches plus CXL 3.0 multi-host sharing anchor distributed pooling and shared memory. The survey also leans on tiered-memory page placement and near-memory processing as the mechanisms that make expansion usable despite CXL's higher latency.

What would settle it

A concrete check: saturate a commercial CXL memory device with random reads and measure sustained latency; if loaded latency exceeds a few microseconds, or if a rack-scale path through CXL switches is slower than an equivalent remote direct memory access path, then the survey's claim that CXL provides near-DRAM, network-beating memory semantics for disaggregation fails.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central claim is that CXL changes what an interconnect is for: instead of shuttling data between disjoint memory domains, a CXL fabric makes memory itself a shared, addressable resource. In a single machine this takes the form of tiered memory systems that keep hot pages in fast DRAM and cold pages in slower CXL memory, near-memory processing that computes beside the data, and a hardware-coherent unified space that lets accelerators read and write host memory without a CPU relay. Across machines, CXL 2.0 memory pooling lets one host claim memory from a pool, while CXL 3.0 memory sharing lets several hosts share one memory module, turning shared CXL memory into a communication channel for remote procedure calls and synchronization. The paper also assembles real-hardware measurements placing CXL memory between DRAM and non-volatile memory in latency (roughly 170–250 nanoseconds) and argues that future systems should be engineered around this memory fabric rather than around individual nodes.

Load-bearing premise

The survey's organizing claim rests on the taxonomy announced in the introduction and used in Sections III and IV: CXL systems divide cleanly into Memory Expansion, Unified Memory, and distributed pooling or sharing, with near-memory processing counted under Memory Expansion; if those categories overlap, as near-memory processing's compute-offload nature suggests, the organization weakens.

Editorial extensions

If this is right

  • If this picture is right, memory expansion is a practical pressure valve for the memory wall: applications that tolerate roughly 170–250 ns access can use terabyte-scale CXL memory, with tiered placement keeping hot pages in DRAM.
  • If this picture is right, hardware cache coherence will let GPUs and other accelerators share memory directly, removing the CPU-mediated data copies that today dominate heterogeneous workloads.
  • If this picture is right, memory pooling decouples memory capacity from individual servers, so data centers can allocate memory where demand actually is rather than where hardware was purchased.
  • If this picture is right, CXL 3.0 shared memory becomes a low-latency communication substrate for remote procedure calls, distributed synchronization, and serverless state, bypassing the network stack inside a rack.
  • If this picture is right, the design center of future systems shifts to memory-centric computing, in which the distance between data and compute is measured in CXL fabric hops and managed explicitly.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper's classification, the line between Memory Expansion and near-memory processing is likely to blur: once CXL devices carry general-purpose compute cores, capacity expansion becomes a compute-placement problem as much as a capacity problem.
  • The survey presents tiering and interleaving as competing policies, but the measurements it cites imply that workload-aware systems will combine both—tiering for latency-sensitive phases and interleaving for bandwidth-hungry phases—because DRAM latency degrades sharply near bandwidth saturation.
  • An extension the survey does not run is a direct benchmark of CXL shared-memory RPC against remote direct memory access (RDMA) at identical intra-rack distances; its own numbers suggest CXL should win inside a rack and lose across racks, which would make CXL-over-Ethernet hybrids the natural scale-out path.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper is a survey of computing systems built around Compute Express Link (CXL). It introduces CXL protocols and device types, reviews evaluation and simulation platforms, proposes a taxonomy for single-machine systems (Memory Expansion and Unified Memory) and for distributed systems (Memory Pooling and Memory Sharing), discusses near-memory processing and heterogeneous accelerator integration, and outlines future research directions toward memory-centric computing. The survey compiles roughly 90 references from 2019 to 2024 and positions itself as a comprehensive review of recent CXL research from single-machine to distributed settings.

Significance. The manuscript is timely and covers a rapidly evolving area. Its main contribution is organizational: it collects and classifies a substantial body of recent CXL systems research. If the taxonomy were consistent, the survey would serve as a useful entry point for newcomers and as a reference for researchers. The paper contains no derivations or fitted models, so circularity is not a concern; its value depends on the accuracy and coherence of its literature organization. The current taxonomy, however, is applied inconsistently, and the unsupported publication-count claim in the Introduction weakens the survey's credibility. These issues are fixable, but they affect the central contribution and require revision.

major comments (3)
  1. [Abstract, Sections III-IV, Table I] The survey's central organizing claim is that single-machine CXL research partitions cleanly into Memory Expansion and Unified Memory based on interconnection type (processor-to-memory vs. processor-to-accelerator). This partition is not consistently applied. Section IV, 'Near-Memory Processing,' is listed under Memory Expansion in Table I and announced in the Section III opening paragraph, yet near-memory processing is a compute-offload mechanism and does not change the interconnect type that allegedly defines the taxonomy. Moreover, ReCXL [14] appears both in Section III.C (Application-Specific Optimizations) and in Section IV.A (Workload-Customized Near-Memory Process), and references [25] and [26] appear in multiple rows of Table I. The categories are therefore neither mutually exclusive nor based on a uniformly applied criterion. The taxonomy should be redefined—for example, by treating near-memory processing as an orthogonal dimension—or the paper should explicitly present the categories as overlapping research themes rather than as a partition.
  2. [Section I (Introduction)] The Introduction states that 'only 10 paper were published during 2019 to 2022, while 40 in 2023, 51 in 2024' with no citation, source, or counting methodology. This quantitative claim is used to justify the survey's timeliness and the need for a new survey. The authors should either provide the bibliographic corpus used for the count, state the inclusion criteria (e.g., which venues and search terms), or soften the claim to avoid an unverifiable statistic that undermines the paper's scholarly framing.
  3. [Section V.A] The subsection 'Memory Expansion with CPU Relay' is classified under Unified Memory, but the works it discusses (e.g., [17], [51]) describe using CXL to expand accelerator memory through a CPU relay. That is memory expansion for accelerators, not the unified-memory scenario defined in Section V's introduction, which emphasizes direct CXL access between coherent accelerators and processors. This placement blurs the boundary between the two headline categories and suggests that the taxonomy's criterion (interconnection type) is not consistently applied across all subsections. The authors should either reclassify these works or revise the category definitions to accommodate relay-based accelerator memory expansion.
minor comments (6)
  1. [Throughout] There are numerous typographical and grammatical errors, including 'Memory Expansion focus' (should be 'focuses'), 'Thet propose' in Section III.A, 'system‘s' with an unbalanced smart quote, and inconsistent British/American spelling ('synchronisation' vs. 'synchronization'). A careful proofreading pass is recommended.
  2. [Section II.B] The affiliation 'UIU' should be 'UIUC' (University of Illinois Urbana-Champaign) when referring to the authors of [24].
  3. [Reference [43]] Reference [43] lists the author as 'Y. T. ea tl.,' which is malformed; the author list should be completed or the entry replaced with a properly formatted citation.
  4. [Table I] Table I lists the same references, such as [25] and [26], in multiple rows (e.g., Latency/Bandwidth, Use Case: Regular Applications, Use Case: Large Language Models). While a single work can support multiple topics, the table should explicitly state that rows are not exclusive, or it will reinforce the taxonomy concerns raised about category disjointness.
  5. [Section VIII.A.1] The claim that memory interleaving makes 'the overall bandwidth can be the sum of CXL memory and DRAM in theory' should be qualified or cited, because the achievable aggregate bandwidth depends on the memory controller, CXL root port, and PCIe topology, and may not reach the theoretical sum.
  6. [Section IX] The 'Memory-Centric Computing in CXL Fabric' section presents a speculative research vision. It would be helpful to explicitly label this section as the authors' position or vision rather than as established results, and to distinguish it from the survey's descriptive contributions.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey is expository and makes no predictive or derived claims that reduce to its own inputs.

full rationale

This is a comprehensive survey paper, not a derivation or modeling paper. It contains no fitted parameters, no equations that transform inputs into outputs, and no prediction that could reduce by construction to its own assumptions. The central claim is a classification of existing CXL research into categories (Memory Expansion, Unified Memory, and distributed pooling/sharing), which is an organizational judgment rather than a derived result. The taxonomy could be criticized for consistency—for example, ReCXL [14] appears both under 'Application-Specific Optimizations' in Section III.C and under 'Workload-Customized Near-Memory Process' in Section IV.A, and near-memory processing is arguably a compute-offload mechanism rather than memory expansion—but such a critique concerns the quality or mutual exclusivity of a survey taxonomy, not circular reasoning. The paper does not define 'Memory Expansion' in terms of the conclusions of the surveyed works, nor does it invoke a self-citation chain to force its classification. Citations such as TPP [9] and Pond [18] provide external latency measurements and system results that are independent of this paper's claims. There are no self-citations by the authors that are load-bearing, and no uniqueness theorem or ansatz is imported from the authors' prior work. Because the survey is self-contained as an expository review and derives no quantitative result from its own categories, the circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey introduces no free parameters and no invented physical entities. It relies on cited measurements of CXL hardware, on a representative literature selection, and on assumptions about future CXL hardware availability.

assumptions (3)
  • domain assumption CXL hardware measurements (latency, bandwidth, coherence behavior) reported in the cited evaluations are accurate and representative.
    Section II builds the survey's background on results from [24], [25], [26], and others; the survey itself does not reproduce these measurements.
  • domain assumption The surveyed papers are a representative sample of CXL research through 2024.
    The survey's classification and future directions depend on this selection, but no systematic search methodology is given.
  • domain assumption Future CXL hardware (CXL 3.0 switches, Type-2 devices) will become available as assumed in the vision sections.
    Sections VIII and IX discuss memory-centric computing as if these components exist; hardware availability remains uncertain.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Next-Gen Computing Systems with Compute Express Link: a Comprehensive Survey." pith.science (2026). https://pith.science/paper/PWK7ABSJ

@misc{pith2026241220249,
  author       = {Pith},
  title        = {Pith review of: Next-Gen Computing Systems with Compute Express Link: a Comprehensive Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PWK7ABSJ}},
  note         = {Machine review of arXiv:2412.20249}
}
read the original abstract

Interconnection is crucial for computing systems. However, the current interconnection performance between processors and devices, such as memory devices and accelerators, significantly lags behind their computing performance, severely limiting the overall performance. To address this challenge, Intel proposes Compute Express Link (CXL), an open industry-standard interconnection. With memory semantics, CXL offers low-latency, scalable, and coherent interconnection between processors and devices. This paper introduces recent advances in CXL-based computing systems from single-machine to distributed. In single-machine systems, we classify existing research into two categories: Memory Expansion and Unified Memory. Memory Expansion focus on processors and memory, aims to address memory wall challenge. Unified memory focus on processors and accelerators, aims to enhance collaboration in heterogeneous computing systems. In distributed systems, we present how to build efficient disaggregation systems based on CXL infrastructure, enabling resource pooling and sharing. Finally, we discuss the future research and envision memory-centric computing with CXL.

Figures

Figures reproduced from arXiv: 2412.20249 by the authors.

Figure 1
Figure 1. Three types of CXL devices. The Type 1 device can directly cache host memory, the Type 2 device can directly cache [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. CXL-based Memory Expansion. It is based on tiered [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Unified Memory. The hardware cache coherence pro [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: CXL Based Disaggregated System [2]. It breaks down the physical isolation between servers and consolidates the data center into a hyper-node through resource pooling. VI. CXL BASED DISAGGREGATED SYSTEM Based on the memory expansion on single-machine, CXL 2.0 proposes a…
Figure 5
Figure 5. Figure 5: Communication via Shared CXL Memory which requires extensive rewriting of networked applications, increasing software complexity. Finally, by analyzing produc￾tion traces from Google and Azure Cloud, it is found that modern servers are large enough for most virtual mac…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving

    cs.DC 2026-06 accept novelty 6.5 of 10

    KV-cache serving systems concentrate into five archetypes under a four-axis taxonomy, with ownership explaining residual distributed design variance and seven measurement gaps blocking next steps.

  2. Bridging the Cognitive Gap: A Unified Memory Paradigm for 6G Agentic AI-RAN

    cs.NI 2026-05 unverdicted novelty 6.0 of 10

    Proposal to replace message-passing interfaces in AI-RAN with zero-copy CXL shared memory, organized as reflexive, contextual, and evolutionary cognitive loops.

Reference graph

Works this paper leans on

93 extracted references · 67 canonical work pages · cited by 2 Pith papers

  1. [2]

    An introduction to the compute express link (cxl) interconnect,

    D. Das Sharma, R. Blankenship, and D. Berger, “An introduction to the compute express link (cxl) interconnect,” ACM Comput. Surv., vol. 56, no. 11, Jul. 2024

  2. [14]

    Enabling efficient large recommendation model training with near cxl memory processing,

    H. Liu, L. Zheng, Y . Huang, J. Zhou, C. Liu, R. Wang, X. Liao, H. Jin, and J. Xue, “Enabling efficient large recommendation model training with near cxl memory processing,” in 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA) , 2024

  3. [25]

    Exploring performance and cost optimization with asic-based cxl memory,

    Y . Tang, P. Zhou, W. Zhang, H. Hu, Q. Yang, H. Xiang, T. Liu, J. Shan, R. Huang, C. Zhao et al., “Exploring performance and cost optimization with asic-based cxl memory,” inProceedings of the Nineteenth European Conference on Computer Systems , 2024, pp. 818–833

  4. [26]

    Ex- ploring and evaluating real-world cxl: Use cases and system adoption,

    J. Liu, X. Wang, J. Wu, S. Yang, J. Ren, B. Shankar, and D. Li, “Ex- ploring and evaluating real-world cxl: Use cases and system adoption,” arXiv preprint arXiv:2405.14209 , 2024

  5. [17]

    Gpu graph processing on cxl- based microsecond-latency external memory,

    S. Sano, Y . Bando, K. Hiwada, H. Kajihara, T. Suzuki, Y . Nakanishi, D. Taki, A. Kaneko, and T. Shiozawa, “Gpu graph processing on cxl- based microsecond-latency external memory,” in Proceedings of the SC ’23 Workshops of The International Conference on High Performance Computing, Network, Storage, and Analysis, ser. SC-W ’23. New York, NY , USA: Associa...

  6. [51]

    Lmb: Augmenting pcie devices with cxl-linked memory buffer,

    J. Wang, X. Zhang, C. Tang, X. Chen, and T. Lu, “Lmb: Augmenting pcie devices with cxl-linked memory buffer,” 2024

  7. [1]

    Ai and memory wall,

    A. Gholami, Z. Yao, S. Kim, C. Hooper, M. W. Mahoney, and K. Keutzer, “Ai and memory wall,” IEEE Micro , vol. 44, no. 3, pp. 33–39, 2024

  8. [3]

    Memory disaggregation: Advances and open challenges,

    H. Al Maruf and M. Chowdhury, “Memory disaggregation: Advances and open challenges,” SIGOPS Oper. Syst. Rev., vol. 57, no. 1, p. 29–37, Jun. 2023. [Online]. Available: https://doi.org/10.1145/3606557.3606562

Show all 93 references
  1. [4]

    Distributed shared memory: A survey of issues and algorithms,

    B. Nitzberg and V . Lo, “Distributed shared memory: A survey of issues and algorithms,” Computer, vol. 24, no. 8, p. 52–60, Aug. 1991. [Online]. Available: https://doi.org/10.1109/2.84877

  2. [5]

    Compute express link 3.0,

    “Compute express link 3.0,” 2022. [Online]. Available: https://www.computeexpresslink.org/ files/ugd/0c1418 a8713008916044ae9604405d10a7773b.pdf

  3. [6]

    Compute express link cxl 3.0 is the exciting building block for dis- aggregation

    “Compute express link cxl 3.0 is the exciting building block for dis- aggregation.” 2022. [Online]. Available: https://www.servethehome.com/ compute-expresslink-cxl-3-0-is-the-exciting-building-block-for-disaggregation/

  4. [7]

    Compute express link™: The breakthrough cpu-to-device intercon- nect

    “Compute express link™: The breakthrough cpu-to-device intercon- nect.” 2022. [Online]. Available: https://www.computeexpresslink.org/ home

  5. [8]

    Compute express link (cxl): Enabling heterogeneous data-centric computing with heterogeneous memory hierarchy,

    D. D. Sharma, “Compute express link (cxl): Enabling heterogeneous data-centric computing with heterogeneous memory hierarchy,” IEEE Micro, vol. 43, no. 2, pp. 99–109, 2023

  6. [9]

    Tpp: Transparent page placement for cxl-enabled tiered-memory,

    H. Maruf, H. Wang, A. Dhanotia, J. Weiner, N. Agarwal, P. Bhat- tacharya, C. Petersen, M. Chowdhury, S. Kanaujia, and P. Chauhan, “Tpp: Transparent page placement for cxl-enabled tiered-memory,” ASPLOS, Jan 2023

  7. [10]

    Tiered memory management: Access latency is the key!

    M. Vuppalapati and R. Agarwal, “Tiered memory management: Access latency is the key!” in Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles, ser. SOSP ’24. New York, NY , USA: Association for Computing Machinery, 2024, p. 79–94

  8. [11]

    CXL- ANNS: Software-Hardware collaborative memory disaggregation and computation for Billion-Scale approximate nearest neighbor search,

    J. Jang, H. Choi, H. Bae, S. Lee, M. Kwon, and M. Jung, “CXL- ANNS: Software-Hardware collaborative memory disaggregation and computation for Billion-Scale approximate nearest neighbor search,” in 2023 USENIX Annual Technical Conference (USENIX ATC 23). Boston, MA: USENIX Asso...

  9. [12]

    Bridging software-hardware for cxl memory disaggregation in billion-scale near- est neighbor search,

    J. Jang, H. Choi, H. Bae, S. Lee, M. Kwon, and M. Jung, “Bridging software-hardware for cxl memory disaggregation in billion-scale near- est neighbor search,” ACM Trans. Storage, vol. 20, no. 2, Feb. 2024

  10. [13]

    Neomem: Hardware/software co-design for cxl-native memory tiering,

    Z. Zhou, Y . Chen, T. Zhang, Y . Wang, R. Shu, S. Xu, P. Cheng, L. Qu, Y . Xiong, J. Zhang, and G. Sun, “Neomem: Hardware/software co-design for cxl-native memory tiering,” in 2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2024

  11. [15]

    Breaking barriers: Expanding gpu memory with sub-two digit nanosecond latency cxl controller,

    D. Gouk, S. Kang, H. Bae, E. Ryu, S. Lee, D. Kim, J. Jang, and M. Jung, “Breaking barriers: Expanding gpu memory with sub-two digit nanosecond latency cxl controller,” in Proceedings of the 16th ACM Workshop on Hot Topics in Storage and File Systems , ser. HotStorage ’24. New ...

  12. [16]

    Failure tolerant training with persistent memory disaggregation over cxl,

    M. Kwon, J. Jang, H. Choi, S. Lee, and M. Jung, “Failure tolerant training with persistent memory disaggregation over cxl,” IEEE Micro, vol. 43, pp. 66–75, 2023

  13. [18]

    Pond: Cxl-based memory pooling systems for cloud platforms,

    H. Li, D. Berger, S. Novakovic, L. Hsu, D. Ernst, P. Zardoshti, M. Shah, S. Rajadnya, S. Lee, I. Agarwal, M. Hill, M. Fontoura, and R. Bian- chini, “Pond: Cxl-based memory pooling systems for cloud platforms,” ASPLOS, Jan 2023

  14. [19]

    Direct access, high- performance memory disaggregation with directcxl,

    D. Gouk, S. Lee, M. Kwon, and M. Jung, “Direct access, high- performance memory disaggregation with directcxl,” ATC, Jan 2022

  15. [20]

    Ddc: A vision for a disaggregated datacenter,

    M. Ewais and P. Chow, “Ddc: A vision for a disaggregated datacenter,” 2024

  16. [21]

    Notnets: Accelerating microservices by bypassing the network,

    P. Alvaro, M. Adiletta, A. Cockroft, F. Hady, R. Illikkal, E. Ramos, J. Tsai, and R. Soul´e, “Notnets: Accelerating microservices by bypassing the network,” 2024

  17. [22]

    Partial failure resilient memory management system for (cxl-based) distributed shared memory,

    M. Zhang, T. Ma, J. Hua, Z. Liu, K. Chen, N. Ding, F. Du, J. Jiang, T. Ma, and Y . Wu, “Partial failure resilient memory management system for (cxl-based) distributed shared memory,” in Proceedings of the 29th Symposium on Operating Systems Principles, ser. SOSP ’23. New York,...

  18. [23]

    Redis ltd

    “Redis ltd.” 2024. [Online]. Available: https://redis.io/

  19. [24]

    Demystifying cxl memory with genuine cxl-ready systems and devices,

    Y . Sun, Y . Yuan, Z. Yu, R. Kuper, C. Song, J. Huang, H. Ji, S. Agarwal, J. Lou, I. Jeong et al., “Demystifying cxl memory with genuine cxl-ready systems and devices,” in Proceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture , 2023, pp. 105–121

  20. [27]

    Evaluating emerging cxl- enabled memory pooling for hpc systems,

    J. Wahlgren, M. Gokhale, and I. B. Peng, “Evaluating emerging cxl- enabled memory pooling for hpc systems,” in 2022 IEEE/ACM Work- shop on Memory Centric High Performance Computing (MCHPC) . IEEE, 2022, pp. 11–20

  21. [28]

    Zero-offload: Democratizing billion-scale model training,

    J. Ren, S. Rajbhandari, R. Aminabadi, O. Ruwase, S. Yang, M. Zhang, L. Dong, and Y . He, “Zero-offload: Democratizing billion-scale model training,” Cornell University - arXiv,Cornell University - arXiv , Jan 2021

  22. [29]

    High-throughput generative inference of large language model with a single gpu

    Y . Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Zhang, I. Stoica, and C. R ´e, “High-throughput generative inference of large language model with a single gpu.”

  23. [30]

    Performance evaluation on cxl-enabled hybrid memory pool,

    Q. Yang, R. Jin, B. Davis, D. Inupakutika, and M. Zhao, “Performance evaluation on cxl-enabled hybrid memory pool,” in 2022 IEEE Inter- national Conference on Networking, Architecture and Storage (NAS) , 2022, pp. 1–5

  24. [31]

    Cxlmemsim: A pure software simulated cxl.mem for performance characterization,

    Y . Yang, P. Safayenikoo, J. Ma, T. A. Khan, and A. Quinn, “Cxlmemsim: A pure software simulated cxl.mem for performance characterization,” 2023

  25. [32]

    emucxl: an emulation framework for cxl- based disaggregated memory applications,

    R. Gond and P. Kulkarni, “emucxl: an emulation framework for cxl- based disaggregated memory applications,” 2024

  26. [33]

    A mess of memory system bench- marking, simulation and application profiling,

    P. Esmaili-Dokht, F. Sgherzi, V . S. Girelli, I. Boixaderas, M. Carmin, A. Monemi, A. Armejach, E. Mercadal, G. Llort, P. Radojkovi ´c, M. Moreto, J. Gim ´enez, X. Martorell, E. Ayguad ´e, J. Labarta, E. Con- falonieri, R. Dubey, and J. Adlard, “A mess of memory system bench- ...

  27. [34]

    Dracksim: Simulating cxl-enabled large-scale disaggre- gated memory systems,

    A. Puri, K. Bellamkonda, K. Narreddy, J. Jose, V . Tamarapalli, and V . Narayanan, “Dracksim: Simulating cxl-enabled large-scale disaggre- gated memory systems,” in Proceedings of the 38th ACM SIGSIM Con- ference on Principles of Advanced Discrete Simulation , ser. SIGSIM- PAD...

  28. [35]

    Nomad: Non-Exclusive memory tiering via transactional page migration,

    L. Xiang, Z. Lin, W. Deng, H. Lu, J. Rao, Y . Yuan, and R. Wang, “Nomad: Non-Exclusive memory tiering via transactional page migration,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . Santa Clara, CA: USENIX Association, Jul. 2024, pp. 19–3...

  29. [36]

    Smt: Software-defined memory tiering for heterogeneous computing systems with cxl memory expander,

    K. Kim, H. Kim, J. So, W. Lee, J. Im, S. Park, J. Cho, and H. Song, “Smt: Software-defined memory tiering for heterogeneous computing systems with cxl memory expander,” IEEE Micro , vol. 43, no. 2, pp. 20–29, 2023

  30. [37]

    Lightweight frequency- based tiering for cxl memory systems,

    K. Song, J. Yang, S. Liu, and G. Pekhimenko, “Lightweight frequency- based tiering for cxl memory systems,” 2023. [Online]. Available: https://arxiv.org/abs/2312.04789

  31. [38]

    Exploring dram cache prefetching for pooled memory,

    C. Tirumalasetty and N. Annapreddy, “Exploring dram cache prefetching for pooled memory,” 2024. [Online]. Available: https: //arxiv.org/abs/2406.14778

  32. [39]

    Overcoming the memory wall with CXL-Enabled SSDs,

    S.-P. Yang, M. Kim, S. Nam, J. Park, J. yong Choi, E. H. Nam, E. Lee, S. Lee, and B. S. Kim, “Overcoming the memory wall with CXL-Enabled SSDs,” in 2023 USENIX Annual Technical Conference (USENIX ATC 23). Boston, MA: USENIX Association, Jul. 2023, pp. 601–617

  33. [40]

    Cache in hand: Expander-driven cxl prefetcher for next generation cxl-ssd,

    M. Kwon, S. Lee, and M. Jung, “Cache in hand: Expander-driven cxl prefetcher for next generation cxl-ssd,” in Proceedings of the 15th ACM Workshop on Hot Topics in Storage and File Systems , ser. HotStorage ’23. New York, NY , USA: Association for Computing Machinery, 2023, p. 24–30

  34. [41]

    Hello bytes, bye blocks: Pcie storage meets compute express link for memory expansion (cxl-ssd),

    M. Jung, “Hello bytes, bye blocks: Pcie storage meets compute express link for memory expansion (cxl-ssd),” in Proceedings of the 14th ACM Workshop on Hot Topics in Storage and File Systems , ser. HotStorage ’22. New York, NY , USA: Association for Computing Machinery, 2022, p. 45–51

  35. [42]

    Exploiting cxl-based memory for distributed deep learning,

    M. Arif, K. Assogba, M. M. Rafique, and S. Vazhkudai, “Exploiting cxl-based memory for distributed deep learning,” in Proceedings of the 51st International Conference on Parallel Processing , ser. ICPP ’22. New York, NY , USA: Association for Computing Machinery, 2023. [Online...

  36. [43]

    Exploring cxl-based kv cache storage for llmserving,

    Y . T. ea tl., “Exploring cxl-based kv cache storage for llmserving,” in Machine Learning for Systems Workshop at NeurIPS 2024 , 2024

  37. [44]

    Clay: Cxl-based scalable ndp architecture accelerating embedding layers,

    S. Yun, H. Nam, K. Kyung, J. Park, B. Kim, Y . Kwon, E. Lee, and J. H. Ahn, “Clay: Cxl-based scalable ndp architecture accelerating embedding layers,” in Proceedings of the 38th ACM International Conference on Supercomputing, ser. ICS ’24. New York, NY , USA: Association for C...

  38. [45]

    Beacon: Scalable near-data-processing accelerators for genome analysis near memory pool with the cxl support,

    W. Huangfu, K. T. Malladi, A. Chang, and Y . Xie, “Beacon: Scalable near-data-processing accelerators for genome analysis near memory pool with the cxl support,” in Proceedings of the 55th Annual IEEE/ACM International Symposium on Microarchitecture, ser. MICRO ’22. IEEE Press...

  39. [46]

    An lpddr-based cxl-pnm platform for tco-efficient inference of transformer-based large language models,

    S.-S. Park, K. Kim, J. So, J. Jung, J. Lee, K. Woo, N. Kim, Y . Lee, H. Kim, Y . Kwon, J. Kim, J. Lee, Y . Cho, Y . Tai, J. Cho, H. Song, J. H. Ahn, and N. S. Kim, “An lpddr-based cxl-pnm platform for tco-efficient inference of transformer-based large language models,” in 2024...

  40. [47]

    Computational cxl-memory solution for accel- erating memory-intensive applications,

    J. Sim, S. Ahn, T. Ahn, S. Lee, M. Rhee, J. Kim, K. Shin, D. Moon, E. Kim, and K. Park, “Computational cxl-memory solution for accel- erating memory-intensive applications,” IEEE Computer Architecture Letters, vol. 22, no. 1, pp. 5–8, 2023

  41. [48]

    Polaris: Enhancing cxl-based memory expanders with memory-side prefetching,

    Z. Zhou, S. Xu, Y . Chen, T. Zhang, R. Shu, L. Qu, P. Cheng, Y . Xiong, and G. Sun, “Polaris: Enhancing cxl-based memory expanders with memory-side prefetching,” in Advanced Parallel Processing Technolo- gies: 15th International Symposium, APPT 2023, Nanchang, China, August 4–...

  42. [49]

    Low-overhead general-purpose near- data processing in cxl memory expanders,

    H. Ham, J. Hong, G. Park, Y . Shin, O. Woo, W. Yang, J. Bae, E. Park, H. Sung, E. Lim, and G. Kim, “Low-overhead general-purpose near- data processing in cxl memory expanders,” in 2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2024

  43. [50]

    Udon: A case for offloading to general purpose compute on cxl memory,

    J. Hermes, J. Minor, M. Wu, A. Patil, and E. V . Hensbergen, “Udon: A case for offloading to general purpose compute on cxl memory,” ArXiv, vol. abs/2404.02868, 2024

  44. [52]

    Design and analysis of cxl performance models for tightly-coupled heterogeneous computing,

    A. M. Cabrera, A. R. Young, and J. S. Vetter, “Design and analysis of cxl performance models for tightly-coupled heterogeneous computing,” in Proceedings of the 1st International Workshop on Extreme Heterogeneity Solutions, ser. ExHET ’22. New York, NY , USA: Association for C...

  45. [53]

    Synergizing cxl with unified memory for scalable gpu memory expansion,

    J. Lee and J. Kim, “Synergizing cxl with unified memory for scalable gpu memory expansion,” in 2024 International Conference on Electron- ics, Information, and Communication (ICEIC) , 2024, pp. 1–4

  46. [54]

    Accelerating performance of gpu-based workloads using cxl,

    M. Arif, A. Maurya, and M. M. Rafique, “Accelerating performance of gpu-based workloads using cxl,” Proceedings of the 13th Workshop on AI and Scientific Computing at Scale using Flexible Computing , 2023

  47. [55]

    Salus: Efficient security support for cxl-expanded gpu memory,

    R. Abdullah, H. Lee, H. Zhou, and A. Awad, “Salus: Efficient security support for cxl-expanded gpu memory,” 2024 IEEE International Sym- posium on High-Performance Computer Architecture (HPCA) , 2024

  48. [56]

    Memory pooling with cxl,

    D. Gouk, M. Kwon, H. Bae, S. Lee, and M. Jung, “Memory pooling with cxl,” IEEE Micro, vol. 43, no. 2, pp. 48–57, 2023

  49. [57]

    Logical memory pools: Flexible and local disaggregated memory,

    E. Amaro, S. Wang, A. Panda, and M. K. Aguilera, “Logical memory pools: Flexible and local disaggregated memory,” Proceedings of the 22nd ACM Workshop on Hot Topics in Networks , 2023. 15

  50. [58]

    Fabric-centric computing,

    M. Liu, “Fabric-centric computing,” in Proceedings of the 19th Work- shop on Hot Topics in Operating Systems, ser. HOTOS ’23. New York, NY , USA: Association for Computing Machinery, 2023, p. 118–126

  51. [59]

    Working with disaggregated systems. what are the challenges and opportunities of rdma and cxl?

    A. M. Geyer, D. Ritter, D. H. Lee, M. Ahn, J. Pietrzyk, A. Krause, D. Habich, and W. Lehner, “Working with disaggregated systems. what are the challenges and opportunities of rdma and cxl?” in Datenbanksys- teme f ¨ur Business, Technologie und Web , 2023

  52. [60]

    Near to far: An evaluation of disaggregated memory for in-memory data processing,

    A. Geyer, J. Pietrzyk, A. Krause, D. Habich, W. Lehner, C. F ¨arber, and T. Willhalm, “Near to far: An evaluation of disaggregated memory for in-memory data processing,” in Proceedings of the 1st Workshop on Disruptive Memory Systems , ser. DIMES ’23. New York, NY , USA: Assoc...

  53. [61]

    Aurelia: Cxl fabric with tentacle,

    S.-T. Wang and W. Wang, “Aurelia: Cxl fabric with tentacle,” Proceed- ings of the 4th Workshop on Resource Disaggregation and Serverless , 2023

  54. [62]

    Cxl over ethernet: A novel fpga-based memory disaggregation design in data centers,

    C. Wang, K. He, R. Fan, X. Wang, W. Wang, and Q. Hao, “Cxl over ethernet: A novel fpga-based memory disaggregation design in data centers,” in 2023 IEEE 31st Annual International Symposium on Field- Programmable Custom Computing Machines (FCCM), 2023, pp. 75–82

  55. [63]

    Rcmp: Reconstructing rdma-based memory disaggregation via cxl,

    Z. Wang, Y . Guo, K. Lu, J. Wan, D. Wang, T. Yao, and H. Wu, “Rcmp: Reconstructing rdma-based memory disaggregation via cxl,”ACM Trans. Archit. Code Optim. , vol. 21, no. 1, Jan. 2024

  56. [64]

    Understanding routable PCIe performance for composable infrastructures,

    W. Hou, J. Zhang, Z. Wang, and M. Liu, “Understanding routable PCIe performance for composable infrastructures,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24). Santa Clara, CA: USENIX Association, Apr. 2024, pp. 297–312

  57. [65]

    A case against cxl memory pooling,

    P. Levis, K. Lin, and A. Tai, “A case against cxl memory pooling,” in Proceedings of the 22nd ACM Workshop on Hot Topics in Networks , ser. HotNets ’23. New York, NY , USA: Association for Computing Machinery, 2023, p. 18–24

  58. [66]

    Design tradeoffs in cxl-based memory pools for public cloud platforms,

    D. S. Berger, D. Ernst, H. Li, P. Zardoshti, M. Shah, S. Rajadnya, S. Lee, L. Hsu, I. Agarwal, M. D. Hill, and R. Bianchini, “Design tradeoffs in cxl-based memory pools for public cloud platforms,” IEEE Micro , vol. 43, no. 2, p. 30–38, Mar. 2023

  59. [67]

    Memory sharing with cxl: Hardware and software design approaches,

    S. Jain, N. Yeleswarapu, H. A. Maruf, and R. Gupta, “Memory sharing with cxl: Hardware and software design approaches,” 2024

  60. [68]

    Trenv: Transparently share serverless execution environments across different functions and nodes,

    J. Huang, M. Zhang, T. Ma, Z. Liu, S. Lin, K. Chen, J. Jiang, X. Liao, Y . Shan, N. Zhang, M. Lu, T. Ma, H. Gong, and Y . Wu, “Trenv: Transparently share serverless execution environments across different functions and nodes,” in Proceedings of the ACM SIGOPS 30th Symposium on...

  61. [69]

    HydraRPC: RPC in the CXL era,

    T. Ma, Z. Liu, C. Wei, J. Huang, Y . Zhuo, H. Li, N. Zhang, Y . Guan, D. Niu, M. Zhang, and T. Ma, “HydraRPC: RPC in the CXL era,” in 2024 USENIX Annual Technical Conference (USENIX ATC 24) . Santa Clara, CA: USENIX Association, Jul. 2024, pp. 387–395

  62. [70]

    Telepathic datacenters: Fast rpcs using shared cxl memory,

    S. Mahar, E. Hajyjasini, S. Lee, Z. Zhang, M. Shen, and S. Swanson, “Telepathic datacenters: Fast rpcs using shared cxl memory,” 2024

  63. [71]

    Cxl shared memory programming: Barely distributed and almost persistent,

    Y . Xu, S. Mahar, Z. Liu, M. Shen, and S. Swanson, “Cxl shared memory programming: Barely distributed and almost persistent,” 2024

  64. [72]

    Dissecting cxl memory performance at scale: Analysis, modeling, and optimization,

    J. Liu, H. Hadian, H. Xu, D. S. Berger, and H. Li, “Dissecting cxl memory performance at scale: Analysis, modeling, and optimization,”

  65. [73]

    The hitchhiker’s guide to programming and optimizing cxl-based heterogeneous systems,

    Z. Wang, S. Mahar, L. Li, J. Park, J. Kim, T. Michailidis, Y . Pan, T. Rosing, D. Tullsen, S. Swanson, K. C. Ryoo, S. Park, and J. Zhao, “The hitchhiker’s guide to programming and optimizing cxl-based heterogeneous systems,” 2024. [Online]. Available: https: //arxiv.org/abs/2411.02814

  66. [74]

    The landscape of parallel computing research: A view from berkeley,

    K. Asanovi ´c, R. Bodik, B. C. Catanzaro, J. J. Gebis, P. Husbands, K. Keutzer, D. A. Patterson, W. L. Plishker, J. Shalf, S. W. Williams, and K. A. Yelick, “The landscape of parallel computing research: A view from berkeley,” Tech. Rep. UCB/EECS-2006-183, Dec 2006

  67. [75]

    Benchmarking cloud serving systems with ycsb,

    B. F. Cooper, A. Silberstein, E. Tam, R. Ramakrishnan, and R. Sears, “Benchmarking cloud serving systems with ycsb,” in Proceedings of the 1st ACM Symposium on Cloud Computing , ser. SoCC ’10. New York, NY , USA: Association for Computing Machinery, 2010, p. 143–154

  68. [76]

    Flexible i/o tester

    “Flexible i/o tester.” 2024. [Online]. Available: https://github.com/ axboe/fio

  69. [77]

    Merci: efficient embedding reduction on commodity hardware via sub-query memoization,

    Y . Lee, S. H. Seo, H. Choi, H. U. Sul, S. Kim, J. W. Lee, and T. J. Ham, “Merci: efficient embedding reduction on commodity hardware via sub-query memoization,” in Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Oper...

  70. [78]

    Managing memory tiers with CXL in virtualized environments,

    Y . Zhong, D. S. Berger, C. Waldspurger, R. Wee, I. Agarwal, R. Agarwal, F. Hady, K. Kumar, M. D. Hill, M. Chowdhury, and A. Cidon, “Managing memory tiers with CXL in virtualized environments,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) ....

  71. [79]

    Characterizing the dilemma of performance and index size in billion-scale vector search and breaking it with second-tier memory,

    R. Cheng, Y . Peng, X. Wei, H. Xie, R. Chen, S. Shen, and H. Chen, “Characterizing the dilemma of performance and index size in billion-scale vector search and breaking it with second-tier memory,”

  72. [80]

    Efficient tensor offloading for large deep-learning model training based on compute express link,

    D. Xu, Y . Feng, K. Shin, D. Kim, H. Jeon, and D. Li, “Efficient tensor offloading for large deep-learning model training based on compute express link,” in SC24: International Conference for High Performance Computing, Networking, Storage and Analysis , 2024, pp. 1–18

  73. [81]

    Available: https://arxiv.org/abs/2405.03267

    [Online]. Available: https://arxiv.org/abs/2405.03267

  74. [82]

    Cxl and the return of scale-up database engines,

    A. Lerner and G. Alonso, “Cxl and the return of scale-up database engines,” Proc. VLDB Endow. , vol. 17, no. 10, p. 2568–2575, Aug

  75. [83]

    Enabling cxl memory expansion for in-memory database management systems,

    M. Ahn, A. Chang, D. Lee, J. Gim, J. Kim, J. Jung, O. Rebholz, V . Pham, K. Malladi, and Y . S. Ki, “Enabling cxl memory expansion for in-memory database management systems,” in Proceedings of the 18th International Workshop on Data Management on New Hardware , ser. DaMoN ’22....

  76. [84]

    Database kernels: Seamless integration of database systems and fast storage via cxl,

    S. Lee, A. Lerner, P. Bonnet, and P. Cudr ´e-Mauroux, “Database kernels: Seamless integration of database systems and fast storage via cxl,” in Conference on Innovative Data Systems Research , 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:266732447

  77. [85]

    Available: https://doi.org/10.14778/3675034.3675047

    [Online]. Available: https://doi.org/10.14778/3675034.3675047

  78. [86]

    A three-tier buffer manager integrating cxl device memory for database systems,

    N. Riekenbrauck, M. Weisgut, D. Lindner, and T. Rabl, “A three-tier buffer manager integrating cxl device memory for database systems,” in 2024 IEEE 40th International Conference on Data Engineering Workshops (ICDEW), 2024, pp. 395–401

  79. [87]

    Understanding and optimizing serverless workloads in cxl-enabled tiered memory,

    Y . Li and S. Yao, “Understanding and optimizing serverless workloads in cxl-enabled tiered memory,” 2023. [Online]. Available: https://arxiv.org/abs/2309.01736

  80. [88]

    Cxl-enabled enhanced memory functions,

    D. Boles, D. Waddington, and D. A. Roberts, “Cxl-enabled enhanced memory functions,” IEEE Micro, vol. 43, no. 2, pp. 58–65, 2023

  81. [89]

    ¯Apta: Fault- tolerant object-granular cxl disaggregated memory for accelerating faas,

    A. Patil, V . Nagarajan, N. Nikoleris, and N. Oswald, “ ¯Apta: Fault- tolerant object-granular cxl disaggregated memory for accelerating faas,” in 2023 53rd Annual IEEE/IFIP International Conference on Depend- able Systems and Networks (DSN) , 2023, pp. 201–215

  82. [90]

    A survey on memory-centric computer architectures,

    A. Gebregiorgis, H. A. Du Nguyen, J. Yu, R. Bishnoi, M. Taouil, F. Catthoor, and S. Hamdioui, “A survey on memory-centric computer architectures,” J. Emerg. Technol. Comput. Syst. , vol. 18, no. 4, Oct. 2022

  83. [91]

    Memory protection keys

    J. Corbet, “Memory protection keys.” 2015. [Online]. Available: https://lwn.net/Articles/643797/

  84. [92]

    Memory-centric computing,

    O. Mutlu, “Memory-centric computing,” 2023. [Online]. Available: https://arxiv.org/abs/2305.20000

  85. [2024]

    Available: https://arxiv.org/abs/2409.14317

    [Online]. Available: https://arxiv.org/abs/2409.14317

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.