Pith. sign in

REVIEW 3 major objections 4 minor 67 references

Mainframe-Style Channel Controllers for Modern Disaggregated Memory Systems

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper proposes memory channel controllers: virtual, virtualizable processors attached to far memory, exposed in an application's address space, that exploit symmetric cache coherence for fine-grained interaction without changing the…

desk verdict A well-argued OS-centric design proposal for NDP; the CP real-time deadline problem is real but openly acknowledged, and the abstraction deserves a careful build and test. read the letter →

arxiv 2506.09758 v2 pith:OEBMXN24 submitted 2025-06-11 cs.OS cs.ARcs.ET

classification cs.OScs.ARcs.ET
keywords near-dataprocessingdisaggregatedmemorychannelcontrollerprogramcachecoherenceCXLoperatingsystemsvirtualization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that near-data processing (NDP) has failed to reach practice because it lacks an OS-centric abstraction, and proposes memory channel controllers (MCCs) to fill that gap. An MCC is a virtual processor located next to disaggregated memory that appears in a process's virtual address space as a memory-mapped region, so applications configure and talk to it with ordinary loads and stores. The key claim is that emerging cache-coherent interconnects, such as CXL.mem 3.0 with back-invalidation or ECI, let the MCC and CPU exchange control and data at cache-line granularity through the coherence protocol itself, enabling fine-grained interaction that older NDP designs cannot offer. If this works, applications get portable, virtualizable access to far-memory accelerators without CPU architectural changes, addressing a main obstacle to NDP adoption.

What carries the argument

The memory channel controller (MCC): a virtual processor on a far-memory node, occupying its own region of an application's virtual address space. Its channel program (CP)—an event-driven program that reacts to coherence messages (e.g., from CPU loads/stores) and completes local DRAM operations—is what turns ordinary memory traffic into computation near data. The design rests on a symmetric-coherence interconnect (CXL.mem 3.0 back-invalidation or ECI) that encodes memory transactions in cache-line granules and lets the MCC actively control cache-line ownership; this coherence fabric is the communication and synchronization substrate that makes fine-grained CP interaction possible.

What would settle it

Run the prototype on the Enzian platform with ECI, or on a CXL.mem 3.0 device if one appears, and measure whether a memory-side processor can generate a reply to every CPU-initiated coherence message (load or store) within the interconnect's timeout window while several MCCs are multiplexed; if any reply exceeds the timeout or requires host bias resolution, the MCC programming model cannot meet its latency and liveness guarantees.

Watch

Extended reading notes

Core claim

The central discovery is that a channel-controller-style abstraction can be built for disaggregated memory by mapping each virtual MCC to a region of the application's virtual address space, split into a control area (MMIO for configuration and downloading channel programs) and a data area where the MCC and CPU interact via cache-coherence transactions. A channel program (CP) runs on the MCC and responds programmatically to coherence messages triggered by the CPU's loads and stores, giving a logical view of data generated at runtime and eliminating per-task setup overhead. The MCC can also DMA to and from host-local memory. This requires only memory-side hardware, no changes to CPU architecture or interconnect protocols, and gives the OS a handle to multiplex, isolate, and virtualize many logical MCCs onto a fixed pool of physical processors.

Load-bearing premise

The whole design assumes an interconnect with symmetric cache coherence where memory transactions are cache-line-sized messages, specifically CXL.mem 3.0's back-invalidation or ECI; the paper states that no CXL.mem 3.0 implementations exist yet, so if such hardware never arrives or behaves differently than assumed, the fine-grained coherence-based MCC cannot be built as described.

Editorial extensions

If this is right

  • Applications such as graph common-neighbor search and in-memory database queries can offload irregular, latency-sensitive traversals to MCCs while CPUs keep data-locality-friendly stages, using coherence streaming instead of task queues.
  • Bulk memory operations like zeroing, copy-on-write, VM migration, and huge-page zeroing can run as simple parameterized channel programs near far memory, removing CPU-side data movement.
  • MCCs can observe CPU memory requests to provide fine-grained access statistics for hot-page migration, garbage collection, and profile-guided optimization without extra hardware counters.
  • The OS can multiplex an unbounded number of virtual MCCs onto a small set of physical MCC processors using cooperative coroutine scheduling, preserving isolation through segmentation-style contiguous mappings rather than full address-space replication.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If CXL.mem 3.0 devices ship with the assumed back-invalidation, the same MCC abstraction could be extended beyond far memory to local DRAM controllers and other devices on the coherence fabric, generalizing the paper's focus.
  • The DataPipes-style safe programming model the paper sketches suggests a concrete research program: compile declarative data-movement specifications into verified channel programs and benchmark them against RDMA-based remote-memory operators such as Farview on identical database workloads.
  • Because the CP sits on the critical path of coherence replies, a direct test of the liveness claim is to measure worst-case and tail latency of CP-generated replies on ECI hardware under MCC multiplexing load and compare against interconnect timeout limits.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes Memory Channel Controllers (MCCs), an OS-centric abstraction for near-data processing in disaggregated memory systems. MCCs are presented as virtual processors that occupy regions of an application's virtual address space, with control and data exchanged through memory operations. The key claimed innovation is exploiting cache coherence provided by emerging interconnects (CXL.mem 3.0, ECI) to enable fine-grained, low-latency interaction between CPU threads and NDP engines, without requiring CPU architecture changes. The paper develops the application model, system software model, and hardware design, and illustrates the approach with graph and database workloads. The channel program (CP) design is explicitly acknowledged as ongoing work, and several aspects are described as open questions or expected to work rather than demonstrated.

Significance. The paper addresses a real gap: the lack of a portable, virtualizable OS abstraction for NDP in CXL-based disaggregated memory. Drawing on mainframe channel controllers, the MCC abstraction is a timely and conceptually clean proposal. Its strengths are the emphasis on system-wide requirements (secure multiplexing, virtualization, scheduling) and the honest enumeration of open problems. The paper also benefits from grounding in the authors' Enzian platform and prior work on coherent interconnects. However, the central claims are design arguments, not validated results. There is no implementation or evaluation, and the fine-grained coherence-based control model rests on a hard real-time assumption that the paper acknowledges but does not resolve. If the indicated issues can be addressed with concrete design details or a validation path, the paper could make a strong contribution; as it stands, it is a promising vision rather than a demonstrated system.

major comments (3)
  1. [Section 5.3] The hard real-time problem for CP response is load-bearing for the central claim of a richer, coherence-based programming model. The paper states that when a reply message is needed, the CP is on the critical path and that a late response can deadlock the interconnect, calling this a hard real-time problem. Yet it provides no worst-case execution time bound for any CP, no scheduler with a liveness or deadline guarantee under multiplexing, and no admission control for overload. The statements that latencies are predictable and that timeouts are millisecond-scale are not sufficient to establish safety: a descheduled or overloaded CP can still miss a deadline. Because this issue affects the ECI platform the authors plan to prototype on, not just future CXL.mem 3.0, it must be addressed before the core abstraction can be accepted as sound. The paper should either provide a schedulability analysis for the cooperative coroutine scheduler, or explicitly reframe the coherence-based control as best-effort with a fallback mechanism.
  2. [Section 5.1] The design depends on interconnects that provide symmetric coherence and encode memory transactions in cache-line granules. The paper acknowledges that CXL.mem 3.0 has this property but that no implementations yet exist, and that ECI is the only currently available platform with this property, on the Enzian research hardware. This makes the claimed portability untestable on commodity CXL systems. The paper should clarify which design elements can be validated on ECI today, which depend on future CXL.mem 3.0 availability, and what mechanisms would be needed if the symmetric-coherence assumption does not hold (e.g., if only bias-based CXL.cache is available). Without such clarification, the portability claim is overstated relative to current hardware reality.
  3. [Section 5.2] The claim that a fixed number of physical processors can multiplex an unbounded number of MCCs via cooperative coroutine scheduling is not substantiated. The paper says that "relatively simple scheduling might provide sufficient guarantees against starvation under load" but does not specify any concrete policy, nor does it prove absence of starvation or bounded response time. Similarly, the claim that maintaining virtual address spaces on the far memory node can be made efficient via segmentation is only a plausibility argument, without analysis of the metadata consistency overhead. Since virtualizability and the absence of arbitrary resource limits are core properties of the proposed abstraction, the paper needs to provide a concrete scheduler design and an analysis of isolation, overhead, and liveness, or these properties must be presented as aspirational rather than guaranteed.
minor comments (4)
  1. [Section 6] Typo: "generous-purpose processor" should be "general-purpose processor."
  2. [Section 6] Typo: "tired memory systems" should be "tiered memory systems."
  3. [Figure 1] The labels 1, 2, 3 in the figure are not fully described in the caption; please add a brief explanation of the numbered operations in the caption or in the running text.
  4. [Section 4] The phrase "the precise semantics for CPs is, at this point, an open question" is a strong caveat that should be reflected in the abstract, where the programming model is presented as a key innovation.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper proposes an OS abstraction and does not derive numeric results from fitted inputs; its self-citations are contextual, not load-bearing.

full rationale

This is a systems design/position paper, not a derivation. The central contribution is a proposed memory channel controller (MCC) abstraction: virtual processors addressed through cache-coherent memory regions. There are no fitted parameters, no empirical predictions, and no uniqueness theorem invoked to force a choice. The design's dependence on symmetric coherent interconnects (CXL.mem 3.0 or ECI) is stated as an assumption (Section 5.1) and is a feasibility/availability risk, not a circular step. The paper explicitly flags open questions: 'The precise semantics for CPs is, at this point, an open question' (Section 4) and the hard-real-time concern 'we believe it is solvable' (Section 5.3). Self-citations to Enzian [14], programmed I/O over coherent interconnects [47], and the NIC-in-OS proposal [57] are used as supporting context and prior-platform references, but the central abstraction does not reduce to these citations. The workload mappings (graph processing, databases, memory copying, access statistics) are illustrative applications of the proposed model, not results derived from it. Therefore no circular step can be exhibited, and the honest finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 2 invented entities

No free parameters appear because the paper makes no numeric claims or fits. The stated axioms are hardware assumptions about symmetric coherent interconnects, hard real-time response of CPs, metadata consistency, and cooperative scheduling. The central invented entity is the MCC, with CP as its programming model; neither has independent evidence yet.

assumptions (4)
  • domain assumption Emerging interconnects provide symmetric cache coherence with memory-transaction semantics.
    Section 5.1 assumes message-based interconnects that encode memory transactions in fixed-size cache line granules and allow symmetric coherency with back-invalidation. CXL.mem 3.0 has the property but no implementations exist yet.
  • domain assumption A channel program can respond to coherence messages within interconnect timeout deadlines to avoid deadlock.
    Section 5.3 acknowledges the CP is on the critical path and calls this a hard real-time problem, claiming it is solvable because DRAM latencies are predictable and CXL/ECI allow millisecond-scale timeouts.
  • domain assumption OS and MCC can keep virtual address space metadata consistent with manageable overhead.
    Section 5.2 argues that contiguous far memory mappings reduce metadata, and cites prior work [57] to suggest scheduling state can be shared over a coherent interconnect; this is asserted, not demonstrated.
  • ad hoc to paper A fixed number of physical processors can multiplex an unbounded number of MCCs via cooperative coroutine scheduling without starvation.
    Section 5.2 proposes cooperative scheduling of CP interpreters as coroutines, with CPU simulation when overloaded; correctness and liveness under load are not proven.
invented entities (2)
  • Memory Channel Controller (MCC)
    purpose: Virtual, dynamically instantiated processors close to far memory that execute channel programs on behalf of applications, providing portable and virtualizable NDP.
    The paper proposes MCC as a new abstraction and hardware/software component. No implementation or external falsifiable evidence is provided; the paper says prototyping on Enzian is expected (Section 7).
  • Channel Program (CP)
    purpose: Program executed by an MCC that interacts with the host via coherence transactions and DMA, replacing task-based offload models.
    CP is the programming model defined by the paper; semantics are still an open question (Section 4), so no independent evidence exists.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Mainframe-Style Channel Controllers for Modern Disaggregated Memory Systems." pith.science (2026). https://pith.science/paper/OEBMXN24

@misc{pith2026250609758,
  author       = {Pith},
  title        = {Pith review of: Mainframe-Style Channel Controllers for Modern Disaggregated Memory Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OEBMXN24}},
  note         = {Machine review of arXiv:2506.09758}
}
read the original abstract

Despite the promise of alleviating the main memory bottleneck, and the existence of commercial hardware implementations, techniques for Near-Data Processing have seen relatively little real-world deployment. The idea has received renewed interest with the appearance of disaggregated or "far" memory, for example in the use of CXL memory pools. However, we argue that the lack of a clear OS-centric abstraction of Near-Data Processing is a major barrier to adoption of the technology. Inspired by the channel controllers which interface the CPU to disk drives in mainframe systems, we propose memory channel controllers as a convenient, portable, and virtualizable abstraction of Near-Data Processing for modern disaggregated memory systems. In addition to providing a clean abstraction that enables OS integration while requiring no changes to CPU architecture, memory channel controllers incorporate another key innovation: they exploit the cache coherence provided by emerging interconnects to provide a much richer programming model, with more fine-grained interaction, than has been possible with existing designs.

Figures

Figures reproduced from arXiv: 2506.09758 by the authors.

Figure 1
Figure 1. The far memory and MCC abstraction We are not the first one to notice those problems. Gao et al. [19] and Ghose et al. [20] observe the gaps in address translation, memory protection and isolation functionality. Barbalace et al. [3] additionally discuss the problems in sched￾uling and the programming model, and call for better run￾time and OS support. More recently, Ham et al. propose M2NDP [22], which focuses on ND… view at source ↗
Figure 2
Figure 2. System architecture (typically, the one where the hardware that it runs on is located), or if should see anything accessible from the appli￾cation’s virtual address space. In the latter case, we argue it is still important that each MCC has an affinity which specifies what memory is local to the MCC. When the system scales to multiple memory nodes, data placement strategies become relevant for reducing data movement… view at source ↗
Figure 3
Figure 3. 𝑛-hop common neighbor pipeline 𝑛-hop common connections between users [54]) [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

67 extracted references · 35 canonical work pages

  1. [1]

    Junwhan Ahn, Sungpack Hong, Sungjoo Yoo, Onur Mutlu, and Kiy- oung Choi. 2015. A Scalable Processing-in-Memory Accelerator for Parallel Graph Processing. In2015 ACM/IEEE 42nd Annual In- ternational Symposium on Computer Architecture (ISCA). 105–117. doi:10.1145/2749469.2750386

  2. [2]

    Arjan van de Ven. 2024. VFIO: Add the SPR_DSA and SPR_IAX De- vices to the Denylist. https://git.kernel.org/pub/scm/linux/kernel/git/ torvalds/linux.git/commit/?id=95feb3160eef

  3. [3]

    Antonio Barbalace, Anthony Iliopoulos, Holm Rauchfuss, and Goetz Brasche. 2017. It’s Time to Think About an Operating System for Near Data Processing Architectures. InProceedings of the 16th Workshop on Hot Topics in Operating Systems(Whistler, BC, Canada)(HotOS ’17). Association for Computing Machinery, New York, NY, USA, 56–61. doi:10.1145/3102980.3102990

  4. [4]

    Alexander Baumstark, Muhammad Attahir Jibril, and Kai-Uwe Sattler

  5. [5]

    Bensoussan, C

    A. Bensoussan, C. T. Clingen, and R. C. Daley. 1972. The Multics Virtual Memory: Concepts and Design.Commun. ACM15, 5 (May 1972), 308–318. doi:10.1145/355602.361306

  6. [6]

    2019.CCIX Base Specification Revision 1.0a Version 1.0 for Evaluation

    CCIX Consortium, Inc. 2019.CCIX Base Specification Revision 1.0a Version 1.0 for Evaluation. Technical Report. 346 pages

  7. [7]

    Dehao Chen, David Xinliang Li, and Tipp Moseley. 2016. AutoFDO: Automatic Feedback-Directed Optimization for Warehouse-Scale Ap- plications. InProceedings of the 2016 International Symposium on Code Generation and Optimization (CGO ’16). Association for Computing Machinery, New York, NY, USA, 12–23. doi:10.1145/2854038.2854044

  8. [8]

    Wen-ke Chen, Sanjay Bhansali, Trishul Chilimbi, Xiaofeng Gao, and Weihaw Chuang. 2006. Profile-Guided Proactive Garbage Collection for Locality Optimization. InProceedings of the 27th ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI ’06). Association for Computing Machinery, New York, NY, USA, 332–

Show all 67 references
  1. [9]

    Audrey Cheng, Xiao Shi, Aaron Kabcenell, Shilpa Lawande, Hamza Qadeer, Jason Chan, Harrison Tin, Ryan Zhao, Peter Bailis, Mahesh Balakrishnan, Nathan Bronson, Natacha Crooks, and Ion Stoica. 2022. TAOBench: An End-to-End Benchmark for Social Network Workloads. Proc. VLDB Endow...

  2. [10]

    Chilimbi and James R

    Trishul M. Chilimbi and James R. Larus. 1998. Using Generational Garbage Collection to Implement Cache-Conscious Data Placement. SIGPLAN Not.34, 3 (Oct. 1998), 37–48. doi:10.1145/301589.286865

  3. [11]

    Tzi-cker Chiueh, Ganesh Venkitachalam, and Prashant Pradhan. 1999. Integrating Segmentation and Paging Protection for Safe, Efficient and Transparent Software Extensions. InProceedings of the Seven- teenth ACM Symposium on Operating Systems Principles (SOSP ’99). Association f...

  4. [12]

    Awasthi, Emmanuel S

    Anita Choudhary, Mahesh Chandra Govil, Girdhari Singh, Lalit K. Awasthi, Emmanuel S. Pilli, and Divya Kapil. 2017. A Critical Survey of Live Virtual Machine Migration Techniques.J. Cloud Comput.6, 1 (Dec. 2017), 92:1–92:41

  5. [13]

    Christopher Clark, Keir Fraser, Steven Hand, Jacob Gorm Hansen, Eric Jul, Christian Limpach, Ian Pratt, and Andrew Warfield. 2005. Live Migration of Virtual Machines. InProceedings of the 2nd Conference on Symposium on Networked Systems Design & Implementation - Volume 2 (NSDI...

  6. [14]

    David Cock, Abishek Ramdas, Daniel Schwyn, Michael Giardino, Adam Turowski, Zhenhao He, Nora Hossle, Dario Korolija, Melissa Liccia- rdello, Kristina Martsenko, Reto Achermann, Gustavo Alonso, and Timothy Roscoe. 2022. Enzian: An Open, General, CPU/FPGA Plat- form for Systems ...

  7. [15]

    2023.Compute Ex- press Link Specification Revision 3.1

    Compute Express Link Consortium, Inc. 2023.Compute Ex- press Link Specification Revision 3.1. Technical Report. 1166 pages. https://computeexpresslink.org/wp-content/uploads/2024/02/ CXL-3.1-Specification.pdf

  8. [16]

    Robert Courts. 1988. Improving Locality of Reference in a Garbage- Collecting Memory Management System.Commun. ACM31, 9 (Sept. 1988), 1128–1138. doi:10.1145/48529.48536

  9. [17]

    P.J. Denning. 1969. Equipment Configuration in Balanced Computer Systems.IEEE Trans. Comput.C-18, 11 (Nov. 1969), 1008–1012. doi:10. 1109/T-C.1969.222571

  10. [19]

    Mingyu Gao, Grant Ayers, and Christos Kozyrakis. 2015. Practical Near- Data Processing for In-Memory Analytics Frameworks. In2015 Inter- national Conference on Parallel Architecture and Compilation (PACT). 113–124. doi:10.1109/PACT.2015.22

  11. [20]

    2019.The Processing-in-Memory Paradigm: Mechanisms to Enable Adoption

    Saugata Ghose, Kevin Hsieh, Amirali Boroumand, Rachata Ausavarungnirun, and Onur Mutlu. 2019.The Processing-in-Memory Paradigm: Mechanisms to Enable Adoption. Springer International Publishing, Cham, 133–194. doi:10.1007/978-3-319-90385-9_5

  12. [21]

    Oliveira, and Onur Mutlu

    Juan Gómez-Luna, Izzat El Hajj, Ivan Fernandez, Christina Gian- noula, Geraldo F. Oliveira, and Onur Mutlu. 2022. Benchmarking a New Paradigm: Experimental Analysis and Characterization of a Real Processing-in-Memory System.IEEE Access10 (2022), 52565–52608. doi:10.1109/ACCESS...

  13. [22]

    Hyungkyu Ham, Jeongmin Hong, Geonwoo Park, Yunseon Shin, Okkyun Woo, Wonhyuk Yang, Jinhoon Bae, Eunhyeok Park, Hyo- jin Sung, Euicheol Lim, and Gwangsun Kim. 2024. Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders. In2024 57th IEEE/ACM International Sym...

  14. [23]

    Malladi, Andrew Chang, and Yuan Xie

    Wenqin Huangfu, Krishna T. Malladi, Andrew Chang, and Yuan Xie

  15. [24]

    Intel. 2024. CVE-2024-21823: Intel DSA and IAA Escalation of Priv- ilege. https://www.intel.com/content/www/us/en/security-center/ advisory/intel-sa-01084.html

  16. [25]

    1964.IBM System/360 Principles of Operation

    International Business Machines Corporation. 1964.IBM System/360 Principles of Operation. IBM Press. https://dl.acm.org/doi/book/10. 5555/1102026

  17. [26]

    International Business Machines Corporation. 1969. IBM System/360 Component Descriptions 2314 Direct Access Storage Facility and 2844 Auxiliary Storage Control. APSys ’25, October 12–13, 2025, Seoul, Republic of Korea Zikai Liu, Jasmin Schult, Pengcheng Xu, and Timothy Roscoe

  18. [27]

    Junhyeok Jang, Hanjin Choi, Hanyeoreum Bae, Seungjun Lee, Miryeong Kwon, and Myoungsoo Jung. 2023. CXL-ANNS: Software- Hardware Collaborative Memory Disaggregation and Computation for Billion-Scale Approximate Nearest Neighbor Search. In2023 USENIX Annual Technical Conference ...

  19. [28]

    Yoon, Jeong-Uk Kang, Sangyeun Cho, Daniel D

    Insoon Jo, Duck-Ho Bae, Andre S. Yoon, Jeong-Uk Kang, Sangyeun Cho, Daniel D. G. Lee, and Jaeheon Jeong. 2016. YourSQL: A High- Performance Database System Leveraging in-Storage Computing.Proc. VLDB Endow.9, 12 (Aug. 2016), 924–935. doi:10.14778/2994509.2994512

  20. [29]

    Aditya K Kamath and Simon Peter. 2024. (MC)2: Lazy MemCopy at the Memory Controller. In2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). 1112–1128. doi:10.1109/ ISCA59077.2024.00084

  21. [30]

    Dario Korolija, Dimitrios Koutsoukos, Kimberly Keeton, Konstantin Taranov, Dejan Milojičić, and Gustavo Alonso. 2021. Farview: Disag- gregated Memory with Operator Off-loading for Database Engines. doi:10.48550/arXiv.2106.07102 arXiv:2106.07102 [cs]

  22. [31]

    Rossbach, and Emmett Witchel

    Youngjin Kwon, Hangchen Yu, Simon Peter, Christopher J. Rossbach, and Emmett Witchel. 2017. Ingens: Huge Page Support for the OS and Hypervisor.SIGOPS Oper. Syst. Rev.51, 1 (Sept. 2017), 83–93. doi:10.1145/3139645.3139659

  23. [32]

    Norman Layer and Edwin D. Reilly. 2003. IBM System 360/370/390 Series. InEncyclopedia of Computer Science. John Wiley and Sons Ltd., GBR, 828–832

  24. [33]

    Berger, Lisa Hsu, Daniel Ernst, Pantea Zar- doshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D

    Huaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst, Pantea Zar- doshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D. Hill, Marcus Fontoura, and Ricardo Bian- chini. 2023. Pond: CXL-Based Memory Pooling Systems for Cloud Platforms. InPro...

  25. [34]

    Berger, Marie Nguyen, Xun Jian, Sam H

    Jinshu Liu, Hamid Hadian, Yuyue Wang, Daniel S. Berger, Marie Nguyen, Xun Jian, Sam H. Noh, and Huaicheng Li. 2025. System- atic CXL Memory Characterization and Performance Analysis at Scale. InProceedings of the 30th ACM International Conference on Architectural Support for P...

  26. [35]

    Elliot Lockerman, Axel Feldmann, Mohammad Bakhshalipour, Alexan- dru Stanescu, Shashwat Gupta, Daniel Sanchez, and Nathan Beckmann

  27. [36]

    Andrew Lumsdaine, Douglas Gregor, Bruce Hendrickson, and Jonathan Berry. 2007. Challenges in Parallel Graph Processing.Parallel Pro- cessing Letters17, 01 (2007), 5–20. doi:10.1142/S0129626407002843 arXiv:https://doi.org/10.1142/S0129626407002843

  28. [37]

    Hasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner, Niket Agarwal, Pallab Bhattacharya, Chris Petersen, Mosharaf Chowd- hury, Shobhit Kanaujia, and Prakash Chauhan. 2023. TPP: Transparent Page Placement for CXL-Enabled Tiered-Memory. InProceedings of the 28th ACM Int...

  29. [38]

    2024.Marvell Structera A 2504 Memory-Expansion Con- troller

    Marvell. 2024.Marvell Structera A 2504 Memory-Expansion Con- troller. Technical Report Marvell_Structera_A MV-SLA25041 _PB. 3 pages. https://www.marvell.com/content/dam/marvell/en/public- collateral/assets/marvell-structera-a-2504-near-memory- accelerator-product-brief.pdf

  30. [39]

    MongoDB. [n. d.]. In-Memory Databases Explained - Mon- goDB. https://www.mongodb.com/resources/basics/databases/in- memory-database

  31. [40]

    August, Hyoun Kyu Cho, Svilen Kanev, Christos Kozyrakis, Trivikram Krishnamurthy, Heiner Litz, Tipp Moseley, and Parthasarathy Ranganathan

    Nayana Prasad Nagendra, Grant Ayers, David I. August, Hyoun Kyu Cho, Svilen Kanev, Christos Kozyrakis, Trivikram Krishnamurthy, Heiner Litz, Tipp Moseley, and Parthasarathy Ranganathan. 2020. As- mDB: Understanding and Mitigating Front-End Stalls in Warehouse- Scale Computers....

  32. [41]

    A. Padegs. 1964. The Structure of SYSTEM/360, Part IV: Channel Design Considerations.IBM Systems Journal3, 2 (1964), 165–179. doi:10.1147/sj.32.0165

  33. [42]

    1999.The PageRank Citation Ranking: Bringing Order to the Web

    Lawrence Page, Sergey Brin, Rajeev Motwani, and Terry Winograd. 1999.The PageRank Citation Ranking: Bringing Order to the Web. Technical Report 1999-66. Stanford InfoLab / Stanford InfoLab. http: //ilpubs.stanford.edu:8090/422/

  34. [43]

    Gopinath

    Ashish Panwar, Sorav Bansal, and K. Gopinath. 2019. HawkEye: Ef- ficient Fine-grained OS Support for Huge Pages. InProceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’19). As- sociation for...

  35. [44]

    Loh, and Abhishek Bhattacharjee

    Binh Pham, Ján Veselý, Gabriel H. Loh, and Abhishek Bhattacharjee

  36. [45]

    2023.CCKit: FPGA Acceleration in Symmetric Coherent Heterogeneous Platforms

    Abishek Ramdas. 2023.CCKit: FPGA Acceleration in Symmetric Coherent Heterogeneous Platforms. Doctoral Thesis. ETH Zurich. doi:10.3929/ethz-b-000642567

  37. [46]

    Redis Development Team. 2024. Redis Documentation: Diagnosing Latency Issues. https://redis.io/docs/latest/operate/oss_and_stack/ management/optimization/latency/

  38. [47]

    Anastasiia Ruzhanskaia, Pengcheng Xu, David Cock, and Timothy Roscoe. 2025. Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent Interconnects. doi:10.48550/arXiv.2409.08141 arXiv:2409.08141 [cs]

  39. [48]

    Joonseop Sim, Soohong Ahn, Taeyoung Ahn, Seungyong Lee, Myunghyun Rhee, Jooyoung Kim, Kwangsik Shin, Donguk Moon, Euiseok Kim, and Kyoung Park. 2022. Computational cxl-memory so- lution for accelerating memory-intensive applications.IEEE Computer Architecture Letters22, 1 (2022), 5–8

  40. [49]

    Solihin, Jaejin Lee, and J

    Y. Solihin, Jaejin Lee, and J. Torrellas. 2002. Using a User-Level Memory Thread for Correlation Prefetching. InProceedings 29th An- nual International Symposium on Computer Architecture. 171–182. doi:10.1109/ISCA.2002.1003576

  41. [50]

    Yan Sun, Jongyul Kim, Zeduo Yu, Jiyuan Zhang, Siyuan Chai, Michael Jaemin Kim, Hwayong Nam, Jaehyun Park, Eojin Na, Yi- fan Yuan, Ren Wang, Jung Ho Ahn, Tianyin Xu, and Nam Sung Kim. 2025. M5: Mastering Page Migration and Memory Manage- ment for CXL-based Tiered Memory Systems...

  42. [51]

    Yan Sun, Yifan Yuan, Zeduo Yu, Reese Kuper, Chihun Song, Jinghan Huang, Houxiang Ji, Siddharth Agarwal, Jiaqi Lou, Ipoom Jeong, Ren Mainframe-Style Channel Controllers for Modern Disaggregated Memory Systems APSys ’25, October 12–13, 2025, Seoul, Republic of Korea Wang, Jung H...

  43. [52]

    Dufy Teguia, Jiaxuan Chen, Stella Bitchebe, Oana Balmau, and Alain Tchana. 2024. vPIM: Processing-in-Memory Virtualization. InProceed- ings of the 25th International Middleware Conference (Middleware ’24). Association for Computing Machinery, New York, NY, USA, 417–430. doi:10...

  44. [53]

    Lukas Vogel, Daniel Ritter, Danica Porobic, Pinar Tözün, Tianzheng Wang, and Alberto Lerner. 2023. Data Pipes: Declarative Control over Data Movement. InConference on Innovative Data Systems Research

  45. [54]

    Rui Wang, Christopher Conrad, and Sam Shah. 2013. Using Set Cover to Optimize a Large-Scale Low Latency Distributed Graph. In5th USENIX Workshop on Hot Topics in Cloud Computing (Hot- Cloud 13). https://www.usenix.org/conference/hotcloud13/workshop- program/presentations/wang

  46. [55]

    Zhao Wang, Yiqi Chen, Cong Li, Yijin Guan, Dimin Niu, Tianchan Guan, Zhaoyang Du, Xingda Wei, and Guangyu Sun. 2025. CTXNL: A Software-Hardware Co-designed Solution for Efficient CXL-Based Transaction Processing. InProceedings of the 30th ACM International Conference on Archit...

  47. [56]

    Louis Woods, Zsolt István, and Gustavo Alonso. 2014. Ibex: An Intelli- gent Storage Engine with Support for Advanced SQL Offloading.Proc. VLDB Endow.7, 11 (July 2014), 963–974. doi:10.14778/2732967.2732972

  48. [57]

    Pengcheng Xu and Timothy Roscoe. 2025. The NIC Should Be Part of the OS(HotOS ’25)

  49. [58]

    Chia-Lin Yang and Alvin R. Lebeck. 2000. Push vs. Pull: Data Movement for Linked Data Structures. InProceedings of the 14th International Conference on Supercomputing (ICS ’00). Association for Computing Machinery, New York, NY, USA, 176–186. doi:10.1145/335231.335248

  50. [59]

    Blackburn, Daniel Frampton, Jennifer B

    Xi Yang, Stephen M. Blackburn, Daniel Frampton, Jennifer B. Sartor, and Kathryn S. Mckinley. 2011. Why Nothing Matters: The Impact of Zeroing. InOOPSLA’11 - Proceedings of the 2011 ACM International Conference on Object Oriented Programming Systems Languages and Applications (...

  51. [60]

    Kaiyuan Zhang, Rong Chen, and Haibo Chen. 2015. NUMA-aware Graph-Structured Analytics.SIGPLAN Not.50, 8 (Jan. 2015), 183–193. doi:10.1145/2858788.2688507

  52. [61]

    Qizhen Zhang, Yifan Cai, Xinyi Chen, Sebastian Angel, Ang Chen, Vin- cent Liu, and Boon Thau Loo. 2020. Understanding the Effect of Data Center Resource Disaggregation on Production DBMSs.Proc. VLDB Endow.13, 9 (May 2020), 1568–1581. doi:10.14778/3397230.3397249

  53. [62]

    Berger, Carl Waldspurger, Ryan Wee, Ishwar Agarwal, Rajat Agarwal, Frank Hady, Karthik Kumar, Mark D

    Yuhong Zhong, Daniel S. Berger, Carl Waldspurger, Ryan Wee, Ishwar Agarwal, Rajat Agarwal, Frank Hady, Karthik Kumar, Mark D. Hill, Mosharaf Chowdhury, and Asaf Cidon. 2024. Managing Memory Tiers with CXL in Virtualized Environments. In18th USENIX Symposium on Operating System...

  54. [63]

    Youwei Zhuo, Chao Wang, Mingxing Zhang, Rui Wang, Dimin Niu, Yanzhi Wang, and Xuehai Qian. 2019. GraphQ: Scalable PIM-Based Graph Processing. InProceedings of the 52nd Annual IEEE/ACM Inter- national Symposium on Microarchitecture. ACM, Columbus OH USA, 712–725. doi:10.1145/33...

  55. [340]

    doi:10.1145/1133981.1134021

  56. [2015]

    InProceedings of the 48th International Symposium on Microarchitecture (MICRO-48)

    Large Pages and Lightweight Memory Management in Virtu- alized Environments: Can You Have It Both Ways?. InProceedings of the 48th International Symposium on Microarchitecture (MICRO-48). Association for Computing Machinery, New York, NY, USA, 1–12. doi:10.1145/2830772.2830773

  57. [2020]

    InProceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’20)

    Livia: Data-Centric Computing Throughout the Memory Hi- erarchy. InProceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’20). Association for Computing Machinery, New York, NY, USA, 417–433. d...

  58. [2022]

    In2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO)

    BEACON: Scalable Near-Data-Processing Accelerators for Genome Analysis near Memory Pool with the CXL Support. In2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO). 727–743. doi:10.1109/MICRO56248.2022.00057

  59. [2023]

    InProceedings of the 19th International Workshop on Data Management on New Hardware (DaMoN ’23)

    Processing-in-Memory for Databases: Query Processing and Data Transfer. InProceedings of the 19th International Workshop on Data Management on New Hardware (DaMoN ’23). Association for Computing Machinery, New York, NY, USA, 107–111. doi:10.1145/ 3592980.3595323

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.