ORCID

0009-0007-7852-5148

Keywords

Persistent Memory, CXL, NVMM, Memory disaggregation

Subject Categories

Computer and Systems Architecture | Computer Engineering | Data Storage Systems

Abstract

Compute express link (CXL) enables persistent memory disaggregation with memory pooling and hardware managed multi-host memory sharing capability, resulting in better resource utilization, increased scalability. Persistency-aware applications need to manage crash consistency across the system which results in significant performance overhead. This dissertation systematically investigates performance overhead to achieve crash consistency in disaggregated persistent memory and proposes solutions to enable persistency-aware application scaling for distributed system. First, we study persistent parallel programming to scale computation capability beyond single processor and determine the underlying hardware limitation to adopt lock-free data structure. We propose hardware support to design durable atomic instruction (DAI) that guarantees crash consistency for atomic operation on shared memory. Our proposed solution utilizes transient state of cache coherence protocol to ensure data durability before any process could read the atomic update. We design durMESI, an extended version of MESI protocol to demonstrate our proposed solution. Second, we explored performance overhead for crash consistency in persistent memory pooling and identified the need for adding persist capability in network level that works in conjunction with existing processor centric optimizations. We introduce distributed persistent domain (DPD), a novel approach of defining persistent domain that enables adding data persist capability at switch level in distributed fashion while guaranteeing correctness and crash consistency. We designed persistent CXL switch (PCS) and analyzed correctness guarantee when adopting DPD to demonstrate validity of DPD. Finally, we evaluate DAIs performance overhead in a disaggregated system. Because disaggregation increases the latency required to make data durable, DAIs cause read stalls and cache-level contention during simultaneous shared memory access. Combining DAIs with DPD’s switch-level persistence successfully reduces this increased contention. Ultimately, this dissertation addresses critical performance bottlenecks in persistent parallel programming and memory pooling, achieving system-wide efficiency for crash-consistent disaggregated persistent memory.

Completion Date

2026

Semester

Summer

Committee Chair

Yan Solihin

Degree

Doctor of Philosophy (Ph.D.)

College

College of Engineering and Computer Science

Department

Computer Science

Format

PDF

Document Type

Dissertation

Language

English

Share

COinS
 

Accessibility Statement

This item was created or digitized prior to April 24, 2027, or is a reproduction of legacy media created before that date. It is preserved in its original, unmodified state specifically for research, reference, or historical recordkeeping. In accordance with the ADA Title II Final Rule, the University Libraries provides accessible versions of archival materials upon request. To request an accommodation for this item, please submit an accessibility request form.