ORCID

0009-0008-2625-9391

Keywords

Sparse Matrix Multiplication, Compressed Sparse Row, Compressed Sparse Column, Non-volatile Memory, Memory Allocation, Graph Convolutional Networks, Sparse General Matrix Multiplication, Compressed Sparse Formats, Memory Optimization, Fault-Tolerance for LLMs, Machine Learning.

Abstract

The rapid advancement of machine learning, from Graph Convolutional Networks (GCNs) to transformer-based Large Language Models (LLMs), continues to expose fundamental limitations in memory efficiency, scalability, and system reliability. GCNs are critical for domains such as biomedical modeling, social networks, and recommendation systems, yet their reliance on sparse general matrix–matrix multiplication (SpGEMM) makes them highly sensitive to GPU memory constraints and irregular access patterns. As graph data scales, out-of-core computation becomes inevitable, but existing systems are hindered by I/O bottlenecks and underutilized GPU resources. Our contribution with AIRES demonstrates how algorithm–system co-design can alleviate these challenges: by introducing block-wise data alignment for sparse formate and a dynamic three-phase scheduling framework with dual-path transfers across GPU memory, host memory, and storage, AIRES significantly reduces I/O latency and sustains up to 1.8x faster end-to-end training. This work highlights how careful coordination between algorithmic structures and system-level memory hierarchies is essential for scaling graph-based learning. LLMs underscore the importance of reliability in large-scale distributed inference. Multi-GPU deployments, while necessary for training and serving massive models, remain vulnerable to hardware failures that can waste compute cycles and disrupt real-time applications. Traditional checkpointing is too slow and resource-intensive for the fast-paced demands of LLM inference, especially during prefill when large KV caches are generated. With GhostServe, we contribute a new approach to fault tolerance in machine learning by embedding lightweight redundancy directly into the inference pipeline. By leveraging blockwise Row–Diagonal Parity (RDP) and custom CUDA kernels for parity encoding and recovery, GhostServe enables near-immediate reconstruction of failed KV cache blocks with negligible overhead, sustaining throughput and preserving accuracy. Evaluations show that GhostServe achieves 2.7x speedup during the checkpointing backup process compared to the state-of-the-art method, and recovers with negligible latency overhead, offering a practical path to high-availability LLM serving for all workloads.

This demonstrates that coding-based redundancy, when carefully integrated with GPU kernels and system-level scheduling, can bridge the gap between high performance and resilience in ML workloads. Together, these contributions reinforce a broader principle: the future of machine learning will depend on co-designing algorithms, kernels, and system architectures that simultaneously maximize efficiency, scalability, and reliability. AIRES shows how memory-conscious design can unlock new levels of performance for graph learning, while GhostServe illustrates how resilience can be achieved without compromising speed in large-scale language model inference. More generally, they point toward an emerging paradigm where high-performance GPU computation kernels are not only optimized for speed, but also for data movement, memory hierarchy, and fault tolerance, critical dimensions for sustaining the continued growth of machine learning at scale.

Completion Date

2025

Semester

Fall

Committee Chair

Jun Wang

Degree

Doctor of Philosophy (Ph.D.)

College

College of Engineering and Computer Science

Department

Electrical and Computer Engineering

Format

PDF

Release Date

6-15-2026

Document Type

Dissertation

Campus Location

Orlando (Main) Campus

Subjects

High performance computing--Research; Neural networks (Computer science)--Design and construction; Machine learning--Graphic methods; Fault-tolerant computing--Mathematical models; Adaptive computing systems

Share

COinS
 

Accessibility Statement

This item was created or digitized prior to April 24, 2027, or is a reproduction of legacy media created before that date. It is preserved in its original, unmodified state specifically for research, reference, or historical recordkeeping. In accordance with the ADA Title II Final Rule, the University Libraries provides accessible versions of archival materials upon request. To request an accommodation for this item, please submit an accessibility request form.