ORCID

0009-0002-6310-2787

Keywords

Hardware/software co-design; domain-specific accelerators; memory-efficient computing; data movement; vision transformers; 3D human mesh recovery; pangenome mapping; graph processing; register-transfer-level optimization; design-space exploration

Subject Categories

Computer and Systems Architecture | Computer Engineering | Hardware Systems

Abstract

Emerging artificial-intelligence and data-intensive scientific workloads increasingly face a memory wall: irregular access patterns and large intermediate data volumes make data movement, rather than arithmetic, the primary constraint on performance and energy efficiency. This dissertation develops a memory-centric hardware/software co-design methodology that jointly reshapes algorithms, architectures, and dataflows to retain frequently reused data on chip. The methodology is demonstrated through three accelerators and an RTL design tool. VITA replaces multi-head attention in vision-transformer-based 3D human mesh recovery with hardware-friendly average pooling and maps the resulting operators to a reconfigurable datapath, achieving 5.05-fold and 69.12-fold speedups over a state-of-the-art GPU and CPU while maintaining competitive accuracy. PanGaea reformulates sequence-to-graph pangenome mapping as a fused, tiled pipeline supported by a tri-mode processing array, improving throughput, energy efficiency, and area efficiency by 1.47-fold, 1.73-fold, and 1.62-fold, respectively, over the strongest prior ASIC. Locus extends this design with bubble-preserving graph partitioning, cross-read subgraph reuse, and variation-aware dataflow; it eliminates 96 percent of DRAM traffic, improves throughput by 2.1-fold, and preserves mapping accuracy within 0.02 percent of the software reference. Finally, KAIROS accelerates RTL timing exploration through a word-level intermediate representation, taxonomy-aware delay models, fan-out management, and retiming. Across 62 designs, it achieves a Spearman rank correlation of 0.96 with post-synthesis frequency and evaluates timing up to 213-fold faster than synthesis. Together, these contributions establish memory-efficient co-design as a repeatable methodology for building emerging-domain accelerators and accelerating their design.

Completion Date

2026

Semester

Summer

Committee Chair

Hao Zheng

Degree

Doctor of Philosophy (Ph.D.)

College

College of Engineering and Computer Science

Department

Department of Electrical and Computer Engineering

Format

PDF

Document Type

Dissertation

Language

English

Share

COinS
 

Accessibility Statement

This item was created or digitized prior to April 24, 2027, or is a reproduction of legacy media created before that date. It is preserved in its original, unmodified state specifically for research, reference, or historical recordkeeping. In accordance with the ADA Title II Final Rule, the University Libraries provides accessible versions of archival materials upon request. To request an accommodation for this item, please submit an accessibility request form.