ORCID
0009-0002-6310-2787
Keywords
Hardware/software co-design; domain-specific accelerators; memory-efficient computing; data movement; vision transformers; 3D human mesh recovery; pangenome mapping; graph processing; register-transfer-level optimization; design-space exploration
Subject Categories
Computer and Systems Architecture | Computer Engineering | Hardware Systems
Abstract
Emerging artificial-intelligence and data-intensive scientific workloads increasingly face a memory wall: irregular access patterns and large intermediate data volumes make data movement, rather than arithmetic, the primary constraint on performance and energy efficiency. This dissertation develops a memory-centric hardware/software co-design methodology that jointly reshapes algorithms, architectures, and dataflows to retain frequently reused data on chip. The methodology is demonstrated through three accelerators and an RTL design tool. VITA replaces multi-head attention in vision-transformer-based 3D human mesh recovery with hardware-friendly average pooling and maps the resulting operators to a reconfigurable datapath, achieving 5.05-fold and 69.12-fold speedups over a state-of-the-art GPU and CPU while maintaining competitive accuracy. PanGaea reformulates sequence-to-graph pangenome mapping as a fused, tiled pipeline supported by a tri-mode processing array, improving throughput, energy efficiency, and area efficiency by 1.47-fold, 1.73-fold, and 1.62-fold, respectively, over the strongest prior ASIC. Locus extends this design with bubble-preserving graph partitioning, cross-read subgraph reuse, and variation-aware dataflow; it eliminates 96 percent of DRAM traffic, improves throughput by 2.1-fold, and preserves mapping accuracy within 0.02 percent of the software reference. Finally, KAIROS accelerates RTL timing exploration through a word-level intermediate representation, taxonomy-aware delay models, fan-out management, and retiming. Across 62 designs, it achieves a Spearman rank correlation of 0.96 with post-synthesis frequency and evaluates timing up to 213-fold faster than synthesis. Together, these contributions establish memory-efficient co-design as a repeatable methodology for building emerging-domain accelerators and accelerating their design.
Completion Date
2026
Semester
Summer
Committee Chair
Hao Zheng
Degree
Doctor of Philosophy (Ph.D.)
College
College of Engineering and Computer Science
Department
Department of Electrical and Computer Engineering
Format
Document Type
Dissertation
Language
English
STARS Citation
Tian, Shilin, "Memory-Efficient Acceleration for Emerging Applications via Hardware/Software Co-Design" (2026). Graduate Studies Theses and Dissertations 2026. 368.
https://stars.library.ucf.edu/gradstudies_etd_2026/368
Accessibility Statement
This item was created or digitized prior to April 24, 2027, or is a reproduction of legacy media created before that date. It is preserved in its original, unmodified state specifically for research, reference, or historical recordkeeping. In accordance with the ADA Title II Final Rule, the University Libraries provides accessible versions of archival materials upon request. To request an accommodation for this item, please submit an accessibility request form.