Keywords

K-Nearest Neighbors, Logistic Regression, Area Under the Receiving Operator Characteristic

Subject Categories

Applied Mathematics | Computer Sciences | Data Science

Abstract

This thesis presents an empirical comparison of two classification methods: Logistic Regression and K Nearest Neighbors (KNN). The primary objective of this research is to evaluate the strengths and limitations of each method when applied to real-world datasets. Several publicly available datasets on diabetes, breast cancer, heart attack risk, and cardiovascular disease, were analyzed. For each dataset, K Nearest Neighbors models were implemented in the same way logistic regression had already been applied. The results demonstrate that while logistic regression offers interpretable parameter estimates and performs well when the underlying predictor and outcome relationship is approximately linear, however KNN can achieve competitive or even improved performance in situations where the decision boundary is nonlinear and more complex. Overall, this thesis showcases the complementary strengths of both classification models and illustrates how KNN can be a valuable alternative or supplementary model in classification problems.

Completion Date

2026

Semester

Summer

Committee Chair

Uddin, Nizam

Degree

Master of Science (M.S.)

College

College of Sciences

Department

School of Data, Mathematical, and Statistical Sciences

Format

PDF

Document Type

Thesis

Language

English

Share

COinS
 

Accessibility Statement

This item was created or digitized prior to April 24, 2027, or is a reproduction of legacy media created before that date. It is preserved in its original, unmodified state specifically for research, reference, or historical recordkeeping. In accordance with the ADA Title II Final Rule, the University Libraries provides accessible versions of archival materials upon request. To request an accommodation for this item, please submit an accessibility request form.