ORCID

0009-0008-8815-4982

Keywords

Privacy, Taxonomy

Subject Categories

Computer Sciences | Information Security | Theory and Algorithms

Abstract

Software applications increasingly rely on user data to provide their functionality, but improper handling of such data can lead to serious privacy noncompliance with applicable regulations and policies. A prominent example is the Facebook–Cambridge Analytica scandal, in which a third-party application collected the personal data of approximately 87 million Facebook users without users' consent. Despite growing attention to privacy compliance, two key challenges hinder the systematic understanding and analysis of privacy noncompliance. First, unlike security vulnerabilities, which have been systematically categorized through taxonomies such as the Common Weakness Enumeration (CWE), privacy noncompliance lacks a technical taxonomy describing how it manifests in real-world software. This gap makes it difficult for researchers, developers, auditors, and regulators to consistently identify, reason about, and assess privacy noncompliance. Second, unlike the availability of well-established benchmark suites for security vulnerabilities, such as the NIST SAMATE Juliet Test Suite, there is no comprehensive benchmark of software samples exhibiting diverse types of privacy noncompliance. The absence of such benchmarks limits the development and evaluation of automated techniques for detecting privacy noncompliance. This thesis presents an initial effort to address these challenges through two exploratory studies. First, we conduct a systematic literature review of 52 papers documenting privacy noncompliance across mobile applications, websites, and IoT devices. Through detailed analysis and manual annotation, we develop a taxonomy of privacy noncompliance organized into six high-level categories derived from established privacy principles. Second, leveraging the technical descriptions captured by the taxonomy, we investigate the feasibility of automatically generating software samples exhibiting privacy noncompliance using large language model (LLM)-based code generation. Our evaluation demonstrates that combining the proposed taxonomy with LLM-based code generation enables the automatic construction of a benchmark containing diverse privacy noncompliance samples, providing a foundation iv for future research on understanding, detecting, and evaluating privacy noncompliance in software systems.

Completion Date

2026

Semester

Summer

Committee Chair

Wang, Xueqiang

Degree

Master of Science in Computer Engineering (M.S.Cp.E.)

College

College of Engineering and Computer Science

Department

Computer Science

Format

PDF

Document Type

Thesis

Language

English

Share

COinS
 

Accessibility Statement

This item was created or digitized prior to April 24, 2027, or is a reproduction of legacy media created before that date. It is preserved in its original, unmodified state specifically for research, reference, or historical recordkeeping. In accordance with the ADA Title II Final Rule, the University Libraries provides accessible versions of archival materials upon request. To request an accommodation for this item, please submit an accessibility request form.