AI Security Compliance and Testing Framework for Large Language Model Systems

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Large Language Models are being integrated into enterprise workflows at a pace that has outrun the security frameworks designed to protect them. Conventional compliance standards such as SOC 2 and ISO 27001 provide no coverage for LLM-specific vulnerabilities including prompt injection, sensitive information disclosure, and system prompt leakage. The OWASP LLM Top 10 defines the relevant threat taxonomy, but no unified automated pipeline exists to translate those controls into repeatable, evidence-based test cases. This research investigates whether automated, systematically constructed adversarial testing can reliably surface exploitable vulnerabilities across LLM applications of varying hardness, and whether evolutionary attack strategies can reach attack surfaces that static prompt libraries cannot anticipate. To address these questions, this research employs a Design Science Research methodology. A 279-prompt library was constructed through three-source triangulation, drawing from CTF competition wins, industry AI red-teaming competition data, and peer-reviewed literature, grounding every technique in documented real-world effectiveness. Three target configurations were designed with isolated independent variables to evaluate detection rates across a baseline system, a hardened system, and a RAG-augmented system. An evolutionary attack engine implementing the SPE-NL genetic algorithm was developed and evaluated across all three configurations. Judge reliability and inter-rater agreement were validated through independent assessment by two practicing cybersecurity professionals. Following this research design, AegisLLM was implemented as an automated security testing suite operationalizing six OWASP LLM Top 10 controls. Empirical evaluation demonstrates that targets resistant to the full static library fall to SPE-NL-evolved payloads within three to five generations, confirming that adaptive evolutionary testing reaches attack surfaces that curated static libraries cannot anticipate. The LLM-as-a-judge classification pipeline achieved 94.5% inter-rater agreement with zero crossover errors between SUCCESS and NO_SUCCESS labels. This thesis contributes to the field in three respects. First, it provides an empirically validated, open-source prompt library mapped explicitly to the OWASP LLM Top 10, sourced from ecologically valid real-world adversarial data. Second, it demonstrates the viability of evolutionary prompt mutation as a structured research method for LLM security evaluation, not merely as an engineering technique. Third, it establishes a replicable evaluation framework combining automated semantic judgment with human inter-rater validation, offering a methodological foundation for future LLM security research.

Description

Thesis (Master's)--University of Washington, 2026

Citation

DOI