SCALES: Dual Information-Theoretic Approaches to Mitigating and Quantifying Prompt Injections in RAG Pipelines

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Large language models deployed within Retrieval-Augmented Generation (RAG) systems are vulnerable to indirect prompt injection, where adversarial instructions embedded in retrieved documents cause the model to override its intended behaviour, exfiltrate confidential context, or execute unauthorised actions. This thesis makes two complementary contributions to securing RAG pipelines. The first is SCALES (Semantic Cascade Architecture for Leakage and Exploitation Security), a modality-aware three-layer defense cascade combining a DeBERTa-based pre-retrieval classifier, a KL-divergence semantic boundary chunker, and an IT-MOC LoRA adapter regularised with Jensen-Shannon Divergence. Evaluated on a 425-sample benchmark spanning explicit text injections, LLM-generated semantic attacks, gradient optimised adaptive attacks, and cross-modal decomposition attacks, SCALES achieves a judge-verified Attack Success Rate of 6.5\% and a Benign Pass Rate of 95.33\%. The second contribution is an information-theoretic evaluation framework that quantifies data leakage severity beyond binary attack success, combining four complementary metrics, lexical, semantic, algorithmic, and distributional similarity, into a Combined Leakage Index validated against LLM-judge harm scores. Per-source decomposition reveals that SCALES suppresses leakage by 99\% on adaptive attacks focused on data leakage demonstrating a level of diagnostic granularity unavailable to single-metric evaluation.

Description

Thesis (Master's)--University of Washington, 2026

Citation

DOI