Optimized Multi-Agent Defense Pipeline Against Advanced Prompt Injection Attacks

dc.contributor.advisorThamilarasu, Geetha G.T
dc.contributor.authorPearson, Ben
dc.date.accessioned2026-09-16T18:24:54Z
dc.date.issued2026-09-16
dc.date.submitted2026
dc.descriptionThesis (Master's)--University of Washington, 2026
dc.description.abstractLarge language models deployed in production remain critically vulnerable to prompt-injection attacks. These attacks embed adversarial inputs that hijack the model into leaking data, taking unauthorized actions, or producing harmful content. Existing multi-agent defenses handle static, known attacks well but leave three gaps. First, they are not evaluated againstadaptive attackers who optimize against the defense. Second, they do not handle indirect injection through external content or multi-turn distributed attacks. Third, they do not measure how the security layer behaves when the inference cache fills up. I designed a multi-agent defense pipeline with three architectural contributions. The first is a white-box adaptive adversary. It attacks all defense components at once through a joint-loss optimization. I evaluated it across five configurations and 24 generations of iterative training. The second is a trust-role-aware cache compression scheme. It partitions cached tokens by trust role. It prevents the loss of security-relevant instructions when the cache fills up. The third is a multi-turn defense that tracks conversation history across turns. It aggregates four detector signals into a per-session trust score, resolved through a learned discriminative classifier. The evaluation produced two headline findings. On an aligned Domain LLM, no configuration produced a verified attack. This held across four evaluation configurations and 24 generations of iterative training. On a vulnerable Domain LLM, the attack succeeded three-of-three times without the pipeline. With the pipeline in place, the attack failed five-of-five times end to end. The multi-turn defense caught every attack pattern on real-user dialogues at near-zero false-positive rate. An early design coupled the false-positive rate and the attack-catch rate through a single parameter. A learned discriminative classifier decoupled them. The cache-layer scheme preserved all protected token classes at full recall. A naive eviction baseline lost them entirely. The perimeter guards reduced attack-success rate by roughly 11×on one benchmark and 70×on another against an undefended baseline. These results reframe what a multi-agent defense is for. On aligned Domain LLMs, the pipeline earns its value through orthogonal coverage of threat surfaces the alignment training does not address. On vulnerable Domain LLMs, the same pipeline becomes the binding direct-injection defense layer. The two roles are complementary rather than competing.
dc.embargo.termsOpen Access
dc.format.mimetypeapplication/pdf
dc.identifier.otherPearson_washington_0250O_30207.pdf
dc.identifier.urihttps://hdl.handle.net/1773/57749
dc.language.isoen_US
dc.rightsnone
dc.subjectLLMs Security
dc.subjectPrompt Injection Attacks
dc.subjectComputer engineering
dc.subjectComputer science
dc.subject.otherComputer science and engineering
dc.titleOptimized Multi-Agent Defense Pipeline Against Advanced Prompt Injection Attacks
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Pearson_washington_0250O_30207.pdf
Size:
2.35 MB
Format:
Adobe Portable Document Format