GuardRAG: Adversarial Preference Training Against Indirect Prompt Injection Attacks

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Retrieval-Augmented Generation (RAG) systems have enhanced the performance of Large Language Models (LLMs) by effectively addressing challenges such as hallucinations and generating irrelevant responses. However, augmenting LLMs with retrieval introduces privacy and data-leakage risks through the retrieval pipeline. Indirect Prompt Injections (IPIs) occur when hidden instructions inside retrieved documents enter a model’s context as trusted evidence, exploiting weak guardrails and leading to unauthorized data leaks, policy overrides, or attacker-controlled behavior. Existing defenses rely on brittle delimiter heuristics or limited retriever adjustments, leaving RAG systems vulnerable to adversarial directives blended seamlessly into retrieved content. This thesis studies how IPIs propagate across the lifecycle of untrusted retrieved content, from retrieval-time exposure to model behavior and persistent agent memory. We first introduce RIPE-II, a corpus-level benchmark that evaluates IPIs under a realistic content-poisoning threat model. RIPE-II contains 32k attacks across six corpora, four domains, and twelve carrier families, and measures retrieval exposure and generation compromise as separate stages. Results show that poisoned passages reach the model on most queries, reranking can amplify poisoning instead of filtering it, and even the strongest evaluated models follow injected directives on roughly half of queries and up to 89\% on some corpora. Cosine-based scoring also underreports semantic compromise by more than 5x, motivating the need for calibrated judge-based evaluation. We then present GuardRAG, an adversarial training framework that converts these observed vulnerabilities into security-aware preference data. GuardRAG teaches the model to produce a span-grounded security report, making each refusal or acceptance decision auditable. On an 8B model that follows about 45\% of stealth injections even when the security policy is included in the prompt, GuardRAG reduces behavioral attack success to 0.3\% while preserving benign-query utility. The resulting 8B model also matches the defense quality of a 70B model nearly nine times its size on a single GPU. However, securing the immediate response does not prevent poisoned content from persisting in tool-using agents with long-term memory. We therefore introduce PRISM-Mem, a provenance-aware memory firewall that screens what may enter persistent memory, cutting agent attack success from 30.3\% to 9.0\% and cross-turn contamination from 98.7\% to zero. Together, RIPE-II, GuardRAG, and PRISM-Mem show that realistic evaluation, security-aware training, and provenance-governed memory can substantially improve RAG and agent security without sacrificing usefulness. This thesis establishes that a threat which grows with model capability can instead be held by a deliberately trained defense across the full retrieval lifecycle.

Description

Thesis (Master's)--University of Washington, 2026

Citation

DOI