Do Medical Language Models Preserve Correct Judgment Under Misleading Context? Construction and Evaluation of an Injection Benchmark
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Clinical and consumer-facing deployments of large language models (LLMs) can now achieve expert-level scores on medical question-answering benchmarks based on licensing exams. However, real-world questions from users rarely resemble the clean, single-turn exam questions in these benchmarks. In reality, interactions are often layered with patient-reported history, retrieved documents, and false claims. This thesis asks whether an LLM that answers a medical question correctly in a clean environment continues to be correct once such misleading context is added. In short, it often does not. To study this question at scale, this thesis presents MedMisBench, a benchmark of 10,932 medical question items paired with 48,889 misleading context option sentences spanning medical-reasoning, agentic-capability, and patient-journey question formats. Each injected sentence is constructed along two independent axes that separately capture the type of medical false claim and its apparent source. This way, failures can be attributed to what is false and who asserts it. Evaluated across 11 configurations of commercial, open-source, and medical-domain LLMs, mean answer accuracy falls from 71.1% on clean questions to 38.0% once an injection sentence is prefixed to one answer option, with a mean attack success rate of 51.5%. Fabricated treatment exceptions and claims attributed to clinical authorities are consistently the most effective types in lowering the accuracy of LLMs, with attack success rates as high as 64.1% and 69.5% respectively. A panel of 14 clinical reviewers spanning 7 countries judged that 38.2% of these injection-induced errors carried serious potential harm to a patient. These findings indicate that strong performance on existing medical benchmarks does not mean that a model will preserve sound clinical judgment once its context is contaminated by misleading information. By releasing MedMisBench as a static, reusable benchmark, this work provides a foundation for evaluating and eventually mitigating this gap between benchmark performance and real-world robustness.
Description
Thesis (Master's)--University of Washington, 2026
