Do Medical Language Models Preserve Correct Judgment Under Misleading Context? Construction and Evaluation of an Injection Benchmark

dc.contributor.advisorShapiro, Linda
dc.contributor.authorZou, Xin Yu
dc.date.accessioned2026-09-16T18:26:05Z
dc.date.issued2026-09-16
dc.date.submitted2026
dc.descriptionThesis (Master's)--University of Washington, 2026
dc.description.abstractClinical and consumer-facing deployments of large language models (LLMs) can now achieve expert-level scores on medical question-answering benchmarks based on licensing exams. However, real-world questions from users rarely resemble the clean, single-turn exam questions in these benchmarks. In reality, interactions are often layered with patient-reported history, retrieved documents, and false claims. This thesis asks whether an LLM that answers a medical question correctly in a clean environment continues to be correct once such misleading context is added. In short, it often does not. To study this question at scale, this thesis presents MedMisBench, a benchmark of 10,932 medical question items paired with 48,889 misleading context option sentences spanning medical-reasoning, agentic-capability, and patient-journey question formats. Each injected sentence is constructed along two independent axes that separately capture the type of medical false claim and its apparent source. This way, failures can be attributed to what is false and who asserts it. Evaluated across 11 configurations of commercial, open-source, and medical-domain LLMs, mean answer accuracy falls from 71.1% on clean questions to 38.0% once an injection sentence is prefixed to one answer option, with a mean attack success rate of 51.5%. Fabricated treatment exceptions and claims attributed to clinical authorities are consistently the most effective types in lowering the accuracy of LLMs, with attack success rates as high as 64.1% and 69.5% respectively. A panel of 14 clinical reviewers spanning 7 countries judged that 38.2% of these injection-induced errors carried serious potential harm to a patient. These findings indicate that strong performance on existing medical benchmarks does not mean that a model will preserve sound clinical judgment once its context is contaminated by misleading information. By releasing MedMisBench as a static, reusable benchmark, this work provides a foundation for evaluating and eventually mitigating this gap between benchmark performance and real-world robustness.
dc.embargo.lift2031-08-21T18:26:05Z
dc.embargo.termsRestrict to UW for 5 years -- then make Open Access
dc.format.mimetypeapplication/pdf
dc.identifier.otherZou_washington_0250O_30313.pdf
dc.identifier.urihttps://hdl.handle.net/1773/57771
dc.language.isoen_US
dc.rightsCC BY-NC-ND
dc.subjectComputer science
dc.subject.otherElectrical and computer engineering
dc.titleDo Medical Language Models Preserve Correct Judgment Under Misleading Context? Construction and Evaluation of an Injection Benchmark
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Zou_washington_0250O_30313.pdf
Size:
1.79 MB
Format:
Adobe Portable Document Format