Spurious Correlations from Supervised Models to Agentic Systems: Detection, Benchmarking, and Mitigation
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Machine learning systems routinely succeed for the wrong reasons, relying on patterns thatare predictive in the training data but incidental to the task, a spurious correlation that
holds until it breaks. This dissertation argues that such failures, though they take a different
form in each training paradigm, can be examined through a single counterfactual lens: keep
the task-inherent features fixed, vary a suspected feature, and ask whether the model’s
behavior changes. What changes across paradigms is not this logic but what it takes to
apply it: whether the suspect feature is known, whether examples with and without it can
be obtained, and at which level of the model’s behavior the shortcut must be measured. Three studies develop this view. Concept Correction (supervised learning) addresses thecase where a spurious attribute is known but unlabeled, inferring the missing group labels
from a small set of out-of-distribution concept images and recovering worst-group accuracy
without manual annotation. SpuriVerse (multimodal pretraining) addresses the case where
the spurious feature is not known in advance, using model-proposed, human-verified can-
didates to build counterfactual image groups; it contributes a benchmark of 124 naturally
occurring correlations on which state-of-the-art vision-language models score only 35%, and
shows that fine-tuning on diverse correlations transfers to unseen ones. Spurious Tool Use
(reinforcement learning) addresses the case where the shortcut hides in an intermediate ac-
tion rather than the final answer, showing that RL-trained agents can invoke tools because
of superficial prompt cues while still answering correctly, and that a dense per-decision
reward suppresses this behavior without sacrificing accuracy. The studies trace a trajectory of growing autonomy, from fixed-label classifiers to agentsthat decide when to act. As systems take more consequential actions, their shortcuts in-
creasingly surface not in what a model predicts but in what it does, in a place where
outcome-based evaluation cannot see them. The contribution of this dissertation is a way
of making that behavior visible: a counterfactual test adapted to the level at which each
system’s shortcuts form, and matched there by mitigation.
Description
Thesis (Ph.D.)--University of Washington, 2026
