Spurious Correlations from Supervised Models to Agentic Systems: Detection, Benchmarking, and Mitigation

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Machine learning systems routinely succeed for the wrong reasons, relying on patterns thatare predictive in the training data but incidental to the task, a spurious correlation that holds until it breaks. This dissertation argues that such failures, though they take a different form in each training paradigm, can be examined through a single counterfactual lens: keep the task-inherent features fixed, vary a suspected feature, and ask whether the model’s behavior changes. What changes across paradigms is not this logic but what it takes to apply it: whether the suspect feature is known, whether examples with and without it can be obtained, and at which level of the model’s behavior the shortcut must be measured. Three studies develop this view. Concept Correction (supervised learning) addresses thecase where a spurious attribute is known but unlabeled, inferring the missing group labels from a small set of out-of-distribution concept images and recovering worst-group accuracy without manual annotation. SpuriVerse (multimodal pretraining) addresses the case where the spurious feature is not known in advance, using model-proposed, human-verified can- didates to build counterfactual image groups; it contributes a benchmark of 124 naturally occurring correlations on which state-of-the-art vision-language models score only 35%, and shows that fine-tuning on diverse correlations transfers to unseen ones. Spurious Tool Use (reinforcement learning) addresses the case where the shortcut hides in an intermediate ac- tion rather than the final answer, showing that RL-trained agents can invoke tools because of superficial prompt cues while still answering correctly, and that a dense per-decision reward suppresses this behavior without sacrificing accuracy. The studies trace a trajectory of growing autonomy, from fixed-label classifiers to agentsthat decide when to act. As systems take more consequential actions, their shortcuts in- creasingly surface not in what a model predicts but in what it does, in a place where outcome-based evaluation cannot see them. The contribution of this dissertation is a way of making that behavior visible: a counterfactual test adapted to the level at which each system’s shortcuts form, and matched there by mitigation.

Description

Thesis (Ph.D.)--University of Washington, 2026

Citation

DOI