Measurement Non-invariance in Complex Psychometric Systems: Structural Regularization-Based DIF Detection and Prompt Effect Modeling
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Educational and psychological assessments play a pivotal role in decision-making across many domains. To ensure their validity, it is essential to maintain measurement invariance across assessment contexts. Measurement invariance requires that individuals sharing the same level of the latent trait have the same probability of any given response, regardless of their demographic characteristics or other group memberships. Differential Item Functioning (DIF) is the item-level manifestation of a measurement invariance violation. When an item exhibits DIF, this principle is violated and observed score differences between groups can no longer be attributed solely to true differences in the underlying trait.Over decades, many models and procedures have been developed to identify items exhibiting DIF. However, existing methods are typically constrained to a small number of groups, often formed by a single demographic variable such as race or sex, and become impractical in more complex settings involving a large number of groups. Such settings arise naturally in large-scale assessments where the number of participating countries or states can easily exceed dozens, as well as in contexts where intersectionality across multiple demographic identities is of interest. These limitations motivate the need for more scalable and structurally informed approaches to DIF detection.
This dissertation advances measurement invariance research in complex settings through two complementary contributions. The first introduces structural regularization-based methods for scalable DIF detection across many groups. The second extends the measurement invariance framework to Large Language Model (LLM) evaluation, where prompt configurations function as the grouping variable and their effects on model performance are explicitly modeled.
The first contribution proposes a new DIF detection framework for multiple-group settings. In this context, DIF detection encompasses two related tasks: identifying items that exhibit DIF at the item level, and determining which specific subgroups are responsible for the detected discrepancy at the item-by-subgroup level. To address both tasks simultaneously, this dissertation introduces the Group Exponential Lasso (GEL), a group-sparse penalty that explicitly accounts for the nested structure inherent in multiple-group settings, allowing item-level and subgroup-level signals to inform each other during estimation. Simulation studies demonstrate that GEL results in better performance compared to classical penalties such as lasso, which treat item-by-group effects as independent and ignore the nested structure. In particular, GEL recovers small but informative subgroup-by-item DIF effects that classical penalties miss, a finding consistent with the theoretical properties of group-sparse regularization.
The second contribution extends measurement invariance analysis to the emerging domain of LLM evaluation. Recent studies have shown that LLM performance varies systematically with prompt configurations, such as the ordering or formatting of response options. These variations closely resemble DIF phenomena. By conceptualizing prompt configurations as groups, the proposed framework enables principled evaluation of prompt sensitivity and supports more stable estimation of latent LLM capabilities. However, this extension is not straightforward. LLM evaluation introduces two challenges absent from traditional DIF settings: (1) the number of prompt-defined groups can reach into the thousands, demanding scalable methods; and (2) unlike traditional grouping variables, prompt configurations simultaneously alter both task difficulty and model capability. To address these challenges, this dissertation introduces PrismEval, a psychometric framework built on a random item-effect model that decomposes performance variation across prompt groups into the main effects of each prompt variable and residual random item effects capturing higher-order interactions. The framework is estimated via a regularized Gaussian Variational Expectation-Maximization (GVEM) algorithm. Applied to the responses of five open-weight LLMs across 6,005 prompt configurations in the DOVE dataset, PrismEval reveals that model scores are strongly prompt-dependent, that prompt sensitivity varies systematically across models, and that a small set of tasks exhibits idiosyncratic difficulty shifts attributable to task-level ambiguity. Simulations further demonstrate that the regularized GVEM estimator accurately recovers model parameters given a sufficient number of prompt configurations.
Together, these two contributions establish a unified psychometric framework for analyzing measurement non-invariance in complex assessment settings. By advancing DIF methodology for human assessment and extending it to LLM evaluation, this dissertation demonstrates that psychometric principles can be both deepened and broadened to meet the demands of modern measurement challenges.
Description
Thesis (Ph.D.)--University of Washington, 2026
