Information-Seeking Agents in the Wild: Approaches to Evaluation in Dynamic and Open Environments
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Information-Seeking Agents (ISAs) stand to reshape longstanding information access practices, ranging from question answering and fact verification to search, synthesis, and decision support. Traditional information access systems have typically left users responsible for formulating queries, inspecting sources, revising assumptions, and assembling evidence into usable knowledge. ISAs shift more of this work to agentic systems designed to fulfill users' informational needs through retrieval, reasoning, summarization, and tool use. Yet the evaluation methods used to measure these systems have not kept pace with this shift. Static QA benchmarks, fixed test items, and task-specific metrics can no longer fully capture systems that operate over changing information environments, use external tools at inference time, and combine parametric memory with retrieved evidence. Without new evaluation methods, measured performance may obscure the difference between reasoning and recall, evidence access and evidence use, or genuine information seeking and retrieval of leaked benchmark artifacts. This dissertation develops approaches for evaluating ISAs in dynamic and open information environments. Concretely, it makes three contributions to the methodological foundations of ISA evaluation. First, it characterizes benchmark leakage as a multi-channel validity threat, showing that static QA evaluation can be compromised both by pretraining contamination and by retrieval-time access to benchmark artifacts. Second, it introduces dynamic benchmark construction methods that generate grounded question-answer instances from explicit evidence structures, reducing memorization risk while preserving comparability across evaluation runs. In structured knowledge settings, this is achieved through regeneration from semantically coherent knowledge-graph subgraphs; in open-web settings, it is achieved through traffic-driven topics, query-conditioned evidence, and auditable intermediate artifacts. Third, it develops an evaluation framework for cross-source sensemaking, where success requires not only retrieving relevant evidence but integrating information across sources, themes, and relationships. Taken together, these contributions argue for a shift from static answer matching toward contamination-aware, evidence-grounded, and dynamically adaptable evaluation. They point toward a future in which ISAs are measured not only by whether they produce correct answers, but by whether they reliably use the right evidence, under the changing conditions in which real information seeking takes place.
Description
Thesis (Ph.D.)--University of Washington, 2026
