AI-assisted video data management for compositional and high-level queries
Date
2026-08-11
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Video data is increasingly used in a variety of domains including drone analytics, robotics, urban traffic management, civil engineering, medical education, and biological research. Users of these datasets often need to find complex events in the videos that involve multiple objects, attributes, relationships, and temporal patterns. The rapid advances in artificial intelligence have significantly enhanced video analytics capabilities, including the ability to extract semantic content from videos, making video data queryable in new and interesting ways. As a result, video data management systems have re-emerged as an active area of research. However, existing systems still face three major challenges: users must often possess database expertise to formulate compositional queries that search for complex events in the video data, suitable visual modules that extract application-specific concepts from videos may not exist, and users may want to directly name domain-specific high-level events whose semantics require inferential reasoning and are difficult to specify exhaustively as predicate compositions. Directly applying large vision-language models (VLMs) to every video can address some of these limitations, but doing so is prohibitively expensive at scale. In this dissertation, we present three systems that make compositional and high-level video queries easier to specify than in existing video data management systems and more efficient to execute than direct VLM-based evaluation. The first system, EQUI-VOCAL, synthesizes declarative video queries from a small number of labeled examples, eliminating the need for users to manually express complex spatio-temporal query logic. The second system, VOCAL-UDF, is a self-enhancing video data management system that automatically generates user-defined functions (UDFs) for new application-specific concepts as it processes queries. VOCAL-UDF uses large language models (LLMs) and VLMs to identify missing modules, implement them as UDFs, and select the most effective ones for query execution. The third system, VOCAL-SIFT, targets high-level queries that name events directly without specifying the full visual logic needed to recognize them. It decomposes these events into necessary predicates, instantiates lightweight filters for those predicates, and invokes expensive VLM reasoning only on candidate videos that pass the filters. Together, these systems show how AI models can be used not only as black-box video query evaluators, but also as collaborators in query synthesis, module generation, semantic decomposition, and query optimization. This dissertation develops an important suite of approaches that make video data management systems more practical, expressive, and scalable.
Description
Thesis (Ph.D.)--University of Washington, 2026
Keywords
Computer science
