AI-assisted video data management for compositional and high-level queries

dc.contributor.advisorBalazinska, Magdalena
dc.contributor.advisorKrishna, Ranjay
dc.contributor.authorZhang, Enhao
dc.date.accessioned2026-08-11T19:26:57Z
dc.date.issued2026-08-11
dc.date.submitted2026
dc.descriptionThesis (Ph.D.)--University of Washington, 2026
dc.description.abstractVideo data is increasingly used in a variety of domains including drone analytics, robotics, urban traffic management, civil engineering, medical education, and biological research. Users of these datasets often need to find complex events in the videos that involve multiple objects, attributes, relationships, and temporal patterns. The rapid advances in artificial intelligence have significantly enhanced video analytics capabilities, including the ability to extract semantic content from videos, making video data queryable in new and interesting ways. As a result, video data management systems have re-emerged as an active area of research. However, existing systems still face three major challenges: users must often possess database expertise to formulate compositional queries that search for complex events in the video data, suitable visual modules that extract application-specific concepts from videos may not exist, and users may want to directly name domain-specific high-level events whose semantics require inferential reasoning and are difficult to specify exhaustively as predicate compositions. Directly applying large vision-language models (VLMs) to every video can address some of these limitations, but doing so is prohibitively expensive at scale. In this dissertation, we present three systems that make compositional and high-level video queries easier to specify than in existing video data management systems and more efficient to execute than direct VLM-based evaluation. The first system, EQUI-VOCAL, synthesizes declarative video queries from a small number of labeled examples, eliminating the need for users to manually express complex spatio-temporal query logic. The second system, VOCAL-UDF, is a self-enhancing video data management system that automatically generates user-defined functions (UDFs) for new application-specific concepts as it processes queries. VOCAL-UDF uses large language models (LLMs) and VLMs to identify missing modules, implement them as UDFs, and select the most effective ones for query execution. The third system, VOCAL-SIFT, targets high-level queries that name events directly without specifying the full visual logic needed to recognize them. It decomposes these events into necessary predicates, instantiates lightweight filters for those predicates, and invokes expensive VLM reasoning only on candidate videos that pass the filters. Together, these systems show how AI models can be used not only as black-box video query evaluators, but also as collaborators in query synthesis, module generation, semantic decomposition, and query optimization. This dissertation develops an important suite of approaches that make video data management systems more practical, expressive, and scalable.
dc.embargo.termsOpen Access
dc.format.mimetypeapplication/pdf
dc.identifier.otherZhang_washington_0250E_29974.pdf
dc.identifier.urihttps://hdl.handle.net/1773/57258
dc.language.isoen_US
dc.rightsnone
dc.subjectComputer science
dc.subject.otherComputer science and engineering
dc.titleAI-assisted video data management for compositional and high-level queries
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Zhang_washington_0250E_29974.pdf
Size:
37.93 MB
Format:
Adobe Portable Document Format