Crossing the Chasm: Bridging Perception and Generative Models for Enhanced Vision-Language Understanding

dc.contributor.advisorHwang, Jenq-Neng
dc.contributor.authorJin, Ying
dc.date.accessioned2026-08-11T19:28:41Z
dc.date.issued2026-08-11
dc.date.submitted2026
dc.descriptionThesis (Ph.D.)--University of Washington, 2026
dc.description.abstractPerception and generative models are usually treated as opposing paradigms in machine learning. Perception models, such as classification, segmentation, and detection models, extract structured labels that integrate easily into downstream software, whereas generative models, such as diffusion models and large language models, synthesize data and offer broader semantic coverage through large-scale, weakly-supervised training. Yet this opposition is mainly one of task formulation. Both learn the mapping between images and labels but differ in its direction. In this dissertation, we show that the two paradigms can be complementary rather than disconnected, and bridging them improves vision-language understanding in both directions: generative models provide richer supervision for perception, while perception models bring verifiability and controllability to generation. To examine the generality of this insight, we study both general and medical domains. For Human-Object Interaction (HOI) recognition, we propose a Heterogeneous Teacher-Student (HTS) framework in which a generative image captioner serves as the teacher and a contrastive-initialized classifier serves as the student. Unlike conventional knowledge distillation, the teacher and student perform different tasks and have complementary strengths. By distilling HOI knowledge from noisy image captions, HTS removes the need for expensive human annotations. Trained with the proposed LogSumExp-Sign loss, HTS achieves 49.6 mAP on HICO without using ground-truth labels, outperforming fully supervised baselines despite being only a fraction of the teacher's size. In addition, it enables a zero-shot HOI detector, derived by pairing with an off-the-shelf object detector without further training, that surpasses prior methods. In the medical domain, we propose DAug for medical image classification and retrieval. Medical image understanding requires meticulous examination of fine, localized abnormalities, and it is challenging for models to learn where to look from limited training data with only image-level supervision, because such supervision does not indicate where abnormalities appear. In DAug, we train a diffusion-based image-to-image translation model to generate abnormality heatmaps, which are appended as additional image channels to direct the model's attention to clinically significant regions. Combined with a novel Image-Text-Class Hybrid Contrastive loss that leverages both text and class labels, DAug achieves state-of-the-art performance on both medical image classification and retrieval benchmarks. To improve radiology report generation, we introduce QRad, which reframes report generation as a self-directed Visual Question Answering process (Auto-VQA). Instead of using an image captioning pipeline, we train the model to first generate a chain of questions and then answer each question, concatenating the answers to form the report. The QA training data is converted from the reports. By isolating sentence-level linguistic variation (such as the omission or ordering of medical topics) from each diagnostic statement, this reformulation lets the model focus on factual accuracy rather than presentation style. As a result, QRad outperforms existing baselines while using only 13% of their parameters. Moreover, by casting classification as a single-token question-answering problem for a target finding, QRad equips a generative model with perception-style capabilities, including class-specific confidence estimation and ROC-based evaluation. Building on this property, we further propose a framework that combines an image classifier with a report generation model to enable threshold-controllable report generation, supporting regulatory validation and allowing clinicians to tune sensitivity-specificity trade-offs for different clinical settings. Taken together, these results establish bridging perception and generation as a general design principle that unites the open-vocabulary generalizability of generative models with the verifiability and controllability of perception models, yielding vision-language models that are at once more capable and more reliable.
dc.embargo.termsOpen Access
dc.format.mimetypeapplication/pdf
dc.identifier.otherJin_washington_0250E_29809.pdf
dc.identifier.urihttps://hdl.handle.net/1773/57321
dc.language.isoen_US
dc.rightsCC BY-ND
dc.subjectGenerative Models
dc.subjectHuman Action Recognition
dc.subjectMedical Imaging
dc.subjectPerception Models
dc.subjectRadiology Report Generation
dc.subjectVision-Language Models
dc.subjectArtificial intelligence
dc.subjectComputer science
dc.subjectMedical imaging
dc.subject.otherElectrical and computer engineering
dc.titleCrossing the Chasm: Bridging Perception and Generative Models for Enhanced Vision-Language Understanding
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Jin_washington_0250E_29809.pdf
Size:
25.03 MB
Format:
Adobe Portable Document Format