Understanding What Deep Vision and Multimodal Models Depend on and Guiding What They Should

The Hong Kong University of Science and Technology
Department of Computer Science and Engineering


PhD Thesis Defence


Title: "Understanding What Deep Vision and Multimodal Models Depend on
and Guiding What They Should"

By

Mr. Weiyan XIE


Abstract:

Over the past decade, visual recognition has advanced from narrow, 
task-specific classifiers to large-scale multimodal large language models 
(MLLMs) supporting open-vocabulary visual understanding and rich multimodal 
reasoning. Yet across this spectrum, real-world reliability remains 
persistently undermined by misaligned visual dependencies: classifiers often 
exploit spurious background correlations rather than causally relevant 
features, and MLLMs frequently hallucinate by leaning on language priors or 
misleading cues in lieu of pertinent visual evidence. This thesis pursues two 
complementary objectives, understanding what deep visual classifiers and 
multimodal models currently depend on and proactively guiding what they 
should depend on, and advances them along three directions. To understand 
what classifiers depend on, ViT-CX estimates the causal effect of semantic 
patches on Vision Transformer predictions, while Contrastive Whole-Output 
Explanation (CWOX) surfaces the discriminative evidence that separates a 
classification model's top-K predictions. To guide classifiers toward 
generalizable causal features, Logit Attribution Matching (LAM) anchors 
decisions to domain-invariant features by aligning per-feature logit 
attributions across image pairs sharing identical core semantics, while Dual 
Risk Minimization (DRM) mitigates robustness degradation when fine-tuning 
vision-language models for visual classification by pairing empirical risk 
minimization with a worst-case proxy derived from LLM-generated core-feature 
descriptions of class labels. Extending both objectives to multimodal models 
for long-document understanding, InSight-doc replaces fixed-resolution, 
single-pass pipelines with an active multi-agent framework that iteratively 
acquires high-resolution crops on demand, guiding the model to depend on 
actively gathered visual evidence while rendering its dependencies fully 
inspectable to users. Collectively, this thesis aims to narrow the gap 
between raw predictive performance and trustworthy models that generalize 
safely and reliably in the real world.


Date:                   Tuesday, 11 August 2026

Time:                   2:00pm - 4:00pm

Venue:                  Room 5501
                        Lifts 25/26

Chairman:               Dr. Laurence L. DELINA (ENVR)

Committee Members:      Prof. Nevin ZHANG (Supervisor)
                        Prof. Fangzhen LIN
                        Dr. Dan XU
                        Dr. Wenhan LUO (AMC)
                        Prof. Antoni Bert CHAN (CityU)