More about HKUST
Understanding What Deep Vision and Multimodal Models Depend on and Guiding What They Should
The Hong Kong University of Science and Technology
Department of Computer Science and Engineering
PhD Thesis Defence
Title: "Understanding What Deep Vision and Multimodal Models Depend on
and Guiding What They Should"
By
Mr. Weiyan XIE
Abstract:
Over the past decade, visual recognition has advanced from narrow,
task-specific classifiers to large-scale multimodal large language models
(MLLMs) supporting open-vocabulary visual understanding and rich multimodal
reasoning. Yet across this spectrum, real-world reliability remains
persistently undermined by misaligned visual dependencies: classifiers often
exploit spurious background correlations rather than causally relevant
features, and MLLMs frequently hallucinate by leaning on language priors or
misleading cues in lieu of pertinent visual evidence. This thesis pursues two
complementary objectives, understanding what deep visual classifiers and
multimodal models currently depend on and proactively guiding what they
should depend on, and advances them along three directions. To understand
what classifiers depend on, ViT-CX estimates the causal effect of semantic
patches on Vision Transformer predictions, while Contrastive Whole-Output
Explanation (CWOX) surfaces the discriminative evidence that separates a
classification model's top-K predictions. To guide classifiers toward
generalizable causal features, Logit Attribution Matching (LAM) anchors
decisions to domain-invariant features by aligning per-feature logit
attributions across image pairs sharing identical core semantics, while Dual
Risk Minimization (DRM) mitigates robustness degradation when fine-tuning
vision-language models for visual classification by pairing empirical risk
minimization with a worst-case proxy derived from LLM-generated core-feature
descriptions of class labels. Extending both objectives to multimodal models
for long-document understanding, InSight-doc replaces fixed-resolution,
single-pass pipelines with an active multi-agent framework that iteratively
acquires high-resolution crops on demand, guiding the model to depend on
actively gathered visual evidence while rendering its dependencies fully
inspectable to users. Collectively, this thesis aims to narrow the gap
between raw predictive performance and trustworthy models that generalize
safely and reliably in the real world.
Date: Tuesday, 11 August 2026
Time: 2:00pm - 4:00pm
Venue: Room 5501
Lifts 25/26
Chairman: Dr. Laurence L. DELINA (ENVR)
Committee Members: Prof. Nevin ZHANG (Supervisor)
Prof. Fangzhen LIN
Dr. Dan XU
Dr. Wenhan LUO (AMC)
Prof. Antoni Bert CHAN (CityU)