More about HKUST
Towards Robust Multimodal AI for Precision Oncology: From Heterogeneous Data to Generalizable Representations
PhD Thesis Proposal Defence
Title: "Towards Robust Multimodal AI for Precision Oncology: From
Heterogeneous Data to Generalizable Representations"
by
Miss Yingxue XU
Abstract:
In precision oncology, pathology enables the characterization of tumor
morphology and the spatial organization of the tumor microenvironment,
serving as a central anchor for connecting these tissue-level features with
molecular, textual, and clinical information. Integrating these tissue-level
features with broader multimodal biomedical data can provide a more
comprehensive basis for individualized cancer diagnosis, treatment
decision-making, and prognosis. However, translating multimodal biomedical
data into robust oncology AI requires addressing a connected progression of
challenges spanning alignment, representation learning, data availability,
generalization, and clinical translation. A first challenge is how to align
heterogeneous modalities with distinct inherent structures. After alignment,
multimodal representations may still contain substantial redundant or
non-discriminative information, requiring models to preserve clinically
relevant signals while suppressing nuisance variation. The subsequent
challenge is practical data availability, as one or more modalities may be
unavailable in real settings. To move beyond task-specific prediction, these
robustness principles must be scaled into generalizable whole-slide
representations and validated within real diagnostic workflows.
This thesis investigates robust multimodal learning methods that
progressively transform heterogeneous oncology data into clinically
meaningful pathology representations for diverse clinical decision-making
tasks in precision oncology. The first part develops task-specific robust
multimodal learning strategies that address cross-modal structural
heterogeneity, information redundancy, and incomplete modality availability
in pathology-genomics learning, using survival prediction as a clinically
relevant testbed. Building on these robustness principles, the second part
moves from task-specific prediction to foundation modeling by learning
generalizable representations through multimodal pretraining. Finally, the
thesis closes this progression by evaluating the clinical utility of these
representations through workflow-aligned validation in breast cancer.
Taken as a whole, the thesis demonstrates that robust multimodal AI for
precision oncology requires a progression from structure-aware alignment, to
compact and missing-modality-tolerant representation learning, to foundation
models that encode clinically meaningful whole-slide multimodal context, and
finally to workflow-aware evaluation in real clinical settings.
Date: Tuesday, 11 August 2026
Time: 10:30am - 12:00noon
Venue: Room 3494
Lift 25/26
Committee Members: Dr. Hao Chen (Supervisor)
Dr. Dan Xu (Chairperson)
Dr. Terence Tsz Wai Wong (CBE)