Towards Robust Multimodal AI for Precision Oncology: From Heterogeneous Data to Generalizable Representations

PhD Thesis Proposal Defence


Title: "Towards Robust Multimodal AI for Precision Oncology: From 
Heterogeneous Data to Generalizable Representations"

by

Miss Yingxue XU


Abstract:

In precision oncology, pathology enables the characterization of tumor 
morphology and the spatial organization of the tumor microenvironment, 
serving as a central anchor for connecting these tissue-level features with 
molecular, textual, and clinical information. Integrating these tissue-level 
features with broader multimodal biomedical data can provide a more 
comprehensive basis for individualized cancer diagnosis, treatment 
decision-making, and prognosis. However, translating multimodal biomedical 
data into robust oncology AI requires addressing a connected progression of 
challenges spanning alignment, representation learning, data availability, 
generalization, and clinical translation. A first challenge is how to align 
heterogeneous modalities with distinct inherent structures. After alignment, 
multimodal representations may still contain substantial redundant or 
non-discriminative information, requiring models to preserve clinically 
relevant signals while suppressing nuisance variation. The subsequent 
challenge is practical data availability, as one or more modalities may be 
unavailable in real settings. To move beyond task-specific prediction, these 
robustness principles must be scaled into generalizable whole-slide 
representations and validated within real diagnostic workflows.

This thesis investigates robust multimodal learning methods that 
progressively transform heterogeneous oncology data into clinically 
meaningful pathology representations for diverse clinical decision-making 
tasks in precision oncology. The first part develops task-specific robust 
multimodal learning strategies that address cross-modal structural 
heterogeneity, information redundancy, and incomplete modality availability 
in pathology-genomics learning, using survival prediction as a clinically 
relevant testbed. Building on these robustness principles, the second part 
moves from task-specific prediction to foundation modeling by learning 
generalizable representations through multimodal pretraining. Finally, the 
thesis closes this progression by evaluating the clinical utility of these 
representations through workflow-aligned validation in breast cancer.

Taken as a whole, the thesis demonstrates that robust multimodal AI for 
precision oncology requires a progression from structure-aware alignment, to 
compact and missing-modality-tolerant representation learning, to foundation 
models that encode clinically meaningful whole-slide multimodal context, and 
finally to workflow-aware evaluation in real clinical settings.


Date:                   Tuesday, 11 August 2026

Time:                   10:30am - 12:00noon

Venue:                  Room 3494
                        Lift 25/26

Committee Members:      Dr. Hao Chen (Supervisor)
                        Dr. Dan Xu (Chairperson)
                        Dr. Terence Tsz Wai Wong (CBE)