Multimodal Graph Intelligence: Reasoning, Prediction, and Generation

The Hong Kong University of Science and Technology
Department of Computer Science and Engineering


PhD Thesis Defence

Title: "Multimodal Graph Intelligence: Reasoning, Prediction, and
Generation"

By

Mr. Yanbin WEI


Abstract:

Graph-structured data plays a fundamental role in modern AI systems, yet 
existing graph methods still face key limitations in flexibility, structural 
expressiveness, and interpretability across heterogeneous tasks. This thesis 
proposal studies a unified direction, termed multimodal-enhanced graph 
intelligence, which integrates visual structural awareness, language 
reasoning, and external knowledge signals to improve graph reasoning, 
prediction, retrieval-augmented generation, and knowledge graph completion. 
The first part presents GITA, a Graph-to-Visual-and-Textual Integration 
framework for instruction-based graph reasoning. GITA converts structural 
graphs into coordinated visual and textual representations, enabling 
vision-language models to perform graph reasoning in a unified and 
user-friendly paradigm. To support systematic evaluation, we introduce GVLQA, 
a large-scale vision-language benchmark for general graph reasoning. The 
second part presents DynamicGTR, a dynamic graph topology representation 
routing framework for zero-shot graph QA. DynamicGTR introduces a Graph 
Response Efficiency objective and a lightweight router that adaptively 
selects topology representation forms for each query, improving both 
correctness and response efficiency. The third part presents GVN and its 
efficient variant E-GVN for link prediction. The core idea is to extract 
visual structural features from local subgraph renderings and fuse them with 
message-passing neural network representations. This design is orthogonal to 
existing structural-feature enhancements and consistently improves link 
prediction performance on both standard and large-scale benchmarks. The 
fourth part presents VizRAG, a retrieval-augmented generation framework 
enhanced by hypergraph visualization. We analyze the advantages and 
feasibility of visual hypergraph cues for knowledge-intensive generation, 
identify key challenges such as visual congestion and rendering bias, and 
introduce HyperViz as a practical toolkit to address them. Extensive 
experiments validate that visualized high-order structure improves retrieval 
quality and downstream response generation. The fifth part presents KICGPTv2, 
an LLM-enhanced framework for knowledge graph completion. It combines a 
structure-aware base KGC model, in-context knowledge prompting, and 
retrieval-augmented reconstruction to improve link prediction, relation 
prediction, and triple classification, with notable gains in long-tail cases 
and low additional training overhead.

Overall, this proposal establishes multimodal integration as a general and 
practical principle for graph AI. By bridging visual reasoning, 
language-based inference, and knowledge-aware graph modeling, the proposed 
research advances unified methodologies and empirical foundations for 
next-generation graph-centric intelligent systems.


Date:                   Tuesday, 18 August 2026

Time:                   3:00pm - 5:00pm

Venue:                  Room 3494
                        Lifts 25/26

Chairman:               

Committee Members:      Prof. James KWOK (Supervisor)
                        Dr. Long CHEN
                        Dr. Dan XU
                        Prof. Can YANG (MATH)
                        Prof. Hau San WONG (CityU)