Structured Scene Representation Learning for End-to-End Autonomous Driving

The Hong Kong University of Science and Technology
Department of Computer Science and Engineering


PhD Thesis Defence


Title: "Structured Scene Representation Learning for End-to-End Autonomous 
Driving"

By

Miss Xiaodong MEI


Abstract:

Autonomous driving is evolving from modular and rule-based systems toward 
end-to-end and data-driven paradigms, which promise greater adaptability in 
complex traffic environments. However, a central challenge remains for 
end-to-end autonomous driving: how to effectively represent dynamic and 
interactive driving scenarios that fundamentally support scene understanding, 
reasoning and planning.

This thesis investigates structured scene representation learning in 
end-to-end autonomous driving. We trace a progressive path to address this 
challenge across three axes: (a) from constrained intersection scenarios to 
diverse urban driving scenarios; (b) from intermediate vectorized inputs to 
raw images and natural language instructions; and (c) from sparse scene 
graphs to dense token sequences and ultimately to unified latent space.

First, we start from explicit interaction modeling of surrounding agents with 
G-CIL, a branched framework that represents the unsignalized intersection 
scenario as a structured graph. We employ graph convolutional networks (GCNs) 
to aggregate scene features, combined with conditional imitation learning to 
generate safe and reactive navigation policies from expert demonstrations.

Second, we move beyond hand-crafted graph structures with HAMF, a hybrid 
Attention-Mamba framework that models various scene elements as a sequence of 
tokens in urban driving scenarios. We jointly encode scene context and future 
motion representations in a unified architecture to generate feasible and 
diverse trajectories without explicit element-relation definitions.

Finally, we extend from vectorized inputs to raw images and natural language 
instructions with LVDrive, a latent visual representation enhanced 
vision-language-action (VLA) model that jointly models future scene 
representations and motion features in a shared latent space, enabling 
future-aware reasoning to refine trajectory generation for end-to-end 
navigation.

Collectively, this thesis traces a pathway from sparse scene graphs to dense 
latent representations, and from constrained intersection navigation with 
vectorized inputs to open-world end-to-end driving with raw sensor data. 
Together, these contributions move toward a more unified and advanced 
foundation for scene representation learning, enhancing scene understanding, 
reasoning and planning in autonomous driving.


Date:                   Monday, 17 August 2026

Time:                   10:00am - 12:00noon

Venue:                  Room 2132B
                        Lift 22

Chairman:               

Committee Members:      Dr. Dan XU (Supervisor)
                        Prof. Qiong LUO
                        Dr. Yangqiu SONG
                        Dr. Jun MA (EMIA)
                        Prof. Rui FAN (Tongji University)