More about HKUST
Structured Scene Representation Learning for End-to-End Autonomous Driving
The Hong Kong University of Science and Technology
Department of Computer Science and Engineering
PhD Thesis Defence
Title: "Structured Scene Representation Learning for End-to-End Autonomous
Driving"
By
Miss Xiaodong MEI
Abstract:
Autonomous driving is evolving from modular and rule-based systems toward
end-to-end and data-driven paradigms, which promise greater adaptability in
complex traffic environments. However, a central challenge remains for
end-to-end autonomous driving: how to effectively represent dynamic and
interactive driving scenarios that fundamentally support scene understanding,
reasoning and planning.
This thesis investigates structured scene representation learning in
end-to-end autonomous driving. We trace a progressive path to address this
challenge across three axes: (a) from constrained intersection scenarios to
diverse urban driving scenarios; (b) from intermediate vectorized inputs to
raw images and natural language instructions; and (c) from sparse scene
graphs to dense token sequences and ultimately to unified latent space.
First, we start from explicit interaction modeling of surrounding agents with
G-CIL, a branched framework that represents the unsignalized intersection
scenario as a structured graph. We employ graph convolutional networks (GCNs)
to aggregate scene features, combined with conditional imitation learning to
generate safe and reactive navigation policies from expert demonstrations.
Second, we move beyond hand-crafted graph structures with HAMF, a hybrid
Attention-Mamba framework that models various scene elements as a sequence of
tokens in urban driving scenarios. We jointly encode scene context and future
motion representations in a unified architecture to generate feasible and
diverse trajectories without explicit element-relation definitions.
Finally, we extend from vectorized inputs to raw images and natural language
instructions with LVDrive, a latent visual representation enhanced
vision-language-action (VLA) model that jointly models future scene
representations and motion features in a shared latent space, enabling
future-aware reasoning to refine trajectory generation for end-to-end
navigation.
Collectively, this thesis traces a pathway from sparse scene graphs to dense
latent representations, and from constrained intersection navigation with
vectorized inputs to open-world end-to-end driving with raw sensor data.
Together, these contributions move toward a more unified and advanced
foundation for scene representation learning, enhancing scene understanding,
reasoning and planning in autonomous driving.
Date: Monday, 17 August 2026
Time: 10:00am - 12:00noon
Venue: Room 2132B
Lift 22
Chairman:
Committee Members: Dr. Dan XU (Supervisor)
Prof. Qiong LUO
Dr. Yangqiu SONG
Dr. Jun MA (EMIA)
Prof. Rui FAN (Tongji University)