More about HKUST
Controlling Factors of Variation in Lip-to-Speech Synthesis: Space, Time, Content, and Speaker
The Hong Kong University of Science and Technology
Department of Computer Science and Engineering
PhD Thesis Defence
Title: "Controlling Factors of Variation in Lip-to-Speech Synthesis: Space,
Time, Content, and Speaker"
By
Mr. Zhe NIU
Abstract:
Lip-to-speech (LTS) synthesis reconstructs intelligible and natural speech
from silent video of a talking face. The visual input does not uniquely
determine the acoustics, and audiovisual training pairs contain nuisance
variation. Mouth crops vary in position and scale, audio and video can have a
consistent temporal offset, similar lip movements can correspond to different
speech content, and training speakers introduce acoustic variation beyond
linguistic content. This thesis organizes these challenges as four factors:
space, time, content, and speaker.
A control mechanism is developed for each factor. The Mouth Alignment Network
(MAN) predicts affine transformations that map face frames to a canonical
configuration. Its training objective combines affine distillation from the
conventional landmark-based alignment pipeline with guidance from a
separately trained Mouth Scoring Network. A training-time synchronization
framework estimates and compensates for the offset between the visual
conditioning and acoustic target. A text-guided model uses hard monotonic
alignment to map a transcript onto the visual frames and conditions the
decoder on the aligned representation, using ground-truth transcripts during
training and lip-reading predictions during inference. Finally, a two-stage
speaker-normalized pipeline first predicts speech in a canonical voice and
then restores the target speaker through voice conversion.
Experiments show that these mechanisms have distinct and complementary roles.
MAN reduces preprocessing cost while matching the recognition performance of
the conventional landmark-based alignment pipeline. Synchronization reduces
sensitivity to training-time audio-visual offsets. Text guidance provides the
largest gain in intelligibility, while the speaker-normalized pipeline
improves acoustic quality and enables target-speaker restoration. The
integrated system combines these benefits, achieves strong intelligibility on
LRS3, and retains similar performance on LRS2 without using LRS2 data to
train its speech-generation stages.
Date: Wednesday, 12 August 2026
Time: 2:00pm - 4:00pm
Venue: Room 3494
Lifts 25/26
Chairman: Dr. Eun Soon IM (CIVL)
Committee Members: Dr. Brian MAK (Supervisor)
Prof. Pedro SANDER
Dr. Dan XU
Prof. Chi Ying TSUI (ECE)
Prof. Tan LEE (CUHK)