Controlling Factors of Variation in Lip-to-Speech Synthesis: Space, Time, Content, and Speaker

The Hong Kong University of Science and Technology
Department of Computer Science and Engineering


PhD Thesis Defence


Title: "Controlling Factors of Variation in Lip-to-Speech Synthesis: Space, 
Time, Content, and Speaker"

By

Mr. Zhe NIU


Abstract:

Lip-to-speech (LTS) synthesis reconstructs intelligible and natural speech 
from silent video of a talking face. The visual input does not uniquely 
determine the acoustics, and audiovisual training pairs contain nuisance 
variation. Mouth crops vary in position and scale, audio and video can have a 
consistent temporal offset, similar lip movements can correspond to different 
speech content, and training speakers introduce acoustic variation beyond 
linguistic content. This thesis organizes these challenges as four factors: 
space, time, content, and speaker.

A control mechanism is developed for each factor. The Mouth Alignment Network 
(MAN) predicts affine transformations that map face frames to a canonical 
configuration. Its training objective combines affine distillation from the 
conventional landmark-based alignment pipeline with guidance from a 
separately trained Mouth Scoring Network. A training-time synchronization 
framework estimates and compensates for the offset between the visual 
conditioning and acoustic target. A text-guided model uses hard monotonic 
alignment to map a transcript onto the visual frames and conditions the 
decoder on the aligned representation, using ground-truth transcripts during 
training and lip-reading predictions during inference. Finally, a two-stage 
speaker-normalized pipeline first predicts speech in a canonical voice and 
then restores the target speaker through voice conversion.

Experiments show that these mechanisms have distinct and complementary roles. 
MAN reduces preprocessing cost while matching the recognition performance of 
the conventional landmark-based alignment pipeline. Synchronization reduces 
sensitivity to training-time audio-visual offsets. Text guidance provides the 
largest gain in intelligibility, while the speaker-normalized pipeline 
improves acoustic quality and enables target-speaker restoration. The 
integrated system combines these benefits, achieves strong intelligibility on 
LRS3, and retains similar performance on LRS2 without using LRS2 data to 
train its speech-generation stages.


Date:                   Wednesday, 12 August 2026

Time:                   2:00pm - 4:00pm

Venue:                  Room 3494
                        Lifts 25/26

Chairman:               Dr. Eun Soon IM (CIVL)

Committee Members:      Dr. Brian MAK (Supervisor)
                        Prof. Pedro SANDER
                        Dr. Dan XU
                        Prof. Chi Ying TSUI (ECE)
                        Prof. Tan LEE (CUHK)