Proactive Response from Video in Human-Centric Embodied Interaction

PhD Qualifying Examination


Title: "Proactive Response from Video in Human-Centric Embodied Interaction"

by

Miss Xiaomeng ZHU


Abstract:

A common formulation in embodied interaction specifies the task before an 
agent begins acting. During human-agent interaction, however, the unfolding 
situation may indicate a need for a response even without an explicit request. 
Research on proactive response therefore examines how agents interpret such 
evidence to decide whether and when to respond and what the response should 
entail. In this survey, we focus on video as the primary perceptual input 
because it can provide rich spatiotemporal signals as an interaction unfolds. 
Existing studies often examine individual stages within task-specific 
settings, whereas the connections among these stages have received less 
systematic treatment. To clarify these connections, we trace how video 
observations inform task-level decisions. We first examine how video supports 
embodied situation understanding across successive semantic levels, from the 
physical scene through human behavior and observable states to task-level 
interaction. We then review how systems infer human goals and intentions from 
observed behavior, identify unexpressed needs from additional context, and 
decide whether and when to respond as evidence accumulates. Building on this 
initiation decision, we finally consider how systems select a language 
response or robot task action, filter candidate actions under task 
constraints, and construct a multistep task plan when several actions are 
required.


Date:                   Tuesday, 8 September 2026

Time:                   2:00pm - 4:00pm

Venue:                  Room 3494
                        Lift 25/26

Committee Members:      Prof. Fangzhen Lin (Supervisor)
                        Dr. Dan Xu (Chairperson)
                        Prof. Nevin Zhang