A Holistic Approach to Human Gaze Understanding: Probabilistic Modeling, Geometric Reasoning, and Multimodal Learning
Event Description
Abstract: Human gaze behavior is a fundamental cue for understanding social intent, human-machine interaction, and cognitive processes. This dissertation addresses the challenges of gaze target estimation (GTE), also known as gaze following, by developing a holistic understanding of gaze across complex environments.
First, we improve GTE performance through Patch-level Distribution Prediction (PDP). Unlike traditional pixel-wise regression, PDP models gaze as a spatial distribution over patches, better accounting for annotation variance and regularizing the pixel-wise heatmap prediction through multi-scale modeling. Second, to mitigate the high cost of data labeling, we present GCDR, the first semi-supervised method for gaze following. By prompting large Visual Question Answering (VQA) models to generate initial Grad-CAM heatmaps and refining them via a diffusion model, GCDR achieves robust performance with minimal human annotation. Third, we expand the applicability of GTE to multi-camera environments. By introducing the Multi-View Gaze Target (MVGT) dataset, along with two novel frameworks for integrating information and predicting gaze targets across views, we explore a new direction that overcomes single-view limitations such as face occlusion and out-of-view targets. Finally, we propose OmniGF, a multi-person gaze following model built on Vision-Language Models (VLMs) that enriches gaze target localization with semantic and social reasoning. By leveraging the semantic capabilities of VLMs alongside structural innovations to ground the model with fine-grained cues for each individual, OmniGF achieves state-of-the-art performance across three gaze following tasks.
Collectively, by tackling the gaze following problem through the distinct yet complementary perspectives of probabilistic modeling, geometric reasoning, and multimodal learning, this dissertation builds a holistic understanding of human gaze, paving the way for more intuitive artificial intelligence systems in downstream applications.
Speaker: Qiaomu Miao
Location: NCS 220