Learning Facial Representations for Emotion Understanding

Event Description

Abstract: Facial emotion understanding aims to recognize, represent, and interpret human affect from facial behavior, and it is important for affective computing, human-computer interaction, digital humans, and mental health assessment. Existing work has represented facial emotion through discrete emotion categories, continuous affective dimensions such as valence and arousal, facial landmarks or geometry, and more recently semantic or language-based emotion descriptors. With the development of deep learning, supervised facial expression recognition has achieved strong performance on benchmark datasets, while self-supervised learning, multimodal large language models, and controllable facial generation have introduced new ways to learn emotion-related facial representations from images, videos, text, audio, and speech. However, many current models still rely heavily on manually annotated emotion labels, third-party perception judgments, or multimodal contextual cues, making it difficult to determine how much emotional information is captured directly from facial behavior itself, especially in naturalistic and clinically meaningful settings. Because these limitations make it challenging to evaluate whether facial representations capture emotionally meaningful behavior in real-world interactions, we propose to study facial emotion understanding through the relationship between facial behavior and language-derived emotional expression in psychiatric interview videos. Using a large dataset of mental health interviews, we extract multiple types of facial representations, including Action Unit features from FaceReader and OpenFace, non-AU facial behavior features from OpenFace, and 3D facial representations from EMOCA and SMIRK. We train segment-aligned transformer regressors to predict language-derived emotional targets, including valence, arousal, and RoBERTa-derived semantic-affective features from transcript segments. The results show that all facial representations achieve meaningful predictive performance across MSE and Pearson correlation metrics in both within-participant and between-participant evaluation settings. This indicates that facial behavior encodes information related to linguistic emotion at multiple levels: moment-to-moment emotional variation within individuals and broader affective differences across individuals. These findings suggest that structured facial representations can support vision-based emotion understanding in naturalistic mental health interviews and motivate future work on self-supervised, personalized, and controllable facial emotion models.

Speaker: Shao-Yu Chang

Zoom: https://stonybrook.zoom.us/j/3679036240?omn=98419305450

Date Start

Date End