Abstract: Facial emotion understanding aims to recognize, represent, and interpret human affect from facial behavior, and it is important for affective computing, human-computer interaction, digital humans, and mental health assessment. Existing work has represented facial emotion through discrete emotion categories, continuous affective dimensions such as valence and arousal, facial landmarks or geometry, and more recently semantic or language-based emotion descriptors. With the development of deep learning, supervised facial expression recognition has achieved strong performance on benchmark datasets, while self-supervised learning, multimodal large language models, and controllable facial generation have introduced new ways to learn emotion-related facial representations from images, videos, text, audio, and speech. However, many current models still rely heavily on manually annotated emotion labels, third-party perception judgments, or multimodal contextual cues, making it difficult to determine how much emotional information is captured directly from facial behavior itself, especially in naturalistic and clinically meaningful settings. Because these limitations make it challenging to evaluate whether facial representations capture emotionally meaningful behavior in real-world interactions, we propose to study facial emotion understanding through the relationship between facial behavior and language-derived emotional expression in psychiatric interview videos. Using a large dataset of mental health interviews, we extract multiple types of facial representations, including Action Unit features from FaceReader and OpenFace, non-AU facial behavior features from OpenFace, and 3D facial representations from EMOCA and SMIRK. We train segment-aligned transformer regressors to predict language-derived emotional targets, including valence, arousal, and RoBERTa-derived semantic-affective features from transcript segments. The results show that all facial representations achieve meaningful predictive performance across MSE and Pearson correlation metrics in both within-participant and between-participant evaluation settings. This indicates that facial behavior encodes information related to linguistic emotion at multiple levels: moment-to-moment emotional variation within individuals and broader affective differences across individuals. These findings suggest that structured facial representations can support vision-based emotion understanding in naturalistic mental health interviews and motivate future work on self-supervised, personalized, and controllable facial emotion models.

Speaker: Shao-Yu Chang

Zoom: https://stonybrook.zoom.us/j/3679036240?omn=98419305450

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Abstract: Designing custom proteins could revolutionize medicine and materials, but it remains an immense scientific challenge. Our work uses large-scale AI foundation models to generate novel proteins tailored to bind specific small molecules. Each AI-generated design is passed through a rigorous, multi-stage validation pipeline to ensure it is biophysically realistic. A key innovation is fine-tuning our model with data from molecular dynamics (MD) simulations, exposing it to the conformational dynamics and energetics of protein-ligand binding. This physics-aware training results in novel protein designs with enhanced stability and more effective binding capabilities.

Bio: Xin Dai is an Assistant Computational Scientist in the Artificial Intelligence Department of the CDS. His work centers on AI for Science with a strong focus on computational biology. He earned his PhD in Physics from Tsinghua University.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Location: CDS, Bldg. 725, Training Room

Join Zoom Meeting: https://bnl.zoomgov.com/j/1604383624?pwd=ffQ5cUPNxTI7nzClKQO6cnsNbhF9Vf.1

Meeting ID: 160 438 3624
Passcode: 558449

Zoom Like a Pro! Unlock Whiteboard, Polls, AI Companion, and more to supercharge student participation. This hands-on workshop explores innovative ways to use Zoom's built-in tools to enhance active learning activities in your classes. Learn how to utilize the Whiteboard feature to make collaborative work more engaging, use Polling and Quizzes for instant feedback, AI Companion for summary, and Breakout Sessions for group activities. Register here: https://stonybrook.zoom.us/meeting/register/tJckf--rpj4pGdRV0ItgTW8Lk7gn_RuykByO#/registration
Abstract: Capturing the spatio-temporal (4D) dynamics of humans has been a long standing research problem in computer vision and graphics. Synthesizing photorealistic human avatars has broad applications, ranging from immersive telepresence in AR/VR and the movie industry, to enriching the education and healthcare systems. Earlier approaches relied on hand-engineered models that use a small amount of data from one or more subjects. With the advent of neural networks, training on large datasets enhanced the output visual quality. Currently, the combination of neural networks with graphics techniques has achieved natural-looking human animation. However, most approaches are identity-specific, trained only on a single identity, and use only one modality.

In this dissertation, we address the problem of learning neural representations of humans in a holistic way. Given that the video data in the real world include multiple modalities (e.g., audio and video) and multiple identities, we develop multi-modal and multi-identity representations. First, we propose to reconstruct the 4D face geometry of humans by leveraging both audio and video information. In this way, the network produces accurate lip shapes and is robust to cases when either modality is insufficient. Next, we introduce a NeRF-based representation for audio-driven human face animation that achieves high-quality lip synchronization for cinematic content. Since humans communicate with their full body, combining body pose, hand gestures, and facial expressions, we extend the network to capture full-body human motion for multiple identities simultaneously. In order to better disentangle identity and non-identity specific information, we subsequently study non-linear interactions between latent factors of variation, and propose a specific multiplicative module. In this way, we learn a multi-identity NeRF that robustly animates human faces under novel expressions and achieves a significant decrease in the total training time. Similarly, we propose a multi-identity Gaussian splatting representation for human bodies, by constructing a high-order tensor. Assuming a low-rank structure, we learn a tensor decomposition that leads to a significant decrease in the total number of learnable parameters, as well as to a robust animation under novel poses. Last but not least, we propose to jointly synthesize audio and visual outputs from just text input. Given the recent rise of large language models, coupling text with natural-looking avatars can enhance the overall interaction between a human and an AI system.

Location: NCS 220 or Zoom

Join a faculty development program to support instructors across campus with navigating/integrating AI in their courses. We're inviting interested faculty to participate in the grant project called Fostering Writing-to-Learn Skills with Critical AI Literacy: A Faculty Development and Student Support Program (funded through the AI3 Institute).

Time commitment and completion requirements :

  • Attend four sessions and a final symposium on the following dates/times:

    • Friday, September 12 from 11am - 12:30pm over Zoom

    • Friday, September 26 from 11am - 12:30pm over Zoom

    • Friday, October 10 from 11am - 12:30pm over Zoom

    • Friday, October 24 from 11am - 12:30pm over Zoom

    • Friday, November 14 from 10am - 1pm in Wang 201 - please note that this is an in person session only

  • Engage with online materials in Brightspace prior to each of the sessions (mainly to update a syllabus, assignment, or teaching strategy that you can share and discuss at the workshop)

Contact: Shyam Sharma, Christine Fena, and Rose Tirotta-Esposito with questions.

https://docs.google.com/document/d/1b51tvfK0HSOkCW7cwYq2nyyeeHtvBZYC7_XHv7Av8wQ/edit?tab=t.0
Abstract: Gaussian Probability Path-based Generative Models (GPPGMs) generate data by reversing a stochastic process that progressively corrupts samples with Gaussian noise. Despite state-of-the-art results in 3D molecular generation, their deployment is hindered by the high cost of long generative trajectories, often requiring hundreds to thousands of steps during training and sampling. In this work, we propose a principled method, named GAGA, to improve generation efficiency without sacrificing training granularity or inference fidelity of GPPGMs. Our key insight is that different data modalities obtain sufficient Gaussianity at markedly different steps during the forward process. Based on this observation, we analytically identify a characteristic step at which molecular data attains sufficient Gaussianity, after which the trajectory can be replaced by a closed-form Gaussian approximation. Unlike existing accelerators that coarsen or reformulate trajectories, our approach preserves full-resolution learning dynamics while avoiding redundant transport through truncated distributional states. Experiments on 3D molecular generation benchmarks demonstrate that our GAGA achieves substantial improvement on both generation quality and computational efficiency.

Speaker: Jingxiang Qu

Location: New Computer Science 220
AI/ML Working Group Seminar

Time/Date: 12:00 PM ET, Tuesday, March 1st, 2022

Seminar Speaker: Yen-Chi (Sam) Chen, CSI, Brookhaven National Laboratory

Title: When reinforcement learning meets quantum computing

Abstract: Recently, reinforcement learning (RL) has demonstrated
various applications with superhuman performance such as mastering the
game of Go.  Meanwhile, the development of quantum computing hardware
shed light on building practical quantum applications to tackle
previously unsolved problems. What will happen if we combine these two
fascinating techniques? In this talk, I will present the recent
progress in quantum RL as well as using classical RL to help certain
tasks in quantum computing.



Host: Meifeng Lin, Computational Science Initiative

_______________________________________________

Nicole Medaglia is inviting you to a scheduled ZoomGov meeting.

Join ZoomGov Meeting
https://bnl.zoomgov.com/j/1619877909?pwd=T041dGl4SURUK0Mwbmp0b1QvVjVtZz09

Meeting ID: 161 987 7909
Passcode: 338057
One tap mobile
+16692545252,,1619877909#,,,,*338057# US (San Jose)
+16468287666,,1619877909#,,,,*338057# US (New York)

Dial by your location
        +1 669 254 5252 US (San Jose)
        +1 646 828 7666 US (New York)
        +1 669 216 1590 US (San Jose)
        +1 551 285 1373 US
Meeting ID: 161 987 7909
Passcode: 338057
Find your local number: https://bnl.zoomgov.com/u/abMDS0zjuq

Join by SIP
1619877909@sip.zoomgov.com

Join by H.323
161.199.138.10 (US West)
161.199.136.10 (US East)
Meeting ID: 161 987 7909
Passcode: 338057
Abstract: Datalog is a powerful language for expressing recursive computations through rules: Horn clauses in first order logic. Although effective at expressing queries over existential properties, Datalog and many of its popular implementations struggle with queries that involve more complex aggregates, requiring users to apply verbose, non-composable, and/or inefficient workarounds. Recent work on lattice-based datalogs addresses many of these concerns for aggregates that can be encoded as lattices (e.g., min or max), but more general aggregates like count remain problematic. In this talk, I will argue that this is not a fundamental limitation of Datalog, but rather from its model of truth: Both datalog semantics and evaluation rules make heavy use of the fact that insertion is both monotone and idempotent. Once a fact is known to be true, it can not be retracted, nor can further discoveries of the same fact alter its truth. Monotonicity is critical for forward progress under Datalog's ``open world'' model, as it allows us to safely assert the truth of a body. Meanwhile, idempotence makes it easier to reason about evaluation, as we need only guarantee that each head atom will be derived at-least-once. Unfortunately, more general aggregates like sum() are neither idempotent, nor monotone. I will introduce Hedgelog, a strict generalization of Datalog that uses general monoids as a basis for truth. I will show that this generalization remains compatible with Datalog's open world model, how it enables cleaner and more composable datalog programs, and how the underlying monoid relations open the door to interesting datastructure-level optimizations.

Bio: Oliver Kennedy is an associate professor at the University at Buffalo. He earned his PhD from Cornell University in 2011 and now leads the Online Data Interactions (ODIn) lab, which operates at the intersection of databases and programming languages. Oliver is the recipient of an NSF CAREER award, an IEEE Region 1 Technological Innovation Award, UB's Exceptional Scholar Award, and several UB SEAS teaching awards. Oliver is also one of the founding board members of Breadcrumb Analytics. Several of Oliver's papers have been invited to Best of compilations from SIGMOD and VLDB. The ODIn lab is currently exploring (i) how we can leverage database techniques like incremental view maintenance to make compilers faster, (ii) how to make it easier for data scientists to track how sources of uncertainty, ambiguity, and/or bias affect analyses, and (iii) how to streamline the interfaces --- both human and software --- between different tools for data science, like python, sql, and spreadsheets.

Location: NCS 120

The Association for Computational Linguistics is the international scientific and professional society for people working on problems involving natural language and computation. Membership includes the ACL quarterly journals, Computational Linguistics and Transactions of the ACL, reduced registration at most ACL-sponsored conferences, discounts on ACL-sponsored publications, and participation in ACL Special Interest Groups.

An annual meeting is held each summer in locations where significant computational linguistics research is carried out.

For more information and registration, visit the official website.