You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes three short talks on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Tuesday, January 7, 2025, 12:00 pm -- CDS, Bldg. 725, Training Room


Speakers

Sanket Jantre
Tao Zhang
Xi Yu


Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1615289117?pwd=Hqkbj9itxWrFnkhZ8rQXHPInO2gxdF.1

Meeting ID: 161 528 9117
Passcode: 991382


Time: Jan 26, 2021 03:00 PM Eastern Time (US and Canada)

All are welcome!

Zoom Meeting:
https://stonybrook.zoom.us/j/93818552212?pwd=ajZkT2x4a2tiaDJUL1h3VFhLZEgwQT09

Meeting ID: 938 1855 2212
Passcode: 802722

Title: Data-Driven Document Unwarping

Abstract: Capturing document images is a common way to digitize and record physical documents due to the ubiquitousness of mobile cameras. To make text recognition easier, it is often desirable to digitally flatten a document image when the physical document sheet is folded or curved. However, unwarping a document from a single image in natural scenes is very challenging due to the complexity of document sheet deformation, document texture, and environmental conditions. Previous model-driven approaches struggle with inefficiency and limited generalizability. In this thesis, I investigate several data-driven approaches to tackle the document unwarping problem.

Data acquisition is the central challenge in data-driven methods. I first design an efficient data synthesis pipeline based on 2D image warping and train DocUNet, the pioneering data-driven document unwarping model, on the synthetic data. A benchmark dataset is also created to facilitate comprehensive evaluation and comparison. To improve the unwarping performance by training on more realistic data, I introduce the Doc3D dataset and DewarpNet. Supervised by 3D shape ground truth in Doc3D, DewarpNet is significantly better than DocUNet. DocUNet and DewarpNet depend on the synthetic data for the ground truth deformation annotation. To exploit the real-world images, I propose PaperEdge, a weakly supervised model trained with in-the-wild document images with easy-to-obtain boundary information. PaperEdge surpasses DewarpNet by utilizing both the synthetic data and weakly annotated real data in the Document In the Wild (DIW) dataset. Finally, I propose to incorporate the 3D physical constraints in training DewarpNet and PaperEdge. The constraints regulate the possible deformations on document papers. I also propose to augment the Doc3D and DIW dataset by introducing an online document segmentation model and better hardware.
Abstract: Facial emotion understanding aims to recognize, represent, and interpret human affect from facial behavior, and it is important for affective computing, human-computer interaction, digital humans, and mental health assessment. Existing work has represented facial emotion through discrete emotion categories, continuous affective dimensions such as valence and arousal, facial landmarks or geometry, and more recently semantic or language-based emotion descriptors. With the development of deep learning, supervised facial expression recognition has achieved strong performance on benchmark datasets, while self-supervised learning, multimodal large language models, and controllable facial generation have introduced new ways to learn emotion-related facial representations from images, videos, text, audio, and speech. However, many current models still rely heavily on manually annotated emotion labels, third-party perception judgments, or multimodal contextual cues, making it difficult to determine how much emotional information is captured directly from facial behavior itself, especially in naturalistic and clinically meaningful settings. Because these limitations make it challenging to evaluate whether facial representations capture emotionally meaningful behavior in real-world interactions, we propose to study facial emotion understanding through the relationship between facial behavior and language-derived emotional expression in psychiatric interview videos. Using a large dataset of mental health interviews, we extract multiple types of facial representations, including Action Unit features from FaceReader and OpenFace, non-AU facial behavior features from OpenFace, and 3D facial representations from EMOCA and SMIRK. We train segment-aligned transformer regressors to predict language-derived emotional targets, including valence, arousal, and RoBERTa-derived semantic-affective features from transcript segments. The results show that all facial representations achieve meaningful predictive performance across MSE and Pearson correlation metrics in both within-participant and between-participant evaluation settings. This indicates that facial behavior encodes information related to linguistic emotion at multiple levels: moment-to-moment emotional variation within individuals and broader affective differences across individuals. These findings suggest that structured facial representations can support vision-based emotion understanding in naturalistic mental health interviews and motivate future work on self-supervised, personalized, and controllable facial emotion models.

Speaker: Shao-Yu Chang

Zoom: https://stonybrook.zoom.us/j/3679036240?omn=98419305450
Abstract: In this talk, we will discuss what a CS PhD entails and the traits and habits that are important for success in PhD programs and future careers. While the talk is targeted to first-year PhD students, PhD students at all levels should derive from it.

Bio: Samir Das is a professor in the Department of Computer Science at Stony Brook
University. He is currently serving as the department chair. He is well recognized in the
community for his research in wireless networks and systems.

Location: NCS120
Abstract: Making large language models usable requires post-training methods that align models to human values, robustly handle underspecified inputs, and generalize to diverse instructions. This talk addresses the challenge of developing responsible AI through post-training from the following angles. First, I address contextual robustness. Preference data, for example, is often underspecified, and I show how underspecification in preference data can lead to diverging preferences. Standard reward models fail to properly handle these disagreements, often making decisive choices even when human annotators are split. I argue that more consequential outputs demand more context, and propose clarification question generation as one solution. Second, I will discuss how we can train models to be better instruction followers. I will show that most models severely overfit on a small set of instruction-following constraints and are not able to generalize well to unseen output constraints. I propose to train models with reinforcement learning from verifiable rewards for verifiable instruction following, and show how this leads to improved generalization on constraint following. Throughout the presentation, I will outline how I have applied these insights into developing open generative models, like Tülu and OLMo, and I will conclude with my research agenda for responsible post-training: critical AI evaluation, broader generalization, and expanding what models can reliably do.


Bio: Valentina Pyatkin is a postdoctoral researcher at the Allen Institute for AI and the University of Washington, advised by Prof. Hanna Hajishirzi and Prof. Yejin Choi. Additionally, she is part- time affiliated with the ETH AI Center, where she mentors students and works on post-training for the Swiss AI Initiative. She obtained her PhD in Computer Science from Bar Ilan University. Her work has been awarded an ACL Outstanding Paper Award and the ACL Best Theme Paper Award, and has been supported by a Schmidt Sciences Postdoctoral Award. During her doctoral studies, she conducted research internships at Google and the Allen Institute for AI, where she received the AI2 Outstanding Intern of the Year Award.

Location: NCS 120

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes three short talks on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Tuesday, January 7, 2025, 12:00 pm -- CDS, Bldg. 725, Training Room


Speakers

Deyu Lu
Mingyuan Ge
Kris Reyes


Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1615289117?pwd=Hqkbj9itxWrFnkhZ8rQXHPInO2gxdF.1

Meeting ID: 161 528 9117
Passcode: 991382

The International Neuroethics Society (INS) Speaker Series on AI & Consciousness

AI has existed as a tool for a long time, performing simple tasks such as sorting documents, suggesting music, and so on. But with the development of new generations of AI, the perception of its value to society has been increasing, as it can bring potential and promising benefits in many areas of human life. AI is known to have errors or biases that result in strange or even dangerous responses, but what happens when in AI-human interaction, the latter have errors or biases? cultural errors or biases? And what could be the implications for human relationships?

Speaker Bio

Dr. Karen Herrera-Ferrá is an independent and global consultant on ethical, medical, psychological, legal, social, cultural, policy-making, human rights and political issues and concerns on the development and use of neuroscience, neurotechnology and AI. She is a former member of the Board of Directors of the International Neuroethics Society.

Register here

https://umaryland.zoom.us/meeting/register/tJMvfuqsqDspG9BKMLfUU49UbuUyP_IEvXRh

CSE 656 Seminars in Computer Vision - Wednesdays 11:30am-12:50pm, Room NCS 120

The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 10 minutes) by multiple people. Students can register for 1 credit for CSE656. Registered students must attend and present a minimum of 2 or 3 talks. Everyone else is welcome to attend. Fill in https://forms.gle/pCVXovgfMfQwGqG38 to subscribe to our mailing list for further announcement.

The first meeting will be Wed Jan 29 at 11.30am, room 120 New CS. The meeting will deal with organizational matters and we will start right away with some presentations. Send David Paredes Merino <dparedesmeri@cs.stonybrook.edu> an email if you are interested but cannot attend the first meeting. Please forward to people outside the CS department that you think might be interested.
Join librarian Christine Fena for an interactive workshop that invites you to explore AI tools firsthand, not just as users, but as critical investigators. Through playful experimentation and collaborative discovery, you'll uncover inherent biases, probe algorithmic flaws, and gain a deeper understanding of AI's limitations and societal impacts.

Location: Melville Library, Central Reading Room, Lab B

https://library.stonybrook.edu/library-events/critiquing-ai/