The Collective Surgical Consciousness: Artificial Intelligence & the Future of Surgery Guest speaker Doctor Ozanan Meireles, the Director of the Surgical AI and Innovation Lab at Massachusetts General Hospital and a faculty member at Harvard Medical School, presents The Collective Surgical Consciousness: Artificial Intelligence & the Future of Surgery. Objectives: * Become familiar with the subfields of AI used in surgery * Understand the importance of a potential paradigm shift in surgical practice, training, and continue medical development * The importance of data acquisition, sharing and ownership, and development of machine learning algorithms
Department of Writing and Rhetoric Lecture

Escaping Human Judgment? The Politics of AI under Authoritarianism

Much of the recent debate around AI-decision making assumes that human involvement is a safeguard: humans bring context, accountability, and fairness to otherwise impersonal algorithmic systems. But whether human or algorithmic judgment is preferred depends on the complex interaction between an individual's social position and the configuration of power in which that judgment takes place.

Speaker: Zheng Fu, PRODIG+ Fellow

Location: Humanities 1008
The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 10 minutes) by multiple people. Students can register for 1 credit for CSE 656. Registered students must attend and present a minimum of 2 or 3 talks. Everyone else is welcome to attend. Fill in https://forms.gle/pCVXovgfMfQwGqG38 to subscribe to our mailing list for further announcement.
Abstract: Making large language models usable requires post-training methods that align models to human values, robustly handle underspecified inputs, and generalize to diverse instructions. This talk addresses the challenge of developing responsible AI through post-training from the following angles. First, I address contextual robustness. Preference data, for example, is often underspecified, and I show how underspecification in preference data can lead to diverging preferences. Standard reward models fail to properly handle these disagreements, often making decisive choices even when human annotators are split. I argue that more consequential outputs demand more context, and propose clarification question generation as one solution. Second, I will discuss how we can train models to be better instruction followers. I will show that most models severely overfit on a small set of instruction-following constraints and are not able to generalize well to unseen output constraints. I propose to train models with reinforcement learning from verifiable rewards for verifiable instruction following, and show how this leads to improved generalization on constraint following. Throughout the presentation, I will outline how I have applied these insights into developing open generative models, like Tülu and OLMo, and I will conclude with my research agenda for responsible post-training: critical AI evaluation, broader generalization, and expanding what models can reliably do.


Bio: Valentina Pyatkin is a postdoctoral researcher at the Allen Institute for AI and the University of Washington, advised by Prof. Hanna Hajishirzi and Prof. Yejin Choi. Additionally, she is part- time affiliated with the ETH AI Center, where she mentors students and works on post-training for the Swiss AI Initiative. She obtained her PhD in Computer Science from Bar Ilan University. Her work has been awarded an ACL Outstanding Paper Award and the ACL Best Theme Paper Award, and has been supported by a Schmidt Sciences Postdoctoral Award. During her doctoral studies, she conducted research internships at Google and the Allen Institute for AI, where she received the AI2 Outstanding Intern of the Year Award.

Location: NCS 120
Abstract: Facial emotion understanding aims to recognize, represent, and interpret human affect from facial behavior, and it is important for affective computing, human-computer interaction, digital humans, and mental health assessment. Existing work has represented facial emotion through discrete emotion categories, continuous affective dimensions such as valence and arousal, facial landmarks or geometry, and more recently semantic or language-based emotion descriptors. With the development of deep learning, supervised facial expression recognition has achieved strong performance on benchmark datasets, while self-supervised learning, multimodal large language models, and controllable facial generation have introduced new ways to learn emotion-related facial representations from images, videos, text, audio, and speech. However, many current models still rely heavily on manually annotated emotion labels, third-party perception judgments, or multimodal contextual cues, making it difficult to determine how much emotional information is captured directly from facial behavior itself, especially in naturalistic and clinically meaningful settings. Because these limitations make it challenging to evaluate whether facial representations capture emotionally meaningful behavior in real-world interactions, we propose to study facial emotion understanding through the relationship between facial behavior and language-derived emotional expression in psychiatric interview videos. Using a large dataset of mental health interviews, we extract multiple types of facial representations, including Action Unit features from FaceReader and OpenFace, non-AU facial behavior features from OpenFace, and 3D facial representations from EMOCA and SMIRK. We train segment-aligned transformer regressors to predict language-derived emotional targets, including valence, arousal, and RoBERTa-derived semantic-affective features from transcript segments. The results show that all facial representations achieve meaningful predictive performance across MSE and Pearson correlation metrics in both within-participant and between-participant evaluation settings. This indicates that facial behavior encodes information related to linguistic emotion at multiple levels: moment-to-moment emotional variation within individuals and broader affective differences across individuals. These findings suggest that structured facial representations can support vision-based emotion understanding in naturalistic mental health interviews and motivate future work on self-supervised, personalized, and controllable facial emotion models.

Speaker: Shao-Yu Chang

Zoom: https://stonybrook.zoom.us/j/3679036240?omn=98419305450
CSE 600 Talk: Haibin Ling - Computer Vision Research and Applications


Abstract: Having been intensively studied over half a decade, computer vision has evolved as a broad research area and become mature in many applications. In this talk, we will summarize our work in computer vision in both core vision topics and application-oriented ones. In particular, for core vision problems, we will report studies on visual tracking, visual matching and visual detection; for applications, we will describe our work on medical image analysis, intelligent transportation, smart projector systems and preliminary work on material property prediction.

Bio: Haibin Ling received the BS and MS degrees from Peking University in 1997 and 2000, respectively, and the PhD degree from the University of Maryland, College Park, in 2006. From 2000 to 2001, he was an assistant researcher at Microsoft Research Asia. From 2006 to 2007, he worked as a postdoctoral scientist at the University of California Los Angeles. In 2007, he joined Siemens Corporate Research as a research scientist. From 2008 to 2019, he worked as a faculty member of Temple University. In fall 2019, he joined the Department of Computer Science of Stony Brook University where he is currently a SUNY Empire Innovation Professor. His research interests include computer vision, augmented reality, medical image analysis, and human computer interaction. He received the Best Student Paper Award at the ACM UIST in 2003, and the NSF CAREER Award in 2014. He serves as Associate Editor for several journals including IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Pattern Recognition (PR), and Computer Vision and Image Understanding (CVIU). He has served or will serve as Area Chair for CVPR 2014, 2016, 2019 and 2020.

Learn how to unlock the power of Image and visuals that will enhance your work by asking the experts questions in-person

No registration required - just stop by!

Location: Frank Melville Jr. Memorial Library Galleria (across from the Central Reading room)

The Fourth Arabic Natural Language Processing Conference (ArabicNLP 2026) is organized by the ACL Special Interest Group on Arabic NLP (SIGARAB).
The research focus of ArabicNLP is, naturally, Arabic, a collection of language varieties, from Classical to Modern Standard Arabic (MSA), and including many living and historical Arabic dialects. Arabic poses many challenges for the field of computational linguistics, including rich morphology, orthographic ambiguity as well as the wide variety of understudied dialects.

Location: Budapest, Hungary

Register here.
The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the Ph.D. program or (2) receive permission from the instructors. Each seminar will consist of multiple short talks (around 15 minutes) by multiple students. Students can register for 1 credit for CSE656. Registered students must attend and present a minimum of 2 talks. Registered students must attend in person. Up to 3 absences will be excused. Everyone else is welcome to attend. The seminar will be taught by Prof. Chao Chen, chao.chen.1@stonybrook.edu.
Abstract: Computational pathology has revolutionized cancer diagnosis and research through the analysis of digitized whole slide images (WSIs). However, their giga-pixel size creates two intertwined bottlenecks: computational inefficiency, as prohibitive GPU memory makes standard end-to-end (E2E) training infeasible, and label inefficiency, as expert annotation is tedious and expensive. This dissertation confronts both challenges through novel architectures, training paradigms, and self-supervised learning methods for efficient WSI analysis.
To improve computational efficiency, this dissertation first introduces a locally supervised learning paradigm that enables E2E training on entire WSIs by partitioning a network into gradient-isolated modules, circumventing the memory bottleneck of backpropagation. Second, it presents Prompt-MIL, a parameter-efficient fine-tuning framework training only a few prompts to guide large pre-trained models, reducing trainable parameters, memory, and training time. Third, this work proposes 2DMamba, the first intrinsic Mamba architecture that preserves the crucial 2D spatial structure of images, overcoming the spatial discrepancy in 1D models. Fourth, it presents Locally Bi-directional Mamba (LBMamba), whose hardware-aware local backward scan integrates bi-directional scanning into a single forward pass, improving the throughput-performance trade-off of Mamba models.
To improve label efficiency, this dissertation proposes a precise location based matching strategy for self-supervised dense contrastive learning, which allows a local patch in one augmented view to match multiple overlapping patches in another, producing more accurate correspondences and superior features for dense prediction tasks like segmentation and detection. Additionally, to better scale multi-channel cell imaging modalities, this dissertation introduces ChannelSFormer, a channel-agnostic vision transformer that disentangles spatial and channel-wise reasoning through divided attention and channel class token, enabling effective representation learning across variable channel configurations in both self-supervised and supervised settings.
In summary, this dissertation presents a holistic investigation into the efficiency bottlenecks in computational pathology. Through these combined contributions in model architecture, training paradigms, and self-supervised learning, this work establishes a more scalable, efficient, and powerful computational framework for analyzing giga-pixel pathology images.

Speaker: Jingwei Zhang

Location: NCS 220

Zoom: https://stonybrook.zoom.us/j/93175806292?pwd=xbtxnQyYGoThz5B1DyJxJxPF9lxiJE.1
Meeting ID: 931 7580 6292
Passcode: 314091