Launching a University-Wide AI Innovation Institute:

Last spring, the Office of the Provost led a group of over 30 faculty, staff, and administrators to consider how we can expand and leverage our strengths in AI research and discovery. The resulting recommendation was to launch a university-wide AI Innovation Institute (AI3), which would expand the Institute for AI-driven Discovery and Innovation established in 2018 from a department-level institute within the College of Engineering and Applied Science (CEAS) to the university-wide AI Innovation Institute reporting to the provost.

As a university-wide enterprise, the AI Innovation Institute (AI3) is intended to accelerate, coordinate, and organize AI innovation and education across Stony Brook. The institute will serve to empower the entire university community and beyond, catalyzing core AI research, curriculum innovation, and societal change in the ever-evolving landscape of knowledge work.

The AI Town Hall, led by AI3 Interim Director Skiena, is an open house event that will provide an overview of the major AI initiatives on campus, including the new AI Seed Grant program and Stony Brook's role in New York State's Empire AI program. The session will include time for questions and discussion about the future of AI at Stony Brook.

Abstract: Human gaze behavior is a fundamental cue for understanding social intent, human-machine interaction, and cognitive processes. This thesis addresses the challenges of gaze target estimation (GTE), also known as gaze following, by developing a holistic understanding of gaze in complex environments.

The first part of this work improves GTE performance by introducing Patch-level Distribution Prediction (PDP). Unlike traditional models that rely on strict pixel-wise regression, PDP models gaze as a distribution over patches, which better accounts for annotation variance and bridges the gap between target location and in/out-of-frame prediction. To address the laborious nature of data labeling, the second part presents GCDR, the first semi-supervised method for gaze following. By prompting large Visual Question Answering (VQA) models to generate initial Grad-CAM heatmaps and refining them with a diffusion model, this method achieves high performance with significantly fewer human annotations. The third part expands the applicability of GTE to multi-camera environments. By introducing the Multi-View Gaze Target (MVGT) dataset, along with two novel frameworks for integrating information between multiple views and predicting the gaze target across views, we explore a new direction that overcomes single-view limitations such as face occlusion and out-of-view targets.

Building on these foundations, the final part of this thesis proposes a new direction toward semantic social gaze understanding using next-generation multimodal Large Language Models (LLMs). Rather than focusing solely on geometric gaze target localization, we aim to enrich gaze prediction with semantic and relational interpretation in complex social scenes. To this end, we will leverage existing gaze following datasets to derive social gaze supervision, including mutual gaze and shared attention, and obtain aligned language descriptions of scene-level gaze behaviors. This proposed work will enable the model to not only locate gaze targets but also predict structured social gaze relations among individuals, meanwhile generating a concise natural-language summary describing the dominant gaze interactions. By integrating spatial gaze estimation, social relation reasoning, and language-based scene understanding within a unified multimodal model, this work takes an important step toward a holistic understanding of human gaze behavior in real-world environments.

Speaker: Qiaomu Miao

The Office for Research and Innovation at Stony Brook University invites you to attend the inaugural Wolf Den, an evening designed to bring together members of the regional innovation and entrepreneurial ecosystem.

Meet investors, researchers, startup founders, and business leaders to exchange ideas, foster collaboration, and strengthen connections that drive technology development and economic growth across Long Island.

Agenda

4:30 - 5:00 PM | Grab some cheer & mingle
5:00 - 5:40 PM | Welcome remarks and AI Panel
5:40 - 6:00PM | Featured lightning pitches
6:00 - 7:00 PM | Food, drinks and great conversations!

Attendees will have the opportunity to learn more about Stony Brook's entrepreneurship ecosystem, hear company pitches from emerging startups, and engage in meaningful networking with innovators, investors and community partners.

Refreshments will be served. Registration is required.

In partnership with Accelerate Long Island.

https://www.stonybrook.edu/commcms/innovation/_events/wolfden.php

How to Succeed in Language Design Without Really Trying presented by Professor Brian Kernighan

ABSTRACT: Why do some languages succeed while others fall by the wayside? I've helped create nearly a dozen languages (mostly small) over the years; a handful are still in widespread use, while others have languished or simply disappeared. I've also been present at the creation of several other languages, including some really major ones. In this talk I'll give my humble, but correct, opinion on factors that affect success and failure, and try to offer some insight into what to do if you're trying to design a new language yourself, and why that might be a good thing.

BIO: Brian Kernighan received a PhD in electrical engineering from Princeton in 1969. He joined the Computer Science department at Princeton in 2000, after many years at Bell Labs. He is a co-creator of several programming languages, including AWK and AMPL, and of a number of tools for document preparation. He is the co-author of a dozen books and some technical papers, and holds 5 patents.
He is a member of the National Academy of Engineering and of the American Academy of Arts and Sciences. His research areas include programming languages, tools and interfaces that make computers easier to use, often for non-specialist users. He has also written two books on technology for
non-technical audiences: Understanding the Digital World in 2017 and Millions, Billions, Zillions: Defending Yourself in a World of Too Many Numbers, published in 2018. His most recent book, Unix: A History and a Memoir, was published in October 2019.
A talk by Jerome Zhengrong Liang entitled, Machine Learning from Original Images to Texture Patterns: A Paradigm Shift from Non-Medical Application to Medical Diagnosis. Abstract: Artificial intelligence (AI) research for medical diagnosis started soon after human began to use computer, initially called artificial neural network (ANN) and now convolutional neural network (CNN). ANN has been mainly explored to classify the experts' handcrafted features from the original (or raw) images, while CNN has been mainly explored directly on the raw images for both tasks of extracting abstract features and classifying the features. Experimental evidences have been shown that CNN can be trained by a large number of the raw images with experts' scores (or labels) to match or even surpass the experts' performance for both non-medical and medical diagnosis applications. However, the performances of the CNN models as well as the experts on medical diagnosis dropped dramatically when the labels of the raw images were replaced by the corresponding medical pathological reports. Accumulated medical knowledge reveals that the lesion heterogeneity is a footprint of lesion evolution and ecology, and the heterogeneity is an indicator of lesion progress and response to medical intervention. The heterogeneity can be reflected by the image contrast distribution (or texture patterns) across the lesion volume. Image textures have been shown as an effective descriptor of the lesion heterogeneity for computer-aided diagnosis. Can we map the raw images into texture patterns (or images) and train CNN to learn from the texture images? This question is the central theme of this presentation with application to CT Colonography or virtual colonoscopy, a game from AlphaGo to PolypGo. Bio: Jerome Zhengrong Liang, PhD, IEEE Fellow Imaging Research and Informatics Laboratory Department of Radiology, Stony Brook University
18th Annual Engineering Ball Flowerfield, St. James, NY Thursday April, 2nd, 7:00 to 10:00 pm Pick up your tickets in 231 Engineering (Monday - Friday, 10:00 am to 4 pm) Presenting Partner: L3Harris
Abstract: Foundation models brought a paradigm shift on representation learning and the deep learning community. In my talk, I will examine the role of foundation models in medical imaging, focusing on their potential to unify diverse tasks through large-scale, generalist architectures. While these models achieve strong performance, their deployment in healthcare raises challenges related to data limitations, privacy, validation, and trust. We will also discuss domain-specific models for imaging, along with efficient adaptation techniques to adapt such models on domains that they have not been trained on. The presentation will also address key issues of reliability and interpretability, highlighting approaches like conformal prediction and counterfactual intervention to improve uncertainty estimation and model transparency. Overall, the talk will emphasize that despite their promise, foundation models require robust evaluation and trustworthy design to ensure safe and effective use in clinical settings.

Speaker: Maria Vakalopoulou is an assistant professor (MCF) in applied mathematics at CentraleSupelec, University Paris Saclay in France and the group leader of the biomathematics group of MICS Laboratory focusing on mathematical modeling in Life Sciences. She is affliated with Inria Saclay in France and Archimedes Unit in Greece. Her main research interest include the development of computational methods for image perception focusing on earth observation and medical applications. Before that, she was a postdoctoral student at CentraleSupelec, where she worked with Nikos Paragios. She completed her PhD at the Remote Sensing Laboratory at the School of Rural, Surveying and Geo-Informatics Engineering of the National Technical University of Athens under the supervision of Konstantinos Karantzalos.

Location: NCS 220
Abstract: Drawing on group-theoretic and information-theoretic foundations, we propose information lattice learning (ILL) as a general framework to learn rules of a signal (e.g., an image or a probability distribution). In our definition, a rule is a coarsened signal used to help us gain one interpretable insight about the original signal. To make full sense of what might govern the signal's intrinsic structure, we seek multiple disentangled rules arranged in a hierarchy, called a lattice. Compared to representation/rule-learning models optimized for a specific task (e.g., classification), ILL focuses on explainability: it is designed to mimic human experiential learning and discover rules akin to those humans can distill and comprehend. We will detail the mathematical foundations and algorithms of ILL, and illustrate how it addresses the fundamental question what makes X an X by creating rule-based explanations designed to help humans understand. Our focus is on explaining X rather than (re)generating it. We show ILL's efficacy and interpretability on benchmarks and assessments, as well as a demonstration of ILL-enhanced classifiers achieving human-level digit recognition using only one or a few MNIST training examples (1-10 per class). We present applications in knowledge discovery, using ILL to distill music theory from scores and chemical laws from molecules and further revealing connections between them. We close with some early work on understanding the principles that govern scattering amplitudes in Super Yang-Mills theory, rather than just predicting them.

Biography: Lav R. Varshney is the Della Pietra Infinity Professor and inaugural director of the AI Innovation Institute at Stony Brook University. He is co-founder and CEO of Kocree, Inc., a startup company building novel human-controllable AI for discovery and creativity, and chief scientist of Ensaras, Inc., a startup company focused on AI and wastewater treatment. He holds appointments at RAND Corporation and at Brookhaven National Laboratory. He was previously on the faculty of the University of Illinois Urbana-Champaign, a visiting scholar at Northwestern's Kellogg School of Management, a principal research scientist at Salesforce Research AI, and a research staff member at IBM Research. He is a former White House staffer, having served on the National Security Council staff as a White House Fellow, where he contributed to national AI and wireless communications policy. His research interests include information theory and artificial intelligence. He received his B.S. degree from Cornell University and his S.M. and Ph.D. degrees from the Massachusetts Institute of Technology.

Location: Room 102

An interactive session to discover how to create ALT text tags from images and create high-impact visuals, from identification to communicating ideas with images.

Discover how to use AI to create ALT text from images as well as identify objects in your environment, and build relatable visuals for high-impact presentations. Images communicate ideas as a way to understand concepts. AI-generated images have helped allow anyone to create these.

In this session, you will

  1. Creating image ALT Tags
  2. Transform ideas into images that are visually appealing
  3. Identify objects from visuals

Register here.