Please join us on Zoom for our next event in the Fall 2025 Stony Brook School of Nursing Research Seminar Series presented by our Office of Research and Innovation.

Topic: Responsible Artificial Intelligence: Promoting Health Equity for All

Speaker: Michael P. Cary, Jr., PhD, RN, FAAN.

Dr. Cary is a tenured Associate Professor at the Duke University School of Nursing. Dually trained as a health services researcher and applied health data scientist, Dr. Cary utilizes AI to investigate health disparities in aging populations, thereby promoting health equity and improving healthcare delivery. He co-directs HUMAINE™, an initiative dedicated to equipping nurses and healthcare professionals with the knowledge and skills necessary for the responsible use of AI in clinical practice.

Register: https://web.cvent.com/event/057978a5-a770-4de5-aca5-ad00287e4902/summary

Abstract:

It is known that models like large language models (LLMs) can often suggest colloquial plans given verbal descriptions of tasks, yet they are unable to reliably provide executable and verifiable plans given formally specified environments. In this talk, I will discuss a strand of efforts to have LLMs generate accurate and explainable plans in textual simulations. Instead of directly generating the plan or actions, LLMs are prompted to generate Planning Domain Definition Language (PDDL) that specifies the environment (domain file) and the task (problem file), which can then be deterministically solved with an off-the-shelf planner. In a 3-phase study, my collaborators and I first observed that it is possible but very challenging for LLMs to generate long-form code such as PDDL domain and problem files given textual specifications. Next, we devise methodologies for LLMs to iteratively generate and refine problem files while exploring a partially-observed, simulated, textual environment. Finally, we show that domain files are even more difficult to generate correctly, even on well-established planning tasks such as BlocksWorld. Finally, I will discuss ongoing efforts to improve said ability of structured generation and promising frontiers to explore.

Bio:
Li Harry Zhang is an assistant professor at Drexel University, focusing on Natural Language Processing (NLP) and artificial intelligence (AI). He obtained his PhD degree from the University of Pennsylvania advised by Prof. Chris Callison-Burch. Prior, he obtained his Bachelor's degree at the University of Michigan mentored by Prof. Rada Mihalcea and Prof. Dragomir Radev. His current research uses large language models (LLMs) to reason and plan via symbolic and structured representations. He has published more than 20 peer-reviewed papers in NLP and AI conferences, such as ACL, EMNLP, and AACL, that have been cited more than 1,000 times. He also consistently serves as Area Chair, Session Chair, and reviewer in those venues. Being a musician, producer, and content creator having over 50,000 subscribers, he is also passionate in the research of AI music and creativity.

Abstract: Graphs are a universal language of science. Molecules, materials, quantum systems, and knowledge bases can all be naturally represented as graphs. This talk explores how graph-based artificial intelligence is emerging as a powerful engine for scientific discovery. Using molecular design as a guiding example, we examine how modern graph AI enables machines not only to analyze complex scientific structures but also to generate new ones. We will discuss graph neural networks for learning predictive models of molecular properties, graph generative models for constructing novel chemical structures, and emerging multimodal graph-language models that support inverse design and synthesis planning. Together, these advances make graph AI more scalable, interpretable, and data-efficient--key capabilities for real-world scientific discovery. As artificial intelligence enters the era of foundation models, the next frontier lies in multimodal reasoning. Scientific knowledge is not purely textual; it is expressed through structures, code, and experimental data. By integrating graph representations with large language models, we move toward AI systems that can reason across multiple modalities and engage with scientific knowledge in its native forms. Looking ahead, we envision AI systems that behave less like tools and more like collaborators in the scientific process--generating hypotheses, designing candidate structures, planning experiments, interpreting results, and iteratively refining ideas through cycles of success and failure. In this vision, multimodal and agentic AI will enable scientists to explore vast and previously inaccessible design spaces, accelerating breakthroughs across domains ranging from drug discovery and materials innovation to software systems and quantum technologies.

Bio: Jie Chen is an interdisciplinary researcher working at the intersection of computing and mathematics, with a current focus on foundation models and AI agents for scientific discovery. His research integrates machine learning, statistics, scientific computing, and numerical linear algebra, with contributions spanning graph neural networks, multimodal graph LLMs, graph structure learning, scalable Gaussian processes, graph coarsening, and matrix functions. He is widely recognized for transformative contributions to graph-based deep learning and large-scale statistical modeling, and for bridging theory with real-world scientific and engineering applications. Dr. Chen has led externally funded, multi-institutional research programs supported by Shell, Evonik, and the U.S. Department of Energy, with applications in materials discovery, financial forensics, and power system resilience. He previously served as a Senior Research Scientist and Manager at IBM Research and the MIT-IBM Watson AI Lab, and as a Postdoctoral Fellow at Argonne National Laboratory. He has published extensively in top-tier AI, statistics, and applied mathematics venues, and his work has been recognized by multiple IBM Outstanding Technical Achievement Awards and the SIAM Student Paper Prize. He earned his Ph.D. in Computer Science from the University of Minnesota and his B.S. in Mathematics with honors from Zhejiang University.

Location: NCS 120

This workshop is intended for researchers, practitioners, students, and industry professionals in AI, robotics, machine learning, human-robot interaction, and related fields.

Workshop Overview:

Instead of learning from data alone, an embodied AI system learns through its movements, sensors, and interactions with the environment. This form of active, experience-based learning, informed by ongoing self-evaluation of its own abilities, enables embodied AI systems to adapt on the fly, understand context rather than just commands, and collaborate with humans in more natural and trustworthy ways.

Workshop Goals:

  1. Foster interdisciplinary dialogue across AI, robotics, and cognitive science.
  2. Identify key challenges and future research directions in embodied intelligence.
  3. Examine the role of embodiment in advancing toward AGI.

This workshop is Invitation-only. Please email Dr. IV Ramakrishnan (ram@cs.stonybrook.edu) to attend.

Read the announcement: https://mcusercontent.com/237207911c0fd4c1f78dd8524/files/070dec2e-a2f5-143e-0fe2-c4ebecdb5193/Embodied_AI_Workshop_Invitation_.pdf

Title: AI-Driven Target Selection Methods for Touch and Gaze Input

Abstract: Accurately selecting targets is an essential aspect of  Human-Computer Interaction. Erroneous selections can cause tedious undo and redo actions. Additionally, some selection errors are non-reversible and can lead to undesirable consequences. However, high-accuracy target selection remains a challenge on touchscreen devices due to the small target size and imprecise touch inputs, and in gaze interaction because of the gaze tracking noise and no easy-to-use selection action. We first propose ReLM, a Reinforcement Learning-based Method for touchscreen target selection. ReLM can automatically show suggestions and require a second touch if the input is ambiguous, and can directly select a target candidate when the input is certain. Our empirical evaluation shows that ReLM reduces the error rate from 6.92% to 1.63%, and the selection time from 2.23s to 1.59s over Shift, an existing suggestion-based method. Compared to BayesianCommand, a direct selection-based method, our ReLM reduces the error rate from 3.64% to 0.89%, while increasing the selection time by only 200 ms. Secondly, we investigate how to improve target selection performance for gaze interaction. We propose BayesGaze, an eye-gaze based target selection method. It accumulates the signal of each gaze point for selecting a target calculated by Bayes Theorem, and uses a threshold mechanism to determine the target selection. Our investigation shows that BayesGaze improves target selection accuracy and speed over a dwell-based selection method, and the Center of Gravity Mapping method.

All are welcome. Here  is the zoom meeting link:
https://stonybrook.zoom.us/j/93130953411?pwd=Rm5IRlVPQ3M0cHJsTXpCVFljUlFGUT09Meeting ID: 931 3095 3411Passcode: 999413

The AI Innovation Institute cordially invites you to the Summer Symposium this Friday, July 31st, from 10:00 AM to 12:00 PM in the Stony Brook Union Ballroom.



Poster Presentations: 10:00 - 11:30 am - Union Ballroom
Closing remarks: (11:25) Dr. Carl Lejuez, Executive Vice President & Provost
Group Photos (11:30-12)

Featuring undergraduate researchers/participants of the following:
  • AI Innovation & Diffusion Research Experience for Undergraduates (REU)
  • Explorations in STEM
  • Frances Velay Fellowship Program
  • Oncology Research and Clinical Learning Experience (ORACLE)
  • SUNY EOP
  • SUNY SOAR
  • URECA REACT/Research Entry Accelerator for College Transfers
  • URECA Summer

Join us in celebrating the hard work and accomplishments of our students.
Abstract: The rapid growth of observational data presents unprecedented opportunities to enhance both the predictability and mechanistic understanding of Earth systems. However, fully harnessing big Earth data needs computational frameworks that bridge the gap between physics-based models and machine learning. In this talk, I will first demonstrate how AI methods can significantly improve the prediction of environmental systems. Despite their predictive accuracy, machine learning models often lack physical interpretability, limiting their ability for scientific inquiry. To address this, I will introduce the developed hybrid, differentiable modeling framework that unifies physical models with machine learning in an end-to-end trainable system. This framework autonomously learns from large observations while maintaining physical clarity. The machine learning components can be seamlessly embedded into physical backbones to assimilate multi-source data, support automatic parameterization, and represent uncertain processes. I will showcase applications of this framework in simulating and understanding the terrestrial water cycle and its interactions with ecosystems at continental and global scales. This talk will highlight how differentiable modeling not only improves the modeling ability in both data-rich and data-scarce scenarios, but also provides a systematic pathway to enhancing model structures, deciphering uncertain physical relations, and facilitating knowledge discovery in Earth system sciences.


IACS Seminar Speaker: Dapeng Feng, Stanford Univeristy

Location: IACS Seminar Room

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Embodied Intelligence at Scientific User Facilities

Abstract: This presentation explores the active work integrating artificial intelligence and robotics at the National Synchrotron Light Source II, and a perspective for the future. Through various case studies, we highlight the optimization of operations, improved experimental outcomes, and the orchestration of distributed multimodal experiments. This ongoing development includes collaborators from across the light and neutron sources in the DOE complex. We will elaborate on the open-source Bluesky project, and its capabilities to support adaptive and autonomous experiments. Additionally, we will discuss how Bluesky can be integrated with open-source robotic control software to unlock new flexible automation for autonomous scientific research, which scales to new experiments and continues to leverage human ingenuity.

Biography: Dr. Phillip M. Maffettone is an Associate Computational Scientist in the Data Science and Systems Integration Division at NSLS-II. His research focuses on accelerating scientific discovery at user facilities through the integration of robotics, artificial intelligence (AI), and advanced experiment orchestration systems. He leads the N3XTware project, constructing the software architecture for the next 12 beamlines to be built at NSLS-II. Prior to this he built the brain on the world's first mobile robotic scientist at the University of Liverpool, and later spearheaded the machine learning platform for a biotechnology start-up, BigHat Biosciences. He holds a DPhil in Inorganic Chemistry from the University of Oxford and a B.S. in Chemical Engineering from the University at Buffalo.

Location: CDS, Bldg. 725, Training Room

Link: https://bnl.zoomgov.com/j/16049713 31?pwd=nc5CV3cOFrdYxordFieP W07tIDmwYb.1

Meeting ID: 160 497 1331
Passcode: 289875

Abstract: Visual generation is a fundamental problem in computer vision and graphics, with applications ranging from 3D capture to content creation and image/video synthesis. Despite rapid progress in neural rendering and generative models, efficiency remains a key obstacle in practice: high-quality 3D reconstruction often depends on dense multi-view supervision; scalable 3D synthesis faces heavy optimization, training, and rendering costs; and modern image/video generators incur substantial computation as token grids grow with spatial resolution and temporal length.
This thesis targets efficient visual world modeling by improving sample efficiency in 3D reconstruction, representation efficiency in 3D generation, and computational efficiency in image/video synthesis. First, we improve sample efficiency for neural implicit surface reconstruction under sparse views by integrating multi-view stereo probability volumes as a geometric regularizer, enabling high-quality reconstruction from as few as three input images. Next, we introduce an explicit 3D representation for 3D generation, built from multi-view depth and RGB predictions with 3D Gaussian features, which enables the use of 2D generative priors while enforcing multi-view consistency via epipolar attention. We then address the computational bottleneck of image and video synthesis with importance-based token merging, using importance signals available during generation to preserve critical information while merging redundant tokens. Finally, we propose efficient mixed-resolution diffusion transformers via cross-resolution phase-aligned attention, aiming to improve attention stability under mixed token grids and support high-fidelity mixed-resolution generation.

Speaker: Haoyu Wu

Location: NCS120