Discover how U.S. Census Bureau Tools can help you find free data for your research projects, community, and more. See how to access the latest American Community Survey and 2020 Census data for various geographies including New York City and Long Island at data.census.gov. Learn about Community Resilience Estimates and how to navigate My Community Explorer; an interactive map-based tool which highlights demographic and socioeconomic data that measure inequality. This session will involve live demonstrations and hands-on exercises for participants. Registrants will receive the Zoom link one day prior to the event.

Please Register for SBU Libraries' AI Club: Exploring Census Data here.
Abstract: Sub-grid turbulence is challenging to resolve in climate models; therefore, it is parameterized. Traditionally, turbulent parameterizations have relied on physics-based and equation-based approaches. However, ad hoc and uncertain components in these parameterizations introduce uncertainty in future climate predictions. Recently, data-driven techniques have emerged as an alternative for modeling sub-grid fluxes. I will demonstrate the use of machine learning to model vertical turbulent fluxes in the ocean surface boundary layer and its impact on reducing biases in NOAA's Geophysical Fluid Dynamics Laboratory ocean climate model.

I will show how neural networks, trained to predict the eddy diffusivity profile from high-fidelity yet computationally expensive turbulence schemes, enhance the vertical mixing scheme in the climate model. These networks replace ad hoc components while maintaining the conservation principles of the standard ocean model equations. The enhanced scheme outperforms its predecessor by reducing biases in the mixed-layer depth and modestly improving tropical upper-ocean stratification in ocean-only global simulations. Furthermore, simplified equations that can replace the neural networks show similar improvements but with lower computational cost and better interpretability. They point to structural deficiencies in the baseline parameterization. This work is one of the first successful applications of machine learning to improve a sub-grid parameterization of turbulent mixing in ocean climate models.

IACS Seminar Speaker: Aakash Sane, Princeton University

Location: IACS Seminar Room or Zoom

Join Zoom Meeting: https://stonybrook.zoom.us/j/97764942108?pwd=MzCWupCe3L9mKdrgfO2bJg3GBbvXuf.1
Meeting ID: 977 6494 2108
Passcode: 519324
Le Hou Dissertation Defense: Deep Learning for Digital Histopathology across Multiple Scales

ABSTRACT: Histopathology is the study of tissue changes caused by diseases such as cancer. It plays a crucial role in disease diagnosis, survival analysis and development of new treatments. Using computer vision techniques, I focus on multiple tasks for automated analysis in digital histopathology images, which are challenging because histopathology images are heterogeneous and complex, due to the large variation of hundreds of cancer types in gigapixel resolution. In this thesis, I show how histopathology image analysis tasks can be viewed in three scales: Whole Slide Image (WSI)-level, patch-level and cellular-level, and present my contributions in each resolution level.

BIO: WSI-level analysis such as classifying WSIs into cancer types is challenging, because conventional classification methods such as off-the-shelf deep learning models cannot be applied directly on gigapixel WSIs due to computational limitations. I contribute a patch-based deep learning method that classifies gigapixel WSIs into cancer types and subtypes with close-to-human performance. This method is useful for computer-aided diagnosis. At patch-level, I contribute a novel method for histopathology image patch classification. On the task of identifying Tumor Infiltrating Lymphocyte (TIL) regions, the prediction result of this method correlates to the survival rate of patients. At cellular-level, I contribute novel methods for nucleus classification and roundness regression, which are interpretable features for histopathology studies. With this method, I generated a large-scale dataset of segmented nuclei, in WSIs from a large publicly available digital histopathology image dataset, to help advance histopathology research.
What AI tools are available to help with the scholarly research process? Are they helpful? What do they do and is it worth the time and energy to try them out? Join librarian Christine Fena to explore and compare established and emerging AI research tools such as Elicit, Scite, Consensus, and Undermind. The online workshop will provide a starting point to understanding what these tools are, the basics of how they work, and how AI research assistants might bring changes to your search process in the future. All are welcome!



Register here via Zoom.

The AI Community at Stony Brook University is proud to announce Datathon 2026.

Dive into data analysis and AI/ML, and get ready to build something big. In this year's underwater-themed event, enjoy a weekend of data analysis, hacking, networking, fun activities, and minigames.

Whether you're a seasoned developer, data scientist, designer, or completely new to hacking, this event is your chance to collaborate, learn data science, and create something impactful with data and AI/ML.

What is Datathon?

AI Community's Datathon is the premier data science competition at Stony Brook University, bringing together students of all skill levels for a weekend of data exploration, analysis, and innovation. Just like a typical hackathon, you will be using your skills to build your dream project.

Unlike a regular hackathon, Datathon is focused on data science. You will be given a set of data to work with, analyze, and apply to your project. You can also find your own data to use. Your project will be presented to a panel of judges consisting of professors and industry professionals!

Who Can Participate

  • Students of all skill levels and majors are welcome.
  • Come with a team or find one at the event or on Discord.
  • This event is open to SBU and non-SBU students.

* Non-SBU Undergraduate Students are ineligible to receive prizes
* You must be 18+ or older (Excludes minors who are active SBU students)

Location: SAC Ballroom B

Register here.

Abstract: Graphs are a universal language of science. Molecules, materials, quantum systems, and knowledge bases can all be naturally represented as graphs. This talk explores how graph-based artificial intelligence is emerging as a powerful engine for scientific discovery. Using molecular design as a guiding example, we examine how modern graph AI enables machines not only to analyze complex scientific structures but also to generate new ones. We will discuss graph neural networks for learning predictive models of molecular properties, graph generative models for constructing novel chemical structures, and emerging multimodal graph-language models that support inverse design and synthesis planning. Together, these advances make graph AI more scalable, interpretable, and data-efficient--key capabilities for real-world scientific discovery. As artificial intelligence enters the era of foundation models, the next frontier lies in multimodal reasoning. Scientific knowledge is not purely textual; it is expressed through structures, code, and experimental data. By integrating graph representations with large language models, we move toward AI systems that can reason across multiple modalities and engage with scientific knowledge in its native forms. Looking ahead, we envision AI systems that behave less like tools and more like collaborators in the scientific process--generating hypotheses, designing candidate structures, planning experiments, interpreting results, and iteratively refining ideas through cycles of success and failure. In this vision, multimodal and agentic AI will enable scientists to explore vast and previously inaccessible design spaces, accelerating breakthroughs across domains ranging from drug discovery and materials innovation to software systems and quantum technologies.

Bio: Jie Chen is an interdisciplinary researcher working at the intersection of computing and mathematics, with a current focus on foundation models and AI agents for scientific discovery. His research integrates machine learning, statistics, scientific computing, and numerical linear algebra, with contributions spanning graph neural networks, multimodal graph LLMs, graph structure learning, scalable Gaussian processes, graph coarsening, and matrix functions. He is widely recognized for transformative contributions to graph-based deep learning and large-scale statistical modeling, and for bridging theory with real-world scientific and engineering applications. Dr. Chen has led externally funded, multi-institutional research programs supported by Shell, Evonik, and the U.S. Department of Energy, with applications in materials discovery, financial forensics, and power system resilience. He previously served as a Senior Research Scientist and Manager at IBM Research and the MIT-IBM Watson AI Lab, and as a Postdoctoral Fellow at Argonne National Laboratory. He has published extensively in top-tier AI, statistics, and applied mathematics venues, and his work has been recognized by multiple IBM Outstanding Technical Achievement Awards and the SIAM Student Paper Prize. He earned his Ph.D. in Computer Science from the University of Minnesota and his B.S. in Mathematics with honors from Zhejiang University.

Location: NCS 120
Fall 2026, Wednesdays 2 to 3:20 pm, NCS 220 and Zoom link to be announced soon.

The seminar will be jointly taught by Prof. Dimitris Samaras (samaras@cs.stonybrook.edu).

The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision.

To enroll in this course, you must either: (1) be in the Ph.D. program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 15 minutes) by multiple students. Students can register for 1 credit for CSE656. Registered students must attend and present a minimum of 2 talks. Registered students must attend in person. Up to 3 absences will be excused. Everyone else is welcome!
AI Institute Seminar Title: A Geometric Understanding of Deep Learning Abstract: This work introduces an optimal transportation (OT) view of generative adversarial networks (GANs). Natural datasets have intrinsic patterns, which can be summarized as the manifold distribution principle: the distribution of a class of data is close to a low-dimensional manifold. GANs mainly accomplish two tasks: manifold learning and probability distribution transformation. The latter can be carried out using the classical OT method. From the OT perspective, the generator computes the OT map, while the discriminator computes the Wasserstein distance between the generated data distribution and the real data distribution; both can be reduced to a convex geometric optimization process. Furthermore, OT theory discovers the intrinsic collaborative--instead of competitive--relation between the generator and the discriminator, and the fundamental reason for mode collapse. We also propose a novel generative model, which uses an autoencoder (AE) for manifold learning and OT map for probability distribution transformation. This AE-OT model improves the theoretical rigor and transparency, as well as the computational stability and efficiency; in particular, it eliminates the mode collapse. The experimental results validate our hypothesis, and demonstrate the advantages of our proposed model.

Abstract: The faster AI automation spreads through the economy, the more profound its potential impacts, both positive (improved productivity) and negative (worker displacement). The previous literature on AI Exposure cannot predict this pace of automation since it attempts to measure an overall potential for AI to affect an area, not the technical feasibility and economic attractiveness of building such systems. In this work, we present a new type of AI task automation model that is end-to-end, estimating: the level of technical performance needed to do a task, the characteristics of an AI system capable of that performance, and the economic choice of whether to build and deploy such a system. The result is a first estimate of which tasks are technically feasible and economically attractive to automate - and which are not. We focus on computer vision, where cost modeling is more developed. We find that at today's costs U.S. businesses would choose not to automate most vision tasks that have AI Exposure, and that only 23% of worker wages being paid for vision tasks would be attractive to automate. This slower roll-out of AI can be accelerated if costs fall rapidly or if it is deployed via AI-as-a-service platforms that have greater scale than individual firms, both of which we quantify. Overall, our findings suggest that AI job displacement will be substantial, but also gradual - and therefore there is room for policy and retraining to mitigate unemployment impacts.

Details of this work can be found here.

Speaker Bio: Neil Thompson is the Director of the FutureTech research project at MIT's Computer Science and Artificial Intelligence Lab and a Principal Investigator at MIT's Initiative on the Digital Economy.

Previously, he was an Assistant Professor of Innovation and Strategy at the MIT Sloan School of Management, where he co-directed the Experimental Innovation Lab (X-Lab), and a Visiting Professor at the Laboratory for Innovation Science at Harvard. He has advised businesses and government on the future of Moore's Law, has been on National Academies panels on transformational technologies and scientific reliability, and is part of the Council on Competitiveness' National Commission on Innovation & Competitiveness Frontiers.

He has a PhD in Business and Public Policy from Berkeley, where he also did Masters degrees in Computer Science and Statistics. He also has a masters in Economics from the London School of Economics, and undergraduate degrees in Physics and International Development. Prior to academia, He worked at organizations such as Lawrence Livermore National Laboratory, Bain and Company, the United Nations, the World Bank, and the Canadian Parliament.

Location: IACS Seminar Room
Virtual Talk: Metadata Matters: Robust Document Classification via Adaptation Methods for Text-driven Public Health by Xiaolei Huang

Zoom link to follow.

Abstract: Document classifiers have been widely applied in solving health-related issues, such as suicide prevention, flu vaccination surveillance and disease diagnosis. However, document metadata including time, gender, age and location has an enormous impact on robustness of 
document classifiers. Language varies across the metadata bringing both challenges and opportunities to build reliable document classifiers. For example, online written language changes over time, and males and females express opinions differently. This talk describes how to use domain adaptation to integrate temporal and user demographic factors into document classifiers. By adapting knowledge of how language varies across the metadata, models can learn generalized representations of language through the metadata-invariant embeddings. 
This approach will lead to metadata-adapted document classifiers and can also extend to personalize classification models by user embedding. 

Bio: Xiaolei Huang is a 4th-year PhD candidate in Information Science at the University of Colorado, Boulder. He is currently a visiting scholar at the Johns Hopkins University. His research interests are in Natural Language Processing, Machine Learning and Public Health. Particularly, he focuses on domain adaptation, cross-lingual transfer learning, user modeling and fairness.