University Libraries Presents the Data Stacks SBU Libraries' workshop series for building practical data skills, led by Ahmad Pratama, the SBU Libraries Data Literacies Lead.

Participants will learn practical techniques for identifying & fixing common data quality problems using Python and/or R. The workshop demonstrates how Generative AI can streamline the data cleaning process, while highlighting potential pitfalls & teaching strategies for accuracy & data integrity.

Register here to attend.











Abstract:
Quantifying similarity is a central notion in science and data analysis, pervading everything from phylogenetic trees to the foundation of clustering. Unfortunately, despite being examined and applied for decades, traditional similarity and distance metrics have fundamental drawbacks. The key problem is that all of them are only defined over pairs of objects, so they scale quadratically when one tries to compare N objects. The present explosion in the amount of data available to us requires new ways to process information, and while some current algorithms can handle millions of points, we need alternatives applicable to billions. This is what motivated us to develop a new framework that can compare any number of objects at the same time. With this, we achieve an unprecedented linear scaling when comparing multiple objects. Here we will discuss the main properties of this formalism, along with its applications in drug design and to the analysis of Molecular Dynamics (MD) simulations. Our indices have proven to be incredibly versatile when applied to chemical space exploration and visualization, allowing us to rigorously quantify the chemical diversity of very large molecular libraries. This has led to the creation of several algorithms to sample important regions in chemical space, including a more efficient way of identifying the prevalence of activity cliffs. Additionally, our indices provide a convenient route to sample complex MD trajectories, allowing to identify representative structures very efficiently. Moreover, we can also cluster biological ensembles in a more robust way than with standard algorithms, which has led to our group's work on MDANCE, a very flexible and efficient open-source clustering module. Drop by if you want to know how we clustered one billion molecules!


Speaker:
Assistant Professor, Department of Chemistry and Quantum Theory Project
University of Florida, Gainesville
Website: https://quintana.chem.ufl.edu/

Location:
Laufer Center Lecture Hall 101

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

We meet once a month at noon in CDSD's Training Room (building 725, room 2-124) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Multilayer machine-learning framework for screening catalytic activity and selectivity

Abstract: Machine learning (ML) studies based on quantum chemical datasets have emerged as a powerful tool for accelerating catalyst discovery. However, their potential applications remain limited by high costs, low data quality, and restricted reliability. Here, we report a multilayer binary classification framework (MLBC) for screening catalytic performance in multistep processes with ML models. Key features include low-cost synthetic data generation via kinetic Monte Carlo simulations, high-quality data that capture catalytic behavior under reaction conditions, robust treatment of imbalanced data distributions, and descriptor selection that enhances reliability and interpretability. Using carbon dioxide (CO 2 ) hydrogenation to methanol (CH 3 OH) on copper (Cu)-based catalysts as a case study, the MLBC framework outperforms conventional ML models by demonstrating high reliability and strong generalization in classifying systems with activity and methanol selectivity exceeding those of Cu. In addition, feature analysis reveals that control over the critical transition steps between competing pathways governs activity and selectivity in CO 2 hydrogenation.

In addition to our speaker, we will have a number of CDS staff in attendance with expertise in AI methods and applications including image analysis, foundation models development, and inverse problem solving.

Biography: Dr. Wenjie Liao is a research associate in the Chemistry Division at Brookhaven National Laboratory, where he develops computational and machine-learning approaches to understand and improve heterogeneous catalysts. He earned his Ph.D. from Stony Brook University. His research focuses on identifying active sites, mapping reaction pathways, and predicting catalyst performance, with the goal of accelerating the discovery of more efficient catalysts for energy and chemical transformations.

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1604383624?pwd=ffQ5cUPNxTI7nzClKQO6cnsNbhF9Vf.1

Meeting ID: 160 438 3624
Passcode: 558449

Please Note: Due to a funding shortfall, we are for the time being no longer able to provide pizza and sodas for these events. We will have coffee though, and all are of course welcome to bring their lunch.

ICB&DD 19th Annual Symposium

Iwao Ojima, Director, ICB&DD
Ivet Bahar Chair, Organizing Committee
Dima KozakovCo-Chair, OrganizingCommittee

There will be poster sessions on projects conducted in the ICB&DD member's laboratories aswell as other laboratories in the area. Awards will be given to the best three posters.

Please see the link for the registration and poster sessions in:
https://www.stonybrook.edu/commcms/icbdd/https://forms.gle/Wh4UzVx9U4HWStXb8
CSE 600 Seminar Series | Fall 2025


Abstract: Large reasoning models have demonstrated capabilities to solve competition-level math problems, answer deep research questions, and address complex coding needs. Much of this progress has been enabled by scaling of data: pre-training data to learn vast knowledge, fine-tuning data to learn natural language reasoning, and RL environments to refine that reasoning. In this talk, I will describe the current LLM reasoning paradigm, its boundaries, and the future of LLM reasoning beyond scaling. First, I will describe the state of reasoning models and where I think scaling can lead to some additional (though perhaps limited) successes. I will then shift to discussing more fundamental issues with models that scale will not resolve in the next few years. I will touch on four current limitations: outdated knowledge, generator-validator gaps, limited creativity, and poor compositional generalization. In all cases, fundamental limitations of LLMs or of supervised learning in general make these problems challenging, inviting future study and novel solutions beyond scaling.

Bio: Greg Durrett is an associate professor in the Department of Computer Science and the Center for Data Science at New York University. His research is broadly in the areas of natural language processing and machine learning. Currently, his group's focus is on reasoning about knowledge in text, verifying correctness of generation methods, and studying how to make progress on problems that defy LLM scaling. He is a 2023 Sloan Research Fellow and a recipient of a 2022 NSF CAREER award. He has served in numerous roles for ACL conferences, recently as a member of the NAACL Board since 2024 and as Senior Area Chair for ACL 2025 and EMNLP 2025. He received his BS in Computer Science and Mathematics from MIT and his PhD in Computer Science from UC Berkeley, where he was advised by Dan Klein.
Join librarian Christine Fena for an interactive workshop that invites you to explore AI tools firsthand, not just as users, but as critical investigators. Through playful experimentation and collaborative discovery, you'll uncover inherent biases, probe algorithmic flaws, and gain a deeper understanding of AI's limitations and societal impacts.

Location: Melville Library, Central Reading Room, Lab B

https://library.stonybrook.edu/library-events/critiquing-ai/
Come learn of the exciting research being done across so many fields using AI! The recipients of AI3's seed awards will present their work in our showcase on November 17, 2025 and we would love to see you there!

The schedule is listed below.

Location: New Computer Science Room 120

Session 1 - 10:30 AM to 11:45

Kevin Reed, PI, Introducing the AI Techniques in Assessing the Future Changes of Extreme Precipitation and Associated Flood Risks
Co-PIs: Tangnyu Song, Ishrat Dollan
Consultant: Jayesh Rathi

Ruwen Qin, PI, AI-Assisted Analysis of Materials in Recycling Streams
Consultant: Vismay Vora

Giuseppe Gazzola, PI, Using AI to Investigate National Literatures: Italy, France, Spain 1733- 1794
Consultant: Jayesh Rathi

Joseph Lemelin, PI, IAE2^3: AI Ecologies
Co-PIs: Katherine Johnston, Aruna Balasubramanian, Matthew Salzano

Niranjan Balasubramanian, Co-PI, Molecular Foundations for Sustainability: Data Analytics for Sustainable Cellulose Scaffolding Modifications to Remediate Diverse Water Contamination Challenges
PI: Benjamin Hsiao, Co-PI: I. V. Ramakrishnan

Owen Rambow, PI,Achieving Common Ground Through Language and Vision in Mixed-Initiative Human-Machine Communication Via zoom
Co-PI Susan Brennan

Session 2 - 12:30 PM to 1:45

Jack McSweeney, PI, Developing Machine Learning Approaches to Classify Internal Waves
Consultant: Vismay Vora

Eric Josephs, PI, Learning Design Rules to Personalize Precision CRISPR Gene Therapies with Interpretable AI
Consultant: Deboparna Banerjee

Shyam Sharma, PI, Fostering Writing-to-Learn Skills through Critical AI Literacy: A Faculty Development and Student Support Program
Co-PIs: Rose Tirotta-Esposito, Christine Fena

Ritwik Banerjee, PI, A Pragmatic Approach to AI for Digital Media Integrity: Combating Complex Misinformation Through Fallacies and Propaganda
Co-PI: Ruobing Li

Ziyu Shu, Co-PI, Novel Clinical Applications of Deep Image Prior-based CT Image Reconstruction
PI: Xin Qian, Co-PIs: Tiezhi Zhang, Zhaozheng Yin

Prateek Prasanna, Co-PI, An Artificial Intelligence-Driven Clinical Decision Support Tool for the Management of Abdominal Aortic Aneurysm
PI: Apostolos Tassiopoulos, Co-PI's: Mary Saltz, Janos Hajagos, Tahsin Kurc



Presenters will give a 5-minute talk with 2 minutes for Q & A.
Presented by Stony Brook University Department of Biomedical Informatics and Long Island Network for Clinical and Translational Science (LINCATS).

The seminar aims to empower participants with the knowledge and skills necessary to harness AI effectively in clinical practice and research. It will equip attendees with practical insights, case studies, and interactive discussions led by experts in both AI and medicine, fostering a collaborative environment where attendee can explore how to overcome barriers and maximize the potential of AI in transforming modern healthcare delivery.

All Stony Brook Audiences Welcome.
Please note: This exciting event is open to all Stony Brook Faculty/Staff/Students. While the overarching theme for this event is the application of AI in medicine, the event is designed to bridge the professional practice gap that exists between cutting-edge AI research and its practical implementation in clinical settings, While AI holds immense promise for transforming healthcare delivery, many physicians and researchers lack the foundational knowledge and practical skills needed to effectively integrate AI into their daily practices.

THIS CONFERENCE IS FOR STONY BROOK UNIVERSITY & HOSPITAL FACULTY/STAFF & STUDENTS ONLY.


Registration link: https://cme.stonybrookmedicine.edu/continuing-medical-education/conferences/235/bench-to-bedside-understanding-the-practical-application-of-ai-in-medicine-2024/10/17/2024

FOR QUESTIONS
joseph.cesaria@stonybrookmedicine.edu
mary.saltz@stonybookmedicine.edu
Climate Uncertainty, Decision Making, and AI for Earth System Predictability Dr. Nathan Urban, Brookhaven National Laboratory

Bio: Nathan Urban is the group leader of the Optimal Experimental Design & Uncertainty Quantification group in the Applied Mathematics Department at Brookhaven National Laboratory's Computing & Data Sciences directorate (CDS). He holds a Ph.D. in theoretical condensed matter physics from Penn State, and has previously held research positions at Los Alamos National Laboratory, Princeton, and Penn State. His research interests include Bayesian inference and spatiotemporal statistics, probabilistic prediction and forecasting, multi-model / model-form / model structural uncertainty quantification, reduced order modeling, scientific machine learning and hybrid physical-data driven modeling, in-situ/streaming data analysis at scale, information fusion, decision making under uncertainty and optimal experimental design, and integrated multiscale computational frameworks for decision support.

Location: IACS Seminar Room

Lunch will be provided
Prepare your Business for the AI-driven future with DocItUSA's Document Management Solutions.
In today's digital world, businesses need to leverage the benefits of the Digital Cloud Age by streamlining Document Organization, Storage, and Accessibility.
Michael Feingold, of Digital Onesource Consulting Solutions/DOCS Consulting, Inc. will show you how DocItUSA equips your company with the tools to efficiently capture, classify, and retrieve documents and enable seamless AI integration.
https://nysbdc.ecenterdirect.com/events/1019400