Abstract: Recent progress in Large Language Models (LLMs) has transformed text and code generation, yet models still falter on scientific reasoning where correctness, constraints, and physical consequences are critical. This talk explores how formal LLM reasoning can advance symbolic scientific modeling. First, our PDE-Controller formalizes informal PDEs (Partial Differential Equations), synthesizes solver-ready code, and plans subgoals to tackle nonconvex control via interactions with external solvers. Second, our Lean Finder accelerates scientific formalization via a semantics-aware search engine for Lean/Mathlib that retrieves relevant theorems, outperforming GPT models and gaining significant traction in the AI-for-math community. Through these efforts, we aim to design a semantics-first LLM that autoformalizes informal scientific problems into machine-checked specifications and synthesizes solver-ready code. This closes the loop between formal analysis and LLM reasoning, ultimately surpassing human heuristics for scientific discovery.

Bio: Dr. Wuyang Chen is a tenure-track Assistant Professor in Computing Science at Simon Fraser University. He is also a visiting research scientist at Microsoft. Previously, he was a postdoctoral researcher in Statistics at the University of California, Berkeley, advised by Professor Michael Mahoney. He obtained his Ph.D. in Electrical and Computer Engineering from the University of Texas at Austin in 2023, advised by Professor Atlas Wang. Dr. Chen's research focuses on integrating AI methods with physical knowledge, scientific machine learning, and theoretical understanding of deep networks. Dr. Chen has published papers at CVPR, ECCV, ICLR, ICML, NeurIPS, and other top conferences. Dr. Chen's research has been recognized by the US NSF newsletter, two Doctoral Dissertation Awards from INNS and iSchools, AAAI New Faculty Highlights, and NVIDIA Academic Grant Award. Dr. Chen also hosted and co-organized many conference workshops at NeurIPS, ICLR, CVPR.

Location: NCS 120

This virtual presentation series is designed to inform the Stony Brook University research community about the Research Funding Landscape of key topic areas. Our Strategic Research Initiatives team will provide insight into the rapidly shifting funding environment using policy briefs, budgetary priorities, and relevant legislation. We will highlight federal and state priorities in the current and upcoming years to help Stony Brook researchers develop strategies for pursuing funding in a rapidly shifting environment. This series is moderated by Mónica Bugallo, Interim Vice President for Research & Innovation.

Join us for the third in the series, focused on the artificial intelligence landscape:


Translating the Funding Landscape for Stony Brook Researchers: Artificial Intelligence
Presented by Catherine Chen, Ph.D., Research Development Associate
Faculty Respondent: Assistant Professor Nav Nidhi Rajput, Department of Materials Science and Chemical Engineering
Wednesday, April 22, 2026 at 2 pm to 3 pm

Registration is Required

Le Hou Dissertation Defense: Deep Learning for Digital Histopathology across Multiple Scales

ABSTRACT: Histopathology is the study of tissue changes caused by diseases such as cancer. It plays a crucial role in disease diagnosis, survival analysis and development of new treatments. Using computer vision techniques, I focus on multiple tasks for automated analysis in digital histopathology images, which are challenging because histopathology images are heterogeneous and complex, due to the large variation of hundreds of cancer types in gigapixel resolution. In this thesis, I show how histopathology image analysis tasks can be viewed in three scales: Whole Slide Image (WSI)-level, patch-level and cellular-level, and present my contributions in each resolution level.

BIO: WSI-level analysis such as classifying WSIs into cancer types is challenging, because conventional classification methods such as off-the-shelf deep learning models cannot be applied directly on gigapixel WSIs due to computational limitations. I contribute a patch-based deep learning method that classifies gigapixel WSIs into cancer types and subtypes with close-to-human performance. This method is useful for computer-aided diagnosis. At patch-level, I contribute a novel method for histopathology image patch classification. On the task of identifying Tumor Infiltrating Lymphocyte (TIL) regions, the prediction result of this method correlates to the survival rate of patients. At cellular-level, I contribute novel methods for nucleus classification and roundness regression, which are interpretable features for histopathology studies. With this method, I generated a large-scale dataset of segmented nuclei, in WSIs from a large publicly available digital histopathology image dataset, to help advance histopathology research.

Abstract: Much like other AI for Science domains, polymer design poses significant challenges. It requires grounding in empirical data and physical laws, precise handling of domain-specific structured representations, and compositional reasoning over multiple interacting constraints--all while working with limited data.

To address these limitations, we introduce PolyBench, a large-scale benchmark comprising over 125K polymer design and analysis tasks grounded in verified experimental and synthetic data. PolyBench includes tasks created from a wide range of data sources and presents diverse structural, property-driven, and synthesis-oriented reasoning problems. Tasks in PolyBench are organized from simple to complex analytical reasoning problems, enabling generalization tests and includes diagnostic probes to evaluate model capabilities. Additionally, to support effective domain alignment, we propose a knowledge-augmented reasoning distillation framework that enriches the dataset with structured chain-of-thought supervision derived from expert-informed reasoning strategies.

Small language models (7B-14B parameters) trained on PolyBench substantially outperform comparably sized baselines and, in many cases, exceed the performance of larger closed-source frontier models on polymer reasoning tasks, while also demonstrating improved transfer to external polymer benchmarks. Last, we conduct a diagnostic study that reveals a compositionality gap: despite strong performance on decomposed sub-questions, models struggle to integrate multiple interacting constraints and intermediate reasoning steps, highlighting fundamental limitations in current scientific language models.

Speaker: Dikshya Mohanty

Location: NCS 115/Online

Zoom: https://stonybrook.zoom.us/j/94746001760?pwd=BCAd8gu7cXLn3PXM6kkbh11V6r0Mr7.1
Meeting ID: 947 4600 1760 Passcode: 987917

Join us as we celebrate this year's Brook & Beyond Challenge finalists.
The Office for Research and Innovation invites you to hear about the two-month journey in which the Brook & Beyond team supported eight cohorts in bringing their bold ideas from the lab to the marketplace. It's an energizing evening that highlights the collaboration, creativity, and entrepreneurial spirit driving discovery across the University.
Meet this year's award recipients, hear pitches from the emerging founders, and applaud their achievements.
Connect, celebrate, and be part of the momentum shaping the future of innovation at
Stony Brook University.
Refreshments will be served. Registration is required.
Register Here.
CG Group member (and SBU faculty) Chao Chen will speak on Fri, March 12, about the use of topological data analysis in machine learning for image analysis.
Chao has shared some of his research with the CG Group previously, and this will be a great opportunity to learn more about this exciting research area related to computational geometry/topology!

Time: Friday, March 12, 2pm-3pm
Place: Zoom
https://stonybrook.zoom.us/my/profweizhu?pwd=RjVIVXg3YUhudzZZQ3pheHUydTJBUT09



Title: Learning with Topological Information - Image Analysis and Label Noise
Speaker: Prof. Chao Chen (SBU)

Abstract: Modern machine learning faces new challenges. We are
analyzing highly complex data with unknown noise. Topology provides
novel structural information to model such data and noise. In this
talk, we discuss two directions in which we are using topological
information in the learning context. In image analysis, we propose a
topological loss to segment and to generate images with not only
per-pixel accuracy, but also topological accuracy. This is necessary
in analysis of images of fine-scale biomedical structures such as
neurons, vessels, etc.  Extracting these structures with correct
topology is essential for the success of downstream
analysis. Meanwhile, we discuss how to use topological information to
train classifiers robust to label noise. This is important in practice
especially when we are using deep neural networks which tend to
overfit noise. These results have been published in NeurIPS, ECCV,
ICML and ICLR.
AI + Music Seminar - The meeting will consist of introductions and organizational discussions, aimed at understanding participants' interests. We'll discuss what the seminars can focus on going forward.

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Experiencing Machine Learning in Collider-Accelerator Control System

Abstract: The Relativistic Heavy Ion Collider (RHIC) at Collider-Accelerator Department (C-AD) of BNL provides the world's only high-energy polarized proton beam. It is in the unique position to study where nuclei obtain their spin. During 25 years of operation at RHIC, the C-AD controls group has developed its own control system to tune the accelerator performance, which contains millions of control points. The successful operation of this system will highly affect the machine performance. RHIC's successor, the Electron-Ion Collider (EIC), will be one of the most complex scientific instruments ever built, with the capability of colliding polarized proton and electron beams. The increasing complexity of instruments will require new, sophisticated control methods/tools to tune and optimize the accelerator performance. In this talk, I will summarize some projects developed in recent years that utilize machine learning in the C-AD controls group.

Biography: Dr. Yuan Gao is an assistant scientist at the Collider-Accelerator Department (C-AD) at Brookhaven, primarily working on developing new machine learning schemes in the control group to enhance system performance. His research interests include game theory, algorithm design, anomaly detection, and simulation modeling.

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1604302440?pwd=0x2I95PIvbkkzIi6rA0MNnon5k2sux.1

Meeting ID: 160 430 2440
Passcode: 478223