Abstract: Sea ice is crucial to Earth's climate, Arctic communities, and ecosystems, yet climate change is driving significant losses, threatening polar stability. Quantifying the long-term impacts of a declining sea ice cover requires tools which improve climate-timescale prediction and bring new understanding of climate interactions. In this talk, I discuss how meeting this challenge requires a multi-disciplinary approach. Climate models, while essential, suffer from systematic biases due to missing or inaccurate physics, leading to uncertainty in future projections. I show how data assimilation (DA) offers a statistical framework for integrating satellite observations with climate models to quantify systematic sea ice model errors. Using convolutional neural networks (CNNs), we can learn these errors based on the model's atmospheric, oceanic, and sea ice conditions--what I term a state-dependent representation of the error. This approach enables real-time corrections to subsequent model simulations, which systematically reduces global sea ice biases. I highlight key successes and challenges in developing this hybrid ML+climate modeling framework, including transfer learning to enhance online generalization of ML models, and new methods for integrating Python-based ML frameworks with Fortran climate model code. Finally, I introduce GPSat, a scalable Gaussian process-based tool for reconstructing complete sea ice fields from sparse satellite altimetry data. Together, the DA+ML framework and GPSat offer future opportunities for improving targeted model physics errors for more robust climate simulation.


IACS Seminar Speaker: William Gregory, Princeton University

Location: IACS Seminar Room

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Uncertainty-Aware Adaptation of LLMs for Protein-Protein Interaction Analysis

Abstract: Identification of protein-protein interactions (PPIs) helps derive cellular mechanistic understanding, particularly in the context of complex conditions such as neurodegenerative disorders, metabolic syndromes, and cancer. Large Language Models (LLMs) have demonstrated remarkable potential in predicting protein structures and interactions via automated mining of vast biomedical literature; yet their inherent uncertainty remains a key challenge for deriving reproducible findings, critical for biomedical applications. In this study, we present an uncertainty-aware adaptation of LLMs for PPI analysis, leveraging fine-tuned LLaMA-3 and BioMedGPT models. To enhance prediction reliability, we integrate LoRA ensembles and Bayesian LoRA models for uncertainty quantification (UQ), ensuring confidence- calibrated insights into protein behavior. Our approach achieves competitive performance in PPI identification across diverse disease contexts while addressing model uncertainty, thereby enhancing trustworthiness and reproducibility in computational biology. These findings underscore the potential of uncertainty-aware LLM adaptation for advancing precision medicine and biomedical research.

Biography: Sanket is a research staff member in the Applied Mathematics department within the Computing and Data Sciences Directorate at Brookhaven National Laboratory. Previously, he was the Amalie Emmy Noether Postdoctoral Fellow in the same department. He earned his Ph.D. in Statistics from Michigan State University.. Sanket's research interests span Bayesian statistics, uncertainty quantification (UQ), deep learning, Markov Chain Monte Carlo (MCMC), variational inference, and sparsity methods. He also focuses on dimensionality reduction, surrogate modeling, hybrid physical-data driven models, and active learning, with applications across climate science, materials science, and life sciences.

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1605691898?pwd=xC7GebG7Kvzxa4AjPSIxJw7e9IZtoY.1

Meeting ID: 160 569 1898
Passcode: 303888

The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors. Each seminar will consist of multiple short talks (around 15 minutes) by multiple students. Students can register for 1 credit for CSE656. Registered students must attend and present a minimum of 2 talks. Everyone else is welcome to attend. Fill in https://forms.gle/q6UG9ygauLp2a8Po8 to subscribe to our mailing list for further announcement.
Abstract: Robot control has evolved from optimization-based controllers---precise but task-specific---through deep reinforcement learning's learned policies, to Vision-Language-Action (VLA) models that leverage pretrained vision-language backbones for language-conditioned manipulation across diverse tasks.
Despite their promise, VLAs exhibit a critical limitation: they function primarily as trajectory learners rather than skill learners. Recent evaluations reveal that VLAs often fail when faced with even minor variations in object initialization or environmental conditions, suggesting they memorize specific trajectories rather than acquiring generalizable manipulation skills. Attempts to address this through 3D spatial representations have shown limited success, indicating that the missing component may be more fundamental than geometric understanding alone.
This work argues that World Models (WMs)---internal representations that predict future states given actions---constitute the missing piece for robust VLA systems. We present one completed contribution and two ongoing investigations.
We developed a dual-layer world model for human-robot interaction that anticipates both physical scene evolution and latent human preferences for assistive tasks. Building on these foundations, we present ongoing work probing VLA internal representations to verify implicit world model existence, and propose a WM-VLA integration approach operating in the native visual domain through embedding prediction and image decoding.
Together, these contributions and investigations establish a foundation for WM-VLA systems, pointing toward robust, generalizable robot policies.
Speaker: Jason Qin
Location: NCS 220

Abstract: Pre-trained diffusion and flow matching models have made visual generation remarkably powerful, enabling high-fidelity synthesis of images and videos from natural language prompts. However, their behavior is still largely dictated by the pre-training data distribution and likelihood objective, which do not directly encode downstream desiderata such as fine-grained semantic alignment, controllability, or realism. This gap motivates post-training: starting from a base generator and further optimizing it with additional supervision signals derived from human or reward model preferences.This work presents post-training for visual generative models through two complementary case studies. First, Hummingbird addresses the problem of fine-grained contextual alignment in image-text-to-image generation. We introduce a multimodal context evaluator that scores the consistency between rich contextual descriptions and generated images, capturing fine-grained alignment beyond global CLIP similarity. By directly backpropagating these differentiable rewards through the diffusion sampler, Hummingbird substantially improves semantic faithfulness while preserving high visual quality.
Second, PISCES tackles post-training for text-to-video generation, where alignment is inherently semantic-spatio-temporal. We show that naive VLM-based rewards suffer from distributional mismatch and token-level misalignment, leading to reward hacking and suboptimal optimization. PISCES introduces a bi-objective, Optimal Transport (OT)-aligned reward module: distributional OT using Neural Optimal Transport to align text and video embedding distributions, and discrete, partial OT over a spatio-temporal cost matrix to capture semantic alignment at the token level. These rewards are integrated into both direct backpropagation and GRPO-style optimization to post-train state-of-the-art text-to-video generators. Together, Hummingbird and PISCES provide a unified view of how carefully designed visual reward models, coupled with OT-based representation alignment, can reliably improve the downstream behavior of pre-trained image and video generators.

Speaker: Minh Quan Le

Location: NCS 220

Zoom: https://stonybrook.zoom.us/j/94798224254?pwd=CFraer25qnpORbJ14aAVHRwaSJOjJM.1

Join the Office of Educational Effectiveness' upcoming workshop on the transformative potential of AI tools to enhance program assessment. Learn how to leverage AI to create targeted learning objectives, detailed rubrics, and precise benchmarks that will elevate the quality and effectiveness of your program assessment process. Join in-person on Oct. 17 at 10:30 am or virtually on Oct. 21 at 12 pm.

Register in advance: https://calendar.stonybrook.edu/site/office-educational-effectiveness/event/leveraging-ai-in-assessment-zoom/

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes three short talks on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Tuesday, January 7, 2025, 12:00 pm -- CDS, Bldg. 725, Training Room


Speakers

Deyu Lu
Mingyuan Ge
Kris Reyes


Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1615289117?pwd=Hqkbj9itxWrFnkhZ8rQXHPInO2gxdF.1

Meeting ID: 161 528 9117
Passcode: 991382