The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 10 minutes) by multiple people. Students can register for 1 credit for CSE 656. Registered students must attend and present a minimum of 2 or 3 talks. Everyone else is welcome to attend. Fill in https://forms.gle/pCVXovgfMfQwGqG38 to subscribe to our mailing list for further announcement.
Join us at the Center for Excellence in Learning and Teaching (CELT) for an interactive Zoom workshop on Generative AI designed for faculty and staff interested in enhancing teaching and assessment practices, increasing student engagement, and navigating the rapidly evolving landscape of AI tools. Participants will be introduced to common AI tools, explore potential instructional uses, and discuss key considerations such as academic integrity, transparency, and equity.

Register now: https://stonybrook.zoom.us/meeting/register/6js1eP64T1ys8tyU57EJ7Q#/registration

Abstract: Pre-trained diffusion and flow matching models have made visual generation remarkably powerful, enabling high-fidelity synthesis of images and videos from natural language prompts. However, their behavior is still largely dictated by the pre-training data distribution and likelihood objective, which do not directly encode downstream desiderata such as fine-grained semantic alignment, controllability, or realism. This gap motivates post-training: starting from a base generator and further optimizing it with additional supervision signals derived from human or reward model preferences.This work presents post-training for visual generative models through two complementary case studies. First, Hummingbird addresses the problem of fine-grained contextual alignment in image-text-to-image generation. We introduce a multimodal context evaluator that scores the consistency between rich contextual descriptions and generated images, capturing fine-grained alignment beyond global CLIP similarity. By directly backpropagating these differentiable rewards through the diffusion sampler, Hummingbird substantially improves semantic faithfulness while preserving high visual quality.
Second, PISCES tackles post-training for text-to-video generation, where alignment is inherently semantic-spatio-temporal. We show that naive VLM-based rewards suffer from distributional mismatch and token-level misalignment, leading to reward hacking and suboptimal optimization. PISCES introduces a bi-objective, Optimal Transport (OT)-aligned reward module: distributional OT using Neural Optimal Transport to align text and video embedding distributions, and discrete, partial OT over a spatio-temporal cost matrix to capture semantic alignment at the token level. These rewards are integrated into both direct backpropagation and GRPO-style optimization to post-train state-of-the-art text-to-video generators. Together, Hummingbird and PISCES provide a unified view of how carefully designed visual reward models, coupled with OT-based representation alignment, can reliably improve the downstream behavior of pre-trained image and video generators.

Speaker: Minh Quan Le

Location: NCS 220

Zoom: https://stonybrook.zoom.us/j/94798224254?pwd=CFraer25qnpORbJ14aAVHRwaSJOjJM.1
Abstract: Sea ice is crucial to Earth's climate, Arctic communities, and ecosystems, yet climate change is driving significant losses, threatening polar stability. Quantifying the long-term impacts of a declining sea ice cover requires tools which improve climate-timescale prediction and bring new understanding of climate interactions. In this talk, I discuss how meeting this challenge requires a multi-disciplinary approach. Climate models, while essential, suffer from systematic biases due to missing or inaccurate physics, leading to uncertainty in future projections. I show how data assimilation (DA) offers a statistical framework for integrating satellite observations with climate models to quantify systematic sea ice model errors. Using convolutional neural networks (CNNs), we can learn these errors based on the model's atmospheric, oceanic, and sea ice conditions--what I term a state-dependent representation of the error. This approach enables real-time corrections to subsequent model simulations, which systematically reduces global sea ice biases. I highlight key successes and challenges in developing this hybrid ML+climate modeling framework, including transfer learning to enhance online generalization of ML models, and new methods for integrating Python-based ML frameworks with Fortran climate model code. Finally, I introduce GPSat, a scalable Gaussian process-based tool for reconstructing complete sea ice fields from sparse satellite altimetry data. Together, the DA+ML framework and GPSat offer future opportunities for improving targeted model physics errors for more robust climate simulation.


IACS Seminar Speaker: William Gregory, Princeton University

Location: IACS Seminar Room
Prof. Eugene A. Feinberg, from the Department of Applied Mathematics and Statistics, presents, Recent Developments in Markov Decision Processes Relevant to AI on April 4 at 4p. The talk discusses recent developments in Markov Decision Processes potentially relevant to artificial intelligence. These developments include complexity estimations for exact and approximate algorithms, decision making with incomplete information and multiple criteria, and continuity properties of optimal values and expectations. Dr. Eugene A. Feinberg is currently Distinguished Professor in the Department of Applied Mathematics and Statistics at Stony Brook University. He is an expert on applied probability, stochastic models of operations research, Markov decision processes, and on industrial applications of operations research and statistics. He has published more than 150 papers and edited the Handbook of Markov Decision Processes. His research has been supported by NSF, DOE, DOD, NYSTAR (New York State Office of Science, Technology, and Academic Research), NYSERDA (New York State Energy Research and Development Authority) and by industry. He is a Fellow of INFORMS (The Institute for Operations Research and Management Sciences) and has received several awards including 2012 IEEE Charles Hirsh Award for developing and implementing smart grid technologies, 2012 IBM Faculty Award, and 2000 Industrial Associates Award from Northrop Grumman. Dr. Feinberg is an Associate Editor for Mathematics of Operations Research and for Applied Mathematics Letters. He is an Area Editor for Operations Research Letters. Refreshments will be provided

The AI Symposium at SUNY Upstate Medical University will bring together researchers, clinicians, data scientists, and industry partners to explore the transformative role of artificial intelligence in personalized healthcare. The event will highlight innovations at the intersection of machine learning, clinical informatics, computational biology, and medical imaging aimed at improving diagnosis, treatment, and patient outcomes.
The program will feature keynote talks from national leaders, panel discussions, and research presentations addressing AI-driven biomarker discovery, multimodal data integration, predictive modeling, and clinical decision support systems. Discussions will also emphasize responsible AI adoption, health equity, and translating computational advances into clinical practice.
The symposium will showcase the growing AI research ecosystem at SUNY Upstate and regional institutions through poster presentations from faculty, trainees, and collaborators. By fostering collaboration across academia, healthcare, and industry, the event aims to accelerate the development of AI technologies that enable data-driven, patient-centered care.

This symposium is free and open to the public; however, registration is required.

Abstract: Autonomous systems, whether on Earth or in space, rely on 3D perception to understand and interact with the world around them. Yet traditional techniques for 3D understanding often depend on human designed features, fixed sensors, and conventional imaging modalities. This constrained approach can limit every stage of perception, from sensing to interpretation to decision making.
In this talk, we'll explore an alternative paradigm for imaging: physically based neural representations for 3D scenes and 3D sensing systems. We will discuss how recent advances in large scale learned representations can be used to jointly optimize both 3D scene models and the design of sensing systems for 3D capture, with the goal of enabling task specific perception systems.
Unlike modern AI models trained on internet scale datasets, these specialized 3D representations typically operate in data sparse regimes and therefore require a different kind of prior. We'll examine how grounding these learned representations in the physics of light transport can improve our understanding of scene structure, and inform imaging system design even with limited data. By connecting physical insights with learned representations, we'll highlight new possibilities for robust, efficient, and adaptive perception in challenging environments.

Speaker: Nikhil Behari is a graduate student in the Camera Culture group at the MIT Media Lab, advised by Professor Ramesh Raskar. His research interests include computational imaging, 3D scene understanding, and multi-agent decision-making under uncertainty, with a focus on automating imaging system design for 3D perception in human and planetary health. His research is supported by the NASA Space Technology Graduate Research Fellowship. He received his bachelor's in Computer Science and Statistics from Harvard University in 2022.

Imagine machines that can see beyond human limitations--drones locating hidden survivors, cameras predicting structural failures, or medical devices detecting tumors beneath the skin. Traditional vision systems are constrained by the boundaries of human perception, missing vast information present in light interactions. This talk explores the development of advanced vision systems that capture underutilized dimensions of light, model intricate light-scene interactions, and extract hidden 3D information--around corners, beneath surfaces, and at high speeds. By jointly developing novel imaging hardware, efficient rendering models, and physics-based learning algorithms, we aim to transcend conventional vision capabilities--unlocking critical applications in autonomous navigation, structural monitoring, and non-invasive medical imaging.

Speaker Bio:


Akshat Dave is a Postdoctoral Associate at MIT Media Lab in the Camera Culture group working with Prof. Ramesh Raskar. He received his Ph.D. from Rice University ECE Department in 2023 where he was advised by Prof. Ashok Veeraraghavan. His research lies at the intersection of applied optics, computer graphics, and computer vision. His research focuses on developing vision systems that go beyond human perception. His work has been recognized by Rice University's Best Thesis Award, OSA Best Paper Prize, and fellowships by Texas Instruments and Qualcomm.
How to Succeed in Language Design Without Really Trying presented by Professor Brian Kernighan

ABSTRACT: Why do some languages succeed while others fall by the wayside? I've helped create nearly a dozen languages (mostly small) over the years; a handful are still in widespread use, while others have languished or simply disappeared. I've also been present at the creation of several other languages, including some really major ones. In this talk I'll give my humble, but correct, opinion on factors that affect success and failure, and try to offer some insight into what to do if you're trying to design a new language yourself, and why that might be a good thing.

BIO: Brian Kernighan received a PhD in electrical engineering from Princeton in 1969. He joined the Computer Science department at Princeton in 2000, after many years at Bell Labs. He is a co-creator of several programming languages, including AWK and AMPL, and of a number of tools for document preparation. He is the co-author of a dozen books and some technical papers, and holds 5 patents.
He is a member of the National Academy of Engineering and of the American Academy of Arts and Sciences. His research areas include programming languages, tools and interfaces that make computers easier to use, often for non-specialist users. He has also written two books on technology for
non-technical audiences: Understanding the Digital World in 2017 and Millions, Billions, Zillions: Defending Yourself in a World of Too Many Numbers, published in 2018. His most recent book, Unix: A History and a Memoir, was published in October 2019.