Abstract: Pre-trained diffusion and flow matching models have made visual generation remarkably powerful, enabling high-fidelity synthesis of images and videos from natural language prompts. However, their behavior is still largely dictated by the pre-training data distribution and likelihood objective, which do not directly encode downstream desiderata such as fine-grained semantic alignment, controllability, or realism. This gap motivates post-training: starting from a base generator and further optimizing it with additional supervision signals derived from human or reward model preferences.This work presents post-training for visual generative models through two complementary case studies. First, Hummingbird addresses the problem of fine-grained contextual alignment in image-text-to-image generation. We introduce a multimodal context evaluator that scores the consistency between rich contextual descriptions and generated images, capturing fine-grained alignment beyond global CLIP similarity. By directly backpropagating these differentiable rewards through the diffusion sampler, Hummingbird substantially improves semantic faithfulness while preserving high visual quality.
Second, PISCES tackles post-training for text-to-video generation, where alignment is inherently semantic-spatio-temporal. We show that naive VLM-based rewards suffer from distributional mismatch and token-level misalignment, leading to reward hacking and suboptimal optimization. PISCES introduces a bi-objective, Optimal Transport (OT)-aligned reward module: distributional OT using Neural Optimal Transport to align text and video embedding distributions, and discrete, partial OT over a spatio-temporal cost matrix to capture semantic alignment at the token level. These rewards are integrated into both direct backpropagation and GRPO-style optimization to post-train state-of-the-art text-to-video generators. Together, Hummingbird and PISCES provide a unified view of how carefully designed visual reward models, coupled with OT-based representation alignment, can reliably improve the downstream behavior of pre-trained image and video generators.

Speaker: Minh Quan Le

Location: NCS 220

Zoom: https://stonybrook.zoom.us/j/94798224254?pwd=CFraer25qnpORbJ14aAVHRwaSJOjJM.1

The Provost's Office is excited to invite you to join in responding to an extraordinary opportunity to enhance our academic and research capabilities in AI at Stony Brook. SUNY recently made funding available to support the creation of departments of AI and Society at its universities. Stony Brook is well-positioned to seize this opportunity to build upon our interdisciplinary strengths in AI.

The office is hosting a forum on Friday, Nov. 15, from 11:30 a.m. to 1:30 p.m., in Ballroom A, SAC. You are invited to attend to learn more about this opportunity and to help us generate ideas to build a compelling proposal for Stony Brook to submit to SUNY. Lunch will be provided.

Please click here to RSVP as soon as possible.

This funding will support innovation in our curriculum, allowing us to create programs that explore the social and societal impact of AI alongside the technological advancements led by researchers in engineering and scientific disciplines.

We believe we can make a significant impact through this SUNY program and look forward to your participation in this initiative.
CSE 600 Talk: Haibin Ling - Computer Vision Research and Applications


Abstract: Having been intensively studied over half a decade, computer vision has evolved as a broad research area and become mature in many applications. In this talk, we will summarize our work in computer vision in both core vision topics and application-oriented ones. In particular, for core vision problems, we will report studies on visual tracking, visual matching and visual detection; for applications, we will describe our work on medical image analysis, intelligent transportation, smart projector systems and preliminary work on material property prediction.

Bio: Haibin Ling received the BS and MS degrees from Peking University in 1997 and 2000, respectively, and the PhD degree from the University of Maryland, College Park, in 2006. From 2000 to 2001, he was an assistant researcher at Microsoft Research Asia. From 2006 to 2007, he worked as a postdoctoral scientist at the University of California Los Angeles. In 2007, he joined Siemens Corporate Research as a research scientist. From 2008 to 2019, he worked as a faculty member of Temple University. In fall 2019, he joined the Department of Computer Science of Stony Brook University where he is currently a SUNY Empire Innovation Professor. His research interests include computer vision, augmented reality, medical image analysis, and human computer interaction. He received the Best Student Paper Award at the ACM UIST in 2003, and the NSF CAREER Award in 2014. He serves as Associate Editor for several journals including IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Pattern Recognition (PR), and Computer Vision and Image Understanding (CVIU). He has served or will serve as Area Chair for CVPR 2014, 2016, 2019 and 2020.

Topic: AI Seminar: Stanley Bak
Time: Monday Nov 1, 2021 12:00 PM Eastern Time (US and Canada)

Join Zoom Meeting
https://stonybrook.zoom.us/j/91227496273?pwd=M3EyUDlzK3Vzd2pDOGpDU1ZjN0k1UT09

Abstract: The field of formal verification has traditionally looked at proving properties about finite state machines or software programs. The surge in deep learning has been accompanied by a surge of progress in trying to apply mathematical and algorithmic techniques to prove things about the function being computed by a neural network.

This talk formalizes the neural network verification problem and describes technical methods for neural network verification based on reachability analysis. Improvements to analysis efficiency will be given, as well as research directions for further exploration. We also include an objective comparison performed this last summer trying to evaluate the best existing verification methods in terms of speed and network size. The competition was performed on common hardware and involved the participation of twelve international teams (the tool authors) on a common set of benchmarks. 

Biography: Stanley Bak is an assistant professor in the Department of Computer Science at Stony Brook University investigating the verification of autonomy, cyber-physical systems, and neural networks. He strives to develop practical formal methods that are both scalable and useful, which demands developing new theory, programming efficient tools and building experimental systems.
Stanley Bak received a Bachelor's degree in Computer Science from Rensselaer Polytechnic Institute (RPI) in 2007 (summa cum laude), and a Master's degree in Computer Science from the University of Illinois at Urbana-Champaign (UIUC) in 2009. He completed his PhD from the Department of Computer Science at UIUC in 2013. He received the Founders Award of Excellence for his undergraduate research at RPI in 2004, the Debra and Ira Cohen Graduate Fellowship from UIUC twice, in 2008 and 2009, and was awarded the Science, Mathematics and Research for Transformation (SMART) Scholarship from 2009 to 2013. From 2013 to 2018, Stanley was a Research Computer Scientist at the US Air Force Research Lab (AFRL), both in the Information Directorate in Rome, NY, and in the Aerospace Systems Directorate in Dayton, OH. He helped run Safe Sky Analytics, a research consulting company investigating verification and autonomous systems, and performed teaching at Georgetown University before joining Stony Brook University as an assistant professor in Fall 2020.
Le Hou Dissertation Defense: Deep Learning for Digital Histopathology across Multiple Scales

ABSTRACT: Histopathology is the study of tissue changes caused by diseases such as cancer. It plays a crucial role in disease diagnosis, survival analysis and development of new treatments. Using computer vision techniques, I focus on multiple tasks for automated analysis in digital histopathology images, which are challenging because histopathology images are heterogeneous and complex, due to the large variation of hundreds of cancer types in gigapixel resolution. In this thesis, I show how histopathology image analysis tasks can be viewed in three scales: Whole Slide Image (WSI)-level, patch-level and cellular-level, and present my contributions in each resolution level.

BIO: WSI-level analysis such as classifying WSIs into cancer types is challenging, because conventional classification methods such as off-the-shelf deep learning models cannot be applied directly on gigapixel WSIs due to computational limitations. I contribute a patch-based deep learning method that classifies gigapixel WSIs into cancer types and subtypes with close-to-human performance. This method is useful for computer-aided diagnosis. At patch-level, I contribute a novel method for histopathology image patch classification. On the task of identifying Tumor Infiltrating Lymphocyte (TIL) regions, the prediction result of this method correlates to the survival rate of patients. At cellular-level, I contribute novel methods for nucleus classification and roundness regression, which are interpretable features for histopathology studies. With this method, I generated a large-scale dataset of segmented nuclei, in WSIs from a large publicly available digital histopathology image dataset, to help advance histopathology research.

AI on Campus: Your Thoughts, Your Future

Join the Conversation: Share Your Thoughts about Learning, Academics, and AI

The world of college is changing fast, and Artificial Intelligence (AI) is at the center of it. We are part of the Institute on AI, Pedagogy, and the Curriculum with AAC&U, and we need to hear from the people AI affects most: you!

This is an open discussion for all students to share their honest experiences, their top concerns, and their best ideas about AI in our academic environment. We'll be diving into these key questions:

  • How can AI actually make learning better or easier? What opportunities do you see for using AI tools to enhance your assignments, research, or skills?

  • What are your biggest worries about AI? Is it about cheating, being graded fairly, or preparing for the job market? How is AI impacting your workload or stress levels?

  • What specific tools, workshops, or policies would help you use AI responsibly and successfully? (Think training, software, or clear rules.)

Dates/Times:

  • Wednesday, 2/4 at 2pm

  • Thursday, 2/5 at 12pm

Please register in advance for the Zoom link.

Can't Make It? Share Your Feedback!

Don't worry if you can't attend! You can still share your thoughts via video in our AI Zoom Room or via email: rose.tirotta-esposito@stonybrook.edu.

Videos will not be shared publicly and comments will only be shared in aggregate.

Your voice matters. Come tell us how AI is affecting your studies, your stress, and your success!

  • Dr. Rose Tirotta-Esposito (Assistant Provost; Director of CELT)

  • Dr. Elizabeth Hewitt (Associate Professor in the Department of Technology and Society (DTS) in the College of Engineering and Applied Sciences)

  • Chris Kretz (Associate Librarian and Head of Academic Engagement at SBU Libraries)

  • Prof. Rajiv Lajmi (Assistant Professor in the School of Health Professions and Chair of Applied Health Informatics)

  • Dr. Matthew Salzano (Assistant Professor in the Department of Communication in the School of Communication and Journalism)

CG Group member (and SBU faculty) Chao Chen will speak on Fri, March 12, about the use of topological data analysis in machine learning for image analysis.
Chao has shared some of his research with the CG Group previously, and this will be a great opportunity to learn more about this exciting research area related to computational geometry/topology!

Time: Friday, March 12, 2pm-3pm
Place: Zoom
https://stonybrook.zoom.us/my/profweizhu?pwd=RjVIVXg3YUhudzZZQ3pheHUydTJBUT09



Title: Learning with Topological Information - Image Analysis and Label Noise
Speaker: Prof. Chao Chen (SBU)

Abstract: Modern machine learning faces new challenges. We are
analyzing highly complex data with unknown noise. Topology provides
novel structural information to model such data and noise. In this
talk, we discuss two directions in which we are using topological
information in the learning context. In image analysis, we propose a
topological loss to segment and to generate images with not only
per-pixel accuracy, but also topological accuracy. This is necessary
in analysis of images of fine-scale biomedical structures such as
neurons, vessels, etc.  Extracting these structures with correct
topology is essential for the success of downstream
analysis. Meanwhile, we discuss how to use topological information to
train classifiers robust to label noise. This is important in practice
especially when we are using deep neural networks which tend to
overfit noise. These results have been published in NeurIPS, ECCV,
ICML and ICLR.
The New York Academy of Sciences Presents AI for Materials: From Discovery to Production - A Virtual Symposium

Event Description: This interdisciplinary symposium covers the application of artificial intelligence (AI) throughout the entire life cycle of new materials -- from materials simulations and synthesis to translating research into high-volume industrial production.

Event Link & Registration: nyas.org/AI4Materials2020
Abstract: Drawing on group-theoretic and information-theoretic foundations, we propose information lattice learning (ILL) as a general framework to learn rules of a signal (e.g., an image or a probability distribution). In our definition, a rule is a coarsened signal used to help us gain one interpretable insight about the original signal. To make full sense of what might govern the signal's intrinsic structure, we seek multiple disentangled rules arranged in a hierarchy, called a lattice. Compared to representation/rule-learning models optimized for a specific task (e.g., classification), ILL focuses on explainability: it is designed to mimic human experiential learning and discover rules akin to those humans can distill and comprehend. We will detail the mathematical foundations and algorithms of ILL, and illustrate how it addresses the fundamental question what makes X an X by creating rule-based explanations designed to help humans understand. Our focus is on explaining X rather than (re)generating it. We show ILL's efficacy and interpretability on benchmarks and assessments, as well as a demonstration of ILL-enhanced classifiers achieving human-level digit recognition using only one or a few MNIST training examples (1-10 per class). We present applications in knowledge discovery, using ILL to distill music theory from scores and chemical laws from molecules and further revealing connections between them. We close with some early work on understanding the principles that govern scattering amplitudes in Super Yang-Mills theory, rather than just predicting them.

Biography: Lav R. Varshney is the Della Pietra Infinity Professor and inaugural director of the AI Innovation Institute at Stony Brook University. He is co-founder and CEO of Kocree, Inc., a startup company building novel human-controllable AI for discovery and creativity, and chief scientist of Ensaras, Inc., a startup company focused on AI and wastewater treatment. He holds appointments at RAND Corporation and at Brookhaven National Laboratory. He was previously on the faculty of the University of Illinois Urbana-Champaign, a visiting scholar at Northwestern's Kellogg School of Management, a principal research scientist at Salesforce Research AI, and a research staff member at IBM Research. He is a former White House staffer, having served on the National Security Council staff as a White House Fellow, where he contributed to national AI and wireless communications policy. His research interests include information theory and artificial intelligence. He received his B.S. degree from Cornell University and his S.M. and Ph.D. degrees from the Massachusetts Institute of Technology.

Location: Room 102