Abstract: Visual generation is a fundamental problem in computer vision and graphics, with applications ranging from 3D capture to content creation and image/video synthesis. Despite rapid progress in neural rendering and generative models, efficiency remains a key obstacle in practice: high-quality 3D reconstruction often depends on dense multi-view supervision; scalable 3D synthesis faces heavy optimization, training, and rendering costs; and modern image/video generators incur substantial computation as token grids grow with spatial resolution and temporal length.
This thesis targets efficient visual world modeling by improving sample efficiency in 3D reconstruction, representation efficiency in 3D generation, and computational efficiency in image/video synthesis. First, we improve sample efficiency for neural implicit surface reconstruction under sparse views by integrating multi-view stereo probability volumes as a geometric regularizer, enabling high-quality reconstruction from as few as three input images. Next, we introduce an explicit 3D representation for 3D generation, built from multi-view depth and RGB predictions with 3D Gaussian features, which enables the use of 2D generative priors while enforcing multi-view consistency via epipolar attention. We then address the computational bottleneck of image and video synthesis with importance-based token merging, using importance signals available during generation to preserve critical information while merging redundant tokens. Finally, we propose efficient mixed-resolution diffusion transformers via cross-resolution phase-aligned attention, aiming to improve attention stability under mixed token grids and support high-fidelity mixed-resolution generation.

Speaker: Haoyu Wu

Location: NCS120

Abstract: Computational pathology has revolutionized cancer diagnosis and research through the analysis of digitized whole slide images (WSIs). However, the giga-pixel size of these images presents profound technical challenges, creating two intertwined bottlenecks: computational inefficiency and label inefficiency. The immense data scale makes standard end-to-end (E2E) training of deep neural networks infeasible due to prohibitive GPU memory requirements, while the reliance on expert pathologists for annotations makes obtaining high-quality labeled data a tedious and expensive process. This proposal confronts these dual challenges by developing a series of novel model architectures, training paradigms, and self-supervised learning methods designed to create a more efficient and effective framework for WSI analysis.

To improve computational efficiency, this proposal first introduces a locally supervised learning paradigm that enables E2E training on entire WSIs by partitioning a network into gradient-isolated modules, circumventing the memory bottleneck of backpropagation. Second, it presents Prompt-MIL, a parameter-efficient fine-tuning framework that reduces the number of trainable parameters, memory consumption, and training time by fine-tuning only few prompts to guide large pre-trained models. Third, this work advances the efficient architecture on WSIs by developing novel State-Space Models (SSMs). It proposes 2DMamba, the first intrinsic Mamba architecture that preserves the crucial 2D spatial structure of images, overcoming the spatial discrepancy inherent in 1D models. Fourth, to address the inefficiency of multi-directional scans in Mamba models, including 2DMamba, it presents Locally Bi-directional Mamba (LBMamba), which introduces a novel, hardware-aware local backward scan that integrates bi-directional scan into a single forward pass, significantly improving throughput performance trade-off. Lastly, it proposes an extension to the LBMamba, warp-level Bi-directional Mamba (WLBMamba) that extends the thread-level bidirectional scan to warp-level bidirectional scan that further improves the throughput performance trade-off.

To improve label efficiency, this proposal proposes a Precise Location-based Matching strategy for self-supervised dense contrastive learning. By allowing a local patch in one augmented view to match multiple overlapping patches in another, creates a more accurate correspondence, leading to superior feature representations for dense prediction tasks like segmentation and detection.

In summary, this proposal presents a holistic investigation into the efficiency bottlenecks in computational pathology. Through these combined contributions in model architecture, training paradigms, and self-supervised learning, this work establishes a more scalable, efficient, and powerful computational framework for analyzing giga-pixel pathology images.

Speaker: Jingwei Zhang

Location: Old Computer Science Room 2114

Zoom: https://stonybrook.zoom.us/j/95187903649?pwd=tV0CNxLu1QKqw7hGmcE1h0rJ2C6n1b.1
Meeting ID: 951 8790 3649 | Passcode: 488916
Join your friends at DoIT for a workshop on Zoom's AI Companion

AI Companion is a meeting summary tool that can capture and summarize what is said in a Zoom meeting transcript, eliminating the need for a notetaker. In addition, meeting participants can use AI Companion to ask questions to a chat bot to get clarity on information within a meeting, and they can also use it to create stylistic virtual backgrounds.

Register here for the online session
AI Seminar: Computational Pathology: Deep Learning, Classification and
Predicting the Future  - Joel Saltz

Abstract:  Pathologists have been looking at tissue through microscopes since the 1800s.  During each pathologist's career,  he or she views slides having  roughly 1,000,000,000,000 cells. Deep learning methods are rapidly being developed to assimilate the huge amount of information walked inside of tissue images and to use this information to predict outcomes and responses to treatments.

Stony Brook is a leader in this type of multi-disciplinary work. I will provide an overview of Stony Brook computational Pathology efforts and articulate how these have the potential to create biomedical advances as well as to drive development of new computer science. 


Bio: Dr. Joel Saltz is a leader in research on advanced information technologies for large scale data science and biomedical/scientific research. He has developed innovative pathology informatics methods, including: the first published whole slide virtual microscope system; pioneering pathology computer-aided diagnosis techniques; and methods for decomposing pathology images into features and linking those features to cancer omics, response to treatment and outcome. He has broken new ground in big data through development of the filter-stream based DataCutter system, the map-reduce style Active Data Repository and the inspector-executor runtime compiler framework. He has also been an active contributor in clinical informatics, having developed
predictive models for hospital readmissions, point of care laboratory testing quality assurance systems, decision support systems for electrophoresis interpretation and graphical user interfaces to support clinical data warehouse queries. Dr. Saltz has been a pioneer in establishing the field of biomedical informatics; he founded and built two highly successful departments of biomedical informatics, one at Ohio State University and one at Emory University. In 2013, he came to Stony Brook as Vice President for Clinical Informatics and Founding Department Chair of Biomedical Informatics - to create a living laboratory for biomedical informatics and to create a third unique biomedical informatics department dually housed in the School of Medicine and the College of Engineering. Dr. Saltz is trained both as a computer scientist and as a physician through the MSTP program at Duke University. He has deep experience in computer science, having served on the computer science faculties at Yale University and the University of Maryland. He completed his residency in clinical
pathology at Johns Hopkins University and he is a practicing, board-certified clinical pathologist. 
Talk by Zhenhua Liu to be followed by AI Institute updates


Abstract: Decision making with uncertainty has been studied in multiple communities extensively. Recently, online optimization has gained popularity partially because of its promising performance guarantees by incorporating predictions. In this talk, I will provide an overview of our work on algorithm designs for online optimization and its applications. Then, I will talk about our recent work in ACM Sigmetrics 2019 on choosing predictions and control algorithms simultaneously and dynamically. Finally, I will discuss some ongoing efforts and collaboration opportunities.

Bio: Zhenhua Liu is currently an assistant professor in the Department of Applied Mathematics and Statistics at Stony Brook University. He is also affiliated with the Department of Computer Science, the AI Institute and the Smart Energy Technology Cluster. He received his PhD degree in Computer Science from California Institute of Technology. His current research interests include cloud computing, online optimization and learning, smart grid, market design and distributed control. His research combines rigorous analysis and system design, and goes from theory, to prototype, and eventually to industry to make real impacts.
Time: May 5, 2022, Thursday, 02:00 PM Eastern Time (US and Canada)
Place: New Computer Science (NCS) Room 220, and Zoom

Zoom link: https://stonybrook.zoom.us/j/95948672934?pwd=d3ZDcUJkK3VweFBDVWhIVDhtaFU2Zz09
Meeting ID:  959 4867 2934
Passcode:  082036

Title:  Generative Adversarial Learning using Optimal Transport

Abstract: 

Generative Adversarial Learning (GAL) aims to learn a target distribution in an adversarial manner. A Generative Adversarial Network (GAN) is a concrete implementation of GAL using a discriminator and a generator that play a min-max game. GANs have been used in many machine learning and computer vision applications. However, GANs are known to be hard to train, mainly because a min-max saddle point optimization problem needs to be solved in GAL. In this thesis, I investigate several methods to improve generative adversarial learning using Optimal Transport (OT). 

Previous Wasserstein GANs (WGANs) do not compute the correct Wasserstein distance to train the discriminator. To address this problem, I propose WGAN-TS that uses the L1 transport cost and computes the correct Wasserstein distance to train the discriminator. To ensure the local convergence of WGANs, I propose WGAN-QC that adopts the quadratic transport cost. I prove that WGAN-QC not only computes the correct Wasserstein distance but also converges to a local equilibrium point. To compute the Wasserstein distance over the whole dataset, I propose to use Semi-Discrete Optimal Transport (SDOT) to match noise points and the real images during GAN training. To measure the quality of an SDOT map, I use the Maximum Relative Error (MRE) and the L_1 distance between the target distribution and the transported distribution obtained by an OT map. I propose statistical methods to estimate the MRE and the L_1 distance. I propose an efficient Epoch Gradient Descent algorithm for SDOT (SDOT-EGD). To deal with the 2D special case of GAL, I propose to use OT to learn 2D distributions. In particular, I adopt OT to match persistent diagrams in training a topology-aware GAN and learn density maps in the crowd counting task. Finally, I use OT and the topological maps of the crowd to improve the crowd counting performance and propose a topology-based metric to measure the quality of the crowd density maps.

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Embodied Intelligence at Scientific User Facilities

Abstract: This presentation explores the active work integrating artificial intelligence and robotics at the National Synchrotron Light Source II, and a perspective for the future. Through various case studies, we highlight the optimization of operations, improved experimental outcomes, and the orchestration of distributed multimodal experiments. This ongoing development includes collaborators from across the light and neutron sources in the DOE complex. We will elaborate on the open-source Bluesky project, and its capabilities to support adaptive and autonomous experiments. Additionally, we will discuss how Bluesky can be integrated with open-source robotic control software to unlock new flexible automation for autonomous scientific research, which scales to new experiments and continues to leverage human ingenuity.

Biography: Dr. Phillip M. Maffettone is an Associate Computational Scientist in the Data Science and Systems Integration Division at NSLS-II. His research focuses on accelerating scientific discovery at user facilities through the integration of robotics, artificial intelligence (AI), and advanced experiment orchestration systems. He leads the N3XTware project, constructing the software architecture for the next 12 beamlines to be built at NSLS-II. Prior to this he built the brain on the world's first mobile robotic scientist at the University of Liverpool, and later spearheaded the machine learning platform for a biotechnology start-up, BigHat Biosciences. He holds a DPhil in Inorganic Chemistry from the University of Oxford and a B.S. in Chemical Engineering from the University at Buffalo.

Location: CDS, Bldg. 725, Training Room

Link: https://bnl.zoomgov.com/j/16049713 31?pwd=nc5CV3cOFrdYxordFieP W07tIDmwYb.1

Meeting ID: 160 497 1331
Passcode: 289875

Title: Cyberinfrastructure for forward prediction and inversion estimation with uncertainty quantification

Seminar Speaker: Dr. Mengyang Gu, Assistant Professor, Department of Statistics and Applied Probability, University of California, Santa Barbara

Abstract: In this talk, we introduce four useful tools for forward prediction and inversion estimation. The first tool is the parallel partial Gaussian process surrogate model for emulating expensive computer simulations with massive coordinates. The tool is implemented in the RobustGaSP package available in R, MATLAB, and Python, for predicting both scalar- and vector-valued outputs with uncertainty assessment. The second tool is implemented in the RobustCalibration package, which handles Bayesian data inversion or model calibration by one or multiple types of experimental observations. A unique feature of the package is the inclusion of fast surrogate models of both scalar- and vector-valued computer simulations that bypass the expensive simulation in one line of code. The third tool is implemented in the AIUQ package, available in both R and MATLAB. In this approach, we show that differential dynamic microscopy, a scattering-based analysis tool that extracts dynamical information from microscopy videos, is equivalent to fitting the temporal auto-covariance in Fourier space, based on a latent factor model we construct. We develop a more efficient estimator and reduce the computational cost to pseudolinear order with respect to the number of observations without approximation, by utilizing the generalized Schur algorithm for the Toeplitz covariance. In the last tool, we developed a new method called the inverse Kalman filter, which enables fast matrix-vector multiplication between a covariance matrix from a dynamic linear model and any real-valued vector with a linear computational cost. These new approaches outline a wide range of applications that include emulating expensive simulation at molecular-, meso- and macro-scales, active learning with error control, nonparametric estimation of particle interaction functions, and data inversion from microscopy and velocity fields.

Join Zoom Meeting: https://bnl.zoomgov.com/j/1606285496?pwd=2yJYSG6lx8gMPiibzgAIBQtKHIjuHV.1
Meeting ID: 160 628 5496
Passcode: 472506