How to Do Spectral Learning at Scale for Science and Engineering

Abstract: Spectral decompositions such as singular value decompositions (SVDs) and eigenvalue decompositions (EVDs) are central tools across a vast swath of scientific computing and machine learning, with abundant engineering applications. Yet many modern methods for learning such decompositions in high dimensions struggle with instability, bias, and poor scalability, even when approximation power is not the limiting factor. I argue that these difficulties are not intrinsic to spectral problems, but instead arise from a shared reliance on Rayleigh-quotient-based constrained optimization, which forces explicit orthogonality handling through penalties, normalization, or whitening.
To address these challenges, I present a reformulation based on unconstrained variational objectives that implicitly encode spectral structure, eliminating the need for orthogonalization and ad-hoc regularization. This perspective leads to a conceptually simpler and scalable parametric framework for learning ordered spectral representations via nested optimization. The resulting framework is well matched to diverse settings in science and engineering. As examples, I demonstrate its effectiveness on eigenvalue problems for linear PDEs such as the Schrödinger equation, spectral (Koopman) analysis of nonlinear dynamical systems such as molecular dynamics, and structured representation learning with deep neural nets. Collectively, these examples illustrate how abandoning Rayleigh-quotient-based formulations resolves long-standing optimization pathologies across domains.

Bio: Jongha (Jon) Ryu is a postdoctoral associate at MIT EECS. He received his Ph.D. in Electrical and Computer Engineering from UC San Diego. His research develops statistical and mathematical foundations for scientific machine learning, with a focus on scalable spectral methods, efficient generative modeling, and reliable uncertainty quantification for scientific and engineering systems.

Location: NCS 120

Join us for an engaging panel discussion featuring researchers who participated in our inaugural AI JAM session on February 26th. Our panelists will share their firsthand experiences using large language models to tackle complex scientific problems, with a special focus on prompt engineering strategies, discussing both breakthroughs and challenges encountered during this collaborative initiative. Learn how these cutting-edge AI tools are being applied to real-world research questions and discover insights that could inform your own scientific endeavors. Attendees are encouraged to come prepared with questions about prompt engineering for the panel discussion.

Moderator: Adolfy Hoisie, Deputy Director, Computing and Data Sciences

Kevin Yager, Group Leader, AI-Accelerated Nanoscience, Center for Functional Nanomaterials
Lingda Li, Associate Computational Scientist, Systems, Architecture and Computing Technologies, Computing and Data Sciences
Liguo Wang, Director of Scientific Operations, Laboratory for BioMolecular Structure (LBMS), National Synchrotron Light Source II
Weiguo Yin, Physicist, Condensed Matter Theory, Condensed Matter Physics and Materials Science Department

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1606837837?pwd=Tc0mwQqLXpDfYOIaoaurmpLD2mMlzS.1 (Meeting ID)

Passcode: 822553

Abstract: The capacity to adapt machine learning models to various contexts, information, and objectives is particularly valuable. In this thesis, I focus on developing Class Conditional Guided Models. These are models that can be adaptively biased towards a class of interest via a conditional input. My primary focus lies in the efficiency of these models. They are constructed to require training only once, with the ability to quickly and conveniently adapt during testing time without necessitating fine-tuning or retraining.
Firstly, I propose RelationVAE, a novel generative model designed for few-shot scenarios, utilizing the prior knowledge of class similarity relationships. RelationVAE is designed to condition on the embeddings of the neighbor classes (i.e. classes with similarity relationships), to generate more reliable samples by making them more similar to the neighbor class. This enables adaptation of the generative model to the provided prior knowledge about class relationships.
As a second focus, I introduce scGAN, a shadow segmentation technique that enables adaptation to varying shadow distributions in different testing environments. scGAN is designed to condition on a sensitivity parameter, a scalar, to control the amount of the shadow detected. In the testing phase, the parameter is set to appropriate values, allowing the model to quickly adapt to specific test environments.
In my third contribution, I propose S-SEG, a methodology for fine-grained counting allowing adaptation to different granularities of fine-grained classes. In fine-grained problems, the distinction between classes is subtle and inconsistent across images, leading to variations in the granularity of the target class from one image to another. S-SEG is designed to be conditioned on an additional input, the sensitivity parameter, to control the granularities of the target class during inference.
My fourth contribution is a text-to-image synthesis method which allows controlling the number of the generated objects of a target class. I propose to generate an intermediate condition, the density map, which reflects the number of objects, together with their layout. This intermediate condition is used to effectively guide the generative model to generate objects with accurate counts.

Speaker: Vu Nguyen

Zoom: https://stonybrook.zoom.us/j/97114455337?pwd=Z4rB9dWcstlahUIs8PRrvQ9b2ZK2Df.1
Meeting ID: 971 1445 5337
Passcode: 272300










Abstract:
Quantifying similarity is a central notion in science and data analysis, pervading everything from phylogenetic trees to the foundation of clustering. Unfortunately, despite being examined and applied for decades, traditional similarity and distance metrics have fundamental drawbacks. The key problem is that all of them are only defined over pairs of objects, so they scale quadratically when one tries to compare N objects. The present explosion in the amount of data available to us requires new ways to process information, and while some current algorithms can handle millions of points, we need alternatives applicable to billions. This is what motivated us to develop a new framework that can compare any number of objects at the same time. With this, we achieve an unprecedented linear scaling when comparing multiple objects. Here we will discuss the main properties of this formalism, along with its applications in drug design and to the analysis of Molecular Dynamics (MD) simulations. Our indices have proven to be incredibly versatile when applied to chemical space exploration and visualization, allowing us to rigorously quantify the chemical diversity of very large molecular libraries. This has led to the creation of several algorithms to sample important regions in chemical space, including a more efficient way of identifying the prevalence of activity cliffs. Additionally, our indices provide a convenient route to sample complex MD trajectories, allowing to identify representative structures very efficiently. Moreover, we can also cluster biological ensembles in a more robust way than with standard algorithms, which has led to our group's work on MDANCE, a very flexible and efficient open-source clustering module. Drop by if you want to know how we clustered one billion molecules!


Speaker:
Assistant Professor, Department of Chemistry and Quantum Theory Project
University of Florida, Gainesville
Website: https://quintana.chem.ufl.edu/

Location:
Laufer Center Lecture Hall 101

Speaker: Gary Kazantsev (Head of Quant Technology Strategy in the Office of the CTO at Bloomberg)

 

Date/Time: Friday, October 15, 2021 10:00AM-11:00AM EST

 

Title: Machine Learning in Finance

Abstract: Machine learning is changing our world at an accelerating pace. In this talk we will discuss the recent developments in how machine learning and artificial intelligence are changing finance, from a perspective of a technology company which is a key  participant in the financial markets. We will give an overview and discuss the evolution of selected flagship Bloomberg ML and AI projects, such as sentiment analysis, question answering, social media analysis, information extraction and prediction of market impact of news stories. We will discuss practical issues in delivering production machine learning solutions to problems of finance, highlighting issues such as interpretability, privacy and nonstationarity. We will also discuss current research directions in machine learning for finance. We will conclude with a Q&A session.

Bio: (https://www.techatbloomberg.com/people/gary-kazantsev/) Gary is the Head of Quant Technology Strategy in the Office of the CTO at Bloomberg. Prior to taking on this role, he created and headed the company's Machine Learning Engineering group, leading projects at the intersection of computational linguistics, machine learning and finance, such as sentiment analysis of financial news, market impact indicators, statistical text classification, social media analytics, question answering, and predictive modeling of financial markets.

Prior to joining Bloomberg in 2007, Gary had earned degrees in physics, mathematics, and computer science from Boston University.

He is engaged in advisory roles with FinTech and Machine Learning startups and has worked at a variety of technology and academic organizations over the last 20 years. In addition to speaking regularly at industry and academic events around the globe, he is a member of the KDD Data Science + Journalism workshop program committee and the advisory board for the AI & Data Science in Trading conference series. He is also a co-organizer of the annual Machine Learning in Finance conference at Columbia University.


Join Zoom Meetinghttps://stonybrook.zoom.us/j/93374426887?pwd=cE9zeW51VXFEN2R0YnNPbHF1WFp0Zz09Meeting ID: 933 7442 6887Passcode: 330347One tap mobile+16468769923,,93374426887# US (New York)+13126266799,,93374426887# US (Chicago)Dial by your location +1 646 876 9923 US (New York) +1 312 626 6799 US (Chicago) +1 301 715 8592 US (Washington DC) +1 346 248 7799 US (Houston) +1 408 638 0968 US (San Jose) +1 669 900 6833 US (San Jose) +1 253 215 8782 US (Tacoma)Meeting ID: 933 7442 6887
























new virtual seminar series on Games, Decisions, and Networks will start this Friday. The series aims at bringing together researchers working on foundations and applications of games theory, decision theory, and networks from computer science, control, economics and operation research. 





The advisory board for the series comprises Asu Ozdaglar (MIT), Christos Papadimitriou (Columbia), Drew Fudenberg (MIT), Eva Tardos (Cornell), Matthew O. Jackson (Stanford), Ramesh Johari (Stanford), and Tamer Başar (UIUC). 
The first talk will be given by Costantinos Daskalakis (MIT) on January 22nd at noon ET, titled Equilibrium Computation and the Foundations of Deep Learning. Upcoming speakers include


- Rakesh Vohra (Upenn)
- Sanjeev Goyal (Cambridge)
- Aaron Roth (Upenn)
- Aislinn Bohren (Upenn)
- Jason Marden (UCSB)

and more to be added!




https://stonybrook.zoom.us/j/91775729097pwd=Qlc5Nks0NmlyKzJwMjR0S0hrdVZ3QT09

Meeting ID: 917 7572 9097
Passcode: 555459


Abstract: As the saying goes, there are many ways to skin a cat.
While we don't want to go around skinning cats, the world of
optimization is rich with different problems, problem formulations,
and methods and approaches, each with different guarantees and
computational benefits. In this talk we will take a tour down the
problem of structured sparsity in sensing to see how one simple
problem can inspire a wide range of analysis and tools. First, I will
present the optimality conditions for a generalized structured sparse
problem, which can be geometrically visualized as alignment of vectors
and matrices. Then I will introduce three approximation methods for
the problem of phase retrieval, which are a twist on stochastic
gradient and coordinate descent methods. These methods leverage
fundamental numerical linear algebra concepts to give fast approximate
solutions to large-scale problems, which then after postprocessing can
produce more reliable sensing results.

Bio: Yifan Sun received her PhD in Electrical Engineering from the
University of California Los Angeles in 2015, with research focusing
on convex optimization and semidefinite programming. She was then
Technicolor Research and Innovation, focusing on machine learning and
data science applications. More recently, she completed two postdocs,
at the University of British Columbia in Vancouver, Canada and
L'Institut National de Recherche en Informatique et Automatique
(INRIA) in Paris, France.
Abstract: As we enter the AI era, domain scientists face a critical question: What can we do to harness AI effectively for scientific discovery? AI has demonstrated remarkable capabilities, from accelerating simulations to uncovering hidden patterns in complex datasets. While these advancements offer unprecedented opportunities, they also raise concerns--AI models often function as black boxes, making it difficult to connect their outputs to established scientific principles. This lack of interpretability can undermine trust and limit adoption, particularly in fields like meteorology where physical understanding is critical.
In this talk, I will explore how interpretable AI can bridge this gap, highlighting its potential to generate explicit, physically meaningful equations rather than opaque neural networks. Through four case studies from my lab, I will showcase how interpretable AI can enhance scientific understanding:
  1. Satellite Precipitation Retrieval: Using AI-based approaches to interpret precipitation retrieval algorithms from AMSU data, we identified critical microwave channels (89 and 150 GHz) that directly link to physical processes in the atmosphere.
  2. Quantitative Precipitation Estimation (QPE): By applying symbolic regression models to polarimetric radar data, we derived mathematical expressions that outperform traditional Z-R relationships and existing QPE algorithms, offering new insights into rainfall microphysics.
  3. Tornado Probability Prediction: Leveraging reinforcement learning-based symbolic deep learning models, we developed interpretable equations that outperform the traditional Significant Tornado Parameter (STP) index, providing a clearer understanding of the relationships between key atmospheric variables and tornado risk.
  4. Domain-Aware Symbolic Regression for Scientific Equations: In our latest work, we introduced a symbolic regression framework that incorporates domain-specific symbol priors extracted from thousands of scientific publications. By encoding common mathematical structures--such as the prevalence of trigonometric functions in physics or logarithmic forms in biology--into a tree-structured reinforcement learning model, we improved both the accuracy and interpretability of discovered equations. This approach accelerates convergence, enforces physical plausibility, and reveals new governing relationships in climate and geophysical data.
Through these examples, I hope to spark discussion on the evolving role of domain scientists in the AI era and inspire new ways to integrate AI with physical understanding in atmospheric research.

IACS Seminar Speaker: Yixin Wen, University of Florida

Location: IACS Seminar Room or Zoom

Join Zoom Meeting: https://stonybrook.zoom.us/j/97596399106?pwd=0PBvElFLqov3biO6OlQxSWLWudkIuH.1
Meeting ID: 975 9639 9106
Passcode: 096213
AI is everywhere and so are the privacy concerns that come with it. At its core, the most common forms of AI we use today are online digital services and thus inherit the usual privacy risks. We'll take a look at indirect prompt injection- a technique that can trick AI tools into revealing or extracting private information as well as techniques being used in academic contexts to manipulate systems and even mislead researchers.

Register here for the online session.