The Provost's Lecture Series features talks by SUNY Distinguished Academy faculty members at Stony Brook University, showcasing the outstanding research and scholarship that is taking place at our institution.

Joe Mitchell

SUNY Distinguished Professor, Applied Mathematics and Statistics
Chair, Department of Applied Mathematics and Statistics, College of Engineering and Applied Sciences

A Case for Algorithms: A Computational Geometer's Perspective

Algorithms are all around us in every smart device and technology that has consumed our daily lives. As a computational geometer, I study algorithms to solve problems that involve a geometric perspective on data. I have observed that practically every technology and field of study has a need for effective algorithms involving geometric data. I reflect on some favorite algorithmic problems that are easy to visualize, but challenging to solve, and argue that the formal study of algorithms remains essential in the age of AI.

Reception to follow immediately after the talks.

Register here.



Abstract: The current approach to materials design, driven by strategic experimentation and supported by physics-based simulation across relevant scales, has been the standard for decades. While the theoretical component in this workflow provides valuable understanding of material behavior, it often fails to deliver actionable guidance for implementation. Advances in artificial intelligence and machine learning (AI/ML), together with high-performance computing (HPC), now offer a viable pathway to close this gap and accelerate both discovery and process optimization. This presentation will outline practical approaches for integrating AI/ML with HPC-enabled, high-throughput computation to explore high-dimensional search spaces. Examples will include the development of engineering alloys for extreme environments, the use of neural networks to rapidly improve computational thermodynamic models, and vapor processing optimization for the manufacturing of ultra-high-temperature ceramics. I will highlight how scientific insight and domain expertise remain essential for translating surrogate model predictions into impactful outcomes. Finally, I will conclude with current challenges and future opportunities for AI/HPC-driven materials research.

Speaker: Dongwon Shin
This seminar will be held in person and online

Join Zoom Meeting: https://stonybrook.zoom.us/j/93730374357?pwd=YDLJ7ELqOQnTZEQhlN8Pa4TuhaiFK8.1
Recently, large-scale language data combined with modern machine learning techniques have shown strong value as means for studying human psychology and behavior. For example, language alone has been shown predictive in mental health, personality, and health behaviors. However, many applications for such language-based assessments have readily available and important data beyond language (i.e. extra-linguistics), such as predicting the subjective well-being of a community using tweets, where one can take into account their age, education, and demographic attributes. Language may capture some characteristics while extra-linguistic variables captures others. We believe that effectively integrating linguistic and extra-linguistic data can yield benefits beyond either independently. In this thesis, we develop methods which effectively integrate extra-linguistic data with language data focused primarily on social scientific applications. The central challenge is dealing with the size and heterogeneity of, often sparse and noisy, language data versus the, often low-dimensional and non-sparse, extra-linguistic variables. First, we consider structured extra-linguistics, like socioeconomic (income and education rates) and demographics (age, gender, etc.), and propose two integration methods, named residualized controls (RC) and residualized factor adaptation (RFA), to be used in county-wise prediction tasks. Demonstrating techniques that integrate information at both the model-level and data-level, we found consistently strong improvement over naively combining features, for example, increasing county level well-being predictions by over 12%. Next, we consider unstructured extra-linguistic data. In the first part, we incorporate social network connections and language over time to propose a novel metric for quantifying the stickiness of words - their ability to spread across friendship connections in a social network over time (or in other words, stick in ones vocabulary after seeing friends use it). We obtain which language features are more probable to disseminate through friendship and show such a metric is useful for predicting who will be friends and what content will spread. In addition, we analyze language content over time by proposing a novel dynamic content-specific topic modeling technique that can help to identify different sub-domains of a thematic scope and can be used to track societal shifts in concerns or views over time.
Join librarian Christine Fena for an interactive workshop that invites you to explore AI tools firsthand, not just as users, but as critical investigators. Through playful experimentation and collaborative discovery, you'll uncover inherent biases, probe algorithmic flaws, and gain a deeper understanding of AI's limitations and societal impacts.

Location: Melville Library, Central Reading Room, Lab B

https://library.stonybrook.edu/library-events/critiquing-ai/
Abstract : Humans reason about everyday situations by making commonsense-based inferences, derived both from explicitly stated information and implicit, unstated knowledge. In this thesis, I investigate whether NLP models have different aspects of causal knowledge about events and how to improve their understanding of narratives and plans.
Answering questions about why people perform actions in a narrative can test whether NLP systems contain and can effectively apply causal knowledge about events. I introduce TellMeWhy, a dataset concerning why characters in short narratives perform the actions described. An evaluation of then SOTA finetuned models show that they are far worse than humans. To improve models, it is important to understand what aspects of causal knowledge they need and how to best use external sources to inject this knowledge. In KnowWhy, I analyze different ways of injecting knowledge into models, which is difficult since we do not know apriori what type of knowledge will be needed to answer a question, hence requiring a ranking model to pick the most important inference. Results show that this retrieved knowledge helps models of all sizes, thereby improving their understanding of narratives.
Next, I study whether models can reason about causal aspects of plans. I focus on testing whether they understand the underlying causal dependencies reflected in the temporal order of a plan's steps. I introduce CAT-Bench, and find that SOTA models are underwhelming, and that model answers are not consistent across questions about the same step pairs. In their current state, these models cannot yet reliably be used for complex user-facing tasks. I then measure contemporary models' ability to perform user-facing and user-centric plan customization. I introduce the use of semi-symbolic edits in large language model (LLM) based agents and test several multi-LLM-agent architectures for plan customization. While LLMs still lack the ability to understand complex customization hints, my results suggest that LLM-based architectures may be worth exploring further for other customization applications. Finally, I distill complex reasoning capabilities into small language models (SLMs) using synthetic data that reflects a decomposition-then-editing process for plan customization. I demonstrate that explicitly teaching this latent causal reasoning significantly improves the quality of SLM-generated customizations. Overall, my work has improved how well NLP models understand complex reasoning associated with events in different contexts.

Speaker: Yash Kumar Lal

Location: NCS 220 or Zoom https://stonybrook.zoom.us/j/95849648243?pwd=dgPpZtDpgwQrK9z1SaPpNbBifaorzk.1

Join University Libraries for an engaging panel discussion where we delve in and learn about the impacts of artificial intelligence on the 2024 US elections! Panelists are Paige Lord, Tom Costello, and Musa al-Gharbi. The discussion will be moderated by Library Dean, Karim Boughida. Co-sponsored by the Office of Diversity, Inclusion, and Intercultural Initiatives.

Please RSVP for Democracy in the Digital Age: AI's Influence on 2024 Elections here.

Submit an abstract celebrating research, new discoveries and achievements in medicine and science!

We encourage faculty, nurse practitioners, post-doctoral fellows, fellows, residents, medical students, graduate students and undergraduate students to submit an abstract. Original research, case reports and case series are welcome.

Abstract submission deadline: FEBRUARY 7, 2025

For more details, visit here.


Please join University Libraries on March 29 at 1:00 via Zoom as we welcome Dr. Zhang, SUNY Empire Innovation Professor at SBU's Power Lab. This lab is pioneering the research of coordinated networked microgrids (NMs) that can possibly help to restore neighboring distribution grids after a major blackout. That these NMs hold promise to significantly enhance the day-to-day reliability of the power grids, we are proud to host Dr. Zhang as a member of our STEM Speaker Series. Registration required.
https://library.stonybrook.edu/library-events/stem-speaker-series-ai-enabled-provably-resilient-networked-microgrids-with-dr-peng-zhang/