Johannes Hachmann, University of Buffalo Assistant Professor of Chemical Engineering presents Making Machine Learning Work in Chemistry

The use of modern machine learning, informatics and data mining approaches is a relatively new development in the chemical and materials domain. These techniques have been exceedingly successful in other application fields, and since there is no fundamental reason why they should not have a similarly transformative impact on chemical and materials research, there is now a concerted effort by the community to introduce data science in this new context. However, adapting techniques from other application domains for the study of chemical and materials systems requires a substantial rethinking and redevelopment of the existing methods.

In this presentation, we will discuss our work on designing advanced, physics-infused neural network architectures, the fusion of unsupervised clustering with supervised regression for local ensemble models, active and transfer learning techniques, bootstrapping approaches to minimize our training data footprint, methods to increase the applicability domain of data-derived models and automated hyperparameter optimization.

Biosketch: Johannes Hachmann is an Assistant Professor of Chemical Engineering at the University at Buffalo (UB), the Director of the Engineering Science in Data Science graduate program, a Core Member of the UB Computational and Data-Enabled Science and Engineering graduate program, and a Faculty Member of the New York State Center of Excellence in Materials Informatics. He earned a Dipl.-Chem. degree (2004) after undergraduate studies at the universities of Jena and Cambridge, M.Sc. (2007) and Ph.D. (2010) degrees in Chemistry from Cornell University, and he conducted postdoctoral research at Harvard University before joining the UB faculty in 2014. The research of the Hachmann Group fuses (first-principles) molecular and materials modeling with virtual high-throughput screening and modern data science (i.e., the use of database technology, machine learning and informatics) to advance a data-driven discovery and rational design paradigm in the chemical and materials disciplines. One of the centerpieces of the group's efforts is the creation of an open, general-purpose software ecosystem for the data-driven design of chemical systems and the exploration of chemical space. This work was recognized with a 2018 NSF CAREER Award.

Description:

Curious about what AI image generation tools are out there and how they work? Come down to the library Galleria space (outside the Central Reading Room) to see some demonstrations and learn more about them.

Librarians Chris Kretz and Ahmad Pratama, along with David Ecker of DoIT, will be hosting Explore AI demos from Monday - Wednesday this week on different topics. Whether you're new to AI or an experienced user, stop by and take a look!

Location: Library Galleria

Abstract: Theory-internal work on opacity in phonology has been focused on the challenges these interactions present for one theory (rules, constraints) versus another. But there has also been interest in studying the formal, invariant properties of opaque and other process interactions (Chandlee et al. 2018; Bakovic and Blumenfeld 2024), though these works crucially differ in their underlying assumptions. In this talk I will recontextualize Chandlee et al. (2018)'s result that opaque maps are ISL in light of Bakovic and Blumenfeld (2024)'s recent formal typology of process interactions, and this recontextualization will provide an answer to an open question about the k-value of an interaction map. I will then discuss the implications of this collective formal understanding of opacity for a recent model of lexicon and phonological grammar learning (i.e., Hua and Jardine 2021, Chandlee and Jardine to appear).


Speaker: Prof. Jane Chandlee, Associate Professor in the Department of Linguistics at Haverford College

Location: IACS Seminar room.
Abstract: Modern technologies enable enhanced integrity and privacy guarantees not just for data, but also for computation. This is perhaps most emphatically demonstrated by the steady rise of zero-knowledge proofs, which are short certificates that attest to the correctness of computations (e.g., an age verification check) without revealing any secret inputs (e.g., the birth date on a digital ID). This subtly powerful technology enables anonymous credentials, privacy-preserving machine learning, anonymous blockchains, and much more--making the question of efficient zero-knowledge proofs fundamental to modern secure systems. Echoing Moore's law for computing, zero-knowledge proofs have improved on this front by ten orders of magnitude in the last two decades. In this talk, I will discuss our work on overcoming a key bottleneck that has emerged in this development: memory efficiency.

Speaker: Abhiram Kothapalli is a postdoctoral scholar at the University of California, Berkeley, hosted by Sanjam Garg. He is a recent graduate of Carnegie Mellon University, where he earned his Ph.D. in Computer Science, advised by Bryan Parno. Previously, he was at the University of Illinois at Urbana-Champaign, where he earned his B.S. in Computer Science and B.S. in Mathematics. Kothapalli's research develops cryptographic techniques aimed at scaling expressive privacy and integrity guarantees across the internet.

Location: NCS 120
What comes after today's large language models and deep neural networks? Join the Computing Community Consortium (CCC) for a virtual 30-min community chat led by David Jensen, CCC Council Member and lead author of the new CCC whitepaper, Envisioning Possible Futures for AI Research. Jensen will explore paradigm-shifting AI Research Futures like Neuro-Symbolic, Embodied, Multi-Agent, and Quantum AI, and then open the floor to the audience for an engaging Q&A discussion.

Register here.

Abstract: Large Language Models (LLMs) have revolutionized how people interact with knowledge, offering unprecedented opportunities to accelerate the pace of scientific discovery. In this talk, I will discuss my research on the synergy between LLMs and scientific knowledge--specifically how these models extract, induce, and verify knowledge to automate the research lifecycle. First, I will cover our work on improving knowledge extraction from vast scientific literature, focusing on enabling models to comprehend long documents in a cost-efficient and comprehensive manner. I will describe a novel paradigm for representing document-level structured information as question-answer pairs and how we address the challenges of long-context understanding by leveraging global context through retrieval-augmented modeling. Next, I present our pioneering work on using LLMs for new scientific hypothesis generation. We introduce a framework employing reinforcement learning with fine-grained reward modeling and adaptive controllers.
This approach balances novelty, feasibility, and effectiveness to generate inspiring and actionable research hypotheses. Finally, I will discuss work on the first LLM Scientist for machine learning research. I will demonstrate how LLMs can move beyond hypothesis generation to participate in the execution and validation of scientific hypotheses, ensuring that the discovered knowledge is not only innovative but also grounded and verified.

Bio: Xinya Du is a tenure-track assistant professor at UT Dallas Computer Science Department. He earned a Ph.D. degree from Cornell University and was a Postdoctoral Research Associate at the University of Illinois (UIUC). He has also worked at Microsoft Research, Google Research, and Allen Institute AI. His research is on large language models, deep learning, and their applications in science.His work has been published in leading NLP and ML conferences (ACL, ICLR, NeurIPS). His research has received multiple recognitions, including a Best Paper Award at AAAI AI for Research and a Best Poster Award at ICML AI for Science workshop. His work was included in the list of Most Influential ACL Papers and has been covered by major media like New Scientist. He was named a Spotlight Rising Star in Data Science by the University of Chicago and is the recipient of several prestigious awards, including the Amazon Research Award, Cisco Research Award, Open Philanthropy Award, and the NSF CAREER Award.

Location: NCS 120
Please join us on Friday for a CSE 600 talk by CS Faculty, Stanley Bak. During this semester, please periodically check the CSE 600 schedule for the latest talk updates.

Title:  Formal Verification Methods for Cyber-Physical Systems and Neural Networks

Time: Friday 4/1, 2:40 PM

Location:  NCS 120

Abstract: Formal verification methods in Computer Science strive to prove properties about all possible executions of a system, and are an alternative development approach to testing when correctness is paramount. Traditionally these have been applied to hardware circuits, state-machine protocols, or software source code. Prof. Stanley Bak will discuss his research on extending formal verification approaches to more complex areas including cyber-physical systems and neural networks.


Speaker Bio: Stanley Bak is an assistant professor in the Department of Computer Science at Stony Brook University investigating the verification of autonomy, cyber-physical systems, and neural networks. He received a PhD from the University of Illinois at Urbana-Champaign (UIUC) in 2013, and worked for four years in the Verification and Validation (V&V) group in the Aerospace Systems Directorate at the Air Force Research Laboratory (AFRL). He received the AFOSR Young Investigator Research Program (YIP) award in 2020.


Abstract: Computer vision seeks to extract semantic and geometric information from images and videos, serving as the perceptual foundation for intelligent systems such as robots and autonomous vehicles. Over the past decade, deep learning has driven remarkable progress in the field, advancing capabilities from 2D recognition to 3D reconstruction. However, the current purely data-driven paradigm faces fundamental challenges, including data inefficiency, curse of high dimensionality, and limited understanding of visual entities beyond individual objects.

In this talk, I will present my recent research on modeling and learning rich visual structures to address these challenges. First, I will introduce a novel framework that integrates explicit visual dependency modeling with deep learning for 2D and 3D dense prediction. Next, I will demonstrate how unfolding the manifold structure of visual data enables unsupervised semantic segmentation. Finally, I will present a recent project that represents, parses, and learns the geometric compositionality of 3D objects to facilitate self-supervised part-whole reconstruction. Through these efforts, I aim to bridge the gap between data-driven deep learning and visual structure modeling, paving the way for more efficient, generalizable, and interpretable computer vision models.

Bio: Dr. Wei Tang is an Assistant Professor in the Department of Computer Science at the University of Illinois Chicago (UIC). He obtained his Ph.D. in Electrical Engineering from Northwestern University, where his dissertation was honored with a Best Dissertation Award. His research interests include computer vision, digital image processing, and machine learning. Dr. Tang has served as an associate editor for several international journals, including Pattern Recognition and Machine Vision and Applications, and as an area chair for leading conferences, including CVPR, ICCV, and WACV. His research has been funded by the National Science Foundation (NSF) and industry partners such as Motorola and Wormpex AI Research.


Location: NCS 115

Zoom: https://stonybrook.zoom.us/j/4624091659?omn=95178138684&jst=3