Abstract:
Large language models (LLMs) have transformed the way humans write code, bringing unprecedented automation to software development. In this talk, I will first provide an overview of my research on enhancing LLMs' code intelligence, optimizing each step of the development pipeline towards more complex software engineering tasks. I will then delve into my key contributions, focusing on how to equip LLMs with a deeper, more comprehensive understanding of software programs. Finally, I will discuss the future of AI-driven software engineering, envisioning a new era of automation that is more reliable, intelligent, and cost-efficient.

Bio:
Yangruibo (Robin) Ding is a Ph.D. candidate in the Department of Computer Science at Columbia University. His research is at the intersection of Software Engineering and Machine Learning, focusing on developing large language models (LLMs) for code. He trains LLMs to generate, analyze, and refine software programs and constructs benchmarks to systematically evaluate LLMs in solving software engineering tasks. He also studies how to improve LLMs' reasoning capability to tackle complex programming tasks, such as debugging and patching. His interdisciplinary research has been published in top-tier conferences of software engineering, programming languages, natural language processing, and machine learning. He won an ACM SIGSOFT Distinguished Paper Award, an IEEE TSE Best Paper Runner-up, and received an IBM Ph.D. Fellowship.
Location:
NCS 120
The Empirical Methods in Natural Language Processing (EMNLP) conference is a premier international academic conference in the field of artificial intelligence and natural language processing (NLP). Organized annually by the Association for Computational Linguistics (ACL) special interest group on linguistic data (SIGDAT), it focuses on research that uses empirical methods to solve language processing problems.

For more information, and registration, visit the official website.

This workshop is intended for researchers, practitioners, students, and industry professionals in AI, robotics, machine learning, human-robot interaction, and related fields.

Workshop Overview:

Instead of learning from data alone, an embodied AI system learns through its movements, sensors, and interactions with the environment. This form of active, experience-based learning, informed by ongoing self-evaluation of its own abilities, enables embodied AI systems to adapt on the fly, understand context rather than just commands, and collaborate with humans in more natural and trustworthy ways.

Workshop Goals:

  1. Foster interdisciplinary dialogue across AI, robotics, and cognitive science.
  2. Identify key challenges and future research directions in embodied intelligence.
  3. Examine the role of embodiment in advancing toward AGI.

This workshop is Invitation-only. Please email Dr. IV Ramakrishnan (ram@cs.stonybrook.edu) to attend.

Read the announcement: https://mcusercontent.com/237207911c0fd4c1f78dd8524/files/070dec2e-a2f5-143e-0fe2-c4ebecdb5193/Embodied_AI_Workshop_Invitation_.pdf


Please join University Libraries on March 29 at 1:00 via Zoom as we welcome Dr. Zhang, SUNY Empire Innovation Professor at SBU's Power Lab. This lab is pioneering the research of coordinated networked microgrids (NMs) that can possibly help to restore neighboring distribution grids after a major blackout. That these NMs hold promise to significantly enhance the day-to-day reliability of the power grids, we are proud to host Dr. Zhang as a member of our STEM Speaker Series. Registration required.
https://library.stonybrook.edu/library-events/stem-speaker-series-ai-enabled-provably-resilient-networked-microgrids-with-dr-peng-zhang/
Communication-Efficient Heterogeneity-Aware Machine Learning System and Architecture by Xuehai Qian

ABSTRACT: The key success of deep learning is the increasing size of models that can achieve high accuracy. At the same time, it is difficult to train the complex models with large data sets. Therefore, it is crucial to accelerate training with distributed systems and architectures, where communication and heterogeneity are two key challenges. In this talk, I will present two heterogeneity-aware decentralized training protocols without communication bottleneck. Specifically, Hop supports arbitrary iteration gap between workers by novel queue-based synchronization which can tolerate heterogeneity with system techniques. Prague uses randomized communication to tolerate heterogeneity with a new training algorithm based on partial reduce -- an efficient communication primitive. If time permits, I will present the systematic tensor partitioning for training on heterogeneous accelerator arrays (e.g., GPU/TPU). We believe that our principled approaches are crucial for achieving high-performance and efficient distributed training.

BIO: Xuehai Qian is an assistant professor at University of Southern California. His research interests include domain-specific systems and architectures, performance tuning and resource management of cloud systems and parallel computer architectures. He received his PhD from the University of Illinois Urbana Champaign and was a postdoc at UC Berkeley. He is the recipient of W.J Poppelbaum Memorial Award at UIUC, NSF CRII and CAREER Award, and the inaugural ACSIC (American Chinese Scholar In Computing) Rising Star Award.
Postmortem Program Analysis from a Conventional Program Analysis Method to an AI-assisted Approach

Abstract: Despite the best efforts of developers, software inevitably contains flaws that may be leveraged as security vulnerabilities. Modern operating systems integrate various security mechanisms to prevent software faults from being exploited. To bypass these defenses and hijack program execution, an attacker needs to constantly mutate an exploit and make many attempts. While in their attempts, the exploit triggers a security vulnerability and makes the running process abnormally terminate.

After a program has crashed and abnormally terminated, it typically leaves behind a snapshot of its crashing state in the form of a core dump. While a core dump carries a large amount of information, which has long been used for software debugging, it barely serves as informative debugging aids in locating software faults, particularly memory corruption vulnerabilities. As such, previous research mainly seeks fully reproducible execution tracing to identify software vulnerabilities in crashes. However, such techniques are usually impractical for complex programs. Even for simple programs, the overhead of fully reproducible tracing may only be acceptable at the time of in-house testing.

In this talk, I will discuss how we tackle this issue by bridging program analysis with artificial intelligence (AI). More specifically, I will first talk about the history of postmortem program analysis, characterizing and disclosing their limitations. Second, I will introduce how we design a new reverse-execution approach for postmortem program analysis. Third, I will discuss how we integrate AI into our reverse-execution method to escalate its analysis efficiency and accuracy. Last but not least, as part of this talk, I will demonstrate the effectiveness of this AI-assisted postmortem program analysis framework by using massive amounts of real-world programs.

Bio: Dr. Xinyu Xing is an Assistant Professor at Pennsylvania State University. His research interests include exploring, designing and developing new program analysis and AI techniques to automate vulnerability discovery, failure reproduction, vulnerability diagnosis (and triage), exploit and security patch generation. His past research has been featured by many mainstream media and received the best paper awards from ACM CCS and ACSAC. Going beyond academic research, he also actively participates and hosts many world-class cybersecurity competitions (such as HITB and XCTF). As the founder of JD-OMEGA, his team has been selected for DEFCON/GeekPwn AI challenge grand final at Las Vegas. Currently, his research is mainly supported by NSF, ONR, NSA and industry partners.
https://stonybrook.zoom.us/j/99820812332?pwd=c05BSTVLNmw3L04yZjdEcG5pem1OZz09 Speaker: Alexei Koulakov of Cold Spring Harbor Laboratory Brain evolution as a machine learning problem We have entered a golden age of artificial intelligence research, driven mainly by the advances in ANNs over the last decade or so. Applications of these techniques--to machine vision, speech recognition, autonomous vehicles, machine translation and many other domains--are coming so quickly that many observers predict that the long-elusive goal of Artificial General Intelligence (AGI) is within our grasp. However, we still cannot build a machine capable of building a nest, stalking prey, or loading a dishwasher. I will describe several projects, ranging from theories of evolution of neural development to the perception of smells, in which we are attempting to understand the algorithms that the nervous system is using to solve some of these challenging problems.
Please join us on Zoom for our next event in the Fall 2025 Stony Brook School of Nursing Research Seminar Series presented by our Office of Research and Innovation.

Topic: Responsible Artificial Intelligence: Promoting Health Equity for All

Speaker: Michael P. Cary, Jr., PhD, RN, FAAN.

Dr. Cary is a tenured Associate Professor at the Duke University School of Nursing. Dually trained as a health services researcher and applied health data scientist, Dr. Cary utilizes AI to investigate health disparities in aging populations, thereby promoting health equity and improving healthcare delivery. He co-directs HUMAINE™, an initiative dedicated to equipping nurses and healthcare professionals with the knowledge and skills necessary for the responsible use of AI in clinical practice.

Register: https://web.cvent.com/event/057978a5-a770-4de5-aca5-ad00287e4902/summary










Abstract:
Quantifying similarity is a central notion in science and data analysis, pervading everything from phylogenetic trees to the foundation of clustering. Unfortunately, despite being examined and applied for decades, traditional similarity and distance metrics have fundamental drawbacks. The key problem is that all of them are only defined over pairs of objects, so they scale quadratically when one tries to compare N objects. The present explosion in the amount of data available to us requires new ways to process information, and while some current algorithms can handle millions of points, we need alternatives applicable to billions. This is what motivated us to develop a new framework that can compare any number of objects at the same time. With this, we achieve an unprecedented linear scaling when comparing multiple objects. Here we will discuss the main properties of this formalism, along with its applications in drug design and to the analysis of Molecular Dynamics (MD) simulations. Our indices have proven to be incredibly versatile when applied to chemical space exploration and visualization, allowing us to rigorously quantify the chemical diversity of very large molecular libraries. This has led to the creation of several algorithms to sample important regions in chemical space, including a more efficient way of identifying the prevalence of activity cliffs. Additionally, our indices provide a convenient route to sample complex MD trajectories, allowing to identify representative structures very efficiently. Moreover, we can also cluster biological ensembles in a more robust way than with standard algorithms, which has led to our group's work on MDANCE, a very flexible and efficient open-source clustering module. Drop by if you want to know how we clustered one billion molecules!


Speaker:
Assistant Professor, Department of Chemistry and Quantum Theory Project
University of Florida, Gainesville
Website: https://quintana.chem.ufl.edu/

Location:
Laufer Center Lecture Hall 101