What Does Learning Mean? presented by Jeffrey Heinz

ABSTRACT
When we develop learning algorithms, what computational problems are we solving? In this talk, I discuss different answers that have been proposed for this question, and discuss some of the consequences for machine learning and artificial intelligence. The main lessons I offer are that (1) feasible solutions to learning problems require careful consideration of a target class C of functions, (2) that such a class C cannot include all functions, or even all computable functions, and so many logically possible functions must be outside of C and (3) class C must have significant structure which the solutions take advantage of. These main ideas are motivated and illustrated from modeling language acquisition and the related problem of grammatical inference from example sequences belonging to formal languages.
Abstract: Many unresolved legal questions over LLMs and copyright center on memorization: whether specific training data have been encoded in the model's weights during training, and whether those memorized data can be extracted in the model's outputs. While many believe that LLMs do not memorize much of their training data, recent work shows that substantial amounts of copyrighted text can be extracted from open-weight models. However, it remains an open question if similar extraction is feasible for production LLMs, given the safety measures these systems implement. We investigate this question using a two-phase procedure: (1) an initial probe to test for extraction feasibility, which sometimes uses a Best-of-N (BoN) jailbreak, followed by (2) iterative continuation prompts to attempt to extract the book. We evaluate our procedure on four production LLMs -- Claude 3.7 Sonnet, GPT-4.1, Gemini 2.5 Pro, and Grok 3 -- and we measure extraction success with a score computed from a block-based approximation of longest common substring (nv-recall). With different per-LLM experimental configurations, we were able to extract varying amounts of text. For the Phase 1 probe, it was unnecessary to jailbreak Gemini 2.5 Pro and Grok 3 to extract text (e.g, nv-recall of 76.8% and 70.3%, respectively, for Harry Potter and the Sorcerer's Stone), while it was necessary for Claude 3.7 Sonnet and GPT-4.1. In some cases, jailbroken Claude 3.7 Sonnet outputs entire books near-verbatim (e.g., nv-recall=95.8%). GPT-4.1 requires significantly more BoN attempts (e.g., 20X), and eventually refuses to continue (e.g., nv-recall=4.0%). Taken together, our work highlights that, even with model- and system-level safeguards, extraction of (in-copyright) training data remains a risk for production LLMs.

Speaker: Xinyue

Location: CS2311
Abstract:
Artificial intelligence (AI)-based methods and computational materials science continue to make inroads into accelerated materials design and development. I will review Al-enabled advances made in the subfield of polymer informatics, with a particular focus on the design of application-specific practical polymeric materials. I will describe exemplar design attempts within a few critical and emerging application spaces, including materials designs for storing, producing, and conserving energy, and those that can prepare us for a sustainable economy powered by recyclable and/or biodegradable polymers. Al- powered workflows help efficiently search the staggeringly large chemical and configurational space of materials, using modern machine-learning (ML) algorithms to solve forward and inverse materials design problems. A practical informatics-based design protocol involves creating a set of application-specific target property criteria, building ML model predictors for those relevant target properties, enumerating or generating a tangible population of viable polymers, and selecting candidates that meet design recommendations. The protocol will be demonstrated for several energy and sustainability-related applications. Finally, I will offer an outlook on the lingering obstacles that must be overcome to achieve widespread adoption of informatics-driven protocols in industrial-scale materials development.

Speaker Bio:
Prof. Ramprasad is the Regents' Entrepreneur, Michael E. Tennenbaum Family Chair and Georgia Research Alliance Eminent Scholar in the School of Materials Science & Engineering at the Georgia Institute of Technology. He is also the CEO and co-founder of Matmerize, Inc. His area of expertise is the development and application of computational and machine learning tools to accelerate sustainable materials development aimed at energy production, storage and utilization. Prof. Ramprasad received his B. Tech. in Metallurgical Engineering at the Indian Institute of Technology, Madras, India, an M.S. degree in Materials Science & Engineering at the Washington State University, and a Ph.D. degree also in Materials Science & Engineering at the University of Illinois, Urbana-Champaign.
Prof. Ramprasad is a Fellow of the Materials Research Society, a Fellow of the American Physical Society, an elected member of the Connecticut Academy of Science and Engineering, and the recipient of the Alexander von Humboldt Fellowship and the Max Planck Society Fellowship for Distinguished Scientists. He has authored or co-authored over 300 peer-reviewed journal articles, 8 book chapters and 8 patents, and has delivered over 300 invited talks at Universities and Conferences worldwide. He is a member of the Editorial Advisory Boards of npj Computational Materials, ACS Materials Letters and Journal of Physical Chemistry A/B/C. He created and chaired the inaugural 2022 Gordon Research Conference on Computational Materials Science and Engineering.

Location: Room 301, Engineering Building
Fall 2025, Mondays 2 to 3:20 pm, NCS 220 and Zoom link to be announced soon.

The seminar will be jointly taught by Prof. Dimitris Samaras samaras@cs.stonybrook.edu.

The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision.

To enroll in this course, you must either: (1) be in the Ph.D. program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 15 minutes) by multiple students. Students can register for 1 credit for CSE656. Registered students must attend and present a minimum of 2 talks. Registered students must attend in person. Up to 3 absences will be excused. Everyone else is welcome to attend.

Abstract: Millions of individuals living in disadvantaged communities are burdened by poverty, illegal drug activities, health concerns, and the lack of reliable and affordable access to facilities (e.g., schools, hospitals, and transit stations). To address these societal problems efficiently with broad support, initiatives have called to engage agents (e.g., residents, community leaders, or stakeholders) and consider their preferences on community improvement decisions to make collective community decisions. In this talk, we will focus on our ongoing AI-empowered collective decision-making approaches to improve the accessibility of individuals to facilities by (a) locating facilities to provide essential services and (b) strengthening existing infrastructures via structural modifications (e.g., constructing new roads, bridges, multi-use paths, or shuttle services) subject to individuals' preferences on the locations of the facilities and which communities to improve access, respectively. In particular, we will discuss our (theoretical and algorithmic) studies on modeling these approaches under several settings (e.g., accounting for fairness and agent preferences) and designing fair, transparent, strategy proof, and (approximately) optimal mechanisms to elicit (true) individual preferences and determine collective community decisions in order to improve facility accessibility. Finally, we will discuss other ongoing and future collective decision-making efforts in urban planning and public health (i.e., our recent studies on substance use research) to improve communities.

Bio: Hau Chan is an assistant professor in the School of Computing at the University of Nebraska-Lincoln. He received his Ph.D. in Computer Science from Stony Brook University in 2015 and completed three years of Postdoctoral Fellowships, including at the Laboratory for Innovation Science at Harvard University in 2018. His main research lies in multi-agent aspects of AI for Society and Social Good, focusing on developing modeling and algorithmic foundations for tackling societal problems involving agents and predicting agent behavior in societal contexts, leveraging AI, game theory, mechanism design, and machine learning to better inform policymaking and (collective) decision-making. His team has been addressing societal challenges and fairness issues in various domains, including security (e.g., reducing vulnerability), public health (e.g., reducing substance use and homelessness), and urban planning (e.g., improving accessibility to public facilities), collaborating with domain experts. His research has been supported by NSF, NIH, and USCYBERCOM. He has received several Best Paper Awards at SDM and AAMAS and distinguished/outstanding SPC/PC member recognitions at IJCAI and WSDM. He has given tutorials and talks on computational game theory and mechanism design at venues such as AAMAS and IJCAI, including an Early Career Spotlight at IJCAI 2022. He has served as co-chairs for the AI and Social Good Track, Demonstration Track, Student Activities, Doctoral Consortium, Job Fair, Scholarships, Finance, and Diversity & Inclusion Activities at AAAI, AAMAS, and IJCAI.

Location: Old Computer Science, room 1310

Defending Software Systems from Cyber Attack Campaigns Presented by R. Sekar The DNC hack of 2016, the Equifax breach of 2017, and the spate of ransomware campaigns in 2019 demonstrate the formidable challenges we face in securing our network and software systems against highly stealthy and sophisticated adversaries. In this talk, I will describe two avenues of research we have been pursuing to help tilt the table against such powerful adversaries. The first is software hardening techniques that make software vulnerabilities harder to exploit. To maximize their applicability and ease of use, our techniques are implemented into compilers, or they directly transform binary code. I will outline some of the exciting new developments we have had in this area over the years, including randomization, memory safety, information-flow tracking, control-flow integrity, and code-pointer integrity. We complement this first line of defense with techniques for analyzing and understanding attack campaigns that manage to slip past all deployed defenses. Our techniques can sift through logs consisting of hundreds of millions of events to zoom in on attack activity that may span just a few hundred events. I will describe our experience in mapping out several DARPA-sponsored red team attack campaigns.
The Provost's Lecture Series features talks by SUNY Distinguished Academy faculty members at Stony Brook University, showcasing the outstanding research and scholarship that is taking place at our institution.

Joe Mitchell

SUNY Distinguished Professor, Applied Mathematics and Statistics
Chair, Department of Applied Mathematics and Statistics, College of Engineering and Applied Sciences

A Case for Algorithms: A Computational Geometer's Perspective

Algorithms are all around us in every smart device and technology that has consumed our daily lives. As a computational geometer, I study algorithms to solve problems that involve a geometric perspective on data. I have observed that practically every technology and field of study has a need for effective algorithms involving geometric data. I reflect on some favorite algorithmic problems that are easy to visualize, but challenging to solve, and argue that the formal study of algorithms remains essential in the age of AI.

Reception to follow immediately after the talks.

Register here.