You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Learning Generalizable Program and Architecture Representations for Performance Modeling

Abstract: Performance modeling is an essential tool in many areas of computer science and engineering. However, existing performance modeling approaches have limitations, such as high computational cost, narrow flexibility, or restricted accuracy/generality. To address these limitations, this talk introduces PerfVec, a novel deep learning-based performance modeling framework that learns high-dimensional and independent/orthogonal program and microarchitecture representations. Once learned, a program representation can be used to predict its performance on any microarchitecture, and likewise, a microarchitecture representation can be applied in the performance prediction of any program. Additionally, PerfVec yields a foundation model that captures the performance essence of instructions, which can be directly used by developers in numerous performance modeling-related tasks without incurring its training cost. The evaluation demonstrates that PerfVec is more general and efficient than previous approaches. This talk will also introduce how PerfVec's design principles can benefit broader research areas.

Biography: Lingda Li is a computer scientist at Brookhaven National Laboratory. He is generally interested in computer architecture and programming model research, with focus on simulation/modeling, memory systems, and machine learning. Before joining BNL, he worked at the Department of Computer Science of Rutgers University as a postdoc to carry out GPGPU research. He obtained a PhD in computer architecture from the Microprocessor Research and Development Center at Peking University.

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1605837856?pwd=kYqJs4bVBt4E0cMCWR6GXH3wxzOoiw.1

Meeting ID: 160 583 7856
Passcode: 161580

https://stonybrook.zoom.us/j/99820812332?pwd=c05BSTVLNmw3L04yZjdEcG5pem1OZz09 Speaker: Alexei Koulakov of Cold Spring Harbor Laboratory Brain evolution as a machine learning problem We have entered a golden age of artificial intelligence research, driven mainly by the advances in ANNs over the last decade or so. Applications of these techniques--to machine vision, speech recognition, autonomous vehicles, machine translation and many other domains--are coming so quickly that many observers predict that the long-elusive goal of Artificial General Intelligence (AGI) is within our grasp. However, we still cannot build a machine capable of building a nest, stalking prey, or loading a dishwasher. I will describe several projects, ranging from theories of evolution of neural development to the perception of smells, in which we are attempting to understand the algorithms that the nervous system is using to solve some of these challenging problems.

Join University Libraries for an engaging panel discussion where we delve in and learn about the impacts of artificial intelligence on the 2024 US elections! Panelists are Paige Lord, Tom Costello, and Musa al-Gharbi. The discussion will be moderated by Library Dean, Karim Boughida. Co-sponsored by the Office of Diversity, Inclusion, and Intercultural Initiatives.

Please RSVP for Democracy in the Digital Age: AI's Influence on 2024 Elections here.
Abstract: Large language models (LLMs) may exhibit unintended or undesirable behaviors. Recent works have concentrated on aligning LLMs to mitigate harmful outputs. Despite these efforts, some anomalies indicate that even a well-conducted alignment process can be easily circumvented, whether intentionally or accidentally. Does alignment fine-tuning yield have robust effects on models, or are its impacts merely superficial? In this work, we make the first exploration of this phenomenon from both theoretical and empirical perspectives. Empirically, we demonstrate the elasticity of post-alignment models, i.e., the tendency to revert to the behavior distribution formed during the pre-training phase upon further fine-tuning. Leveraging compression theory, we formally deduce that fine-tuning disproportionately undermines alignment relative to pre-training, potentially by orders of magnitude. We validate the presence of elasticity through experiments on models of varying types and scales. Specifically, we find that model performance declines rapidly before reverting to the pre-training distribution, after which the rate of decline drops significantly. Furthermore, we further reveal that elasticity positively correlates with the increased model size and the expansion of pre-training data. Our findings underscore the need to address the inherent elasticity of LLMs to mitigate their resistance to alignment.

Speaker: Huajian Zhang

Location: CS2311
Join the Department of Computer Science as we welcome Lyle Ungar, University of Pennsylvania, who will be delivering a lecture on 'Measuring Cultural Variation using Natural Language Processing.' When: 11/08/24 @ 2:30 PM Where: New Computer Science Building, Room 120. Reception to follow. Abstract: Cultures vary widely in how they view the world, for example being more individualist or collectivist. Such cultural differences are, of course, reflected in the words that people use. We first show a variety of ways in which multilingual language models are not multicultural; they speak Hindi or Mandarin, but still think like Americans. In contrast, we then present a scalable method that uses embedding-derived lexica to successfully measure regional variation in culture. Bio: Lyle Ungar is a Professor of Computer and Information Science at the University of Pennsylvania, where he also holds secondary appointments in Psychology, Bioengineering, Genomics and Computational Biology, and Operations, Information and Decisions. His group uses natural language processing and explainable AI for psychological research, including analyzing social media and cell phone sensor data to better understand the drivers of physical and mental well-being. They are currently building socio-emotionally sensitive GPT-based tutors and coaches.
Abstract:

What is the nature of linguistic knowledge, and how is it acquired from limited data? In recent years, the program of subregular linguistics has identified formal language classes expressive enough to account for most phenomena in natural language but also sufficiently limited to be efficiently learned from positive data. An advantage to these formal learning algorithms is that they come with mathematically proven guarantees about their performance, and it is easy to reason about how and why they behave the way they do.

In this talk, I discuss the Multi Tier-based 2-Strictly Local Inference Algorithm (MT2SLIA), which probably learns the syntactically relevant class of 2-Factor Muti Tier-based Strictly Local (2FMSTL) tree languages. This algorithm efficiently learns from a polynomially-sized sample of positive data by identifying missing substructures and generalizing these as constraints over tiers in a principled manner.

I will introduce a working prototype implementation of this algorithm and demonstrate its behavior on a curated sample of natural language data to show how it can learn relevant syntactic patterns.

Bio:

Logan Swanson is a third year PhD student in the Department of Linguistics at Stony Brook University. He is advised by Dr. Jefferey Heinz and Dr. Thomas Graf. His interests include learning theory, computational syntax, and language change. His current research focuses on understanding the learning-theoretic elements of natural language by designing, implementing, and testing learning algorithms for linguistically relevant formal language classes.

*Please note: this seminar will be held in person (IACS Seminar Room w/ food provided) and online.

Join Zoom Meeting
https://stonybrook.zoom.us/j/95707958315?pwd=6ITUJ0ffCXjRJb4wpt0KMDTApfSLZ0.1

Meeting ID: 957 0795 8315
Passcode: 920473
Join a faculty development program to support instructors across campus with navigating/integrating AI in their courses. We're inviting interested faculty to participate in the grant project called Fostering Writing-to-Learn Skills with Critical AI Literacy: A Faculty Development and Student Support Program (funded through the AI3 Institute).

Time commitment and completion requirements :

  • Attend four sessions and a final symposium on the following dates/times:

    • Friday, September 12 from 11am - 12:30pm over Zoom

    • Friday, September 26 from 11am - 12:30pm over Zoom

    • Friday, October 10 from 11am - 12:30pm over Zoom

    • Friday, October 24 from 11am - 12:30pm over Zoom

    • Friday, November 14 from 10am - 1pm in Wang 201 - please note that this is an in person session only

  • Engage with online materials in Brightspace prior to each of the sessions (mainly to update a syllabus, assignment, or teaching strategy that you can share and discuss at the workshop)

Contact: Shyam Sharma, Christine Fena, and Rose Tirotta-Esposito with questions.

https://docs.google.com/document/d/1b51tvfK0HSOkCW7cwYq2nyyeeHtvBZYC7_XHv7Av8wQ/edit?tab=t.0
The Hudson River Estuary (HRE) and New York Bight (NYB) are closely connected, with HRE acting as crucial areas where many NYB marine species spawn and grow. Understanding how these biotic and abiotic environments interact, especially with rapid climate change, is key to better managing fisheries and conserving ecosystems. To better understand the HRE-NYB ecosystem, we develop a comprehensive ecosystem model that links physical and biological processes. Using data from long-term monitoring programs, we analyze ecological patterns and identify key factors regulating the ecosystem. We use this information to develop a model that mimics the food web from tiny plankton to large predators in the ecosystem. This model can help us better understand how changes in the environment, like rising temperatures, and human activities such as fishing affect marine lives and ecosystem over time. The insights from this model can support smarter fisheries management and efforts to conserve marine ecosystems in the HRE-NYB region.

IACS Student Seminar Speaker: Xiangyan Yang, Dept. of Applied Math & Statistics

Location: IACS Seminar Room or Zoom

Join Zoom Meeting: https://stonybrook.zoom.us/j/91650247483?pwd=fvAGEwadplJh7jFC5RWcdvZ5NWPJth.1
Meeting ID: 916 5024 7483
Passcode: 631055