Stony Brook University Libraries invites students, faculty, & staff to join a conversation about how AI is transforming the private sector workforce. As AI tools move from experimentation to everyday business use, companies are rethinking roles, skill sets, leadership, and long-term strategy. This discussion-based event will focus on the fast-paced changes and directions at tech companies and their possible impact. This event will be particularly relevant for students preparing for an AI influenced job market and how to position themselves for opportunities in a rapidly evolving professional landscape.

The discussion will be led by Tariq Khan, Senior Director of Private Cloud Solutions at Hewlett Packard Enterprise. Tariq is a technology leader and architect with experience across private cloud, hybrid cloud, and data center platforms. He is responsible for shaping the technology architecture and strategic direction of HPE's Private Cloud offerings across on premises and cloud integrated environments.

Light refreshments will be served.


Location: Melville Library, NRR, Learning Lab
Abstract: Many unresolved legal questions over LLMs and copyright center on memorization: whether specific training data have been encoded in the model's weights during training, and whether those memorized data can be extracted in the model's outputs. While many believe that LLMs do not memorize much of their training data, recent work shows that substantial amounts of copyrighted text can be extracted from open-weight models. However, it remains an open question if similar extraction is feasible for production LLMs, given the safety measures these systems implement. We investigate this question using a two-phase procedure: (1) an initial probe to test for extraction feasibility, which sometimes uses a Best-of-N (BoN) jailbreak, followed by (2) iterative continuation prompts to attempt to extract the book. We evaluate our procedure on four production LLMs -- Claude 3.7 Sonnet, GPT-4.1, Gemini 2.5 Pro, and Grok 3 -- and we measure extraction success with a score computed from a block-based approximation of longest common substring (nv-recall). With different per-LLM experimental configurations, we were able to extract varying amounts of text. For the Phase 1 probe, it was unnecessary to jailbreak Gemini 2.5 Pro and Grok 3 to extract text (e.g, nv-recall of 76.8% and 70.3%, respectively, for Harry Potter and the Sorcerer's Stone), while it was necessary for Claude 3.7 Sonnet and GPT-4.1. In some cases, jailbroken Claude 3.7 Sonnet outputs entire books near-verbatim (e.g., nv-recall=95.8%). GPT-4.1 requires significantly more BoN attempts (e.g., 20X), and eventually refuses to continue (e.g., nv-recall=4.0%). Taken together, our work highlights that, even with model- and system-level safeguards, extraction of (in-copyright) training data remains a risk for production LLMs.

Speaker: Xinyue

Location: CS2311
Hosted by the College of Business faculty and staff, this virtual session will explore how our MS in Business Analytics & Intelligence program can enable you to become a future leader with AI expertise. This session will also cover the program's benefits, curriculum, cost, and admission requirements.

Register here.

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

We meet once a month at noon in CDSD's Training Room (building 725, room 2-124) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Multilayer machine-learning framework for screening catalytic activity and selectivity

Abstract: Machine learning (ML) studies based on quantum chemical datasets have emerged as a powerful tool for accelerating catalyst discovery. However, their potential applications remain limited by high costs, low data quality, and restricted reliability. Here, we report a multilayer binary classification framework (MLBC) for screening catalytic performance in multistep processes with ML models. Key features include low-cost synthetic data generation via kinetic Monte Carlo simulations, high-quality data that capture catalytic behavior under reaction conditions, robust treatment of imbalanced data distributions, and descriptor selection that enhances reliability and interpretability. Using carbon dioxide (CO 2 ) hydrogenation to methanol (CH 3 OH) on copper (Cu)-based catalysts as a case study, the MLBC framework outperforms conventional ML models by demonstrating high reliability and strong generalization in classifying systems with activity and methanol selectivity exceeding those of Cu. In addition, feature analysis reveals that control over the critical transition steps between competing pathways governs activity and selectivity in CO 2 hydrogenation.

In addition to our speaker, we will have a number of CDS staff in attendance with expertise in AI methods and applications including image analysis, foundation models development, and inverse problem solving.

Biography: Dr. Wenjie Liao is a research associate in the Chemistry Division at Brookhaven National Laboratory, where he develops computational and machine-learning approaches to understand and improve heterogeneous catalysts. He earned his Ph.D. from Stony Brook University. His research focuses on identifying active sites, mapping reaction pathways, and predicting catalyst performance, with the goal of accelerating the discovery of more efficient catalysts for energy and chemical transformations.

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1604383624?pwd=ffQ5cUPNxTI7nzClKQO6cnsNbhF9Vf.1

Meeting ID: 160 438 3624
Passcode: 558449

Please Note: Due to a funding shortfall, we are for the time being no longer able to provide pizza and sodas for these events. We will have coffee though, and all are of course welcome to bring their lunch.

Abstract: Large Language Model (LLM) agents have demonstrated remarkable generalization capabilities across multi-domain tasks. Existing agent tuning approaches typically employ supervised finetuning on entire expert trajectories. However, behavior-cloning of full trajectories can introduce expert bias and weaken generalization to states not covered by the expert data. Additionally, critical steps--such as planning, complex reasoning for intermediate subtasks, and strategic decision-making--are essential to success in agent tasks, so learning these steps is the key to improving LLM agents. For more effective and efficient agent tuning, we propose ATLAS that identifies the critical steps in expert trajectories and finetunes LLMs solely on these steps with reduced costs. By steering the training's focus to a few critical steps, our method mitigates the risk of overfitting entire trajectories and promotes generalization across different environments and tasks. In extensive experiments, an LLM finetuned on only 30% critical steps selected by ATLAS outperforms the LLM finetuned on all steps and recent open-source LLM agents. ATLAS maintains and improves base LLM skills as generalist agents interacting with diverse environments.

Speaker: Bijoy

Location: Old Computer Science Building - CS2311
Making sense of Twitter @ Bloomberg presented by Daniel Preotiuc-Pietro

ABSTRACT: The Bloomberg Terminal has provided ways for investors and journalists to sift through and understand the immense volume of tweets and discover financially-relevant content ever since the SEC approved the use of Twitter for company disclosures back in 2013.

In the first part of the talk, I will showcase how tweets impact financial markets and how Bloomberg is using Natural Language Processing methods to identify financially relevant tweets that move the markets. Our processing pipeline feeds directly to clients, journalists in the newsroom and powers several news analytic products offered by the company including trending companies and consumer sentiment for publicly traded equities.

However, understanding user pragmatic intent in individual tweets would allow us to gain deeper insights and enable new applications. I will present several recent research studies focused on understanding intent including identifying complaints and the roles with which vulgarity is used in social media and how these can help improve applications such as sentiment analysis and hate speech detection.

BIO: Daniel Preotiuc-Pietro is a Senior Research Engineer and Team Lead at Bloomberg LP, where he works on analyzing and building models for real-world large scale social media and news mining and information extraction. His research interests are focused on understanding the social and temporal aspects of text, especially from social media, with applications in domains such as Social Psychology, Law, Political Science and Journalism. Several of his research studies were featured in popular press including the Washington Post, BBC, New Scientist, Scientific American or FiveThirtyEight. He is a co-organizer of the Natural Legal Language Processing workshop series. Prior to joining Bloomberg LP, Daniel was a postdoctoral researcher at the University of Pennsylvania with the interdisciplinary World Well Being Project and obtained his PhD in Natural Language Processing and Machine Learning at the University of Sheffield, UK.
The Natural Language Processing Reading Group at Stony Brook University meets weekly to discuss recent research papers in NLP and related fields.
Join the Google Group here.

Title: J-Space

Abstract: The rise of the term mechanistic interpretability has accompanied increasing interest in understanding neural models -- particularly language models. However, this jargon has also led to a fair amount of confusion. So, what does it mean to be mechanistic? We describe four uses of the term in interpretability research. The most narrow technical definition requires a claim of causality, while a broader technical definition allows for any exploration of a model's internals. However, the term also has a narrow cultural definition describing a cultural movement. To understand this semantic drift, we present a history of the NLP interpretability community and the formation of the separate, parallel mechanistic interpretability community. Finally, we discuss the broad cultural definition -- encompassing the entire field of interpretability -- and why the traditional NLP interpretability community has come to embrace it. We argue that the polysemy of mechanistic is the product of a critical divide within the interpretability community.

Location: NCS 220

Zoom Link: https://stonybrook.zoom.us/j/92942833294?pwd=ysKMaQx7Hp0lkq5nOb3zJ9ivNZVvLv.1&jst=2
Abstract: Modern language agents often need to solve tasks requiring long-horizon, multi-turn interactions, where they retrieve external information, adapt to observations, and answer interdependent queries. Yet, most LLM systems rely on full-context prompting, appending all past turns regardless of their relevance. This leads to un-bounded memory growth, increased computational costs, and degraded reasoning performance on out-of-distribution input lengths due to LLM forgetting the context. We introduce MEM1, an end-to-end reinforcement learning framework that enables agents to operate with constant context size when solving long multi-turn tasks. At each turn, MEM1 updates a compact shared internal state that jointly supports memory consolidation and reasoning. Leveraging reinforcement learning (RL) and rollout trajectory truncation, we train a MEM1 agent to develop internal states that integrate prior memory with new observations from the environment while strategically discarding irrelevant or redundant information. Experiments across three domains, including internal retrieval QA, open-domain web QA, and multi-turn web shopping, show that MEM1-7B improves performance by 3.5x while reducing memory usage by 3.7x compared to Qwen2.5-14B-Instruct on an augmented multi-hop QA dataset with 16 objectives in each task, and generalizes beyond the training horizon. Our results demonstrate the promise of reasoning-driven memory consolidation as a scalable alternative to existing solutions for training long-horizon task-solving agents that involve multiple interactions, where both efficiency and performance are optimized.

Speaker: Yiyang Feng

Location: CS2311

Scaling the NY AI Innovation Ecosystem

The State University of New York at Stony Brook will bring together leading AI experts to promote a future where AI drives responsible progress. This two-day event will provide a significant opportunity to explore the future of AI, exchange ideas, and connect with those at the forefront of research and deployment. We invite faculty, staff, and students from all SUNY institutions and beyond, as well as industry AI practitioners and policymakers to attend.

Recognized AI experts from academia, industry, and government will present on topics such as AI applications, innovative developments in research and technology, workforce development, as well as ethical and societal impacts.

A 90-minute poster session is included in the schedule. If you would like to submit an abstract for consideration, please see the Call for Abstracts. The poster session segment of the symposium will be held in honor of the Inauguration of Dr. Andrea Goldsmith, the State University of New York at Stony Brook's seventh President. Poster printing for all participants will be covered by the Inauguration Planning Committee. SUNY students presenting posters are also eligible for travel reimbursement.

We kindly ask faculty to encourage their students to attend and to submit their work for presentation.

For additional information and to register, visit the symposium website. Please direct any questions to suny-ai-symposium-sbu@stonybrook.edu.

Register.

Abstract: Self-supervised representation learning (SRL) has emerged as a pivotal advancement in machine learning, offering high-quality data representations without the need for labeled datasets. While SRL has demonstrated enhanced adversarial robustness compared to supervised learning, its resilience against other attack types, particularly backdoor attacks, remains an open question. Recent studies have revealed potential vulnerabilities in SRL, underscoring the necessity for a comprehensive security analysis. However, existing research often extrapolates attacks from supervised learning paradigms, neglecting the unique challenges and opportunities inherent to self-supervised mechanisms.

This thesis proposal aims to address three critical objectives in the realm of self-supervised learning: (1) exploring novel attack vectors, (2) implementing and evaluating practical attacks, and (3) developing robust countermeasures. We focus on two key SRL paradigms: Contrastive Learning and Diffusion Models. For Contrastive Learning, we synthesize existing security vulnerabilities and introduce innovative attack vectors, such as CTRL, to uncover distinctive risks. We conduct a comparative analysis of contrastive and supervised learning approaches in their defense against these threats, exploring potential safeguards and highlighting the limitations of current protective measures in self-supervised contexts. Regarding Diffusion Models, we demonstrate inherent vulnerabilities in their application to adversarial purification.

Our research aims to illuminate the unique challenges posed by emerging attack vectors in self-supervised learning, fostering technical advancements to address underlying security risks in real-world applications. By contributing to the development of more resilient and secure self-supervised representation learning systems, we seek to enhance their reliability and trustworthiness in practical scenarios. This comprehensive examination of SRL's security landscape will provide valuable insights for the broader machine-learning community and pave the way for more robust AI systems.

Join here.