What AI tools are available to help with the scholarly research process? Are they helpful? What do they do and is it worth the time and energy to try them out? Join librarian Christine Fena to explore and compare established and emerging AI research tools such as Elicit, Scite, Consensus, and Undermind. The online workshop will provide a starting point to understanding what these tools are, the basics of how they work, and how AI research assistants might bring changes to your search process in the future. All are welcome!



Register here via Zoom.
TITLE: Towards a Theory of Encode/Decoder Architectures by Andrej Risteski of CMU

ABSTRACT: A common choice of architecture in representation learning (i.e., learning a good embedding of the data) is an encoder/decoder architecture, which tries to map a part of the input into a good latent representation (via an encoder), and predict the remaining part of the input (via a decoder). Two common examples are universal machine translation: where one tries to learn to translate between any pair of a set of languages via a common latent language, given paired up corpora for only a part of the pairs; and contextual encoders -- where one tries to predict a part of the image, given the rest of the image.
 
We will give a framework for analyzing the sample complexity of such architectures -- i.e., how many pairs of languages do we need to have paired up corpora for? How many image prediction tasks do we have to solve to get a good representation?
As artificial intelligence continues to transform higher education and the world beyond, how are students engaging with this change? Join us for a student-led discussion that explores how AI is influencing academic integrity, learning practices, and students' perspectives on its role in future workplaces.

Our panelists will share their experiences and reflections on questions such as:
1. What counts as appropriate and inappropriate use of AI in coursework?
2. How do faculty approach AI and talk about its implications in class?
3. What does AI mean for students' learning and ethical decision-making?
4. How are students building their understanding of AI tools and their potential uses in professional contexts?

This conversation offers an authentic look at how students are navigating the promises and challenges of AI--both in their studies and as they look ahead to applying these technologies responsibly in their fields.

Register here.
Postmortem Program Analysis from a Conventional Program Analysis Method to an AI-assisted Approach

Abstract: Despite the best efforts of developers, software inevitably contains flaws that may be leveraged as security vulnerabilities. Modern operating systems integrate various security mechanisms to prevent software faults from being exploited. To bypass these defenses and hijack program execution, an attacker needs to constantly mutate an exploit and make many attempts. While in their attempts, the exploit triggers a security vulnerability and makes the running process abnormally terminate.

After a program has crashed and abnormally terminated, it typically leaves behind a snapshot of its crashing state in the form of a core dump. While a core dump carries a large amount of information, which has long been used for software debugging, it barely serves as informative debugging aids in locating software faults, particularly memory corruption vulnerabilities. As such, previous research mainly seeks fully reproducible execution tracing to identify software vulnerabilities in crashes. However, such techniques are usually impractical for complex programs. Even for simple programs, the overhead of fully reproducible tracing may only be acceptable at the time of in-house testing.

In this talk, I will discuss how we tackle this issue by bridging program analysis with artificial intelligence (AI). More specifically, I will first talk about the history of postmortem program analysis, characterizing and disclosing their limitations. Second, I will introduce how we design a new reverse-execution approach for postmortem program analysis. Third, I will discuss how we integrate AI into our reverse-execution method to escalate its analysis efficiency and accuracy. Last but not least, as part of this talk, I will demonstrate the effectiveness of this AI-assisted postmortem program analysis framework by using massive amounts of real-world programs.

Bio: Dr. Xinyu Xing is an Assistant Professor at Pennsylvania State University. His research interests include exploring, designing and developing new program analysis and AI techniques to automate vulnerability discovery, failure reproduction, vulnerability diagnosis (and triage), exploit and security patch generation. His past research has been featured by many mainstream media and received the best paper awards from ACM CCS and ACSAC. Going beyond academic research, he also actively participates and hosts many world-class cybersecurity competitions (such as HITB and XCTF). As the founder of JD-OMEGA, his team has been selected for DEFCON/GeekPwn AI challenge grand final at Las Vegas. Currently, his research is mainly supported by NSF, ONR, NSA and industry partners.
The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors. Each seminar will consist of multiple short talks (around 15 minutes) by multiple students. Students can register for 1 credit for CSE656. Registered students must attend and present a minimum of 2 talks. Everyone else is welcome to attend. Fill in https://forms.gle/q6UG9ygauLp2a8Po8 to subscribe to our mailing list for further announcement.
Discover how U.S. Census Bureau Tools can help you find free data for your research projects, community, and more. See how to access the latest American Community Survey and 2020 Census data for various geographies including New York City and Long Island at data.census.gov. Learn about Community Resilience Estimates and how to navigate My Community Explorer; an interactive map-based tool which highlights demographic and socioeconomic data that measure inequality. This session will involve live demonstrations and hands-on exercises for participants. Registrants will receive the Zoom link one day prior to the event.

Please Register for SBU Libraries' AI Club: Exploring Census Data here.

Abstract: Much like other AI for Science domains, polymer design poses significant challenges. It requires grounding in empirical data and physical laws, precise handling of domain-specific structured representations, and compositional reasoning over multiple interacting constraints--all while working with limited data.

To address these limitations, we introduce PolyBench, a large-scale benchmark comprising over 125K polymer design and analysis tasks grounded in verified experimental and synthetic data. PolyBench includes tasks created from a wide range of data sources and presents diverse structural, property-driven, and synthesis-oriented reasoning problems. Tasks in PolyBench are organized from simple to complex analytical reasoning problems, enabling generalization tests and includes diagnostic probes to evaluate model capabilities. Additionally, to support effective domain alignment, we propose a knowledge-augmented reasoning distillation framework that enriches the dataset with structured chain-of-thought supervision derived from expert-informed reasoning strategies.

Small language models (7B-14B parameters) trained on PolyBench substantially outperform comparably sized baselines and, in many cases, exceed the performance of larger closed-source frontier models on polymer reasoning tasks, while also demonstrating improved transfer to external polymer benchmarks. Last, we conduct a diagnostic study that reveals a compositionality gap: despite strong performance on decomposed sub-questions, models struggle to integrate multiple interacting constraints and intermediate reasoning steps, highlighting fundamental limitations in current scientific language models.

Speaker: Dikshya Mohanty

Location: NCS 115/Online

Zoom: https://stonybrook.zoom.us/j/94746001760?pwd=BCAd8gu7cXLn3PXM6kkbh11V6r0Mr7.1
Meeting ID: 947 4600 1760 Passcode: 987917

CSE 656 Seminars in Computer Vision - Wednesdays 11:30am-12:50pm, Room NCS 120

The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 10 minutes) by multiple people. Students can register for 1 credit for CSE656. Registered students must attend and present a minimum of 2 or 3 talks. Everyone else is welcome to attend. Fill in https://forms.gle/pCVXovgfMfQwGqG38 to subscribe to our mailing list for further announcement.

The first meeting will be Wed Jan 29 at 11.30am, room 120 New CS. The meeting will deal with organizational matters and we will start right away with some presentations. Send David Paredes Merino <dparedesmeri@cs.stonybrook.edu> an email if you are interested but cannot attend the first meeting. Please forward to people outside the CS department that you think might be interested.