University Libraries Presents

What new tools and extra powers are available to you through the library's subscription databases? Join faculty librarian Chris Kretz, Head of Academic Engagement, on a tour of the AI landscape from our database vendors.

Location: Melville Library, Central Reading Room, Lab A, or Online.

Register here to join.

CSE 600 Seminar Series | Fall 2025


Abstract: Large reasoning models have demonstrated capabilities to solve competition-level math problems, answer deep research questions, and address complex coding needs. Much of this progress has been enabled by scaling of data: pre-training data to learn vast knowledge, fine-tuning data to learn natural language reasoning, and RL environments to refine that reasoning. In this talk, I will describe the current LLM reasoning paradigm, its boundaries, and the future of LLM reasoning beyond scaling. First, I will describe the state of reasoning models and where I think scaling can lead to some additional (though perhaps limited) successes. I will then shift to discussing more fundamental issues with models that scale will not resolve in the next few years. I will touch on four current limitations: outdated knowledge, generator-validator gaps, limited creativity, and poor compositional generalization. In all cases, fundamental limitations of LLMs or of supervised learning in general make these problems challenging, inviting future study and novel solutions beyond scaling.

Bio: Greg Durrett is an associate professor in the Department of Computer Science and the Center for Data Science at New York University. His research is broadly in the areas of natural language processing and machine learning. Currently, his group's focus is on reasoning about knowledge in text, verifying correctness of generation methods, and studying how to make progress on problems that defy LLM scaling. He is a 2023 Sloan Research Fellow and a recipient of a 2022 NSF CAREER award. He has served in numerous roles for ACL conferences, recently as a member of the NAACL Board since 2024 and as Senior Area Chair for ACL 2025 and EMNLP 2025. He received his BS in Computer Science and Mathematics from MIT and his PhD in Computer Science from UC Berkeley, where he was advised by Dan Klein.

Abstract: Traditional questionnaires remain the primary method for assessing psychological outcomes and beliefs, capturing individuals' and populations' inner states. This dissertation presents an alternative computational method that overcomes key limitations in current mental health monitoring, particularly in spatiotemporal resolution, responses to major events, and automatic belief identification. By analyzing ∼1 billion Tweets from 2 million geo-located users, we created a big data pipeline for estimating depression and anxiety at the county-week level. These Language-Based Mental Health Assessments (LBMHA) demonstrated higher reliability and validity than traditional survey measures. Our approach effectively captured mental health trends and highlighted significant increases in mental illness following major events. Using the LBMHA pipeline, we conducted quasi-experiments, research designs that simulate randomized control trials, to generate explanations for mental health changes due to COVID-19 incidence/death. Utilizing these time-series analyses, we conducted discontinuity forecasting for community-specific anxiety shifts using statistical learning via ensemble and contextual models. To likewise investigate individual internal states, we created a novel task and annotated dataset for self belief language identification. Our fine-tuned language model for self-belief classification, despite its relatively small scale, outperformed GPT-4o. The self belief topics identified by our model successfully predicted depression, anxiety, and stress, offering insights into the relationship between self-conceptualization and mental health. The adoption of scalable language-based assessments with modern distributed computation presents a promising avenue for advancing community and individual mental health research.

Speaker: Siddharth Mangalik

https://stonybrook.zoom.us/j/91251321639?pwd=faggV5jZ7ByFDCFmnLXD3HiYxjQ1Eb.1&jst=2
TITLE: Towards a Theory of Encode/Decoder Architectures by Andrej Risteski of CMU

ABSTRACT: A common choice of architecture in representation learning (i.e., learning a good embedding of the data) is an encoder/decoder architecture, which tries to map a part of the input into a good latent representation (via an encoder), and predict the remaining part of the input (via a decoder). Two common examples are universal machine translation: where one tries to learn to translate between any pair of a set of languages via a common latent language, given paired up corpora for only a part of the pairs; and contextual encoders -- where one tries to predict a part of the image, given the rest of the image.
 
We will give a framework for analyzing the sample complexity of such architectures -- i.e., how many pairs of languages do we need to have paired up corpora for? How many image prediction tasks do we have to solve to get a good representation?

AI on Campus: Your Thoughts, Your Future

Join the Conversation: Share Your Thoughts about Learning, Academics, and AI

The world of college is changing fast, and Artificial Intelligence (AI) is at the center of it. We are part of the Institute on AI, Pedagogy, and the Curriculum with AAC&U, and we need to hear from the people AI affects most: you!

This is an open discussion for all students to share their honest experiences, their top concerns, and their best ideas about AI in our academic environment. We'll be diving into these key questions:

  • How can AI actually make learning better or easier? What opportunities do you see for using AI tools to enhance your assignments, research, or skills?

  • What are your biggest worries about AI? Is it about cheating, being graded fairly, or preparing for the job market? How is AI impacting your workload or stress levels?

  • What specific tools, workshops, or policies would help you use AI responsibly and successfully? (Think training, software, or clear rules.)

Dates/Times:

  • Wednesday, 2/4 at 2pm

  • Thursday, 2/5 at 12pm

Please register in advance for the Zoom link.

Can't Make It? Share Your Feedback!

Don't worry if you can't attend! You can still share your thoughts via video in our AI Zoom Room or via email: rose.tirotta-esposito@stonybrook.edu.

Videos will not be shared publicly and comments will only be shared in aggregate.

Your voice matters. Come tell us how AI is affecting your studies, your stress, and your success!

  • Dr. Rose Tirotta-Esposito (Assistant Provost; Director of CELT)

  • Dr. Elizabeth Hewitt (Associate Professor in the Department of Technology and Society (DTS) in the College of Engineering and Applied Sciences)

  • Chris Kretz (Associate Librarian and Head of Academic Engagement at SBU Libraries)

  • Prof. Rajiv Lajmi (Assistant Professor in the School of Health Professions and Chair of Applied Health Informatics)

  • Dr. Matthew Salzano (Assistant Professor in the Department of Communication in the School of Communication and Journalism)

Biomedical Informatics Grand Rounds Talk

Educational objectives:
  1. Explain how generative AI can write data-analysis code from natural-language prompts
  2. Describe how virtual computational laboratories can be built to explore clinicalquestions
  3. Identify the key challenges of applying AI models to protected healthcare data

Speaker: Dr. Janos Hajagos, Ph.D Chief Data Analytics and Research Assistant Professor, Biomedical Informatics, Stony Brook University

Remote Access: Join Zoom meeting
Meeting ID: 95617197636 Passcode: 924293


Time:
Sep 7, Tue, 11:00am EDT

Place:
NCS 220 or on Zoom (info below)

Title: Data-Driven Document Unwarping


Abstract:
Capturing document images is a common way to digitize and record physical documents due to the ubiquitousness of mobile cameras. To make text recognition easier, it is often desirable to digitally flatten a document image when the physical document sheet is folded or curved. However, unwarping a document from a single image in natural scenes is very challenging due to the complexity of document sheet deformation, document texture, and environmental conditions. Previous model-driven approaches struggle with inefficiency and limited generalizability. In this thesis, I investigate several data-driven approaches to tackle the document unwarping problem.

Data acquisition is the central challenge in data-driven methods. I first design an efficient data synthesis pipeline based on 2D image warping and train DocUNet, the pioneering data-driven document unwarping model, on the synthetic data. A benchmark dataset is also created to facilitate comprehensive evaluation and comparison. To improve the unwarping performance by training on more realistic data, I introduce the Doc3D dataset and DewarpNet. Supervised by 3D shape ground truth in Doc3D, DewarpNet is significantly better than DocUNet. DocUNet and DewarpNet depend on the synthetic data for the ground truth deformation annotation. To exploit the real-world images, I propose PaperEdge, a weakly supervised model trained with in-the-wild document images with easy-to-obtain boundary information. PaperEdge surpasses DewarpNet by utilizing both the synthetic data and weakly annotated real data in the Document In the Wild (DIW) dataset. Finally, I propose directly predicting the $uv$ parameterized 3D mesh of the document with 3D constraints and using the accessible 3D presentations like depth maps as training targets. Predicting the 3D mesh of the document solves the unwarping task and also benefits VR/AR applications.

Join Zoom Meeting
https://stonybrook.zoom.us/j/96440592912?pwd=ZU5waTdyUzRFNW5SRHM5ME84TWdFQT09

Meeting ID: 964 4059 2912
Passcode: 793149
One tap mobile
+16468769923,,96440592912# US (New York)
+13017158592,,96440592912# US (Washington DC)

Dial by your location
        +1 646 876 9923 US (New York)
        +1 301 715 8592 US (Washington DC)
        +1 312 626 6799 US (Chicago)
        +1 253 215 8782 US (Tacoma)
        +1 346 248 7799 US (Houston)
        +1 408 638 0968 US (San Jose)
        +1 669 900 6833 US (San Jose)
Meeting ID: 964 4059 2912
Find your local number: https://stonybrook.zoom.us/u/adxTt9ZbuJ

The SUNY AI Symposium brings together AI experts from across the state, in Western New York and around the country.


This two-day event showcases AI thought leaders, SUNY researchers, students and companies of all sizes who leverage AI to produce positive outcomes--with scientific discovery, business innovation and economic impact. Come curious, explore the fascinating world of AI and leave with connections to those at the forefront of innovation.

As artificial intelligence continues to transform higher education and the world beyond, how are students engaging with this change? Join us for a student-led discussion that explores how AI is influencing academic integrity, learning practices, and students' perspectives on its role in future workplaces.

Our panelists will share their experiences and reflections on questions such as:
1. What counts as appropriate and inappropriate use of AI in coursework?
2. How do faculty approach AI and talk about its implications in class?
3. What does AI mean for students' learning and ethical decision-making?
4. How are students building their understanding of AI tools and their potential uses in professional contexts?

This conversation offers an authentic look at how students are navigating the promises and challenges of AI--both in their studies and as they look ahead to applying these technologies responsibly in their fields.

Register here.
Spring 2025, Mondays 3.30 to 4.50 pm, NCS 220.

The seminar will be jointly taught by Prof. Chao Chen, chao.chen.1@stonybrook.edu and Prof. Dimitris Samaras samaras@cs.stonybrook.edu

The overall purpose of this seminar is to bring together people with interests
in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision.

To enroll in this course, you must either: (1) be in the Ph.D. program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 15 minutes) by multiple students. Students can register for 1 credit for CSE656. Registered students must attend and present a minimum of 2 talks. Registered students must attend in person. Up to 3 absences will be excused. Everyone else is welcome to attend.

Join here. Meeting ID: 927 2069 8658. Passcode: 130934.
.