The IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025 will be held from June 11th to June 15th, 2025, at the Music City Center, Nashville, TN. The IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR) is the premier annual computer vision event comprising the main conference and several co-located workshops and short courses. With its high quality and low cost, it provides an exceptional value for students, academics and industry researchers. Register here.
CSE 600 Seminar Series | Fall 2025


Abstract: Large reasoning models have demonstrated capabilities to solve competition-level math problems, answer deep research questions, and address complex coding needs. Much of this progress has been enabled by scaling of data: pre-training data to learn vast knowledge, fine-tuning data to learn natural language reasoning, and RL environments to refine that reasoning. In this talk, I will describe the current LLM reasoning paradigm, its boundaries, and the future of LLM reasoning beyond scaling. First, I will describe the state of reasoning models and where I think scaling can lead to some additional (though perhaps limited) successes. I will then shift to discussing more fundamental issues with models that scale will not resolve in the next few years. I will touch on four current limitations: outdated knowledge, generator-validator gaps, limited creativity, and poor compositional generalization. In all cases, fundamental limitations of LLMs or of supervised learning in general make these problems challenging, inviting future study and novel solutions beyond scaling.

Bio: Greg Durrett is an associate professor in the Department of Computer Science and the Center for Data Science at New York University. His research is broadly in the areas of natural language processing and machine learning. Currently, his group's focus is on reasoning about knowledge in text, verifying correctness of generation methods, and studying how to make progress on problems that defy LLM scaling. He is a 2023 Sloan Research Fellow and a recipient of a 2022 NSF CAREER award. He has served in numerous roles for ACL conferences, recently as a member of the NAACL Board since 2024 and as Senior Area Chair for ACL 2025 and EMNLP 2025. He received his BS in Computer Science and Mathematics from MIT and his PhD in Computer Science from UC Berkeley, where he was advised by Dan Klein.
Looking to learn about a new topic or skill? Look no further! Gemini's Guided Learning feature acts as your own personal tutor, teaching you about a particular subject through an engaging back and forth conversation. This AI tool helps users develop their knowledge and skills on a wide variety of topics, acting as a patient mentor, breaking down complex topics step-by-step. This session will take place on 2/24 at 11 AM. Please register using the link below!
https://stonybrookuniversity.co1.qualtrics.com/jfe/form/SV_a9PVlBw0E1Bal1A?

The Office for Research and Innovation at Stony Brook University invites you to attend the inaugural Wolf Den, an evening designed to bring together members of the regional innovation and entrepreneurial ecosystem.

Meet investors, researchers, startup founders, and business leaders to exchange ideas, foster collaboration, and strengthen connections that drive technology development and economic growth across Long Island.

Agenda

4:30 - 5:00 PM | Grab some cheer & mingle
5:00 - 5:40 PM | Welcome remarks and AI Panel
5:40 - 6:00PM | Featured lightning pitches
6:00 - 7:00 PM | Food, drinks and great conversations!

Attendees will have the opportunity to learn more about Stony Brook's entrepreneurship ecosystem, hear company pitches from emerging startups, and engage in meaningful networking with innovators, investors and community partners.

Refreshments will be served. Registration is required.

In partnership with Accelerate Long Island.

https://www.stonybrook.edu/commcms/innovation/_events/wolfden.php



Abstract: In high-dimensional data spaces, vast empty regions often exist where no known data points are present. These empty spaces are not merely gaps but hold untapped potential for discovering novel configurations, optimizing parameters, and improving decision-making processes. However, traditional exploration techniques struggle to identify and leverage these regions due to the curse of dimensionality. To address this, we introduce the Empty Space Search Algorithm (ESA), a scalable, physics-inspired method that systematically identifies and explores these uncharted voids. ESA operates by modeling the data space as a dynamic system, using a repulsion-attraction mechanism to locate optimal empty space configurations (ESCs) without requiring exhaustive search. Building upon ESA, we present GapMiner, a visual analytics system that integrates human-in-the-loop AI to iteratively refine and validate ESCs. GapMiner combines parallel coordinate visualization, interactive optimization, and deep learning-based predictive modeling to enhance the efficiency of empty space exploration. This methodology has broad applications, including accelerating convergence in evolutionary algorithms through a more diverse initial population, optimizing adversarial learning strategies, and discovering novel parameter configurations in reinforcement learning. Our approach demonstrates that empty space is not just an absence of data but a frontier for new possibilities in high-dimensional problem-solving.
Bio: Xinyu Zhang received his B.E. in Computer Science from Shandong University, Taishan College, in 2019. He is currently a final-year Ph.D. candidate in the Department of Computer Science at Stony Brook University, advised by Prof. Klaus Mueller. His research focuses on multivariate data analysis, scientific visualization, and reinforcement learning. He has published multiple papers in top-tier journals and conferences, including IEEE TVCG and NeurIPS.
*this seminar will be held in person (food provided on a first come, first serve basis), and online (zoom link below)!
Topic: IACS Student Seminar Speaker: Xinyu Zhang
Time: Feb 26, 2025 12:00 PM Eastern Time (US and Canada)
Join Zoom Meeting
https://stonybrook.zoom.us/j/91848218975?pwd=lfITFa61GaXZ2Wsa1B1OnbLQMmXvOE.1

Meeting ID: 918 4821 8975
Passcode: 027337
Abstract: In today's digital era, language functions not only as a medium of information transmission but also as a mechanism of persuasion, framing, and control. The proliferation of online platforms has amplified this dual role: while enabling unprecedented access to knowledge, it has also exacerbated challenges such as misinformation, rhetorical manipulation, and cultural or linguistic disparities in information access. As a result, pragmatic language understanding and information integrity have emerged as central concerns for both computational linguistics and society at large. This research follows how claims are produced, reframed, and contested online through three interconnected threads. First, it models pragmatic deflection in discourse by investigating whataboutism, a rhetorical device that deflects criticism by redirecting discourse, and introduced novel datasets from Twitter (now X) and YouTube. This work underscores how subtle pragmatic maneuvers can erode discourse integrity without relying on outright falsehoods. Second, it advances retrieval and alignment for information integrity in health and news communication. These systems trace claims and narratives across genres (e.g., social posts and news reports) and languages (Chinese and English), linking social posts with journalistic reporting and aligning Chinese news with English biomedical evidence. By accounting for cultural context, assertions can be linked to reliable evidence and organized for systematic comparison. This work surfaces the risks of missing sources, unverifiable claims, and framing disparities in global health discourse, and demonstrates computational solutions that enhance both the credibility and accessibility of information. Third, the methodological centerpiece is Class Distillation (ClaD), a geometry-aware training paradigm for distilling a small, well-defined target class from a large, heterogeneous background. ClaD couples a distribution-aware contrastive loss (instantiated here in a Mahalanobis form when its assumptions fit the data) with an interpretable decision algorithm tuned for class separation. Evaluated on sarcasm, metaphor, and sexism detection, ClaD delivers strong efficiency and robustness, matching or surpassing larger models while using fewer computational resources, making these pipelines practical by learning reliably from small, sharply defined classes. In sum, this research presents an integrated account of language understanding in the digital age. It exposes how integrity falters through pragmatic deflection, cross-genre drift, and cross-lingual misalignment, and translates these insights to move pragmatic language understanding to systems for evidence retrieval, alignment, and verification; and it sheds light on where and how integrity is threatened, and delivers methods that leverage pragmatic language use.

Speaker: Chenlu Wang

Location: (Old) Computer Science Building, Room 2311
Abstract: Many unresolved legal questions over LLMs and copyright center on memorization: whether specific training data have been encoded in the model's weights during training, and whether those memorized data can be extracted in the model's outputs. While many believe that LLMs do not memorize much of their training data, recent work shows that substantial amounts of copyrighted text can be extracted from open-weight models. However, it remains an open question if similar extraction is feasible for production LLMs, given the safety measures these systems implement. We investigate this question using a two-phase procedure: (1) an initial probe to test for extraction feasibility, which sometimes uses a Best-of-N (BoN) jailbreak, followed by (2) iterative continuation prompts to attempt to extract the book. We evaluate our procedure on four production LLMs -- Claude 3.7 Sonnet, GPT-4.1, Gemini 2.5 Pro, and Grok 3 -- and we measure extraction success with a score computed from a block-based approximation of longest common substring (nv-recall). With different per-LLM experimental configurations, we were able to extract varying amounts of text. For the Phase 1 probe, it was unnecessary to jailbreak Gemini 2.5 Pro and Grok 3 to extract text (e.g, nv-recall of 76.8% and 70.3%, respectively, for Harry Potter and the Sorcerer's Stone), while it was necessary for Claude 3.7 Sonnet and GPT-4.1. In some cases, jailbroken Claude 3.7 Sonnet outputs entire books near-verbatim (e.g., nv-recall=95.8%). GPT-4.1 requires significantly more BoN attempts (e.g., 20X), and eventually refuses to continue (e.g., nv-recall=4.0%). Taken together, our work highlights that, even with model- and system-level safeguards, extraction of (in-copyright) training data remains a risk for production LLMs.

Speaker: Xinyue

Location: CS2311