CSE 600 Seminar Series | Fall 2025


Abstract: Large reasoning models have demonstrated capabilities to solve competition-level math problems, answer deep research questions, and address complex coding needs. Much of this progress has been enabled by scaling of data: pre-training data to learn vast knowledge, fine-tuning data to learn natural language reasoning, and RL environments to refine that reasoning. In this talk, I will describe the current LLM reasoning paradigm, its boundaries, and the future of LLM reasoning beyond scaling. First, I will describe the state of reasoning models and where I think scaling can lead to some additional (though perhaps limited) successes. I will then shift to discussing more fundamental issues with models that scale will not resolve in the next few years. I will touch on four current limitations: outdated knowledge, generator-validator gaps, limited creativity, and poor compositional generalization. In all cases, fundamental limitations of LLMs or of supervised learning in general make these problems challenging, inviting future study and novel solutions beyond scaling.

Bio: Greg Durrett is an associate professor in the Department of Computer Science and the Center for Data Science at New York University. His research is broadly in the areas of natural language processing and machine learning. Currently, his group's focus is on reasoning about knowledge in text, verifying correctness of generation methods, and studying how to make progress on problems that defy LLM scaling. He is a 2023 Sloan Research Fellow and a recipient of a 2022 NSF CAREER award. He has served in numerous roles for ACL conferences, recently as a member of the NAACL Board since 2024 and as Senior Area Chair for ACL 2025 and EMNLP 2025. He received his BS in Computer Science and Mathematics from MIT and his PhD in Computer Science from UC Berkeley, where he was advised by Dan Klein.
The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors. Each seminar will consist of multiple short talks (around 15 minutes) by multiple students. Students can register for 1 credit for CSE656. Registered students must attend and present a minimum of 2 talks. Everyone else is welcome to attend. Fill in https://forms.gle/q6UG9ygauLp2a8Po8 to subscribe to our mailing list for further announcement.
CSE 600 Seminar Series | Fall 2025


Abstract: Virtual worlds are prevalent in applications ranging from entertainment, healthcare, retail, to workforce training. With the demand for virtual content growing exponentially, the market for such content is valued at over $200 Billion, which is accelerating the need for advanced computational solutions. In this talk, I will focus on a key challenge in virtual content creation: simulating autonomous agents.
I begin by overviewing this problem domain, through the lens of a physics-based dynamics simulation, which enables the simulation of thousands of agents at interactive rates with GPU programming, achieving a level of performance previously unattainable.
Next, I'll present our recent results in Deep Reinforcement Learning for multi-agent navigation, which enable refined, reward-based strategies to control agent movement. We demonstrate how these techniques can simulate realistic crowds, with broad applications in pedestrians, robots, and swarms. Lastly, I conclude my talk by discussing our lab's work-at-large and the wide range of research opportunities in this emerging area.

Speaker: Tomer Weiss is a professor with New Jersey Institute of Technology since 2020. He received the best student, presentation, and best paper awards in various ACM SIGGRAPH conferences for his work on simulating multi-agent crowds. He was also a finalist in both ACM SIGGRAPH Thesis Fast Forward, and the ACM SIGGRAPH Asia Doctoral Symposium in 2018. He received his PhD in computer science from UCLA in 2018. His research interests include multi-agent dynamics, scene understanding, and interactive visual computing.
AI for Conservation: AI and Humans Combating Extinction Together by Daniel I. Rubenstein of Princeton University

ABSTRACT: The state of our planet is not good. We have lost more than 60% of the world's wildlife. Stopping the decline remains a challenge, especially since acquiring appropriate knowledge is expensive, time consuming and risky. Visual observations following the fates of a few individuals was the currency of the realm. But GPS technology and now machine learning provide a non-invasive scalable alternative. Photographs, taken by field scientists, tourists, automated cameras and incidental photographers, are the most abundant source of data on wildlife today. Wildbook, a project of tech for conservation coordinated by a non-profit Wild Me, is an autonomous computational system that starts from massive collections of images and, by detecting various species of animals and identifying individuals, combined with sophisticated data management, turns them into high-resolution information databases, enabling scientific inquiry, conservation and citizen science.

BIO: Dan Rubenstein is the Class of 1877 Professor of Zoology. He is currently Director of Princeton's Environmental Studies Program and is former Chair of Princeton University's Department of Ecology and Evolutionary Biology and Director of Princeton's Program in African Studies. He is a behavioral ecologist who studies how environmental variation and individual differences shape social behavior, social structure, sex
roles and the dynamics of populations. He has special interests in all species of wild horses, zebras and asses, and has done field work on them throughout the world identifying rules governing decision-making, the emergence of complex behavioral patterns and how these understandings influence their management
and conservation. In Kenya he also works with pastoral communities to develop and assess impacts of various grazing strategies on rangeland quality, wildlife use and livelihoods. He has also developed a scout program for gathering data on Grevy's zebras and created curricular modules for local schools to raise awareness about the plight of this endangered species. He engages people as 'Citizen Scientists' and has recently extended his work to measuring the effects of environmental change, including issues pertaining to the global commons
and changes wrought by management and by global warming, on behavior.

The Vedanta Forum is devoted to one of humanity's oldest and most profound pursuits -- thinking. Thinking about who we truly are: the one that remains constant through childhood and old age, through waking, dream, and deep sleep. Thinking about the source and cause of creation, and its relationship to what inheres in us.

Across history, such thinking, both meditative and scientific, has been aimed at these questions. The ancient Upanishads proclaimed, Tat Tvam Asi -- Thou Art That -- revealing the non-dual identity of the individual and the ultimate reality. Centuries later, modern scientists such as Schrödinger and Bohr echoed similar intuitions about the unity of existence.

Over time, many philosophical approaches, traditions, and interpretive schools have arisen from such inquiry, each offering unique perspectives. The Forum will:

  • Focus on universal approaches and traditions and examine their teachings,

  • Foster comparative studies, and

  • Explore the practical benefits to society from such thinking,

through scholarly studies, dialogue, and debate also promoting accessibility to all qualified seekers. Additionally, the Forum will explore how these reflections can enrich life, education, and even technology.

Location: NCS 120 (New Computer Science), Engineering Dr, Stony Brook, NY 11794.

The program is available at: https://www.vedantaforum.org/events/program

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Abstract: Two-dimensional (2D) materials such as graphene, hBN, and TMDs offer atomically sharp interfaces and unprecedented tunability when vertically assembled into van der Waals heterostructures. These stacks have enabled discoveries ranging from moiré superconductivity and correlated insulators to quantum emitters and next-generation nanoelectronic devices. Yet constructing high-quality heterostructures remains largely artisanal: researchers manually identify exfoliated flakes, align a polymer stamp by eye, and finely adjust temperature and contact geometry through tacit skill. This manual workflow is difficult to reproduce, scales poorly, and prevents systematic exploration of the enormous combinatorial space of materials, twist angles, and interfacial conditions. AutoLab is an autonomous platform that translates this tacit human expertise into programmable, feedback-driven control. Instead of pressing flakes with predefined trajectories, AutoLab uses machine vision to detect polymer-wafer contact, dynamically regulates contact evolution through closed-loop actuation and temperature control, and captures high-quality flakes with the cleanliness and precision of expert manual fabrication. The system integrates perception, decision making, and motion planning into a single robotic framework, enabling reproducible stacking, wafer-level coverage, and accelerated discovery. Beyond 2D materials, AutoLab illustrates a broader paradigm for AI-native scientific automation: codifying human experimental reasoning into algorithms that interrogate data in real time, adaptively adjust instrumentation, and generate scalable, high-fidelity datasets. Such platforms could generalize to diverse research domains--quantum device fabrication, optical alignment, surface science, autonomous microscopy, and other workflows where expert intuition currently limits throughput and reproducibility. By bridging artisanal manipulation and robotic autonomy, AutoLab points toward a future where scientific discovery is accelerated by machines that not only execute instructions, but learn, respond, and collaborate with human scientists.

Biography: Dr. Yutao Li is a research associate from Department of Condensed Matter Physics and Material Science, Brookhaven National Laboratory. He has 8 years of experience in 2D material sample fabrication, and investigation in their electronic transport, optical and mechanical properties.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1604383624?pwd=ffQ5cUPNxTI7nzClKQO6cnsNbhF9Vf.1

Meeting ID: 160 438 3624
Passcode: 558449

The Provost's Lecture Series features talks by SUNY Distinguished Academy faculty members at Stony Brook University, showcasing the outstanding research and scholarship that is taking place at our institution.

Joe Mitchell

SUNY Distinguished Professor, Applied Mathematics and Statistics
Chair, Department of Applied Mathematics and Statistics, College of Engineering and Applied Sciences

A Case for Algorithms: A Computational Geometer's Perspective

Algorithms are all around us in every smart device and technology that has consumed our daily lives. As a computational geometer, I study algorithms to solve problems that involve a geometric perspective on data. I have observed that practically every technology and field of study has a need for effective algorithms involving geometric data. I reflect on some favorite algorithmic problems that are easy to visualize, but challenging to solve, and argue that the formal study of algorithms remains essential in the age of AI.

Reception to follow immediately after the talks.

Register here.
Hidden Biases. Ethical Issues in NLP, and What to Do about Them presented by Dirk Hovy of Bocconi University

ABSTRACT: Through language, we fundamentally express who we are as humans. This property makes text a fantastic resource for research into the complexity of the human mind, from social sciences to humanities. However, it is exactly that property that also creates some ethical problems. Texts reflect the authors' biases, which get magnified by statistical models. This has unintended consequences for our analysis: If our data is not reflective of the population as a whole, if we do not pay attention to the biases contained, we can easily draw the wrong conclusions, and create disadvantages for our users.

In this talk, I will discuss several types of biases that affect NLP models, their sources, and potential counter measures: (1) Bias stemming from data, i.e., selection bias (if our texts do not adequately reflect the population we want to study), label bias (if the labels we use are skewed) and semantic bias (the latent stereotypes encoded in embeddings); (2) Biases deriving from the models themselves, i.e., their tendency to amplify any imbalances that are present in the data; (3) Design bias, i.e., the biases arising from our (the researchers) decisions which topics to analyze, which data sets to use, and what to do with them. For each bias, I will provide examples and discuss the possible ramifications for a wide range of applications, and various ways to address and counteract these biases, ranging from simple labeling considerations to new types of models.

BIO: Dirk Hovey is an associate professor of Computer Science in the department of marketing at Bocconi University. He received his PhD from the University of Southern California in Los Angeles, where he worked as a research assistant at the Information Sciences Institute. 

He works in Natural Language Processing (NLP), a subfield of artificial intelligence. His research focuses on computational social science. His interests include integrating sociolinguistic knowledge into NLP models, using large-scale statistics to model the interaction between people's socio-demographic profile and their language use, and ethics for data science and algorithmic fairness.

Abstract: Pretraining vision encoders with self-supervision (SSL) leads to stronger representations that excel across diverse downstream tasks. One of the key factors enabling self-supervision is extracting multiple views of the same scene to formulate either: 1) View-invariant pretraining (DINO, SimCLR, iBOT), where the objective is predicting the same representation for different views of the scene; or 2) Cross-view pretraining (cross-view Masked Autoencoders), where the objective is predicting missing parts of one view using other views. For extracting multiple views, view-invariant methods rely on a combination of handcrafted augmentations (random cropping, color jittering, gaussian blur, etc.) of the same image, whereas cross-view pretraining methods rely on image cropping or video frames. In this work, we present methods to effectively incorporate synthetic views from diffusion models into SSL training.
For view-invariant pretraining, we introduce Gen-SIS, a method that leverages the ability of diffusion models to generate interpolated images through interpolation in conditioning space. We introduce a disentanglement pretext task: disentangling two source images from an interpolated synthetic image. This disentanglement task, in addition to vanilla single-source generative augmentation for view extraction, improves visual pretraining of various view-invariant methods (DINO, SimCLR, iBOT).
For cross-view pretraining, we introduce CDG-MAE, a novel cross-view masked autoencoder (MAE) based method that uses diverse synthetic views generated from static images via an image-conditioned diffusion model to learn dense correspondences. We present a quantitative method to evaluate the local and global consistency of the generated views to choose the right diffusion model for cross-view pretraining. These generated views exhibit substantial changes in pose and perspective, providing a rich training signal that overcomes the limitations of video (expensive) and crop-based (less variation) methods. CDG-MAE substantially narrows the gap to video-based MAE methods on video label propagation tasks while maintaining the data advantages of image-only MAEs.

Speaker: Varun Belagali

Location: NCS 120
Zoom: https://stonybrook.zoom.us/j/93647452432?pwd=hZaX7LXCAD8KPHWYE1Afw2sDI3owpv.1
The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 10 minutes) by multiple people. Students can register for 1 credit for CSE 656. Registered students must attend and present a minimum of 2 or 3 talks. Everyone else is welcome to attend. Fill in https://forms.gle/pCVXovgfMfQwGqG38 to subscribe to our mailing list for further announcement.