Event Website: bnl.gov/nysds
Dates: September 28-29, 2026
Location: SUNY Global Center, New York, NY
Co-hosts: Brookhaven National Laboratory, the Institute for Advanced Computational Science (IACS), and the AI Innovation Institute at Stony Brook University.

Join us for this premier annual conference that brings together researchers and thought leaders from academia, national labs, and industry to exchange ideas and foster cross-disciplinary collaboration centered on data-driven science and technology.

The theme of NYSDS 2026 is Transformational AI from Science to Society. This year's conference will focus on the ways in which artificial intelligence (AI), machine learning (ML) and robotics are impacting everything from our everyday lives to scientific endeavors. NYSDS2026 will feature the following main tracks:

  • Robotics and Embodied AI: advances in perception, control, learning and interaction for autonomous physical agents in real-world and scientific environments.

  • AI for Science: innovative applications ranging from the physical sciences to biology and medicine, and discussion of important ways in which AI is changing the nature of science.

  • AI for Energy: how energy is discovered and produced, how hazards are mitigated to ensure reliability of the power grid, and our search for new sources of critical minerals and materials.

  • AI for Risk Assessment: inventive uses of large amounts of data and the AI tools to assess and mitigate risks in national security, climate, and financial markets.

  • AI for Education: the ways in which education is being reimagined in the age of AI and the roadblocks to training the next generation of students.

Each track will include invited presentations, contributed talks, posters, and panel discussions.

Are you tired of drowning in a sea of resumes and losing top talent in the hiring whirlwind? Transform your hiring process through a different lens and learn about AI in the Workplace and the Applicant Tracking System (ATS). Whether you're a recent graduate seeking your first job or an undergraduate student looking to delve into more career-oriented opportunities, this workshop by SBU Career Center is designed to equip you with the knowledge and strategies needed to succeed.

Register here: https://stonybrook.joinhandshake.com/stu/events/1568133?

Speaker Petar Djuric Refreshments will be provided Deep Gaussian processes: Theory and applications Petar M. Djurić Department of Electrical and Computer Engineering Stony Brook University Abstract: Gaussian processes are an infinite-dimensional generalization of multivariate normal distributions. They provide a principled approach to learning with kernel machines and they have found wide applications in many fields. More recently, with the advance of deep learning, the concept of deep Gaussian processes has emerged. Deep Gaussian processes can be viewed as multilayer hierarchical organizations of Gaussian processes that are equivalent to infinitely wide multiple layer neural networks. Deep Gaussian processes have improved capacity for prediction and classification over standard Gaussian processes, while models based on them continue to allow for full Bayesian treatment and for applications when the amount of available data is limited. The theory of recent progress in deep Gaussian processes will be presented and some applications will be provided. Biosketch: Petar M. Djurić received the B.S. and M.S. degrees in electrical engineering from the University of Belgrade, Belgrade, Yugoslavia, respectively, and the Ph.D. degree in electrical engineering from the University of Rhode Island, Kingston, RI, USA. He is a SUNY Distinguished Professor and currently, he is a Chair of the Department of Electrical and Computer Engineering, Stony Brook University, Stony Brook, NY, USA. Djurić was a recipient of the IEEE Signal Processing Magazine Best Paper Award in 2007 and the EURASIP Technical Achievement Award in 2012. From 2008 to 2009, he was a Distinguished Lecturer of the IEEE Signal Processing Society. He was the Editor-in-Chief of the IEEE Transactions on Signal and Information Processing over Networks (2015-2018). Djurić is a Fellow of IEEE and EURASIP

Abstract: Pretraining vision encoders with self-supervision (SSL) leads to stronger representations that excel across diverse downstream tasks. One of the key factors enabling self-supervision is extracting multiple views of the same scene to formulate either: 1) View-invariant pretraining (DINO, SimCLR, iBOT), where the objective is predicting the same representation for different views of the scene; or 2) Cross-view pretraining (cross-view Masked Autoencoders), where the objective is predicting missing parts of one view using other views. For extracting multiple views, view-invariant methods rely on a combination of handcrafted augmentations (random cropping, color jittering, gaussian blur, etc.) of the same image, whereas cross-view pretraining methods rely on image cropping or video frames. In this work, we present methods to effectively incorporate synthetic views from diffusion models into SSL training.
For view-invariant pretraining, we introduce Gen-SIS, a method that leverages the ability of diffusion models to generate interpolated images through interpolation in conditioning space. We introduce a disentanglement pretext task: disentangling two source images from an interpolated synthetic image. This disentanglement task, in addition to vanilla single-source generative augmentation for view extraction, improves visual pretraining of various view-invariant methods (DINO, SimCLR, iBOT).
For cross-view pretraining, we introduce CDG-MAE, a novel cross-view masked autoencoder (MAE) based method that uses diverse synthetic views generated from static images via an image-conditioned diffusion model to learn dense correspondences. We present a quantitative method to evaluate the local and global consistency of the generated views to choose the right diffusion model for cross-view pretraining. These generated views exhibit substantial changes in pose and perspective, providing a rich training signal that overcomes the limitations of video (expensive) and crop-based (less variation) methods. CDG-MAE substantially narrows the gap to video-based MAE methods on video label propagation tasks while maintaining the data advantages of image-only MAEs.

Speaker: Varun Belagali

Location: NCS 120
Zoom: https://stonybrook.zoom.us/j/93647452432?pwd=hZaX7LXCAD8KPHWYE1Afw2sDI3owpv.1
Towards Saving Lives with Natural Language Processing Andrew Schwartz Dept. of Computer Science Stony Brook Analyzing language use patterns is proving to be a valuable and unique approach to understanding the psychological, social, and health factors of people. On the individual level, Facebook and Twitter have been found predictive of mental health, personality, demographics, and occupational class (among others). At the community or county-level, Twitter has been found predictive of flu and allergy outbreaks, life satisfaction, atherosclerotic heart disease mortality, health behavioral risk factors, excessive drinking, and HIV prevalence. While these techniques have shown robust links over a plethora of important aspects of human life, it is not clear whether any lives have been saved, at least directly, by such work. At their core, some barriers to improving health care and saving lives are likely not NLP or even AI problems, but others are perhaps technical in nature and suggest changing the way we model data. This seminar will have two parts: a presentation and a discussion. I will start by going over recent and on-going work toward predicting mental health outcomes --- depression, addiction relapse, future psychological distress --- from human language use patterns. Then, I will present an imperfect vision of a future where NLP helps to save lives and open the floor for discussion of technical barriers and whether such a vision is practical. Biography: Andrew Schwartz received his PhD in Computer Science from the University of Central Florida in 2011 with research on acquiring lexical semantic knowledge from the Web. He then joined the University of Pennsylvania where he was a Postdoctoral Research Fellow and later Visiting Assistant Professor in Computer & Information Science. He is Lead Research Scientist for the World Well-Being Project, a multidisciplinary group of Computer Scientists and Psychologists studying physical and psychological well-being based on language in social media.
Abstract: Modern technologies enable enhanced integrity and privacy guarantees not just for data, but also for computation. This is perhaps most emphatically demonstrated by the steady rise of zero-knowledge proofs, which are short certificates that attest to the correctness of computations (e.g., an age verification check) without revealing any secret inputs (e.g., the birth date on a digital ID). This subtly powerful technology enables anonymous credentials, privacy-preserving machine learning, anonymous blockchains, and much more--making the question of efficient zero-knowledge proofs fundamental to modern secure systems. Echoing Moore's law for computing, zero-knowledge proofs have improved on this front by ten orders of magnitude in the last two decades. In this talk, I will discuss our work on overcoming a key bottleneck that has emerged in this development: memory efficiency.

Speaker: Abhiram Kothapalli is a postdoctoral scholar at the University of California, Berkeley, hosted by Sanjam Garg. He is a recent graduate of Carnegie Mellon University, where he earned his Ph.D. in Computer Science, advised by Bryan Parno. Previously, he was at the University of Illinois at Urbana-Champaign, where he earned his B.S. in Computer Science and B.S. in Mathematics. Kothapalli's research develops cryptographic techniques aimed at scaling expressive privacy and integrity guarantees across the internet.

Location: NCS 120
Abstract: This dissertation addresses the methodological disconnect between Natural Language Processing (NLP) and human-centric analysis by shifting the unit of analysis from document to human behavior in two broad respects: (i) time-ordering: modeling documents as sequential person-indexed behavioral observations, and (ii) person-level semantics: evaluation and explainability of models by their latent structure of psychological constructs rather than just its predictive accuracy against narrow proxy measures. First, we consider the most basic implication of language as a person's behaviors when measuring their psychological constructs: relationship between language sample size and model's predictive performance. We empirically show that the state-of-the-art transformers are often over-parameterized for typical NLP dataset sizes and can be reduced in dimensionality without performance loss. Establishing the author as the unit of analysis naturally allows us to treat their behavior as a time-ordered sequence. Second, we introduce a longitudinal evaluation framework that establishes ecologically valid evaluation settings, namely, cross-sectional and prospective generalization, and separates error measurement of the model into within-person dynamics and between-person differences. We demonstrate that traditional NLP evaluations based on random document splits can yield reversed conclusions under ecologically valid generalization settings. To address this, we develop models that capture the trajectory of mental states (e.g., mood shifts) rather than static traits. Third, moving into person-level semantics, we evaluate the latent structure of large language models using a novel machine behavior analytic framework. We find that while GPT-4 achieves high predictive correlation with self-reports, its latent symptoms structure diverges from clinical understanding. Finally, we propose a method for modeling multidimensional behaviors, embedding concurrent behavioral signals alongside language to predict future states. Taken together, this work suggests that operationalizing language as behavior advances NLP methods into a rigorous instrument for valid psychological inquiry.

Speaker: Adithya Ganesan

Location: Join Zoom Meeting (ID: 99021939129, Passcode: 569493)
Are you concerned about AI issues with your asynchronous online courses? Is your fully online course vulnerable to AI plagiarism? Do you want to engage your online students using AI? Discover the future of education with our AI-powered solutions designed specifically for online asynchronous courses. This innovative approach uses artificial intelligence to transform the way courses are delivered, making learning more personalized, engaging, and effective.

https://stonybrook.zoom.us/meeting/register/tJMvd-irqTotGtQONZqerPf_TnhXcx8t2sA1