The AI Community at Stony Brook University is proud to announce Datathon 2026.

Dive into data analysis and AI/ML, and get ready to build something big. In this year's underwater-themed event, enjoy a weekend of data analysis, hacking, networking, fun activities, and minigames.

Whether you're a seasoned developer, data scientist, designer, or completely new to hacking, this event is your chance to collaborate, learn data science, and create something impactful with data and AI/ML.

What is Datathon?

AI Community's Datathon is the premier data science competition at Stony Brook University, bringing together students of all skill levels for a weekend of data exploration, analysis, and innovation. Just like a typical hackathon, you will be using your skills to build your dream project.

Unlike a regular hackathon, Datathon is focused on data science. You will be given a set of data to work with, analyze, and apply to your project. You can also find your own data to use. Your project will be presented to a panel of judges consisting of professors and industry professionals!

Who Can Participate

  • Students of all skill levels and majors are welcome.
  • Come with a team or find one at the event or on Discord.
  • This event is open to SBU and non-SBU students.

* Non-SBU Undergraduate Students are ineligible to receive prizes
* You must be 18+ or older (Excludes minors who are active SBU students)

Location: SAC Ballroom B

Register here.

Abstract: Autonomous systems, whether on Earth or in space, rely on 3D perception to understand and interact with the world around them. Yet traditional techniques for 3D understanding often depend on human designed features, fixed sensors, and conventional imaging modalities. This constrained approach can limit every stage of perception, from sensing to interpretation to decision making.
In this talk, we'll explore an alternative paradigm for imaging: physically based neural representations for 3D scenes and 3D sensing systems. We will discuss how recent advances in large scale learned representations can be used to jointly optimize both 3D scene models and the design of sensing systems for 3D capture, with the goal of enabling task specific perception systems.
Unlike modern AI models trained on internet scale datasets, these specialized 3D representations typically operate in data sparse regimes and therefore require a different kind of prior. We'll examine how grounding these learned representations in the physics of light transport can improve our understanding of scene structure, and inform imaging system design even with limited data. By connecting physical insights with learned representations, we'll highlight new possibilities for robust, efficient, and adaptive perception in challenging environments.

Speaker: Nikhil Behari is a graduate student in the Camera Culture group at the MIT Media Lab, advised by Professor Ramesh Raskar. His research interests include computational imaging, 3D scene understanding, and multi-agent decision-making under uncertainty, with a focus on automating imaging system design for 3D perception in human and planetary health. His research is supported by the NASA Space Technology Graduate Research Fellowship. He received his bachelor's in Computer Science and Statistics from Harvard University in 2022.
Hieu Le presents Incorporating Physical Illumination Constraints into Deep Learning Shadow Detection and Removal (PhD Proposal)

Shadows provide useful cues to analyze the scene but also hamper many computer vision algorithms such as image segmentation, object detection or tracking. For those reasons, shadow detection and shadow removal have been well studied topics in computer vision. Early approaches for shadow detection and removal focus on physical illumination models of shadows. These methods can express, identify, and remove shadows in a physically plausible manner. However, these models are often hard to optimize and slow in inference due to reliance on hand-designed image features. On the other hand, recent deep-learning approaches have achieved breakthroughs in performances for both shadow detection and removal. They learn to extract useful features automatically through training while being extremely efficient in computation. However, these models are data-dependent, opaque and ignore the physical aspects of shadows.

We propose to incorporate physical illumination constraints into deep-learning frameworks. Thus the mapping learned by the deep-network closely follows the physics of shadows, enabling the network to systematically and realistically modify shadows in images. For shadow detection, we present a novel GAN framework in which the generator can generate realistic images with attenuated shadows that can be used to train a shadow detector. For shadow removal, we propose a method that uses deep-networks to estimate the unknown parameters for a shadow image formation model that removes shadows. The system outputs shadow-free images in high-quality with no image artifacts and achieves state-of-the-art shadow removal performance. Lastly, we propose a system trained without the need for any shadow-free images in which physical constraints play pivotal roles that enable training the networks.

For Zoom information, please email events@cs.stonybrook.edu.

Visual Analytics and Machine Learning for Biomedical Imaging Diagnosis

 

Arie Kaufman

 

We present an integrated approach using visual analytics and machine learning (ML) to diagnose abnormalities in 3D radiological imaging and biological microscopes. The primary example will involve 3D virtual pancreatography (VP), a novel visualization-ML procedure and application for non-invasive diagnosis and classification of pancreatic lesions, the precursors of pancreatic cancer. Currently, non-invasive screening of patients is performed through visual inspection of 2D axis-aligned CT images, though the relevant features are often not clearly visible nor automatically detected. VP is an end-to-end visual diagnosis system that includes an ML-based automatic segmentation of the pancreatic gland and the lesions, a semi-automatic approach to extract the primary pancreatic duct, an ML-based automatic classification of lesions into four prominent types, and specialized 3D and 2D exploratory visualizations of the pancreas, lesions and surrounding anatomy. We combine volume rendering with pancreas- and lesion-centric visualizations and measurements for effective diagnosis. We designed VP through close collaboration and feedback from expert radiologists, and evaluated it on multiple real-world CT datasets with various pancreatic lesions and case studies examined by the expert radiologists. Other applications include virtual colonoscopy, COVID-19, pathology, brain neurites, etc.


Biography: Arie Kaufman is Distinguished Professor and formerChair of the Department of Computer Science at Stony Brook University, where he is also Director of the Center for Visual Computing (CVC), and Chief Scientist at the Center of Excellence in Wireless and Information Technology (CEWIT). 

He received his PhD in Computer Science at Ben-Gurion University of the Negev in 1977.   He is known for his work in visualization, graphics, virtual reality, user interfaces, multimedia, and their applications, especially in bio-medicine. He is especially well known for his work on the 3-dimensional virtual colonoscopy, a revolutionary low-risk technique for colon cancer screening, and for pioneering the use of Graphics Processing Units (GPUs) and GPU-clusters. In 2012, he presided over the development and opening of the Reality Deck, the largest virtual reality display in the world, at Stony Brook University.

Kaufman was the founding Editor in Chief of IEEE Transactions on Visualization and Computer Graphics (TVCG), co-founded the IEEE Visualization Conference and Volume Graphics series, and is currently the director of IEEE Computer Society Technical Committee on Visualization and Graphics. He is an IEEE Fellow, ACM Fellow, winner of many awards, including the IEEE Visualization Career Award, and member of the European Academy of Sciences.



Steven Skiena is inviting you to a scheduled Zoom meeting.

Topic: AI Seminar: Arie Kaufman
Time: Apr 21, 2021 10:00 AM Eastern Time (US and Canada)

Join Zoom Meeting
https://stonybrook.zoom.us/j/96017498640?pwd=SE0rdHB6ZVlCM2ZpY2RnRUxyVnR3Zz09

Title: Cyberinfrastructure for forward prediction and inversion estimation with uncertainty quantification

Seminar Speaker: Dr. Mengyang Gu, Assistant Professor, Department of Statistics and Applied Probability, University of California, Santa Barbara

Abstract: In this talk, we introduce four useful tools for forward prediction and inversion estimation. The first tool is the parallel partial Gaussian process surrogate model for emulating expensive computer simulations with massive coordinates. The tool is implemented in the RobustGaSP package available in R, MATLAB, and Python, for predicting both scalar- and vector-valued outputs with uncertainty assessment. The second tool is implemented in the RobustCalibration package, which handles Bayesian data inversion or model calibration by one or multiple types of experimental observations. A unique feature of the package is the inclusion of fast surrogate models of both scalar- and vector-valued computer simulations that bypass the expensive simulation in one line of code. The third tool is implemented in the AIUQ package, available in both R and MATLAB. In this approach, we show that differential dynamic microscopy, a scattering-based analysis tool that extracts dynamical information from microscopy videos, is equivalent to fitting the temporal auto-covariance in Fourier space, based on a latent factor model we construct. We develop a more efficient estimator and reduce the computational cost to pseudolinear order with respect to the number of observations without approximation, by utilizing the generalized Schur algorithm for the Toeplitz covariance. In the last tool, we developed a new method called the inverse Kalman filter, which enables fast matrix-vector multiplication between a covariance matrix from a dynamic linear model and any real-valued vector with a linear computational cost. These new approaches outline a wide range of applications that include emulating expensive simulation at molecular-, meso- and macro-scales, active learning with error control, nonparametric estimation of particle interaction functions, and data inversion from microscopy and velocity fields.

Join Zoom Meeting: https://bnl.zoomgov.com/j/1606285496?pwd=2yJYSG6lx8gMPiibzgAIBQtKHIjuHV.1
Meeting ID: 160 628 5496
Passcode: 472506
Communication-Efficient Heterogeneity-Aware Machine Learning System and Architecture by Xuehai Qian

ABSTRACT: The key success of deep learning is the increasing size of models that can achieve high accuracy. At the same time, it is difficult to train the complex models with large data sets. Therefore, it is crucial to accelerate training with distributed systems and architectures, where communication and heterogeneity are two key challenges. In this talk, I will present two heterogeneity-aware decentralized training protocols without communication bottleneck. Specifically, Hop supports arbitrary iteration gap between workers by novel queue-based synchronization which can tolerate heterogeneity with system techniques. Prague uses randomized communication to tolerate heterogeneity with a new training algorithm based on partial reduce -- an efficient communication primitive. If time permits, I will present the systematic tensor partitioning for training on heterogeneous accelerator arrays (e.g., GPU/TPU). We believe that our principled approaches are crucial for achieving high-performance and efficient distributed training.

BIO: Xuehai Qian is an assistant professor at University of Southern California. His research interests include domain-specific systems and architectures, performance tuning and resource management of cloud systems and parallel computer architectures. He received his PhD from the University of Illinois Urbana Champaign and was a postdoc at UC Berkeley. He is the recipient of W.J Poppelbaum Memorial Award at UIUC, NSF CRII and CAREER Award, and the inaugural ACSIC (American Chinese Scholar In Computing) Rising Star Award.

Learn which AI models are available at SBU and how to access them.

As AI becomes part of our everyday work, it's important to know which AI models are available at Stony Brook and how to access them. In the first half of this session, we'll show you how to access Stony Brook's Gemini and Microsoft Copilot models and how to make sure you're using the university-approved environment. In the second half, we'll review Stony Brook's Data Classification policies so you know what institutional data can be used with these AI tools and what information should remain protected, including sensitive data. This session will provide the latest campus guidance so you can use AI confidently and responsibly.

Register here.

Hyperscale Verification in Microsoft Azure talk by Nikolaj Bjorner

Abstract: Cloud providers are increasingly embracing network verification for managing complex datacenter network infrastructure. Microsoft's Azure cloud infrastructure integrates the SecGuru tool, which leverages the Z3 Satisfiability Modulo Theories solver, for checking network access
control lists. It also integrates a verifier that uses both custom verification algorithms and Z3 that checks correctness of forwarding tables in Azure data-centers. These tools assure that the network is configured to preserve desired intent over hundreds of thousands of network devices. We describe our experiences building and running SecGuru for network verification in Azure.

Finally we mention recent advances in Z3, including a distributed version of Z3 that scales with Azure's elastic cloud. It integrates recent advances in lookahead and distributed SAT solving for Z3's
engines for SMT. A different recent advance includes integration of DNNs to learn variable branching strategies for high-performance SAT solvers, including MiniSAT, Glucose and Z3's SAT solver.

Bio: Nikolaj Bjorner is a Principal Researcher at Microsoft Research, Redmond, working in the area of Automated Theorem Proving and Software Engineering. His current main line of work is around the state-of-the art theorem prover Z3, which is used as a foundation of several software engineering tools. Z3 received the 2015 ACM SIGPLAN Software System award and most influential tool paper in the first 20 years of TACAS in 2014, and test of time award at ETAPS 2018. Together with Leonardo de Moura received the CADE 2019 Herbrand award for contributions to SMT and applications. Previously, he developed the DFSR, Distributed File System - Replication, and Remote Differential
Compression protocols, RDC, part of Windows Server since 2005 and before that worked on distributed file sharing systems at a startup, and program synthesis and transformation systems at the Kestrel Institute. He received his Master's and PhD degrees in computer science from Stanford University.

Abstract: Recent progress in Large Language Models (LLMs) has transformed text and code generation, yet models still falter on scientific reasoning where correctness, constraints, and physical consequences are critical. This talk explores how formal LLM reasoning can advance symbolic scientific modeling. First, our PDE-Controller formalizes informal PDEs (Partial Differential Equations), synthesizes solver-ready code, and plans subgoals to tackle nonconvex control via interactions with external solvers. Second, our Lean Finder accelerates scientific formalization via a semantics-aware search engine for Lean/Mathlib that retrieves relevant theorems, outperforming GPT models and gaining significant traction in the AI-for-math community. Through these efforts, we aim to design a semantics-first LLM that autoformalizes informal scientific problems into machine-checked specifications and synthesizes solver-ready code. This closes the loop between formal analysis and LLM reasoning, ultimately surpassing human heuristics for scientific discovery.

Bio: Dr. Wuyang Chen is a tenure-track Assistant Professor in Computing Science at Simon Fraser University. He is also a visiting research scientist at Microsoft. Previously, he was a postdoctoral researcher in Statistics at the University of California, Berkeley, advised by Professor Michael Mahoney. He obtained his Ph.D. in Electrical and Computer Engineering from the University of Texas at Austin in 2023, advised by Professor Atlas Wang. Dr. Chen's research focuses on integrating AI methods with physical knowledge, scientific machine learning, and theoretical understanding of deep networks. Dr. Chen has published papers at CVPR, ECCV, ICLR, ICML, NeurIPS, and other top conferences. Dr. Chen's research has been recognized by the US NSF newsletter, two Doctoral Dissertation Awards from INNS and iSchools, AAAI New Faculty Highlights, and NVIDIA Academic Grant Award. Dr. Chen also hosted and co-organized many conference workshops at NeurIPS, ICLR, CVPR.

Location: NCS 120