The Pittsburgh Supercomputing Center is pleased to present a Machine Learning and Big Data workshop.

This workshop will focus on topics including big data analytics and machine learning with Spark, as well as deep learning.

This will be an IN PERSON event hosted by various satellite sites, there WILL NOT be a direct to desktop option for this event. SBU's Institute for Advanced Computational Science (IACS) is one of those satellite sites!

Location: IACS Conference Room #2

Interested applicants must first have an ACCESS ID. If you don't have the ID, please visit this page to create one: ACCESS USER REGISTRATION.


Once you have an ACCESS ID, please login (see top right here) then register here.
Abstract: Recent progress in large language and vision models demonstrates how far we can go by scaling with vast internet-scale data. In contrast, physical AI, agents that perceive and act in the real world, still lags far behind. Today, both academia and industry primarily pursue generalizable physical AI by scaling up: collecting large-scale action-video datasets or training world models that enable interaction through learned environments. However, this paradigm is inherently inefficient and will soon reach a data ceiling. In this talk, I argue for a shift from scaling up to scaling out. I introduce reality world simulators, a new paradigm that converts real-world videos into diverse, interactive simulation environments. Instead of relying on more data collection, this approach expands data through structured reconstruction and recomposition, enabling both higher data efficiency and physically grounded interaction. I will present a three-pronged approach: 1) Scaling out via Digital Twins: reconstructing controllable, interactive environments from monocular videos to support diverse agent exploration. 2) Scaling out via Digital Cousins: disentangling scene structure into compositional elements to generate large-scale variations of real-world environments. 3) Scaling out via Embodied Humans: incorporating realistic human dynamics to improve safety and social compliance in robot learning. Finally, I will outline a roadmap toward building generalizable and safe physical AI systems for open-world deployment.

Bio: Dr. Wayne Wu is a postdoctoral researcher at UCLA Computer Science, working closely with Bolei Zhou, and collaborating with Trevor Darrell (UC Berkeley EECS) and Jiaqi Ma (UCLA CEE). He received his Ph.D. in Computer Science and Technology from Tsinghua University in June 2022 and was previously a visiting Ph.D. student at Nanyang Technological University. He also spent seven years in industry, where he led the research and development of products that reached more than 10 million end users worldwide. His research lies at the intersection of computer vision, robotics, and computer graphics. He focuses on developing infrastructure and methods to scale physical AI, enabling robots to work reliably and safely in the open world. He has published over 50 papers at top-tier venues including CVPR, ICCV, ICLR, NeurIPS, and ICRA, with over 9,500 citations and 10,000 GitHub stars. His work has received a CVPR Best Paper Candidate and multiple Oral, Spotlight, and Highlight presentations. He was also honored with the 2025 UCLA Chancellor's Award for Postdoctoral Research, recognizing the best postdocs at UCLA, and he was the only awardee from the School of Engineering. He serves as an Area Chair at CVPR 2026.

Location: NCS 120



Abstract: Trustworthy AI deployment in high-stakes domains requires systems that are fair, private, robust, and controllable as they scale. Yet these demands are often pursued through ad-hoc approaches, lacking a systematic understanding of the inherent trade-offs between competing objectives. We add fairness regularizers and hope bias decreases. We train on massive datasets and hope the model learns the underlying logic of how concepts combine, rather than memorizing statistical shortcuts. We encrypt data and hope the resulting computational overhead remains manageable. But hope isnot a science.
In this talk, I argue that what trustworthy AI lacks is not better heuristics but a deeper science of what these properties fundamentally cost and what is achievable. Before we can fix a system, we must map the terrain: what trade-offs are unavoidable, what regions of performance areunreachable, and how far current methods fall from what is actually achievable. My research builds this map across fairness, privacy, robustness, and controllability, following a common methodology: diagnose where models fail, characterize the fundamental limits any method must obey, and design systems that approach those limits. I will present this framework, its extension to scientific applications where we replace statistical constraints with physical laws to ensure AI systems remain grounded in reality, and a vision for scaling these principles to the rapidly expanding ecosystem of composed and interacting AI systems.


Bio: Dr. Vishnu Boddeti is an Associate Professor in the Department of Computer Science and Engineering at Michigan State University, where he leads the Human Analysis Lab (HAL). His research develops mathematical frameworks for trustworthy AI, spanning fairness, privacy, robustness, and physics-informed learning, with an emphasis on characterizing fundamental limits and building systems that achieve them. His work has been supported by NSF, NIST, DARPA, ONR, Ford, and others, and recognized with a Meta Research Award (2021). His research has been featured on the cover of Nature, recognized as an Editor's Highlight in Nature Communications, and received multiple best paper awards, including the 2024 IEEE-CCF Cloud Computing Best Paper Award and the TMLR Outstanding Certification Finalist (2023). He serves as Senior Area Editor for IEEE Transactions on Information Forensics and Security and completed his PhD in ECE from Carnegie Mellon University in 2012.

Location: NCS 120
Abstract: Large language models are prone to memorizing some of their training data. Memorized (and possibly sensitive) samples can then be extracted at generation time by adversarial or benign users. There is hope that model alignment---a standard training process that tunes a model to harmlessly follow user instructions---would mitigate the risk of extraction. However, we develop two novel attacks that undo a language model's alignment and recover thousands of training examples from popular proprietary aligned models such as OpenAI's ChatGPT. Our work highlights the limitations of existing safeguards to prevent training data leakage in production language models.

Speaker: Pegah Alipoormolabashi

Location: CS2311
AI for Conservation: AI and Humans Combating Extinction Together by Daniel I. Rubenstein of Princeton University

ABSTRACT: The state of our planet is not good. We have lost more than 60% of the world's wildlife. Stopping the decline remains a challenge, especially since acquiring appropriate knowledge is expensive, time consuming and risky. Visual observations following the fates of a few individuals was the currency of the realm. But GPS technology and now machine learning provide a non-invasive scalable alternative. Photographs, taken by field scientists, tourists, automated cameras and incidental photographers, are the most abundant source of data on wildlife today. Wildbook, a project of tech for conservation coordinated by a non-profit Wild Me, is an autonomous computational system that starts from massive collections of images and, by detecting various species of animals and identifying individuals, combined with sophisticated data management, turns them into high-resolution information databases, enabling scientific inquiry, conservation and citizen science.

BIO: Dan Rubenstein is the Class of 1877 Professor of Zoology. He is currently Director of Princeton's Environmental Studies Program and is former Chair of Princeton University's Department of Ecology and Evolutionary Biology and Director of Princeton's Program in African Studies. He is a behavioral ecologist who studies how environmental variation and individual differences shape social behavior, social structure, sex
roles and the dynamics of populations. He has special interests in all species of wild horses, zebras and asses, and has done field work on them throughout the world identifying rules governing decision-making, the emergence of complex behavioral patterns and how these understandings influence their management
and conservation. In Kenya he also works with pastoral communities to develop and assess impacts of various grazing strategies on rangeland quality, wildlife use and livelihoods. He has also developed a scout program for gathering data on Grevy's zebras and created curricular modules for local schools to raise awareness about the plight of this endangered species. He engages people as 'Citizen Scientists' and has recently extended his work to measuring the effects of environmental change, including issues pertaining to the global commons
and changes wrought by management and by global warming, on behavior.

George Em Karniadakis received his SM and PhD from Massachusetts Institute of Technology. He was appointed lecturer in the Department of Mechanical Engineering at MIT in 1987 and subsequently he joined the Center for Turbulence Research at Stanford/Nasa Ames. He joined Princeton University as assistant professor in the Department of Mechanical and Aerospace Engineering and as associate faculty in the program of applied and computational mathematics. He was a visiting professor at Caltech in 1993 in the Aeronautics Department and joined Brown University as associate professor of applied mathematics in the Center for Fluid Mechanics in 1994. After becoming a full professor in 1996, he continues to be a visiting professor and senior lecturer of Ocean/Mechanical Engineering at MIT. He is an AAAS fellow (2018), fellow of the Society for Industrial and Applied Mathematics (2010), fellow of the American Physical Society (2004), fellow of the American Society of Mechanical Engineers (2003) and associate fellow of the American Institute of Aeronautics and Astronautics (2006). He received the Alexander von Humboldt award in 2017, the Ralf E Kleinman award (2015), the J. Tinsley Oden Medal (2013), and the CFD award (2007) from the US Association in Computational Mechanics. His h-index is 103, and he has been cited over 52,000 times.


Abstract:
Karniadakis will present a new approach to develop a data-driven, learning-based framework for predicting outcomes of physical and biological systems, governed by PDEs, and for discovering hidden physics from noisy data. He will introduce a deep learning approach based on neural networks (NNs) and generative adversarial networks (GANs). He will also introduce new NNs that learn functionals and nonlinear operators from functions and corresponding responses for system identification. Unlike other approaches that rely on big data, here we learn from small data by exploiting the information provided by the physical conservation laws, which are used to obtain informative priors or regularize the neural networks. He will demonstrate the power of PINNs for several inverse problems in fluid mechanics, solid mechanics and biomedicine including wake flows, shock tube problems, material characterization, brain aneurysms, etc., where traditional methods fail due to lack of boundary and initial conditions or material properties. He will also present a new NN, DeepM&Mnet, which uses DeepOnets as building blocks for multiphysics problems, and he will demonstrate its unique capability in a 7-field hypersonics application.  

To register and for more information, click here 

Scaling the NY AI Innovation Ecosystem

The State University of New York at Stony Brook will bring together leading AI experts to promote a future where AI drives responsible progress. This two-day event will provide a significant opportunity to explore the future of AI, exchange ideas, and connect with those at the forefront of research and deployment. We invite faculty, staff, and students from all SUNY institutions and beyond, as well as industry AI practitioners and policymakers to attend.

Recognized AI experts from academia, industry, and government will present on topics such as AI applications, innovative developments in research and technology, workforce development, as well as ethical and societal impacts.

A 90-minute poster session is included in the schedule. If you would like to submit an abstract for consideration, please see the Call for Abstracts. The poster session segment of the symposium will be held in honor of the Inauguration of Dr. Andrea Goldsmith, the State University of New York at Stony Brook's seventh President. Poster printing for all participants will be covered by the Inauguration Planning Committee. SUNY students presenting posters are also eligible for travel reimbursement.

We kindly ask faculty to encourage their students to attend and to submit their work for presentation.

For additional information and to register, visit the symposium website. Please direct any questions to suny-ai-symposium-sbu@stonybrook.edu.

Register.

Abstract: Sub-grid turbulence is challenging to resolve in climate models; therefore, it is parameterized. Traditionally, turbulent parameterizations have relied on physics-based and equation-based approaches. However, ad hoc and uncertain components in these parameterizations introduce uncertainty in future climate predictions. Recently, data-driven techniques have emerged as an alternative for modeling sub-grid fluxes. I will demonstrate the use of machine learning to model vertical turbulent fluxes in the ocean surface boundary layer and its impact on reducing biases in NOAA's Geophysical Fluid Dynamics Laboratory ocean climate model.

I will show how neural networks, trained to predict the eddy diffusivity profile from high-fidelity yet computationally expensive turbulence schemes, enhance the vertical mixing scheme in the climate model. These networks replace ad hoc components while maintaining the conservation principles of the standard ocean model equations. The enhanced scheme outperforms its predecessor by reducing biases in the mixed-layer depth and modestly improving tropical upper-ocean stratification in ocean-only global simulations. Furthermore, simplified equations that can replace the neural networks show similar improvements but with lower computational cost and better interpretability. They point to structural deficiencies in the baseline parameterization. This work is one of the first successful applications of machine learning to improve a sub-grid parameterization of turbulent mixing in ocean climate models.

IACS Seminar Speaker: Aakash Sane, Princeton University

Location: IACS Seminar Room or Zoom

Join Zoom Meeting: https://stonybrook.zoom.us/j/97764942108?pwd=MzCWupCe3L9mKdrgfO2bJg3GBbvXuf.1
Meeting ID: 977 6494 2108
Passcode: 519324