Predictable Autonomy for Cyber-Physical Systems by Stanley Bak, Safe Sky Analytics

ABSTRACT: Cyber-physical systems combine complex physics with complex software. Although these systems offer significant potential in fields such as smart grid design, autonomous robotics and medical systems, verification of CPS designs remains challenging. Model-based design permits simulations to be used to explore potential system behaviors, but individual simulations do not provide full coverage of what the system can do. In particular, simulations cannot guarantee the absence of unsafe behaviors, which is unsettling as many CPS are safety-critical systems.

The goal of set-based analysis methods is to explore a system's behaviors using sets of states, rather than individual states. The usual downside of this approach is that set-based analysis methods are limited in scalability, working only for very small models. This talk describes our recent process on improving the scalability of set-based reachability computation for LTI hybrid automaton models, some of which can apply to very large systems (up to one billion continuous state variables!). Lastly, we'll discuss the significant overlap of techniques used for our scalable reachability analysis methods with set-based input/output analysis of neural networks.

BIO: Stanley Bak is a computer scientist investigating the predictable design of autonomous cyber-physical systems. He strives to develop practical formal methods that are both scalable and useful, which demands developing new theory, programming efficient tools and building experimental systems. He received a Bachelor's degree in Computer Science from Rensselaer Polytechnic Institute (RPI) in 2007 (summa cum laude), and a Master's degree in Computer Science from the University of Illinois at Urbana-Champaign (UIUC) in 2009. He completed his PhD from the Department of Computer Science at UIUC in 2013. He received the Founders Award of Excellence for his undergraduate research at RPI in 2004, the Debra and Ira Cohen Graduate Fellowship from UIUC twice, in 2008 and 2009, and was awarded the Science, Mathematics and Research for Transformation (SMART) Scholarship from 2009 to 2013. From 2013 to 2018, Stanley was a Research Computer Scientist at the US Air Force Research Lab (AFRL), both in the Information Directorate in Rome, NY, and in the Aerospace Systems Directorate in Dayton, OH. He currently helps run Safe Sky Analytics, a research consulting company investigating verification and autonomous systems, and performs teaching as an Adjunct Professor at Georgetown University.
The Hudson River Estuary (HRE) and New York Bight (NYB) are closely connected, with HRE acting as crucial areas where many NYB marine species spawn and grow. Understanding how these biotic and abiotic environments interact, especially with rapid climate change, is key to better managing fisheries and conserving ecosystems. To better understand the HRE-NYB ecosystem, we develop a comprehensive ecosystem model that links physical and biological processes. Using data from long-term monitoring programs, we analyze ecological patterns and identify key factors regulating the ecosystem. We use this information to develop a model that mimics the food web from tiny plankton to large predators in the ecosystem. This model can help us better understand how changes in the environment, like rising temperatures, and human activities such as fishing affect marine lives and ecosystem over time. The insights from this model can support smarter fisheries management and efforts to conserve marine ecosystems in the HRE-NYB region.

IACS Student Seminar Speaker: Xiangyan Yang, Dept. of Applied Math & Statistics

Location: IACS Seminar Room or Zoom

Join Zoom Meeting: https://stonybrook.zoom.us/j/91650247483?pwd=fvAGEwadplJh7jFC5RWcdvZ5NWPJth.1
Meeting ID: 916 5024 7483
Passcode: 631055
Abstract: Self-supervised representation learning (SRL) has emerged as a pivotal advancement in machine learning, offering high-quality data representations without the need for labeled datasets. While SRL has demonstrated enhanced adversarial robustness compared to supervised learning, its resilience against other attack types, particularly backdoor attacks, remains an open question. Recent studies have revealed potential vulnerabilities in SRL, underscoring the necessity for a comprehensive security analysis. However, existing research often extrapolates attacks from supervised learning paradigms, neglecting the unique challenges and opportunities inherent to self-supervised mechanisms.

This thesis proposal aims to address three critical objectives in the realm of self-supervised learning: (1) exploring novel attack vectors, (2) implementing and evaluating practical attacks, and (3) developing robust countermeasures. We focus on two key SRL paradigms: Contrastive Learning and Diffusion Models. For Contrastive Learning, we synthesize existing security vulnerabilities and introduce innovative attack vectors, such as CTRL, to uncover distinctive risks. We conduct a comparative analysis of contrastive and supervised learning approaches in their defense against these threats, exploring potential safeguards and highlighting the limitations of current protective measures in self-supervised contexts. Regarding Diffusion Models, we demonstrate inherent vulnerabilities in their application to adversarial purification.

Our research aims to illuminate the unique challenges posed by emerging attack vectors in self-supervised learning, fostering technical advancements to address underlying security risks in real-world applications. By contributing to the development of more resilient and secure self-supervised representation learning systems, we seek to enhance their reliability and trustworthiness in practical scenarios. This comprehensive examination of SRL's security landscape will provide valuable insights for the broader machine-learning community and pave the way for more robust AI systems.

Join here.

The Della Pietra Lecture Series is pleased to present this lecture by Scott Aaronson:

Abstract: I'll survey some areas where I think theoretical computer science, math, and statistics can potentially contribute to the urgent quest to align powerful AI with humane values. These areas include: the watermarking of AI outputs, mechanistic interpretability (including Paul Christiano's No-Coincidence Principle, and succinct digests of the training process to aid interpretability), and theoretical guarantees for out-of-distribution generalization.

Speaker: Scott Aaronson is Schlumberger Chair of Computer Science at the University of Texas at Austin, and founding director of its Quantum Information Center. He received his bachelor's from Cornell University and his PhD from UC Berkeley. Aaronson's research has focused mainly on the capabilities and limits of quantum computers. His first book, Quantum Computing Since Democritus, was published in 2013 by Cambridge University Press. He received the National Science Foundation's Alan T. Waterman Award, the United States PECASE Award, the Tomassoni-Chisesi Prize in Physics, and the ACM Prize in Computing, and is a Fellow of the ACM and the AAAS and a member of the National Academy of Sciences. He blogs at Shtetl-Optimized, https://www.scottaaronson.com/blog.

Location: Della Pietra Family Auditorium (SCGP 103)
Abstract: The recent expansion of online sport wagering and igaming has led to higher rates of problem gambling, particularly among emerging adults and other population subgroups. The Center for Gambling Studies (CGS) at the Rutgers University, School of Social Work, is using big data analysis, machine learning and GIS mapping to identify geographic locations with populations most at risk to guide the development of targeted interventions. This presentation will review the GIS StoryMap for the State of New Jersey, including a blueprint for the highest risk target service areas in the state. It will also present findings from a machine learning model that identifies the key risk factors for high-intensity online casino bettors. Implications for prevention, treatment and policy initiatives will be discussed.

Bio: Lia Nower, J.D., Ph.D., is a Distinguished Professor, Associate Dean for Research, and Director of the Center for Gambling Studies at Rutgers University. A clinician and attorney, her research focuses on big data analysis and machine learning models for online gambling and sports wagering; gambling and video gaming among emerging adults; policy initiatives around harm reduction and responsible gambling, and etiology and treatment of problem gambling. Dr. Nower serves as a senior editor for Addiction. She has received both the Research (2019) and the Lifetime Research Award (2022) from the National Council on Problem Gambling and the Board of Trustees Award for Research (2022) from Rutgers University.

Join Zoom Meeting: https://stonybrook.zoom.us/j/95617197636?pwd=KytzZ2pVRG9SZGpKZUtpNXJISjNjZz09
Meeting ID: 956 1719 7636 Passcode: 924293
Are you concerned about AI issues with your asynchronous online courses? Is your fully online course vulnerable to AI plagiarism? Do you want to engage your online students using AI? Discover the future of education with our AI-powered solutions designed specifically for online asynchronous courses. This innovative approach uses artificial intelligence to transform the way courses are delivered, making learning more personalized, engaging, and effective.

Register here: https://stonybrook.zoom.us/meeting/register/RD94cHiHRwCj6xNkCZqNEg

Abstract: Traditional questionnaires remain the primary method for assessing psychological outcomes and beliefs, capturing individuals' and populations' inner states. This dissertation presents an alternative computational method that overcomes key limitations in current mental health monitoring, particularly in spatiotemporal resolution, responses to major events, and automatic belief identification. By analyzing ∼1 billion Tweets from 2 million geo-located users, we created a big data pipeline for estimating depression and anxiety at the county-week level. These Language-Based Mental Health Assessments (LBMHA) demonstrated higher reliability and validity than traditional survey measures. Our approach effectively captured mental health trends and highlighted significant increases in mental illness following major events. Using the LBMHA pipeline, we conducted quasi-experiments, research designs that simulate randomized control trials, to generate explanations for mental health changes due to COVID-19 incidence/death. Utilizing these time-series analyses, we conducted discontinuity forecasting for community-specific anxiety shifts using statistical learning via ensemble and contextual models. To likewise investigate individual internal states, we created a novel task and annotated dataset for self belief language identification. Our fine-tuned language model for self-belief classification, despite its relatively small scale, outperformed GPT-4o. The self belief topics identified by our model successfully predicted depression, anxiety, and stress, offering insights into the relationship between self-conceptualization and mental health. The adoption of scalable language-based assessments with modern distributed computation presents a promising avenue for advancing community and individual mental health research.

Speaker: Siddharth Mangalik

https://stonybrook.zoom.us/j/91251321639?pwd=faggV5jZ7ByFDCFmnLXD3HiYxjQ1Eb.1&jst=2
Abstract: Humans perceive the world around them by recognizing global patterns and structures such as object parts, branches, their spatial arrangement, and so on. Most deep learning models, however, take a fundamentally local approach. They process images pixel-by-pixel rather than focusing on structures as a whole. While these models indeed perform well on many tasks, the local (pixel-level) versus global (structure-level) disconnect makes them harder to interpret and control.

Topology, in a general sense, is a mathematical language for describing structure. It delineates how different parts of an image relate to one another, capturing both individual structures and their overall layout. Preserving topology enforces structural correctness and, by extension, semantic validity.

In this thesis, we investigate how topological constraints can be used to bridge the gap between local and global understanding. We use topology to inform the design of deep learning models that are explicitly structure-aware. Our thesis focuses on dense prediction tasks, which include image segmentation, uncertainty estimation, and generative modeling. First, we introduce a topological interaction module for semantic segmentation that encodes containment and exclusion constraints directly into the learning process. This preserves anatomical hierarchies and improves multi-class consistency. Next, since segmentation models can never be truly perfect, we address the need for reliable uncertainty estimation to identify error-prone regions. Unlike conventional pixel-wise uncertainty maps, which tend to be noisy and difficult to interpret, we propose reasoning at the level of structural units--branches and connections--which are more visually discernible and actionable. Finally, we leverage topology for generative modeling. We propose a topology-guided diffusion framework that can be controlled using structural attributes like object count and connectivity.

Together, these contributions establish a unified approach to topology-informed, structure-preserving dense prediction models. By integrating topological reasoning with deep networks, this thesis advances models that are not only accurate, but also structurally consistent, interpretable, and controllable. The results from this thesis have been published in ECCV, NeurIPS, and ICLR.

Speaker: Saumya Gupta

Location: New Computer Science (NCS) 120


Zoom: https://stonybrook.zoom.us/j/93643318604?pwd=kv8DagpbayzizivU29UCYItnlzlYRM.1&jst=2