Scaling the NY AI Innovation Ecosystem

The State University of New York at Stony Brook will bring together leading AI experts to promote a future where AI drives responsible progress. This two-day event will provide a significant opportunity to explore the future of AI, exchange ideas, and connect with those at the forefront of research and deployment. We invite faculty, staff, and students from all SUNY institutions and beyond, as well as industry AI practitioners and policymakers to attend.

Recognized AI experts from academia, industry, and government will present on topics such as AI applications, innovative developments in research and technology, workforce development, as well as ethical and societal impacts.

A 90-minute poster session is included in the schedule. If you would like to submit an abstract for consideration, please see the Call for Abstracts. The poster session segment of the symposium will be held in honor of the Inauguration of Dr. Andrea Goldsmith, the State University of New York at Stony Brook's seventh President. Poster printing for all participants will be covered by the Inauguration Planning Committee. SUNY students presenting posters are also eligible for travel reimbursement.

We kindly ask faculty to encourage their students to attend and to submit their work for presentation.

For additional information and to register, visit the symposium website. Please direct any questions to suny-ai-symposium-sbu@stonybrook.edu.

Register.

Abstract: Robot control has evolved from optimization-based controllers---precise but task-specific---through deep reinforcement learning's learned policies, to Vision-Language-Action (VLA) models that leverage pretrained vision-language backbones for language-conditioned manipulation across diverse tasks.
Despite their promise, VLAs exhibit a critical limitation: they function primarily as trajectory learners rather than skill learners. Recent evaluations reveal that VLAs often fail when faced with even minor variations in object initialization or environmental conditions, suggesting they memorize specific trajectories rather than acquiring generalizable manipulation skills. Attempts to address this through 3D spatial representations have shown limited success, indicating that the missing component may be more fundamental than geometric understanding alone.
This work argues that World Models (WMs)---internal representations that predict future states given actions---constitute the missing piece for robust VLA systems. We present one completed contribution and two ongoing investigations.
We developed a dual-layer world model for human-robot interaction that anticipates both physical scene evolution and latent human preferences for assistive tasks. Building on these foundations, we present ongoing work probing VLA internal representations to verify implicit world model existence, and propose a WM-VLA integration approach operating in the native visual domain through embedding prediction and image decoding.
Together, these contributions and investigations establish a foundation for WM-VLA systems, pointing toward robust, generalizable robot policies.
Speaker: Jason Qin
Location: NCS 220
Abstract: Large Language Models (LLMs) have transitioned from standalone prediction interfaces into integrated systems that incorporate content protection, external knowledge retrieval, and multi-step reasoning. While these functional layers expand model capabilities, they also introduce complex, inter-component dependencies that create novel and systemic security risks. This research provides a systematic deconstruction of the structural vulnerabilities emerging across these functional layers.

In this proposal, we evaluate the security boundaries of LLM systems through three pivotal dimensions:
The Content Layer: We present Watermark under Fire, revealing the inherent fragility of content-based tracing mechanisms under adaptive perturbations and highlighting the limitations of surface-level safety measures.
The Retrieval Layer: We introduce GraphRAG under Fire to examine the security of topology-aware knowledge integration. We reveal how graph-based indexing can be exploited as a structural lever for high-success poisoning attacks.
The Reasoning Layer: We detail AutoRAN, the first framework demonstrating the hijacking of internal safety reasoning in Large Reasoning Models (LRMs). This work proves that the transparency of the reasoning process itself creates a critical and exploitable attack surface.

Collectively, these studies demonstrate a systemic failure of add-on safety mechanisms in securing the broader LLM ecosystem. By identifying recurring patterns of exploitation across different system layers, this research provides the necessary foundation for transitioning from reactive patching to a more unified and architecturally-grounded approach to AI trustworthiness.

Speaker: Jiacheng Liang

Zoom: https://stonybrook.zoom.us/j/6669990420?pwd=dkY0eEw5YXpPSWo3RUE4OE1oVW90UT09&omn=97367037382
Meeting ID: 666 999 0420
Passcode: 075299
Abstract: Recent advances in Spatial Transcriptomics (ST) pair histology images with spatially resolved gene expression profiles, enabling predictions of gene expression across different tissue locations based on image patches. This opens up new possibilities for enhancing whole slide image (WSI) prediction tasks with localized gene expression. However, existing methods do not fully leverage the interactions between different tissue locations, which are crucial for accurate joint prediction. To address this, we introduce MERGE (Multi-faceted hiErarchical gRaph for Gene Expressions), which combines a multi-faceted hierarchical graph construction strategy with graph neural networks (GNN) to improve gene expression predictions from WSIs. By clustering tissue image patches based on both spatial and morphological features, and incorporating intra- and inter-cluster edges, our approach fosters interactions between distant tissue locations during GNN learning. As an additional contribution, we evaluate different data smoothing techniques that are necessary to mitigate artifacts in ST data, often caused by technical imperfections. We advocate for adopting gene-aware smoothing methods that are more biologically justified. Experimental results on gene expression prediction show that our GNN method outperforms state-of-the-art techniques on multiple metrics such as mean squared error (MSE), mean absolute error (MAE), and pearson correlation coefficient (PCC). Qualitative analysis establishes the effectiveness of MERGE in capturing cancer marker genes, thus consolidating its utility in diagnostics. As an extension of this work, we use MERGE in a setting with an uncertainty calibration branch to perform robust gene expression smoothing. We show that using patch-wise uncertainty from an uncertainty calibration model and the gene expression predictions from MERGE to enrich the ground truth gene expression matrix, results in better alignment with pathologist annotations, thus establishing that the smoothing is biologically informed.

Speaker: Aniruddha Ganguly

Location: Virtual Zoom Meeting


https://stonybrook.zoom.us/j/5474847973?pwd=Sng0Q2h1c1d3cm9sbFBmYUczMHZNdz09
Meeting ID: 547 484 7973
Passcode: 206739
Abstract: Survival prediction using whole slide images (WSIs) can be formulated as a multiple instance learning (MIL) problem. However, existing MIL methods often fail to explicitly capture pathological heterogeneity within WSIs, both globally -- through long-tailed morphological distributions, and locally through -- tile-level prediction uncertainty. Optimal transport (OT) provides a principled way of modeling such heterogeneity by incorporating marginal distribution constraints. Building on this insight, we propose OTSurv, a novel MIL framework from an optimal transport perspective. Specifically, OTSurv formulates survival predictions as a heterogeneity-aware OT problem with two constraints: (1) global long-tail constraint that models prior morphological distributions to avert both mode collapse and excessive uniformity by regulating transport mass allocation, and (2) local uncertainty-aware constraint that prioritizes high-confidence patches while suppressing noise by progressively raising the total transport mass. We then recast the initial OT problem, augmented by these constraints, into an unbalanced OT formulation that can be solved with an efficient, hardware-friendly matrix scaling algorithm. Empirically, OTSurv sets new state-of-the-art results across six popular benchmarks, achieving an absolute 3.6% improvement in average C-index. In addition, OTSurv achieves statistical significance in log-rank tests and offers high interpretability, making it a powerful tool for survival prediction in digital pathology.

Speaker: Qin Ren

Location: https://stonybrook.zoom.us/j/97458849045?pwd=SREklaxa7EFhKaK6TeRXgQPH64vrQJ.1&jst=2
The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 10 minutes) by multiple people. Students can register for 1 credit for CSE 656. Registered students must attend and present a minimum of 2 or 3 talks. Everyone else is welcome to attend. Fill in https://forms.gle/pCVXovgfMfQwGqG38 to subscribe to our mailing list for further announcement.

Abstract: This talk shows how machine learning can address challenges in Astrophysics. We specifically focus on black hole simulations and supernova observations. First, we present a super-resolution technique for black hole simulations that avoids the need for high-resolution labels by leveraging the Hamiltonian and momentum constraints from general relativity. This method reduces constraint violations by one to two orders of magnitude. Next, we introduce Maven, a multimodal foundation model for supernova science. Using contrastive learning to align photometric and spectroscopic data, Maven achieves state-of-the-art results in classification and redshift estimation by pre-training on synthetic data and fine-tuning on real observations.

Bio: Thomas Helfer is a computational physicist specializing in deep learning and physics. Currently based at the Institute for Advanced Computational Science at Stony Brook University, Thomas was previously a postdoctoral fellow at Johns Hopkins and did his PhD with Eugene Lim at King's College in London. In his work, he looks to bridge topics; in his PhD, he bridged theoretical particle physics and gravitational waves. Now, in his postdoctoral work, he aims to find novel applications of deep learning in astrophysics.

*please note: this seminar will be held in a hybrid format*


Location: IACS Seminar Room OR Join Zoom Meeting
https://stonybrook.zoom.us/j/98617630652?pwd=tb4hplPgb3bTTifPCJTCcsn3P9vX8y.1

Meeting ID: 986 1763 0652
Passcode: 882994
Abstract:
People shift their visual attention to gather and prioritize information from their surroundings, helping them navigate complex environments. Understanding these attentional shifts involves decoding the features that guide where attention is directed (spatial areas of focus) and when attention shifts (timing). Decoding these processes can aid applications from interface design to medical diagnosis. However, prior models have not fully explored the underlying factors addressing these aspects. In this dissertation, we study the factors that guide visual attention across diverse image types, spanning natural images, graphic design documents, and whole slide images (WSIs) of cancer tissues, while also predicting visual attention based on these factors.
First, we propose a method to quantify object recognition uncertainty as a factor influencing spatio-temporal attention (where and when) in natural images. We found that it plays a larger role than bottom-up saliency in guiding visual attention. Second, we analyze graphic design documents such as webpages, comics, posters, mobile UIs, etc., which differ from natural images in that they are designed to convey specific messages or elicit desired viewer response. We propose a unified and interpretable deep learning model that predicts both static and dynamic visual attention behavior (addressing where and when) by integrating document layout and content saliency as factors, enhancing attention prediction performance. Finally, in the domain of digital pathology, we investigate pathologists' attention during their examination of giga-pixel WSIs of prostate cancer with an objective to aid in the development of computer-assisted pathology training and clinical decision support systems. Using a digital microscope interface, we collected the largest known dataset of pathologist attention, which allows us to study the factors that guide their spatial and temporal attention patterns (where and when) and develop predictive models. Our study explores key factors guiding their attention, including magnification, slide staining, the nature of the diagnostic task, and their expertise. Motivated by this analysis, we propose deep learning models to solve two tasks: 1) predicting pathologist attention via spatial (heatmaps) and spatio-temporal (scanpaths) models, and 2) inferring pathologist expertise level, both essential technical components towards developing an AI-assisted pathology training pipeline.

Speaker:
Souradeep Chakraborty

Location: New Computer Science Bldg., Room 220

Zoom Link: https://stonybrook.zoom.us/j/9755288447?pwd=TW95T2xqOUZjRnlqcnVFcUQvN0JMdz09
Meeting ID: 975 528 8447
Passcode: 338037