Abstract: Computational pathology has revolutionized cancer diagnosis and research through the analysis of digitized whole slide images (WSIs). However, the giga-pixel size of these images presents profound technical challenges, creating two intertwined bottlenecks: computational inefficiency and label inefficiency. The immense data scale makes standard end-to-end (E2E) training of deep neural networks infeasible due to prohibitive GPU memory requirements, while the reliance on expert pathologists for annotations makes obtaining high-quality labeled data a tedious and expensive process. This proposal confronts these dual challenges by developing a series of novel model architectures, training paradigms, and self-supervised learning methods designed to create a more efficient and effective framework for WSI analysis.

To improve computational efficiency, this proposal first introduces a locally supervised learning paradigm that enables E2E training on entire WSIs by partitioning a network into gradient-isolated modules, circumventing the memory bottleneck of backpropagation. Second, it presents Prompt-MIL, a parameter-efficient fine-tuning framework that reduces the number of trainable parameters, memory consumption, and training time by fine-tuning only few prompts to guide large pre-trained models. Third, this work advances the efficient architecture on WSIs by developing novel State-Space Models (SSMs). It proposes 2DMamba, the first intrinsic Mamba architecture that preserves the crucial 2D spatial structure of images, overcoming the spatial discrepancy inherent in 1D models. Fourth, to address the inefficiency of multi-directional scans in Mamba models, including 2DMamba, it presents Locally Bi-directional Mamba (LBMamba), which introduces a novel, hardware-aware local backward scan that integrates bi-directional scan into a single forward pass, significantly improving throughput performance trade-off. Lastly, it proposes an extension to the LBMamba, warp-level Bi-directional Mamba (WLBMamba) that extends the thread-level bidirectional scan to warp-level bidirectional scan that further improves the throughput performance trade-off.

To improve label efficiency, this proposal proposes a Precise Location-based Matching strategy for self-supervised dense contrastive learning. By allowing a local patch in one augmented view to match multiple overlapping patches in another, creates a more accurate correspondence, leading to superior feature representations for dense prediction tasks like segmentation and detection.

In summary, this proposal presents a holistic investigation into the efficiency bottlenecks in computational pathology. Through these combined contributions in model architecture, training paradigms, and self-supervised learning, this work establishes a more scalable, efficient, and powerful computational framework for analyzing giga-pixel pathology images.

Speaker: Jingwei Zhang

Location: Old Computer Science Room 2114

Zoom: https://stonybrook.zoom.us/j/95187903649?pwd=tV0CNxLu1QKqw7hGmcE1h0rJ2C6n1b.1
Meeting ID: 951 8790 3649 | Passcode: 488916
Abstract: Artificial Intelligence for Science (AI4Sci) has become a transformative approach in modeling and understanding complex physical systems, encompassing different scales such as atomistic systems and continuum systems. In atomistic systems, AI has shown potential in accelerating simulations, optimizing molecular dynamics, and predicting material and molecular properties through data-driven approaches, enhancing computational efficiency while preserving accuracy. For continuum systems, AI provides powerful tools for solving partial differential equations (PDEs) and learning physical patterns from data, capturing intricate dynamics that govern physical and engineering processes. This work explores AI methods--particularly equivariance for neural networks and neural operators--bridging atomistic and continuum representations. We analyze the implications of incorporating symmetries to improve model robustness and learning efficiency, providing a cohesive AI- driven framework for advancing scientific discovery. The findings aim to underscore the role of AI in enhancing accuracy, applicability, scalability, interopretability, and generalization across scales, from molecular simulations to physical modeling, opening pathways for next- generation applications in computational science. Biography: Wenhan Gao is a third-year Ph.D. student in Applied Mathematics at Stony Brook University, where he works under the supervision of Professor Yi Liu. Wenhan's research focuses on equivariant neural networks, graph neural networks, and AI for partial differential equations. Wenhan's work seeks to leverage the power of symmetries to aid AI models, particularly in fields such as computer vision (image and video generation), physical simulation (modeling climate change), and computational chemistry (drug discovery). He has published papers on the aforementioned topics in leading venues like NeurIPS, Transactions on Machine Learning Research (TMLR), and Journal of Computational Physics (JCP). He also has several preprints under review in leading venues like ICLR and CVPR. In addition to his research, Wenhan has served as a reviewer for top-tier conferences, including ICLR, NeurIPS, ICML, and KDD, and as a lecturer for undergraduate and graduate courses at Stony Brook University. Wenhan was awarded the NeurIPS Travel Award and Excellence in Teaching for Fall 2023.
Please join us on Zoom for our next event in the Fall 2025 Stony Brook School of Nursing Research Seminar Series presented by our Office of Research and Innovation.

Topic: Responsible Artificial Intelligence: Promoting Health Equity for All

Speaker: Michael P. Cary, Jr., PhD, RN, FAAN.

Dr. Cary is a tenured Associate Professor at the Duke University School of Nursing. Dually trained as a health services researcher and applied health data scientist, Dr. Cary utilizes AI to investigate health disparities in aging populations, thereby promoting health equity and improving healthcare delivery. He co-directs HUMAINE™, an initiative dedicated to equipping nurses and healthcare professionals with the knowledge and skills necessary for the responsible use of AI in clinical practice.

Register: https://web.cvent.com/event/057978a5-a770-4de5-aca5-ad00287e4902/summary



Abstract: The current approach to materials design, driven by strategic experimentation and supported by physics-based simulation across relevant scales, has been the standard for decades. While the theoretical component in this workflow provides valuable understanding of material behavior, it often fails to deliver actionable guidance for implementation. Advances in artificial intelligence and machine learning (AI/ML), together with high-performance computing (HPC), now offer a viable pathway to close this gap and accelerate both discovery and process optimization. This presentation will outline practical approaches for integrating AI/ML with HPC-enabled, high-throughput computation to explore high-dimensional search spaces. Examples will include the development of engineering alloys for extreme environments, the use of neural networks to rapidly improve computational thermodynamic models, and vapor processing optimization for the manufacturing of ultra-high-temperature ceramics. I will highlight how scientific insight and domain expertise remain essential for translating surrogate model predictions into impactful outcomes. Finally, I will conclude with current challenges and future opportunities for AI/HPC-driven materials research.

Speaker: Dongwon Shin
This seminar will be held in person and online

Join Zoom Meeting: https://stonybrook.zoom.us/j/93730374357?pwd=YDLJ7ELqOQnTZEQhlN8Pa4TuhaiFK8.1
AI3, SBU Libraries and IACS present
at International Love Data Week
sponsored by The Office of the Provost and
Educational and Institutional Effectiveness (EIE)

Special Talk and Panel Discussion

How I Learned to Stop Worrying and Love AI (For Now)


with Paul Fain from The Job and Work Shift

A reporter's take on what we know--and what we don't know--about AI's emerging impacts on the labor market. The discussion will include the latest research from economists and the AI labs themselves about how workers are using AI, and current thinking among experts on how the tech's rapid deployment will play out across job roles, industries, and regions.

Panel discussion to follow with:

  • Lav Varshney, Della Pietra Infinity Professor and inaugural director of the AI Innovation Institute
  • Nicholas Johnson, Director of AI, SBU Libraries
  • Marianna Savoca, Associate Vice President for Career Readiness and Experiential Education
Paul Fain is co-founder of Work Shift, editor of the must-read newsletter, The Job, and host of The Cusp podcast. A veteran higher education reporter, Paul is perhaps the nation's top journalist focused on connections between education and work. He started Work Shift after a decade as a senior reporter and then news editor at Inside Higher Ed, where he led the outlet's coverage of low-income and first-generation students, college completion, community colleges, federal policy, and emerging models of higher education. He also was the founding host of the successful podcast, The Key with Inside Higher Ed, and has contributed chapters for books on innovation in higher education, published by the Harvard University Press and the Stanford University Press. Earlier in his career, Paul was a senior reporter at The Chronicle of Higher Education.

Limited Seats!

Registration is required.

Abstract:

Many real world complex problems are multi-step reasoning tasks. These range from analytic tasks such as answering questions to automation tasks where agents complete tasks on behalf of users.. Evaluation, datasets, and models for such tasks can be unreliable for multiple reasons. (i) Datasets often have annotation artifacts and biases, allowing models to take reasoning shortcuts. Such shortcuts can allow models to make effective guesses -- or, in a sense, cheat -- to achieve high performance without any multi-step reasoning. This issue is further exacerbated for complex tasks because as the number of the required reasoning steps increases, so do the avenues for bypassing those steps. (ii) Models trained on such dataset/s learn to solve the task by taking reasoning shortcuts instead of proper multi-step reasoning. As a result, these models are not robust (reliable) when evaluated in an out-of-distribution evaluation setting. (iii) Lastly, recent works have shown that language models can solve complex multi-step tasks by producing a step-by-step explanation without any training. However, these methods often hallucinate factually incorrect (i.e., unreliable) explanations when posed with knowledge-intensive tasks.

I address these challenges by carefully characterizing the requirements of robust multi-step reasoning and designing reliable evaluation datasets and training methods that necessitate thorough multi-step reasoning. In DiRe, I first formalize and introduce Disconnected Reasoning, i.e., reasoning that allows models to arrive at the correct answer by bypassing necessary reasoning steps, and use this formalization to measure how much multi-step reasoning a model does on a dataset. In MuSiQue, I built a multi-step reasoning dataset for QA from scratch that avoids cheatability via disconnected reasoning, providing a more reliable evaluation. In TeaBReaC, I developed a synthetically generated multi-step QA pretraining dataset designed to force models to avoid disconnected reasoning and learn reliable multi-step reasoning. In IRCoT, I address the reliability of model-generated multi-step reasoning chains by interleaving models' step-by-step reasoning with a step-by-step retrieval from an external corpus, resulting in more factually correct reasoning. Finally, in AppWorld, I built a multi-step reasoning dataset that requires highly interactive problem-solving in an environment carefully designed to ensure models need thorough reasoning to succeed.
Speaker: Harsh Trivedi

Location: NCS 220 or Zoom

https://stonybrook.zoom.us/j/99096379762?pwd=zYCJZQVxRuZd9BboscO4nlodCwsKBr.1
Le Hou Dissertation Defense: Deep Learning for Digital Histopathology across Multiple Scales

ABSTRACT: Histopathology is the study of tissue changes caused by diseases such as cancer. It plays a crucial role in disease diagnosis, survival analysis and development of new treatments. Using computer vision techniques, I focus on multiple tasks for automated analysis in digital histopathology images, which are challenging because histopathology images are heterogeneous and complex, due to the large variation of hundreds of cancer types in gigapixel resolution. In this thesis, I show how histopathology image analysis tasks can be viewed in three scales: Whole Slide Image (WSI)-level, patch-level and cellular-level, and present my contributions in each resolution level.

BIO: WSI-level analysis such as classifying WSIs into cancer types is challenging, because conventional classification methods such as off-the-shelf deep learning models cannot be applied directly on gigapixel WSIs due to computational limitations. I contribute a patch-based deep learning method that classifies gigapixel WSIs into cancer types and subtypes with close-to-human performance. This method is useful for computer-aided diagnosis. At patch-level, I contribute a novel method for histopathology image patch classification. On the task of identifying Tumor Infiltrating Lymphocyte (TIL) regions, the prediction result of this method correlates to the survival rate of patients. At cellular-level, I contribute novel methods for nucleus classification and roundness regression, which are interpretable features for histopathology studies. With this method, I generated a large-scale dataset of segmented nuclei, in WSIs from a large publicly available digital histopathology image dataset, to help advance histopathology research.