Abstract: Recent advances in Spatial Transcriptomics (ST) pair histology images with spatially resolved gene expression profiles, enabling predictions of gene expression across different tissue locations based on image patches. This opens up new possibilities for enhancing whole slide image (WSI) prediction tasks with localized gene expression. However, existing methods do not fully leverage the interactions between different tissue locations, which are crucial for accurate joint prediction. To address this, we introduce MERGE (Multi-faceted hiErarchical gRaph for Gene Expressions), which combines a multi-faceted hierarchical graph construction strategy with graph neural networks (GNN) to improve gene expression predictions from WSIs. By clustering tissue image patches based on both spatial and morphological features, and incorporating intra- and inter-cluster edges, our approach fosters interactions between distant tissue locations during GNN learning. As an additional contribution, we evaluate different data smoothing techniques that are necessary to mitigate artifacts in ST data, often caused by technical imperfections. We advocate for adopting gene-aware smoothing methods that are more biologically justified. Experimental results on gene expression prediction show that our GNN method outperforms state-of-the-art techniques on multiple metrics such as mean squared error (MSE), mean absolute error (MAE), and pearson correlation coefficient (PCC). Qualitative analysis establishes the effectiveness of MERGE in capturing cancer marker genes, thus consolidating its utility in diagnostics. As an extension of this work, we use MERGE in a setting with an uncertainty calibration branch to perform robust gene expression smoothing. We show that using patch-wise uncertainty from an uncertainty calibration model and the gene expression predictions from MERGE to enrich the ground truth gene expression matrix, results in better alignment with pathologist annotations, thus establishing that the smoothing is biologically informed.

Speaker: Aniruddha Ganguly

Location: Virtual Zoom Meeting


https://stonybrook.zoom.us/j/5474847973?pwd=Sng0Q2h1c1d3cm9sbFBmYUczMHZNdz09
Meeting ID: 547 484 7973
Passcode: 206739
Hidden Biases. Ethical Issues in NLP, and What to Do about Them presented by Dirk Hovy of Bocconi University

ABSTRACT: Through language, we fundamentally express who we are as humans. This property makes text a fantastic resource for research into the complexity of the human mind, from social sciences to humanities. However, it is exactly that property that also creates some ethical problems. Texts reflect the authors' biases, which get magnified by statistical models. This has unintended consequences for our analysis: If our data is not reflective of the population as a whole, if we do not pay attention to the biases contained, we can easily draw the wrong conclusions, and create disadvantages for our users.

In this talk, I will discuss several types of biases that affect NLP models, their sources, and potential counter measures: (1) Bias stemming from data, i.e., selection bias (if our texts do not adequately reflect the population we want to study), label bias (if the labels we use are skewed) and semantic bias (the latent stereotypes encoded in embeddings); (2) Biases deriving from the models themselves, i.e., their tendency to amplify any imbalances that are present in the data; (3) Design bias, i.e., the biases arising from our (the researchers) decisions which topics to analyze, which data sets to use, and what to do with them. For each bias, I will provide examples and discuss the possible ramifications for a wide range of applications, and various ways to address and counteract these biases, ranging from simple labeling considerations to new types of models.

BIO: Dirk Hovey is an associate professor of Computer Science in the department of marketing at Bocconi University. He received his PhD from the University of Southern California in Los Angeles, where he worked as a research assistant at the Information Sciences Institute. 

He works in Natural Language Processing (NLP), a subfield of artificial intelligence. His research focuses on computational social science. His interests include integrating sociolinguistic knowledge into NLP models, using large-scale statistics to model the interaction between people's socio-demographic profile and their language use, and ethics for data science and algorithmic fairness.
Date: March 11, 2022
Time: 2:40PM EST

Title: Towards Scalable and Efficient Machine Learning as a Service (MLaaS)

Abstract:
Driven by the explosive growth of big data, the sustained advances of
Machine Learning (ML), and the fast evolving of computer system
techniques, the past few years have witnessed a surging demand for
Machine Learning as a Service (MLaaS). MLaaS is an emerging computing
paradigm that facilitates ML model design, training, inference serving
and provides optimized executions of ML tasks in an automated,
scalable, and efficient manner. In this talk, I will demonstrate how
to integrate ML algorithm research and system research in synergy to
address the pressing challenges in MLaaS. I will first share a story
about how our system experience led to a novel large batching
algorithm design that revolutionizes large-scale training. Then I will
tell another story about how our gradient compression algorithm
research helped us to discover overlooked critical features of modern
ML systems and thereby build a compression-aware distributed ML
system. I will also briefly discuss a promising future of harnessing
serverless computing for MLaaS model inference serving. I will
conclude my talk with a discussion of interdisciplinary research and
future plans.

Bio:
Dr. Feng Yan is an Assistant Professor of Computer Science and
Engineering at University of Nevada, Reno (UNR) and director of the
Intelligent Data and Systems Lab (IDS Lab). Dr. Yan received M.S. and
Ph.D. degrees in Computer Science from the College of William and Mary
and worked at Microsoft Research and HP Labs. Dr. Yan's research
bridges the fields of big data, machine learning, and systems. The
focus of his research is on developing methodologies and building
systems that are automated, high-performing, efficient, robust, and
user-centric. Some of his recent research topics include large-scale
distributed deep learning, machine learning as a service (MLaaS),
federated learning, AutoML, serverless computing, and broad topics in
cloud and HPC. Dr. Yan is also dedicated to interdisciplinary research
and has established fruitful collaborations with domain experts in
areas such as health, physics, geography, material science, mechanical
engineering, civil engineering, and innovated big data and AI-driven
approaches for these domains. Dr. Yan and his team are actively
publishing at the most prestigious venues in computer system area
(such as SOSP, SC, HPDC, USENIX ATC, EuroSys, FAST, VLDB, etc.) and
machine learning area (such as NIPS/NeurIPS, KDD, AAAI, etc.). Dr. Yan
and his students are the recipients of the Best Student Paper Award of
IEEE CLOUD 2018, the Best Paper Award of CLOUD 2019, and the Best
Student Paper Award of ITNG 2021. Dr. Yan is the recipient of the NSF
CAREER Award, the NSF CRII Award, the CSE Best Researcher Award, and
has been nominated for the Regents' Rising Researcher Award. Dr. Yan
serves as Social Media Chair of ACM SIGMETRICS. To learn more
information, please visit Dr. Yan's homepage:
https://www.cse.unr.edu/~fyan.
All are welcome to attend BMI grand rounds talk by Dr. Le Lu on 04/14. 

Le Lu, Ph.D 
Executive Director, PAII Inc 
Johns Hopkins University
IEEE Fellow, MICCAI Board Member


Time: Wednesday, April 14, 2021 3:00 pm - 4:00 pm 

Zoom Meeting 
https://stonybrook.zoom.us/j/95617197636?pwd=KytzZ2pVRG9SZGpKZUtpNXJISjNjZz09 
Meeting ID: 956 1719 7636 Passcode: 924293

Title: 
In Search of Effective and Reproducible Clinical Imaging Biomarkers for Population Health and Oncology Applications of Screening, Diagnosis and Prognosis

Bio: 
Le Lu received a PhD in 2007 from Johns Hopkins University. During his first six years at Siemens, he made significant contributions to the company's CT colonography and Lung CAD product lines. From 2013 to 2017, Dr. Lu served as a staff scientist in the Radiology and Imaging Sciences department of the National Institutes of Health Clinical Center. He then went on to found Nvidia's medical image analysis group and he held the position of senior research manager until June 2018. Since then, he has been the Executive Director at PAII Inc., Bethesda Research lab, Maryland, USA which has become one of the leading industrial research labs in medical imaging. He was the main technical leader for two of the most-impactful public radiology image dataset releases (NIH ChestXray14, NIH DeepLesion 2018). He won NIH Clinical Center Director Award in 2017, NIH Mentor of the year award in 2015, and won numerous best paper awards in MICCAI and RSNA from 2016 to 2020 (over 10000 citations). In 2021, He was elected into IEEE Fellow class cited for his contribution to machine learning for cancer detection and diagnosis, and MICCAI society board member (MICCAI-Industry Workgroup Chair). He is currently an Associate Editor for IEEE Trans. Pattern Analysis and Machine Intelligence and IEEE Signal Processing Letters. He has served as an Area Chair for recent MICCAI, AAAI, CVPR, WACV, ICIP and ICHI conferences for 14 times.

Abstract: 
This talk will first give an overall on the work of employing deep learning to permit novel clinical workflows in two population health tasks, namely using conventional ultrasound for liver steatosis screening and quantitative reporting; osteoporosis screening via conventional X-ray imaging and AI readers. These two tasks were generally considered as infeasible tasks for human readers, but as proved by our scientific and clinical studies and peer-reviewed publications, they are suitable for AI readers. AI can be a supplementary and useful tool to assist physicians for cheaper and more convenient/precision patient management. Next, the main part of this talk describes a roadmap on three key problems in pancreatic cancer imaging solution: early screening, precision differential diagnosis, and deep prognosis on patient survival prediction. (1) Based on a new self- learning framework, we train the pancreatic ductal adenocarcinoma (PDAC) segmentation model using a larger quantity of patients (≈1,000, four institutions), with a mix of annotated/unannotated venous or multi-phase CT images. Pseudo annotations are generated by combining two teacher models with different PDAC segmentation specialties on unannotated images, and can be further refined by a teaching assistant model that identifies associated vessels around the pancreas. Our approach makes it technically feasible for robust large-scale PDAC screening from multi-institutional multi-phase partially-annotated CT scans. (2) We propose a holistic segmentation-mesh classification network (SMCN) to provide patient-level diagnosis, by fully utilizing the geometry and location information. SMCN learns the pancreas and mass segmentation task and builds an anatomical correspondence-aware organ mesh model by progressively deforming a pancreas prototype on the raw segmentation mask. Our results are comparable to a multimodality clinical test that combines clinical, imaging, and molecular testing for clinical management of patients with cysts. (3) Accurate preoperative prognosis of resectable PDACs for personalized treatment is highly desired in clinical practice. We present a novel deep neural network for the survival prediction of resectable PDAC patients, 3D Contrast-Enhanced Convolutional Long Short-Term Memory network (CE- ConvLSTM), to derive the tumor attenuation signatures from CE-CT imaging studies. Our framework can significantly improve the prediction performances upon existing state-of-the-art survival analysis methods. This deep tumor signature has evidently added values (as a predictive biomarker) to be combined with the existing clinical staging system.

More information can be found at:
https://bmi.stonybrookmedicine.edu/sites/default/files/Lu_le_04_14.pdf
Zoom Like a Pro! Unlock Whiteboard, Polls, AI Companion, and more to supercharge student participation. This hands-on workshop explores innovative ways to use Zoom's built-in tools to enhance active learning activities in your classes. Learn how to utilize the Whiteboard feature to make collaborative work more engaging, use Polling and Quizzes for instant feedback, AI Companion for summary, and Breakout Sessions for group activities. Register here: https://stonybrook.zoom.us/meeting/register/tJckf--rpj4pGdRV0ItgTW8Lk7gn_RuykByO#/registration
The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 10 minutes) by multiple people. Students can register for 1 credit for CSE 656. Registered students must attend and present a minimum of 2 or 3 talks. Everyone else is welcome to attend. Fill in https://forms.gle/pCVXovgfMfQwGqG38 to subscribe to our mailing list for further announcement.

Get hands-on with data cleaning techniques using Python and AI tools. Join SBU Libraries' Data Literacies Lead, Ahmad Pratama, to learn how to identify and rectify errors, handle missing data, and prepare your dataset for analysis. This workshop introduces you to powerful yet easy-to-use tools and techniques that make data cleaning efficient and effective, turning chaotic data into valuable insights.

Please register for the Data Cleaning with Python and AI here.