Over the past decade, Artificial Intelligence (AI) has made stunning advances, from mastering language to solving the structure of proteins. These breakthroughs arise from more than forty years of work in neural networks, where ideas from neuroscience have inspired solutions in AI. In this lecture, Anthony Zador, MD, PhD, will explore how reverse engineering the brain's computations has driven progress in both fields, and how this back-and-forth between neuroscience and AI is set to grow even stronger -- with brain-inspired designs driving new AI advances while AI tools transform our understanding of how the brain works.

Speaker:
Dr. Zador works at the intersection of neuroscience and artificial intelligence. He is the Alle Davis Harris Professor of Biology at Cold Spring Harbor Laboratory, where he served as Chair of Neuroscience. He was named one of Foreign Policy's 100 Leading Global Thinkers and is a recipient of the Brain Research Foundation Fellowship, the Gill Symposium Transformative Investigator Award, and the Allen Distinguished Investigator Award.

Watch online at stonybrook.edu/live

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes three short talks on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Tuesday, November 26, 2024, 12:00 pm -- CDS, Bldg. 725, Training Room

Speakers

Hanfei Yan, NSLS-II

David Park, CDS, AI Dept

Xihaier Luo, CDS, AI Dept

Join Zoom Meeting

https://bnl.zoomgov.com/j/1601052863?pwd=eIX9qZKPGNtQ11uwbK8JP5hIdIxA3V.1

Meeting ID: 160 105 2863

Passcode: 442980


Nam Nguyen

4-5pm, Dec 17 2020

https://stonybrook.zoom.us/j/94214254415?pwd=K1VoQml4cFdlVW51VW41dWtid2tJdz09



The molecular mechanisms and functions in complex biological systems
currently remain elusive. Recent high-throughput techniques, such as
next-generation sequencing, have generated a wide variety of
multiomics datasets that enable the identification of biological
functions and mechanisms via multiple facets. However, integrating
these large-scale multiomics data and discovering functional insights
are, nevertheless, challenging tasks. To address these challenges,
machine learning has been broadly applied to analyze multiomics. In
particular, multiview learning is more effective than previous
integrative methods for learning data's heterogeneity and revealing
cross-talk patterns. Although it has been applied to various contexts,
such as computer vision and speech recognition, multiview learning has
not yet been widely applied to biological data--specifically,
multiomics data. Therefore, we have developed a framework called
multiview empirical risk minimization (MV-ERM) for unifying multiview
learning methods (Nguyen, et al., PLoS Computational Biology, 2020).
MV-ERM enables potential applications to understand multiomics
including genomics, transcriptomics, and epigenomics, in an aim to
discover the functional and mechanistic interpretations across omics.
Based on MV-ERM, we have developed the following methods:
ManiNetCluster, Varmole and ECMarker.



(1) ManiNetCluster (Nguyen, et al., BMC Genomics, 2019) is a manifold
learning method which simultaneously aligns and clusters gene networks
(e.g., co-expression) to systematically reveal the links of genomic
function between different phenotypes. Specifically, ManiNetCluster
employs manifold alignment to uncover and match local and non-linear
structures among networks, and identifies cross-network functional
links. We demonstrated that ManiNetCluster better aligns the
orthologous genes from their developmental expression profiles across
model organisms than state-of-the-art methods. This indicates the
potential non-linear interactions of evolutionarily conserved genes
across species in development. Furthermore, we applied ManiNetCluster
to time series transcriptome data measured in the green alga
Chlamydomonas reinhardtii to discover the genomic functions linking
various metabolic processes between the light and dark periods of a
diurnally cycling culture;



(2) Varmole (Nguyen, et al., Bioinformatics, 2020) is an interpretable
deep learning method that simultaneously reveals genomic functions and
mechanisms while predicting phenotype from genotype. In particular,
Varmole embeds multi-omic networks into a deep neural network
architecture and prioritizes variants, genes and regulatory linkages
via biological drop-connect without needing prior feature selections.
With an application to schizophonia, we demonstrate that Varmole
provides an effective alternative for recent statistical methods that
associate functional omic data (e.g. gene expression) with genotype
and phenotype and that link variants to individual genes in population
studies such as genome-wide association study;



(3) ECMarker (Jin*, Nguyen*, et al., Bioinformatics, 2020) is an
interpretable and scalable machine learning model that predicts gene
expression biomarkers for disease phenotypes and simultaneously
reveals underlying regulatory mechanisms. Particularly, ECMarker is
built on the integration of semi- and discriminative- restricted
Boltzmann machines, a neural network model for classification allowing
lateral connections at the input gene layer. With application to the
gene expression data of non-small cell lung cancer (NSCLC) patients,
we found that ECMarker not only achieved a relatively high accuracy
for predicting cancer stages but also identified the biomarker genes
and gene networks implying the regulatory mechanisms in lung cancer
development.



Finally, we propose a novel multiview learning method, Malignomics, to
predict phenotypes from heterogeneous multi-omic features. Malignomics
will first align multi-omic features by deep manifold alignment onto a
common latent space, better predicting nonlinear relationships across
omics. This deep alignment aims to preserve both global consistency
and local smoothness across omics and reveal higher-order nonlinear
interactions (i.e., manifolds) among cross-omic features. Second, it
uses these manifold structures to regularize the classifiers for
predicting phenotypes. This manifold-regularization allows
highlighting cross-omic feature manifolds and prioritizing the
features and interactions for the phenotypes. The prioritized
multi-omic features will further reveal underlying phenotypic
functions and mechanisms and thus enhance the biological
interpretation of Malignomics. We will apply Malignomics to
multi-omics data in neuropsychiatric disorders, and prioritize gene
regulatory networks linking risk variants, regulatory elements, and
genes for the disorders. We will also compare Malignomics with the
state-of-the-arts, and investigate how the manifold regulation will
potentially improve understanding of multi-omics functions and
predicting diseases.

The AI Community at Stony Brook University is proud to announce Datathon 2026.

Dive into data analysis and AI/ML, and get ready to build something big. In this year's underwater-themed event, enjoy a weekend of data analysis, hacking, networking, fun activities, and minigames.

Whether you're a seasoned developer, data scientist, designer, or completely new to hacking, this event is your chance to collaborate, learn data science, and create something impactful with data and AI/ML.

What is Datathon?

AI Community's Datathon is the premier data science competition at Stony Brook University, bringing together students of all skill levels for a weekend of data exploration, analysis, and innovation. Just like a typical hackathon, you will be using your skills to build your dream project.

Unlike a regular hackathon, Datathon is focused on data science. You will be given a set of data to work with, analyze, and apply to your project. You can also find your own data to use. Your project will be presented to a panel of judges consisting of professors and industry professionals!

Who Can Participate

  • Students of all skill levels and majors are welcome.
  • Come with a team or find one at the event or on Discord.
  • This event is open to SBU and non-SBU students.

* Non-SBU Undergraduate Students are ineligible to receive prizes
* You must be 18+ or older (Excludes minors who are active SBU students)

Location: SAC Ballroom B

Register here.

Virtual Talk: Contextual Modeling for Natural Language Understanding, Generation and Grounding by Rui Zhang

Zoom link to come.

Abstract: Natural language is a fundamental form of information and communication. In both human-human and human-computer communication, people reason about the context of text and world state to understand language and produce language response. In this talk, I present 
several deep-neural-network-based systems that first understand the meaning of language grounded in various contexts where the language is used, and then generate effective language responses in different forms for information access and human-computer communication. First, 
I will introduce Speaker Interaction RNNs for addressee and response selection in multi-party conversations based on explicit representations for different discourse participants. Then, I will 
present a text summarization approach for generating email subject lines by optimizing quality scores in a reinforcement learning framework. Finally, I will show an editing-based multi-turn SQL query generation system towards intelligent natural language interfaces to databases. 

Bio: Rui Zhang is a final-year PhD student at Yale University advised by Professor Dragomir Radev. His research interest lies in various natural language processing problems in understanding, generation, and grounding. He has been working on (1) End-to-End Neural Modeling for Entities, Sentences, Documents and Multi-party Multi-turn Dialogues, (2) Text Summarization for Emails, News and Scientific Articles, (3) Cross-lingual Information Retrieval for Low-Resource Languages, (4) Context-Dependent Text-to-SQL Semantic Parsing in Human-Computer Interaction. Rui Zhang has published papers and served as Program Committee members at top-tier NLP and AI conferences including ACL, NAACL, EMNLP, AAAI and CoNLL. During his PhD, he has done research internships at IBM Thomas J. Watson Research Center, Grammarly Research and Google AI. He was a graduate student at the University of Michigan and got his Bachelor's degrees at both the University of Michigan and Shanghai Jiao Tong University from the UM-SJTU Joint Institute.
Educational objectives:

1. Explain how AI represents a Cognitive Revolution in academic medicine, redefining thefundamental limits of human cognition and knowledge work.
2. Differentiate between superficial AI adoption (innovation theatre) and truetransformationthrough AI-native institutional design.
3. Analyze how AI fundamentally reshapes clinical, research, and educational work-from dataentry to verification, recall to recognition, and hypothesis generation to evaluation-andidentify implications for redesigning academic health systems.

Speaker: Jiajie Zhang, Ph.D.,Dean, Professor, and Glassell Family FoundationDistinguished Chair in Informatics Excellence,D. Bradley McWilliams School of BiomedicalInformatics, UTHealth Houston

Location: MART Building, Room: 7M-0602 (7th Floor)
Optimization and Machine Learning - presented by Yifan Sun

Abstract: Optimization is a growing topic of interest in the machine learning community. It starts out as an option to check in Tensorflow (SGD? Adam? Adagrad?), but as we get more into the how and why of these options, we uncover many fundamental principles relating to operations research, control theory, and dynamical systems, dating back as far as the Cold World era. 

In this talk I will give a broad overview of some of the important optimization themes in machine learning. I will try to give connections between tools we are used to seeing in popular packages 
and fundamental optimization concepts like duality, convexity, contractive operators, etc. While we cannot hope to completely cover this diverse research area, I hope to provide a glimpse of this exciting research area that is permeating more and more into the machine learning world. 

Bio: Yifan Sun received her PhD in Electrical Engineering from the University of California Los Angeles in 2015, with research focusing on convex optimization and semidefinite programming. She was then Technicolor Research and Innovation, focusing on machine learning and 
data science applications. More recently, she completed two postdocs focusing on optimization, at the University of British Columbia in Vancouver, Canada and INRIA, in Paris, France.

Abstract: Pretraining vision encoders with self-supervision (SSL) leads to stronger representations that excel across diverse downstream tasks. One of the key factors enabling self-supervision is extracting multiple views of the same scene to formulate either: 1) View-invariant pretraining (DINO, SimCLR, iBOT), where the objective is predicting the same representation for different views of the scene; or 2) Cross-view pretraining (cross-view Masked Autoencoders), where the objective is predicting missing parts of one view using other views. For extracting multiple views, view-invariant methods rely on a combination of handcrafted augmentations (random cropping, color jittering, gaussian blur, etc.) of the same image, whereas cross-view pretraining methods rely on image cropping or video frames. In this work, we present methods to effectively incorporate synthetic views from diffusion models into SSL training.
For view-invariant pretraining, we introduce Gen-SIS, a method that leverages the ability of diffusion models to generate interpolated images through interpolation in conditioning space. We introduce a disentanglement pretext task: disentangling two source images from an interpolated synthetic image. This disentanglement task, in addition to vanilla single-source generative augmentation for view extraction, improves visual pretraining of various view-invariant methods (DINO, SimCLR, iBOT).
For cross-view pretraining, we introduce CDG-MAE, a novel cross-view masked autoencoder (MAE) based method that uses diverse synthetic views generated from static images via an image-conditioned diffusion model to learn dense correspondences. We present a quantitative method to evaluate the local and global consistency of the generated views to choose the right diffusion model for cross-view pretraining. These generated views exhibit substantial changes in pose and perspective, providing a rich training signal that overcomes the limitations of video (expensive) and crop-based (less variation) methods. CDG-MAE substantially narrows the gap to video-based MAE methods on video label propagation tasks while maintaining the data advantages of image-only MAEs.

Speaker: Varun Belagali

Location: NCS 120
Zoom: https://stonybrook.zoom.us/j/93647452432?pwd=hZaX7LXCAD8KPHWYE1Afw2sDI3owpv.1

Google Docs has been updated to integration Gemini AI. Learn about the integrations that gives you more features within Google Docs. There are also AI Automations that will keep you updated when things happen with documents.

In this session, you will:

  1. Learn about the AI features within Google Docs
  2. Learn about workflow to be notified of changes to Docs
  3. Understand the features that allow AI to work

Register here.