Each seminar consists of multiple short talks (around 15 minutes) by several students.
Join Zoom Meeting:
https://stonybrook.zoom.us/j/93547152068?pwd=WVpoRVgzelBXeloxdXVEakNSb2M5UT09
Meeting ID: 935 4715 2068 | Passcode: 481832
https://stonybrook.zoom.us/j/
Meeting ID: 980 7952 6509
Passcode: 949941
supervised learning algorithms. However, in stark contrast to young animals (including humans), training such
networks requires enormous numbers of labeled examples, leading to the belief that animals must rely instead
mainly on unsupervised learning. The reason is that most animal behavior is not the result of clever learning
algorithms--supervised or unsupervised-- but is encoded in the genome. Specifically, animals are born with
highly structured brain connectivity, which enables them to learn very rapidly. Because the wiring diagram is
far too complex to be specified explicitly in the genome, it must be compressed through a genomic
bottleneck. I will describe results showing how the genomic bottleneck algorithm can lead to dramatic
compression of networks and better generalization, particularly for rapid transfer in supervised and
reinforcement learning.
Anthony Zador is professor of neuroscience at CSHL.
ABSTRACT: A common choice of architecture in representation learning (i.e., learning a good embedding of the data) is an encoder/decoder architecture, which tries to map a part of the input into a good latent representation (via an encoder), and predict the remaining part of the input (via a decoder). Two common examples are universal machine translation: where one tries to learn to translate between any pair of a set of languages via a common latent language, given paired up corpora for only a part of the pairs; and contextual encoders -- where one tries to predict a part of the image, given the rest of the image.
We will give a framework for analyzing the sample complexity of such architectures -- i.e., how many pairs of languages do we need to have paired up corpora for? How many image prediction tasks do we have to solve to get a good representation?
Recent advances in deep learning have significantly enhanced the capabilities of Natural Language Processing (NLP) and Vision-Language Models (VLMs). However, these advancements come with increased vulnerabilities, notably through backdoor attacks that pose severe security threats. This thesis addresses two critical dimensions of Trustworthy AI and Efficient Multimodal Representation Learning: (1) security through analyzing, detecting, and designing backdoor attacks in NLP and VLMs, and (2) efficiency through advanced multimodal representation methods tailored for clinical and medical imaging applications.
In the first dimension, we explore the internal mechanisms exploited by backdoor attacks, identifying the distinctive phenomenon of attention focus drifting in compromised transformer models, where trigger tokens consistently hijack attention. Leveraging these insights, we propose robust detection frameworks, including the attention-based Trojan detector (AttenTD) and a task-agnostic logit-based detection method (TABDet), achieving effective identification of backdoored NLP models across diverse tasks. We further introduce novel backdoor attack methodologies: the Trojan Attention Loss (TAL), enhancing attack efficiency and stealth through direct attention manipulation, and BadCLM, demonstrating critical vulnerabilities in clinical decision-support systems by effectively compromising clinical language models.
Extending our security exploration to multimodal settings, we investigate backdoor attacks on Vision-Language Models (VLMs), particularly in complex image-to-text generation tasks, proposing innovative techniques (TrojVLM, VLOOD) capable of embedding backdoors without direct access to original training data, thus showcasing practical risks in real-world scenarios.
In the second dimension, we address efficiency and interpretability challenges in clinical and pathology applications. We introduce TCP-LLaVA, the first multimodal large language model (MLLM) designed explicitly for Whole Slide Image (WSI) Visual Question Answering (VQA). Utilizing a novel token compression mechanism inspired by transformer-based models, TCP-LLaVA substantially reduces computational resource consumption while maintaining superior VQA performance across multiple tumor subtypes. Additionally, we present a multimodal transformer model integrating structured Electronic Health Records (EHR) with clinical notes, demonstrating enhanced predictive accuracy and interpretability for in-hospital mortality prediction through integrated gradient-based interpretability methods.
Together, these contributions present a comprehensive approach to ensuring AI models are not only secure against malicious manipulation but also efficient and interpretable for critical clinical applications, underscoring the essential need for trustworthy and effective AI systems.
Speaker: Weimin Lyu
Zoom: https://stonybrook.zoom.us/j/
Meeting ID: 239 232 6575
Passcode: 436192
The program will fund projects for up to a one-year period, depending on the availability of funds. AI^3 anticipates making at least six awards on this call. A one-year, no-cost extension can be requested in the final 6 months of a project with approval subject to progress towards project goals and active participation in research themes.
Competitive applications will actively incorporate modern AI technologies into the work; integrate students; document significant potential for future funding or other growth-oriented outcomes; and highlight innovations.
The 2024 application deadline will be October 15, at 11:59 PM EST. Recipients will be notified by December 20, and projects are anticipated to commence at the start of the Spring 2025 semester.
The seminar will be jointly taught by Prof. Dimitris Samaras (samaras@cs.stonybrook.edu).
The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision.
To enroll in this course, you must either: (1) be in the Ph.D. program or (2) receive permission from the instructors.
Each seminar will consist of multiple short talks (around 15 minutes) by multiple students. Students can register for 1 credit for CSE656. Registered students must attend and present a minimum of 2 talks. Registered students must attend in person. Up to 3 absences will be excused. Everyone else is welcome to attend.
Please note: Exceptionally, the first meeting on 1/28 will be in NCS 120.
You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.
We meet every other Tuesday at noon in CDSD's Training Room (building 725, room 2-124) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.
Abstract: Identifying model Hamiltonians is a vital step toward creating predictive models of materials. We combined Bayesian optimization with the EDRIXS numerical package to infer Hamiltonian parameters from resonant inelastic X-ray scattering (RIXS) spectra within the single atom approximation. To evaluate the efficacy of our method, we tested it on experimental RIXS spectra for several materials and demonstrated that it can reproduce results obtained from hand-fitted parameters to a precision similar to expert human analysis while providing a more systematic mapping of parameter space. Our work provides a key first step toward solving the inverse scattering problem to extract effective multi- orbital models from information-dense RIXS measurements, which can be applied to a host of quantum materials.
Biography: Marton Lajer is a postdoctoral researcher at the Condensed Matter Physics and Materials Science Department, Brookhaven National Laboratory. Marton obtained his PhD in theoretical physics at the Eotvos Lorand University, Hungary, in 2021. He was a junior research fellow at the Wigner Research Centre for Physics in Budapest before joining BNL in September 2022. His background spans various analytical and performance-critical numerical methods, mostly in the context of low- dimensional quantum field theories and quantum many-body systems. His research currently focuses on incorporating AI-enhanced methods to various problems in inelastic spectroscopy.
In addition to our speaker, we will have a number of CDS staff in attendance with expertise in AI methods and applications including image analysis, foundation models development, and inverse problem solving.
Location: CDS, Bldg. 725, Training Room
Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1604383624?pwd=ffQ5cUPNxTI7nzClKQO6cnsNbhF9Vf.1
Meeting ID: 160 438 3624
Passcode: 558449
Please Note: Due to a funding shortfall, we are for the time being no longer able to provide pizza and sodas for these events. We will have coffee though, and all are of course welcome to bring their lunch.