Abstract: Implicit functions have long been a fundamental representation for both 2D and 3D objects in computer graphics, playing a significant role in the field's early development. With the rise of 3D deep learning and the rapid advancement of neural rendering techniques, implicit representations of 3D shapes have regained significant attention in recent years. In this talk, I will present several recent research projects focusing on implicit function-based 3D reconstruction and neural rendering. Furthermore, I will discuss potential future developments in this dynamic and rapidly evolving field.

Biography: Ying He is an Associate Professor at the College of Computing and Data Science, Nanyang Technological University, where he also serves as the Director of the Centre for Augmented and Virtual Reality. His research interests lie in geometric computation and analysis, with applications spanning computer graphics, 3D vision, computer-aided design, multimedia, and wireless sensor networks. Dr. He is an active member of the technical program committees for major conferences on geometric modeling and has served on the editorial boards of IEEE Transactions on Visualization and Computer Graphics, Computer Graphics Forum, and Computational Visual Media. He has also taken on key leadership roles as General/Program Co-Chair for several conferences, including Shape Modeling International (SMI) 2022, Solid and Physical Modeling (SPM) 2022 & 2023, Geometric Modeling and Processing (GMP) 2014 & 2021, and Computational Visual Media (CVM) 2020. For more information, please visit https://personal.ntu.edu.sg/yhe/

Location: NCS 115

The Institute for AI-Driven Discovery and Innovation hosts Dr. Mary
Simoni for a talk on her music and its intersection with AI, as part
of the Music and AI Seminars series.

The event will be held on Thursday, December 10, 2020, at 3:00 PM.

Abstract: Mary Simoni, Dean of Humanities, Arts & Social Sciences at
Rensselaer Polytechnic Institute will discuss her research in the use
of computer algorithms and technology in the composition and
performance of music. The talk will feature compositions inspired by
Augmented Transition Networks (ATNs), employ motion tracking to
control synthesis parameters, and a work in progress that employs
machine learning using training data that juxtaposes classical music
with COVID-19. During this talk, participants will be introduced to
several technologies that support music information retrieval, machine
learning, and algorithmic composition such as jSymbolic, Weka, and
Common Music.

Zoom details below:
https://stonybrook.zoom.us/j/98236706900?pwd=bDFEZFZtaHBWU0cyL0wxK3UrdUpIdz09
Meeting ID: 982 3670 6900
Passcode: 133945  
The Empirical Methods in Natural Language Processing (EMNLP) conference is a premier international academic conference in the field of artificial intelligence and natural language processing (NLP). Organized annually by the Association for Computational Linguistics (ACL) special interest group on linguistic data (SIGDAT), it focuses on research that uses empirical methods to solve language processing problems.

For more information, and registration, visit the official website.
The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the Ph.D. program or (2) receive permission from the instructors. Each seminar will consist of multiple short talks (around 15 minutes) by multiple students. Students can register for 1 credit for CSE656. Registered students must attend and present a minimum of 2 talks. Registered students must attend in person. Up to 3 absences will be excused. Everyone else is welcome to attend. The seminar will be taught by Prof. Chao Chen, chao.chen.1@stonybrook.edu.
Abstract: Generative image models trained on massive datasets encode the statistics of our visual world. An off-the-shelf diffusion or flow-matching model can therefore serve as a general-purpose image prior, providing information about images across a wide range of tasks. Harnessing these capabilities, however, requires combining the model at inference time with additional signals or constraints unknown during training. This thesis will develop training-free methods for inference under a generative denoising prior. I first show how sampling from a trained denoiser can be formulated as an optimization problem, which combined with a constraint, can solve tasks ranging from conditional generation and weakly supervised segmentation to combinatorial optimization. I then demonstrate how the inherent properties of denoisers can accelerate inference under such constraints. Replacing gradient descent with an inexact Newton update, based on the symmetry of the denoiser's Jacobian, substantially reduces inference costs without tradeoffs. I also explore a middle ground between model adaptation and fully training-free inference by using the denoiser's robust internal representations to learn constraints from limited labeled data. These methods are applied to gigapixel image domains such as digital histopathology and remote sensing, where generative models can only be trained at a patch scale. To synthesize arbitrarily large images at resolutions unseen during training, I introduce inference-time algorithms that enforce consistency across spatially overlapping patches and image scales. Finally, I propose repurposing inference-time algorithms from sampling tools, to mechanisms for understanding the prior learned by a denoising model. Extending the previously developed techniques, I analyze the Jacobians of generative denoisers, where preliminary results indicate that Jacobian spectra correlates with generative quality. This motivates a Jacobian-spectrum regularization as a way to improve model performance using insights derived from inference-time algorithms.

Speaker: Alexandros Graikos

Location: NCS 220
Abstract:

Recent advances in deep learning have significantly enhanced the capabilities of Natural Language Processing (NLP) and Vision-Language Models (VLMs). However, these advancements come with increased vulnerabilities, notably through backdoor attacks that pose severe security threats. This thesis addresses two critical dimensions of Trustworthy AI and Efficient Multimodal Representation Learning: (1) security through analyzing, detecting, and designing backdoor attacks in NLP and VLMs, and (2) efficiency through advanced multimodal representation methods tailored for clinical and medical imaging applications.

In the first dimension, we explore the internal mechanisms exploited by backdoor attacks, identifying the distinctive phenomenon of attention focus drifting in compromised transformer models, where trigger tokens consistently hijack attention. Leveraging these insights, we propose robust detection frameworks, including the attention-based Trojan detector (AttenTD) and a task-agnostic logit-based detection method (TABDet), achieving effective identification of backdoored NLP models across diverse tasks. We further introduce novel backdoor attack methodologies: the Trojan Attention Loss (TAL), enhancing attack efficiency and stealth through direct attention manipulation, and BadCLM, demonstrating critical vulnerabilities in clinical decision-support systems by effectively compromising clinical language models.

Extending our security exploration to multimodal settings, we investigate backdoor attacks on Vision-Language Models (VLMs), particularly in complex image-to-text generation tasks, proposing innovative techniques (TrojVLM, VLOOD) capable of embedding backdoors without direct access to original training data, thus showcasing practical risks in real-world scenarios.

In the second dimension, we address efficiency and interpretability challenges in clinical and pathology applications. We introduce TCP-LLaVA, the first multimodal large language model (MLLM) designed explicitly for Whole Slide Image (WSI) Visual Question Answering (VQA). Utilizing a novel token compression mechanism inspired by transformer-based models, TCP-LLaVA substantially reduces computational resource consumption while maintaining superior VQA performance across multiple tumor subtypes. Additionally, we present a multimodal transformer model integrating structured Electronic Health Records (EHR) with clinical notes, demonstrating enhanced predictive accuracy and interpretability for in-hospital mortality prediction through integrated gradient-based interpretability methods.

Together, these contributions present a comprehensive approach to ensuring AI models are not only secure against malicious manipulation but also efficient and interpretable for critical clinical applications, underscoring the essential need for trustworthy and effective AI systems.

Speaker: Weimin Lyu

Zoom: https://stonybrook.zoom.us/j/2392326575?pwd=SVQ2VkFXTnZZYmJUMXgvTXBuZWM3UT09

Meeting ID: 239 232 6575
Passcode: 436192
Abstract: Language is not just something we generate, but something humans use to interact with the world around them. Indeed, today's conversational AI agents speak fluently, but often treat language as prediction rather than interaction, producing responses that sound correct while failing to recognize when requests are ungrounded, impossible, or misunderstood. My research asks what it would take for multimodal agents to take the next step and use language effectively to take actions in grounded contexts and in this talk, I argue that many challenges in multimodal LLM design, including alignment, hallucination, and adaptation, are due to a lack of pragmatics: an understanding of the implicit context behind the implied actions of the words in the query. From visual understanding, to automatic speech recognition, to hallucination detection, I will demonstrate that incorporating pragmatic/contextual reasoning substantially improves agent behavior, and that pragmatic reasoning will drive a necessary shift in how we build multimodal conversational agents that can see, listen, act, and speak in context.


Bio: David M. Chan, Ph.D., is a postdoctoral scholar at the University of California, Berkeley, specializing in multimodal conversational AI. His research focuses on developing scalable AI systems that move beyond language prediction toward grounded language use, integrating vision, audio, and language to enable pragmatic interaction, improve AI-human collaboration, and reduce hallucinations in generative models. Beyond academia, he has worked with leading organizations including Amazon, Google, and NASA, to build and deploy safe, efficient, and accessible machine learning systems. He is also the developer and maintainer of TSNE-CUDA, an open-source tool for high-dimensional data visualization, used by over 40,000 researchers in fields ranging from biomedical technology to industrial manufacturing. David holds a Ph.D. and M.Sc. in Computer Science from UC Berkeley, where he was a graduate fellow with the Center for Technology, Society & Policy (CTSP), and dual B.Sc. degrees in Computer Science and Mathematics from the University of Denver.

Location: NCS 120

Join Stony CELT for a focused, one-hour overview on how to redesign and future-proof assessments in the age of AI. This session will cover three key areas: leveraging AI as a co-pilot for developing effective exam questions, designing authentic assessments, and exploring how AI can strategically support active learning structures like Team-Based Learning (TBL), Project-Based Learning (PBL), and Scenario-Based Learning (SBL).

Register here to attend.

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes three short talks on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Speakers

Kriti Chopra, Computing & Data Sciences (CDS)
Thomas Flynn, Computing & Data Sciences (CDS)
Wenjie Liao, Chemistry Division

Tuesday, January 7, 2025, 12:00 pm -- CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1615289117?pwd=Hqkbj9itxWrFnkhZ8rQXHPInO2gxdF.1

Meeting ID: 161 528 9117
Passcode: 991382

Looking to learn about a new topic or skill? Look no further! Gemini's Guided Learning feature acts as your own personal tutor, teaching you about a particular subject through an engaging back and forth conversation. This AI tool helps users develop their knowledge and skills on a wide variety of topics, acting as a patient mentor, breaking down complex topics step-by-step. This session will take place on 2/24 at 11 AM. Please register using the link below!
https://stonybrookuniversity.co1.qualtrics.com/jfe/form/SV_a9PVlBw0E1Bal1A?