Abstract: Recent studies have highlighted the vulnerability of Natural Language Processing (NLP) and Vision-Language Models (VLMs) to backdoor attacks, posing significant security risks. Understanding these attack strategies is crucial for assessing model robustness and developing effective defenses. This thesis proposal aims to investigate the vulnerability of language and vision-language models, analyze abnormal behaviors in backdoor-attacked models, and develop defense methods to enhance safety of modern machine learning models at deployment.


We investigate the internal mechanisms of backdoored NLP models, identifying a distinct attention focus drifting phenomenon, where trigger tokens hijack attention regardless of the input context. Through comprehensive qualitative and quantitative analysis, we provide insights into the underlying mechanisms that enable backdoor attacks. Building on these insights, we propose detection methods to differentiate backdoored models from clean ones, through inspecting both the attention distribution and the model predictions. To better understand the vulnerability, we develop advanced backdoor attack strategies targeting language models in classification tasks. For BERT variants, we introduce Trojan Attention Loss (TAL), a novel method that directly manipulates attention patterns to enhance backdoor effectiveness, ensuring stealth and robustness. Vision-Language Models have demonstrated strong performance in recent years. Yet their vulnerability is largely underexplored. We investigate advanced backdoor attack strategies on Vision-Language Models, focusing on image-to-text generation tasks. We demonstrate how backdoors can be embedded in complex multimodal tasks while maintaining semantic integrity under poisoned inputs. Additionally, we propose innovative techniques for injecting backdoors without requiring access to the original training data, expanding the feasibility of real-world attacks.

This proposal provides novel insights into the internal mechanisms of backdoored models, propose effective detection strategies, and develop advanced attack techniques that expose critical vulnerabilities. These findings underscore the urgent need for robust security measures to defend against emerging backdoor threats in deep learning models. The results have been published in top venues including ICLR, ECCV, NAACL, EMNLP, etc.

Speaker: Weimin Lyu


Zoom link: https://stonybrook.zoom.us/j/99880605139?pwd=cfWbRG6n9v3GXEa7OqvXa5cOp5eLBv.1
Meeting ID: 998 8060 5139
Passcode: 843302
The IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025 will be held from June 11th to June 15th, 2025, at the Music City Center, Nashville, TN. The IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR) is the premier annual computer vision event comprising the main conference and several co-located workshops and short courses. With its high quality and low cost, it provides an exceptional value for students, academics and industry researchers. Register here.
Abstract: Modern decision-making increasingly relies on complex data, imperfect models, and limited domain expertise--yet decisions must still be made with confidence and accountability. This talk presents a research perspective on visual analytics as a bridge between data, models, and human judgment. Through three case studies spanning public-health risk analysis, multivariate scientific visualization, and causal model auditing with large language models, I will show how interactive visualization can reveal structure in high-dimensional data, support reasoning under uncertainty, and help humans critically assess both statistical and AI-generated explanations. Together, these examples illustrate how visual analytics enables users not only to explore data, but to form, challenge, and refine beliefs that underpin scientific and societal decisions.

Bio: Klaus Mueller received his Ph.D. in Computer Science from The Ohio State University in 1998. He is a Professor in the Department of Computer Science at Stony Brook University and a Senior Scientist at the Computational Science Initiative at Brookhaven National Laboratory. He currently serves as the Acting Chair of the Department of Technology and Society at Stony Brook. From 2012 to 2015, he was the Founding Chair of the Computer Science Department at SUNY Korea, where he also served as Vice President for Academic Affairs and Finance for two years.
His research interests span visual analytics, explainable AI, machine learning and data science, human-centered responsible AI, fairness, belief modeling and personalized communication, virtual and augmented reality, and computational and medical imaging. Dr. Mueller received the U.S. National Science Foundation Early Career Award in 2001, the SUNY Chancellor's Award for Excellence in Scholarship and Creative Activity in 2011, and the Meritorious Service Certificate and Golden Core Award of the IEEE Computer Society in 2016. In 2018, he was inducted into the U.S. National Academy of Inventors.
To date, he has authored more than 300 peer-reviewed journal and conference papers, which have been cited over 15,000 times. He is a frequent speaker at international conferences, has organized or participated in 18 tutorials, chaired the IEEE Visualization Conference in 2009, served as elected Chair of the IEEE Technical Committee on Visualization and Computer Graphics (VGTC) from 2012-2015, and was Editor-in-Chief of IEEE Transactions on Visualization and Computer Graphics from 2019-2022. He is a Fellow of the IEEE.

Location: NCS 120
Abstract: Humans perceive the world through global structures such as parts, branches, and their spatial arrangement. Most deep learning models, however, operate mainly at the pixel level. This disconnect between local and global understanding limits interpretability and control. In this thesis, we explore topology as a mathematical framework for bridging local predictions and global structure in dense prediction and generation tasks. We first incorporate topological constraints into semantic segmentation to preserve anatomical relationships and improve multi-class consistency. We next develop structure-level uncertainty estimation, producing more interpretable and actionable measures of model error over branches and connections rather than isolated pixels. Then, we introduce a topology-guided diffusion framework for controllable image generation using structural attributes such as object count and connectivity. Finally, we extend image generation to the longitudinal task, where we aim to capture structural changes across timepoints. All these contributions together establish topology as a unifying interface for building dense prediction models that are structurally aware, interpretable, and controllable.

Speaker: Saumya Gupta

Location: NCS 220

Zoom: https://stonybrook.zoom.us/j/97950688136?pwd=NCa3XOsgIaMIsTVlQBQJ11n27NzL8s.1
Meeting ID: 979 5068 8136
Passcode: 941798

The AI Community will be hosting our very first Datathon๐Ÿ’ก๐Ÿ“Š

Ready to turn data into groundbreaking insights? ๐Ÿง 

Compete in our Datathon, where you'll analyze real-world data ๐Ÿ“ˆ and share innovate solutions in these tracks:

๐Ÿซ Student Life

๐ŸŒฑ Environment & Sustainability

๐Ÿ’‰ Health & Wellness

๐Ÿ’ฐ Finance & Economics

Whether you're a data pro or just starting out, this is your chance to network, learn, and win exciting prizes! ๐Ÿ†๐ŸŽ‰ Bring your creativity ๐Ÿงฉ collaborate with fellow students ๐Ÿง‘โ€๐Ÿคโ€๐Ÿง‘ and gain hands-on experience showcasing your analytical skills ๐Ÿ’ป

Submissions will be judged by professors ๐Ÿง‘โ€๐Ÿซ so take this chance to impress them!

There will be free food โ˜• and games ๐ŸŽฒ to fuel your brain and imagination! Don't miss out--register now and unleash the power of data! ๐Ÿ”ฅโœจ

Registration Form: https://forms.gle/6XYMfmhyAByzFpxz5

Time: Friday (4/4) 10:30am - 5pm โฐ

Location: Bauman Center ๐Ÿ“

The annual conference on Neural Information Processing Systems is a multi-track interdisciplinary annual meeting that includes invited talks, demonstrations, symposia, and oral and poster presentations of refereed papers. Along with the conference is a professional exposition focusing on machine learning in practice, a series of tutorials, and topical workshops that provide a less formal setting for the exchange of ideas.

For more information and registration, visit the official website.

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Machine Learning for Seismic Low Frequency Extrapolation

Abstract: The cycle skipping problem that plagues seismic inversion can be mitigated by utilizing low-frequency seismic data, which captures the kinematics of wave propagation, in conjunction with a reasonable initial velocity model. However, seismic sources and receivers are band-limited and cannot provide signals down to 0 Hz. To improve solution of the seismic inverse problem one can synthesize the missing low-frequency content by solving a regression problem using machine learning (ML). The recorded high-frequency (HF) seismic data is the input and the ML models are trained to predict the missing low-frequency (LF) seismic data. Deep learning models utilizing convolutional neural networks (CNNs) and generative adversarial networks (GANs) demonstrate important capabilities for LF extrapolation. However, such models require powerful hardware and careful training. We explore the feasibility of using less costly ML models such as a random forest, Gaussian process surrogates, and gradient boosting as alternatives to computationally expensive deep learning models.

Biography: Sue Minkoff is Chair of Applied Mathematics at Brookhaven National Laboratory. From 2012-2024 she was a Professor of Mathematical Sciences and an Affiliated Professor in the Departments of Sustainable Earth Systems Sciences and Science and Mathematics Education at the University of Texas at Dallas. From 2000-2012 she served on the faculty in the Department of Mathematics and Statistics at the University of Maryland, Baltimore County. She received her doctorate in Computational and Applied Mathematics from Rice University. From 1995-1997 she was a National Science Foundation-Industrial postdoc joint with the University of Texas at Austin and British Petroleum, and from 1997-2000 she held the von Neumann Fellowship in the Mathematics Department at Sandia National Labs. In 2000 Minkoff was promoted to Senior Member of the Technical Staff in Sandia's Geophysics Department. Minkoff's research interests include scientific computing, inverse problems, uncertainty quantification and digital twins modeling, Earth science, and photonics.

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1606848158?pwd=miUtq7OkYL5SNkjbgVb19teZPNennd.1

Meeting ID: 160 684 8158
Passcode: 068399

Abstract: Visual generation is a fundamental problem in computer vision and graphics, with applications ranging from 3D capture to content creation and image/video synthesis. Despite rapid progress in neural rendering and generative models, efficiency remains a key obstacle in practice: high-quality 3D reconstruction often depends on dense multi-view supervision; scalable 3D synthesis faces heavy optimization, training, and rendering costs; and modern image/video generators incur substantial computation as token grids grow with spatial resolution and temporal length.
This thesis targets efficient visual world modeling by improving sample efficiency in 3D reconstruction, representation efficiency in 3D generation, and computational efficiency in image/video synthesis. First, we improve sample efficiency for neural implicit surface reconstruction under sparse views by integrating multi-view stereo probability volumes as a geometric regularizer, enabling high-quality reconstruction from as few as three input images. Next, we introduce an explicit 3D representation for 3D generation, built from multi-view depth and RGB predictions with 3D Gaussian features, which enables the use of 2D generative priors while enforcing multi-view consistency via epipolar attention. We then address the computational bottleneck of image and video synthesis with importance-based token merging, using importance signals available during generation to preserve critical information while merging redundant tokens. Finally, we propose efficient mixed-resolution diffusion transformers via cross-resolution phase-aligned attention, aiming to improve attention stability under mixed token grids and support high-fidelity mixed-resolution generation.

Speaker: Haoyu Wu

Location: NCS120
Abstract: In today's digital era, language functions not only as a medium of information transmission but also as a mechanism of persuasion, framing, and control. The proliferation of online platforms has amplified this dual role: while enabling unprecedented access to knowledge, it has also exacerbated challenges such as misinformation, rhetorical manipulation, and cultural or linguistic disparities in information access. As a result, pragmatic language understanding and information integrity have emerged as central concerns for both computational linguistics and society at large. This research follows how claims are produced, reframed, and contested online through three interconnected threads. First, it models pragmatic deflection in discourse by investigating whataboutism, a rhetorical device that deflects criticism by redirecting discourse, and introduced novel datasets from Twitter (now X) and YouTube. This work underscores how subtle pragmatic maneuvers can erode discourse integrity without relying on outright falsehoods. Second, it advances retrieval and alignment for information integrity in health and news communication. These systems trace claims and narratives across genres (e.g., social posts and news reports) and languages (Chinese and English), linking social posts with journalistic reporting and aligning Chinese news with English biomedical evidence. By accounting for cultural context, assertions can be linked to reliable evidence and organized for systematic comparison. This work surfaces the risks of missing sources, unverifiable claims, and framing disparities in global health discourse, and demonstrates computational solutions that enhance both the credibility and accessibility of information. Third, the methodological centerpiece is Class Distillation (ClaD), a geometry-aware training paradigm for distilling a small, well-defined target class from a large, heterogeneous background. ClaD couples a distribution-aware contrastive loss (instantiated here in a Mahalanobis form when its assumptions fit the data) with an interpretable decision algorithm tuned for class separation. Evaluated on sarcasm, metaphor, and sexism detection, ClaD delivers strong efficiency and robustness, matching or surpassing larger models while using fewer computational resources, making these pipelines practical by learning reliably from small, sharply defined classes. In sum, this research presents an integrated account of language understanding in the digital age. It exposes how integrity falters through pragmatic deflection, cross-genre drift, and cross-lingual misalignment, and translates these insights to move pragmatic language understanding to systems for evidence retrieval, alignment, and verification; and it sheds light on where and how integrity is threatened, and delivers methods that leverage pragmatic language use.

Speaker: Chenlu Wang

Location: (Old) Computer Science Building, Room 2311