Abstract: AI has achieved remarkable advancements in image recognition and natural language processing. However, its applications in Earth and environmental sciences are still emerging. Unprecedented data from satellites, sensors, and in-situ measurements oIers new opportunities to improve physics-based models and forecasts of environmental systems with AI and to gain deeper insights into these phenomena. Extreme systems, such as weather and climate events, pose distinct challenges for AI, such as limited sampling of rare events, non-trivial data augmentation, errors-in-variables, and complexities of transfer learning across diverse tasks. In this talk, we will explore some of these challenges and showcase AI architectures designed to address them. We will use specific examples of forecasting dust storms, precipitation extremes, flash floods, and drought events in the Middle East. Finally, we will discuss a different AI approach for studying sinkhole formation in the Dead Sea.

Speaker: Prof. Yinon Rudich, Department of Earth and Planetary Sciences, Weizmann Institute, Israel


Join Zoom Meeting
ID: 98731258879
Passcode: cJjGQJqP
AI Seminar: Video Architecture Search - Michael Ryoo Abstract: Video understanding is a challenging problem. Because a video contains spatio-temporal data, its feature representation is required to abstract both appearance and motion information. This is not only essential for automated understanding of the semantic content of videos, such as Web-video classification or sport activity recognition, but is also crucial for robot perception and learning. Previously, convolutional neural networks (CNNs) for videos were normally built by manually extending known 2D architectures such as Inception and ResNet to 3D or by carefully designing two-stream CNN architectures that fuse together both appearance and motion information. However, designing an optimal video architecture to best take advantage of spatio-temporal information in videos still remains an open problem. In this talk, we discuss recent progress in neural architecture search for videos, obtaining more optimal network architectures for video understanding.
DeepMath Conference on the Mathematical Theory of Deep Neural Networks Recent advances in deep neural networks (DNNs), combined with open, easily-accessible implementations, have made DNNs a powerful, versatile method used widely in both machine learning and neuroscience. These advances in practical results, however, have far outpaced a formal understanding of these networks and their training. The dearth of rigorous analysis for these techniques limits their usefulness in addressing scientific questions and, more broadly, hinders systematic design of the next generation of networks. Recently, long-past-due theoretical results have begun to emerge from researchers in a number of fields. The purpose of this conference is to give visibility to these results, and those that will follow in their wake, to shed light on the properties of large, adaptive, distributed learning architectures, and to revolutionize our understanding of these systems.​​​
Abstract:

Photorealistic editing of human facial expressions and head articulations remains a long-standing topic in the computer graphics and computer vision community. Methods enabling such control have great potential in AR/VR applications where a 3D immersive experience is valuable, especially when this control extends to novel views of the scene in which the human subject appears. Traditionally, 3D Morphable Face Models (3DMMs) have been used to control the facial expressions and head pose of a human head. However, the PCA-based shape and expression spaces of 3DMMs lack the expressivity. They cannot model essential elements of the human head such as hair, skin details, and accessories such as glasses that are paramount for realistic reanimation. In this thesis, we present a set of methods that enables facial reanimation, starting from editing expressions in still face images to creating fully controllable neural 3D portraits with control over facial expressions, head pose, and viewing direction of the scene using only casually captured monocular videos from a smartphone to finally achieving studio-like quality from the said monocular captures.
First, we propose a method for editing facial expressions in near-frontal facial images through the unsupervised disentangling of expression-induced deformations and texture changes. Next, we extend facial expression editing to human subjects in 3D scenes. We represent the scene and the subject in it using a semantically guided neural field. This enables control over the subject's facial expressions and the viewing direction of the scene they're in. We then present a method that learns, in an unsupervised manner, to deform static 3D neural fields using facial expression and head-pose dependent deformations, enabling control over facial expressions and head pose of the subject along with the viewing direction of the 3D scene they're in. Next, we propose a method that makes the learning of the aforementioned deformation field robust to strong illumination effects, which adversely impact the registration of the deformation. We then propose an extension of this unsupervised deformation model to 3D Gaussian splatting by constraining it using a 3D morphable model, resulting in a rendering speed of 18 FPS--a 100x speed improvement over prior work. Finally, we propose a method that bridges the quality gap between 3D portraits created using in-the-wild monocular data and multi-view studio capture data. We accomplish this using a two-stage method. First, we train a StyleGAN to relight and inpaint in-the-wild face texture maps (with strong illumination effects and incompletely captured regions). Next, we both reconstruct and generate identity-specific facial details that may be poorly captured in the in-the-wild captures. Once trained, we can generate studio-like complete avatars from monocular phone captures.

Speaker: Shahrukh Athar

Zoom Link:
https://stonybrook.zoom.us/j/94228500743?pwd=RqOBgG6tbJkKaFBlWFwBkYFX0VRovV.1

Meeting ID: 94228500743
Passcode: 661599
IACS Research Theme: Human Centered Computing Seminar

Abstract: The AI art platform Artbreeder hosts daily remix parties where users build on each other's work, creating transparent evolutionary chains of images from a single seed. This study analyzes 130,882 images from 368 remix parties to identify the drivers of novelty, complexity, and competitive success. The results reveal an interesting tension: while more novel parent images produce more novel and complex children and attract more likes, users paradoxically prefer to remix images that are less novel and complex. At the group level, larger remix parties produce more novelty at the cost of lower complexity. Additionally, images tend to converge towards common thematic attractors (e.g., steampunk scenes, alien architecture, furries) over the course of remix parties. These results provide quantitative insights into collective creativity--the production of novelty by groups of people--a typically opaque aspect of human cultural evolution.

Speaker: Dr. Mason Youngblood

Location: Institute for Advanced Computational Science, Seminar Room
Abstract: Recent studies have highlighted the vulnerability of Natural Language Processing (NLP) and Vision-Language Models (VLMs) to backdoor attacks, posing significant security risks. Understanding these attack strategies is crucial for assessing model robustness and developing effective defenses. This thesis proposal aims to investigate the vulnerability of language and vision-language models, analyze abnormal behaviors in backdoor-attacked models, and develop defense methods to enhance safety of modern machine learning models at deployment.


We investigate the internal mechanisms of backdoored NLP models, identifying a distinct attention focus drifting phenomenon, where trigger tokens hijack attention regardless of the input context. Through comprehensive qualitative and quantitative analysis, we provide insights into the underlying mechanisms that enable backdoor attacks. Building on these insights, we propose detection methods to differentiate backdoored models from clean ones, through inspecting both the attention distribution and the model predictions. To better understand the vulnerability, we develop advanced backdoor attack strategies targeting language models in classification tasks. For BERT variants, we introduce Trojan Attention Loss (TAL), a novel method that directly manipulates attention patterns to enhance backdoor effectiveness, ensuring stealth and robustness. Vision-Language Models have demonstrated strong performance in recent years. Yet their vulnerability is largely underexplored. We investigate advanced backdoor attack strategies on Vision-Language Models, focusing on image-to-text generation tasks. We demonstrate how backdoors can be embedded in complex multimodal tasks while maintaining semantic integrity under poisoned inputs. Additionally, we propose innovative techniques for injecting backdoors without requiring access to the original training data, expanding the feasibility of real-world attacks.

This proposal provides novel insights into the internal mechanisms of backdoored models, propose effective detection strategies, and develop advanced attack techniques that expose critical vulnerabilities. These findings underscore the urgent need for robust security measures to defend against emerging backdoor threats in deep learning models. The results have been published in top venues including ICLR, ECCV, NAACL, EMNLP, etc.

Speaker: Weimin Lyu


Zoom link: https://stonybrook.zoom.us/j/99880605139?pwd=cfWbRG6n9v3GXEa7OqvXa5cOp5eLBv.1
Meeting ID: 998 8060 5139
Passcode: 843302