AI Seminar: Video Architecture Search - Michael Ryoo Abstract: Video understanding is a challenging problem. Because a video contains spatio-temporal data, its feature representation is required to abstract both appearance and motion information. This is not only essential for automated understanding of the semantic content of videos, such as Web-video classification or sport activity recognition, but is also crucial for robot perception and learning. Previously, convolutional neural networks (CNNs) for videos were normally built by manually extending known 2D architectures such as Inception and ResNet to 3D or by carefully designing two-stream CNN architectures that fuse together both appearance and motion information. However, designing an optimal video architecture to best take advantage of spatio-temporal information in videos still remains an open problem. In this talk, we discuss recent progress in neural architecture search for videos, obtaining more optimal network architectures for video understanding.

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes three short talks on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Tuesday, January 7, 2025, 12:00 pm -- CDS, Bldg. 725, Training Room

Speakers

Chuntian Cao, CDS AID - Neural Network Potential (NNP) for Battery Electrolytes

Yeonju Go, NPP Physics - Generative AI for High-Energy Nuclear Physics

Gilchan Park, CDS AID - Graph RAG: Indexing, Retrieval and Generation

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1615289117?pwd=Hqkbj9itxWrFnkhZ8rQXHPInO2gxdF.1

Meeting ID: 161 528 9117
Passcode: 991382

Come learn of the exciting research being done across so many fields using AI! The recipients of AI3's seed awards will present their work in our showcase on November 17, 2025 and we would love to see you there!

The schedule is listed below.

Location: New Computer Science Room 120

Session 1 - 10:30 AM to 11:45

Kevin Reed, PI, Introducing the AI Techniques in Assessing the Future Changes of Extreme Precipitation and Associated Flood Risks
Co-PIs: Tangnyu Song, Ishrat Dollan
Consultant: Jayesh Rathi

Ruwen Qin, PI, AI-Assisted Analysis of Materials in Recycling Streams
Consultant: Vismay Vora

Giuseppe Gazzola, PI, Using AI to Investigate National Literatures: Italy, France, Spain 1733- 1794
Consultant: Jayesh Rathi

Joseph Lemelin, PI, IAE2^3: AI Ecologies
Co-PIs: Katherine Johnston, Aruna Balasubramanian, Matthew Salzano

Niranjan Balasubramanian, Co-PI, Molecular Foundations for Sustainability: Data Analytics for Sustainable Cellulose Scaffolding Modifications to Remediate Diverse Water Contamination Challenges
PI: Benjamin Hsiao, Co-PI: I. V. Ramakrishnan

Owen Rambow, PI,Achieving Common Ground Through Language and Vision in Mixed-Initiative Human-Machine Communication Via zoom
Co-PI Susan Brennan

Session 2 - 12:30 PM to 1:45

Jack McSweeney, PI, Developing Machine Learning Approaches to Classify Internal Waves
Consultant: Vismay Vora

Eric Josephs, PI, Learning Design Rules to Personalize Precision CRISPR Gene Therapies with Interpretable AI
Consultant: Deboparna Banerjee

Shyam Sharma, PI, Fostering Writing-to-Learn Skills through Critical AI Literacy: A Faculty Development and Student Support Program
Co-PIs: Rose Tirotta-Esposito, Christine Fena

Ritwik Banerjee, PI, A Pragmatic Approach to AI for Digital Media Integrity: Combating Complex Misinformation Through Fallacies and Propaganda
Co-PI: Ruobing Li

Ziyu Shu, Co-PI, Novel Clinical Applications of Deep Image Prior-based CT Image Reconstruction
PI: Xin Qian, Co-PIs: Tiezhi Zhang, Zhaozheng Yin

Prateek Prasanna, Co-PI, An Artificial Intelligence-Driven Clinical Decision Support Tool for the Management of Abdominal Aortic Aneurysm
PI: Apostolos Tassiopoulos, Co-PI's: Mary Saltz, Janos Hajagos, Tahsin Kurc



Presenters will give a 5-minute talk with 2 minutes for Q & A.

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

We meet every other Tuesday at noon in CDSD's Training Room (building 725, room 2-124) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

In addition to our speaker, we will have a number of CDS staff in attendance with expertise in AI methods and applications including image analysis, foundation models development, and inverse problem solving.

AI-Driven Physics-Informed Phase Retrieval from a Single X-ray

Abstract: X-ray phase-contrast imaging enables the visualization of weakly absorbing or low-contrast structures and plays an important role in materials, biological, and energy research. Conventional X-ray holography and phase-retrieval techniques typically require multiple intensity measurements acquired at different propagation distances to recover phase information, increasing acquisition time, radiation dose, and experimental complexity. In this work, we present an AI-driven, physics-informed approach for phase retrieval using only a single X-ray intensity measurement. The method adapted a generative neural network as an inverse reconstruction engine, with physical models of X-ray wave propagation embedded directly into the optimization process. This allows phase and absorption information to be recovered from a single hologram without relying on paired, unpaired, or simulated training datasets. By combining physical constraints with self-supervised AI reconstruction, the approach achieves stable and quantitative results across a wide range of imaging conditions. The results demonstrate how physics-informed AI can reduce experimental requirements and enable data-efficient, automated phase retrieval for next-generation X-ray imaging workflows.

Biography: Xiaogang Yang is a computational scientist in the Data Analysis & Workflow Integration group at NSLS-II, focusing on AI development for X-ray imaging, data analysis, and automated workflows. He earned his PhD from Delft University of Technology, completed his postdoctoral research at Argonne National Laboratory, and previously held a tenured position at PETRA III (DESY).

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1604383624?pwd=ffQ5cUPNxTI7nzClKQO6cnsNbhF9Vf.1

Meeting ID: 160 438 3624
Passcode: 558449

Please Note: Due to a funding shortfall, we are for the time being no longer able to provide pizza and sodas for these events. We will have coffee though, and all are of course welcome to bring their lunch.

Abstract: The capacity to adapt machine learning models to various contexts, information, and objectives is particularly valuable. In this thesis, I focus on developing Class Conditional Guided Models. These are models that can be adaptively biased towards a class of interest via a conditional input. My primary focus lies in the efficiency of these models. They are constructed to require training only once, with the ability to quickly and conveniently adapt during testing time without necessitating fine-tuning or retraining.
Firstly, I propose RelationVAE, a novel generative model designed for few-shot scenarios, utilizing the prior knowledge of class similarity relationships. RelationVAE is designed to condition on the embeddings of the neighbor classes (i.e. classes with similarity relationships), to generate more reliable samples by making them more similar to the neighbor class. This enables adaptation of the generative model to the provided prior knowledge about class relationships.
As a second focus, I introduce scGAN, a shadow segmentation technique that enables adaptation to varying shadow distributions in different testing environments. scGAN is designed to condition on a sensitivity parameter, a scalar, to control the amount of the shadow detected. In the testing phase, the parameter is set to appropriate values, allowing the model to quickly adapt to specific test environments.
In my third contribution, I propose S-SEG, a methodology for fine-grained counting allowing adaptation to different granularities of fine-grained classes. In fine-grained problems, the distinction between classes is subtle and inconsistent across images, leading to variations in the granularity of the target class from one image to another. S-SEG is designed to be conditioned on an additional input, the sensitivity parameter, to control the granularities of the target class during inference.
My fourth contribution is a text-to-image synthesis method which allows controlling the number of the generated objects of a target class. I propose to generate an intermediate condition, the density map, which reflects the number of objects, together with their layout. This intermediate condition is used to effectively guide the generative model to generate objects with accurate counts.

Speaker: Vu Nguyen

Zoom: https://stonybrook.zoom.us/j/97114455337?pwd=Z4rB9dWcstlahUIs8PRrvQ9b2ZK2Df.1
Meeting ID: 971 1445 5337
Passcode: 272300

Time: Jan 26, 2021 03:00 PM Eastern Time (US and Canada)

All are welcome!

Zoom Meeting:
https://stonybrook.zoom.us/j/93818552212?pwd=ajZkT2x4a2tiaDJUL1h3VFhLZEgwQT09

Meeting ID: 938 1855 2212
Passcode: 802722

Title: Data-Driven Document Unwarping

Abstract: Capturing document images is a common way to digitize and record physical documents due to the ubiquitousness of mobile cameras. To make text recognition easier, it is often desirable to digitally flatten a document image when the physical document sheet is folded or curved. However, unwarping a document from a single image in natural scenes is very challenging due to the complexity of document sheet deformation, document texture, and environmental conditions. Previous model-driven approaches struggle with inefficiency and limited generalizability. In this thesis, I investigate several data-driven approaches to tackle the document unwarping problem.

Data acquisition is the central challenge in data-driven methods. I first design an efficient data synthesis pipeline based on 2D image warping and train DocUNet, the pioneering data-driven document unwarping model, on the synthetic data. A benchmark dataset is also created to facilitate comprehensive evaluation and comparison. To improve the unwarping performance by training on more realistic data, I introduce the Doc3D dataset and DewarpNet. Supervised by 3D shape ground truth in Doc3D, DewarpNet is significantly better than DocUNet. DocUNet and DewarpNet depend on the synthetic data for the ground truth deformation annotation. To exploit the real-world images, I propose PaperEdge, a weakly supervised model trained with in-the-wild document images with easy-to-obtain boundary information. PaperEdge surpasses DewarpNet by utilizing both the synthetic data and weakly annotated real data in the Document In the Wild (DIW) dataset. Finally, I propose to incorporate the 3D physical constraints in training DewarpNet and PaperEdge. The constraints regulate the possible deformations on document papers. I also propose to augment the Doc3D and DIW dataset by introducing an online document segmentation model and better hardware.
Abstract: Foundation models brought a paradigm shift on representation learning and the deep learning community. In my talk, I will examine the role of foundation models in medical imaging, focusing on their potential to unify diverse tasks through large-scale, generalist architectures. While these models achieve strong performance, their deployment in healthcare raises challenges related to data limitations, privacy, validation, and trust. We will also discuss domain-specific models for imaging, along with efficient adaptation techniques to adapt such models on domains that they have not been trained on. The presentation will also address key issues of reliability and interpretability, highlighting approaches like conformal prediction and counterfactual intervention to improve uncertainty estimation and model transparency. Overall, the talk will emphasize that despite their promise, foundation models require robust evaluation and trustworthy design to ensure safe and effective use in clinical settings.

Speaker: Maria Vakalopoulou is an assistant professor (MCF) in applied mathematics at CentraleSupelec, University Paris Saclay in France and the group leader of the biomathematics group of MICS Laboratory focusing on mathematical modeling in Life Sciences. She is affliated with Inria Saclay in France and Archimedes Unit in Greece. Her main research interest include the development of computational methods for image perception focusing on earth observation and medical applications. Before that, she was a postdoctoral student at CentraleSupelec, where she worked with Nikos Paragios. She completed her PhD at the Remote Sensing Laboratory at the School of Rural, Surveying and Geo-Informatics Engineering of the National Technical University of Athens under the supervision of Konstantinos Karantzalos.

Location: NCS 220

Abstract: Much like other AI for Science domains, polymer design poses significant challenges. It requires grounding in empirical data and physical laws, precise handling of domain-specific structured representations, and compositional reasoning over multiple interacting constraints--all while working with limited data.

To address these limitations, we introduce PolyBench, a large-scale benchmark comprising over 125K polymer design and analysis tasks grounded in verified experimental and synthetic data. PolyBench includes tasks created from a wide range of data sources and presents diverse structural, property-driven, and synthesis-oriented reasoning problems. Tasks in PolyBench are organized from simple to complex analytical reasoning problems, enabling generalization tests and includes diagnostic probes to evaluate model capabilities. Additionally, to support effective domain alignment, we propose a knowledge-augmented reasoning distillation framework that enriches the dataset with structured chain-of-thought supervision derived from expert-informed reasoning strategies.

Small language models (7B-14B parameters) trained on PolyBench substantially outperform comparably sized baselines and, in many cases, exceed the performance of larger closed-source frontier models on polymer reasoning tasks, while also demonstrating improved transfer to external polymer benchmarks. Last, we conduct a diagnostic study that reveals a compositionality gap: despite strong performance on decomposed sub-questions, models struggle to integrate multiple interacting constraints and intermediate reasoning steps, highlighting fundamental limitations in current scientific language models.

Speaker: Dikshya Mohanty

Location: NCS 115/Online

Zoom: https://stonybrook.zoom.us/j/94746001760?pwd=BCAd8gu7cXLn3PXM6kkbh11V6r0Mr7.1
Meeting ID: 947 4600 1760 Passcode: 987917

Predicting Subjective Attributes in Visual Data - Zijun Wei

ABSTRACT: Recent progress in deep neural networks has revolutionized many computer vision tasks such as image classification, detection and segmentation. However, in addition to excelling in tasks that predict well-defined objective information, human-centered artificial intelligence systems should also be able to model subjective attributes, as defined by human perceptual behavior, that goes beyond the pure physical content of visual data. Example subjective tasks are the prediction of spatial or temporal regions that are interesting to humans (e.g., attract attention or are visually pleasing) and the recognition of subjective attributes (e.g., visually elicited sentiments). Better models for these tasks will improve the human-computer interaction experience in various applications. This thesis investigates several approaches to address the challenges in predicting those subjective attributes in visual data over a diverse set of tasks. I first present a novel framework for real-time automatic photo composition. The framework consists of a cost-effective data collection workflow, an efficient model training pipeline and a lightweight module to account for personalized preferences. Then I develop a novel and general algorithm to detect interesting segments in sequential data, which can be naturally applied to video summarization tasks. Furthermore, I propose methods that learn to represent sentiments elicited by images, in an unsupervised manner, using linguistic features extracted from large scale Web data. To conclude this thesis, I introduce a human-vision-inspired image classification algorithm that also predicts spatial visual attention even though no attention data was used for training it.