Abstract: Gaussian Probability Path-based Generative Models (GPPGMs) generate data by reversing a stochastic process that progressively corrupts samples with Gaussian noise. Despite state-of-the-art results in 3D molecular generation, their deployment is hindered by the high cost of long generative trajectories, often requiring hundreds to thousands of steps during training and sampling. In this work, we propose a principled method, named GAGA, to improve generation efficiency without sacrificing training granularity or inference fidelity of GPPGMs. Our key insight is that different data modalities obtain sufficient Gaussianity at markedly different steps during the forward process. Based on this observation, we analytically identify a characteristic step at which molecular data attains sufficient Gaussianity, after which the trajectory can be replaced by a closed-form Gaussian approximation. Unlike existing accelerators that coarsen or reformulate trajectories, our approach preserves full-resolution learning dynamics while avoiding redundant transport through truncated distributional states. Experiments on 3D molecular generation benchmarks demonstrate that our GAGA achieves substantial improvement on both generation quality and computational efficiency.

Speaker: Jingxiang Qu

Location: New Computer Science 220

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes three short talks on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Tuesday, January 7, 2025, 12:00 pm -- CDS, Bldg. 725, Training Room


Speakers

Sanket Jantre
Tao Zhang
Xi Yu


Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1615289117?pwd=Hqkbj9itxWrFnkhZ8rQXHPInO2gxdF.1

Meeting ID: 161 528 9117
Passcode: 991382

Abstract: The development of embodied AI has largely focused on scaling data and computational power, often at the cost of energy efficiency. In contrast, biological intelligence achieves remarkable adaptability with minimal resources, inspiring a shift toward neuromorphic AI, an approach that mimics the structure and dynamics of biological neural systems. In this talk, I will explore the promises and challenges of neuromorphic computer vision from three key perspectives: algorithms, robot actions, and data. First, I will discuss algorithmic advances, including continuous visual hull reconstruction, continuous-time human motion field estimation, and unsupervised independent motion segmentation. Next, I will illustrate how neuromorphic vision enables agile robotic actions by leveraging event-based perception for real-time decision-making. Finally, I will address challenges in training data-driven models with event data, highlighting strategies to enhance data availability and efficiency. By integrating these elements, neuromorphic AI paves the way for energy-efficient, high-performance embodied intelligence in dynamic real-world environments.

Speaker Bio: Ziyun (Claude) Wang is a fifth-year Ph.D. student in the General Robotics, Automation, Sensing & Perception (GRASP) Lab at the University of Pennsylvania, advised by Professor Kostas Daniilidis. His research focuses on developing algorithms for neuromorphic computer vision and integrating them with real hardware to enable agile perception in embodied AI systems. Prior to his Ph.D., he worked at the Samsung AI Center New York, where he developed 3D reconstruction techniques for robotic applications and earned three patents. He also contributed to the Apple Vision Pro team, enhancing user comfort for AR glasses. His research work has been recognized at major computer vision, robotics, and machine learning venues including the AAAI Conference on Artificial Intelligence (AAAI), European Conference on Computer Vision (ECCV), International Conference on Learning Representations (ICLR), Conference on Computer Vision and Pattern Recognition (CVPR) workshops, and IEEE Robotics and Automation Letters (R-AL), with an oral presentation at ECCV placing in the top 2.7%. His research aims to drive the development of next-generation bio-inspired AI systems, enabling more efficient, adaptive, and intelligent embodied perception.
Abstract: Visual generation is a fundamental problem in computer vision and graphics, with applications ranging from 3D capture to content creation and image/video synthesis. Despite rapid progress in neural rendering and generative models, efficiency remains a key obstacle in practice: high-quality 3D reconstruction often depends on dense multi-view supervision; scalable 3D synthesis faces heavy optimization, training, and rendering costs; and modern image/video generators incur substantial computation as token grids grow with spatial resolution and temporal length.
This thesis targets efficient visual world modeling by improving sample efficiency in 3D reconstruction, representation efficiency in 3D generation, and computational efficiency in image/video synthesis. First, we improve sample efficiency for neural implicit surface reconstruction under sparse views by integrating multi-view stereo probability volumes as a geometric regularizer, enabling high-quality reconstruction from as few as three input images. Next, we introduce an explicit 3D representation for 3D generation, built from multi-view depth and RGB predictions with 3D Gaussian features, which enables the use of 2D generative priors while enforcing multi-view consistency via epipolar attention. We then address the computational bottleneck of image and video synthesis with importance-based token merging, using importance signals available during generation to preserve critical information while merging redundant tokens. Finally, we propose efficient mixed-resolution diffusion transformers via cross-resolution phase-aligned attention, aiming to improve attention stability under mixed token grids and support high-fidelity mixed-resolution generation.

Speaker: Haoyu Wu

Location: NCS120
Join the Department of Computer Science as we welcome Lyle Ungar, University of Pennsylvania, who will be delivering a lecture on 'Measuring Cultural Variation using Natural Language Processing.' When: 11/08/24 @ 2:30 PM Where: New Computer Science Building, Room 120. Reception to follow. Abstract: Cultures vary widely in how they view the world, for example being more individualist or collectivist. Such cultural differences are, of course, reflected in the words that people use. We first show a variety of ways in which multilingual language models are not multicultural; they speak Hindi or Mandarin, but still think like Americans. In contrast, we then present a scalable method that uses embedding-derived lexica to successfully measure regional variation in culture. Bio: Lyle Ungar is a Professor of Computer and Information Science at the University of Pennsylvania, where he also holds secondary appointments in Psychology, Bioengineering, Genomics and Computational Biology, and Operations, Information and Decisions. His group uses natural language processing and explainable AI for psychological research, including analyzing social media and cell phone sensor data to better understand the drivers of physical and mental well-being. They are currently building socio-emotionally sensitive GPT-based tutors and coaches.

Abstract: Spectroscopy and imaging are two primary tools for probing material structures. However, the discovery of trends that guide the design of improved materials is often hindered by intertwined physical interactions or significant experimental noise. In this talk, I will present machine learning approaches that address both challenges. The first part focuses on the interpretation of X-ray absorption spectroscopy (XAS). We developed a controlled projection algorithm, RankAAE, which disentangles coupled structural descriptors in complex datasets and reveals analysis rules for inferring new structural information visually from spectra. The second part targets transmission electron microscopy (TEM) imaging of material structures. We developed a machine learning model capable of denoising extremely noisy images, while demonstrating strong out-of-distribution generalization. I will describe the construction of these models and demonstrate their effectiveness through representative scientific case studies.

Bio: Dr. Xiaohui Qu is a Staff Scientist in the Theory and Computation Group at the Center for Functional Nanomaterials (CFN), Brookhaven National Laboratory. His research focuses on developing interpretable machine learning and data analytics methods for materials science, with an emphasis on extracting structural insights from X-ray absorption spectroscopy and transmission electron microscopy. Dr. Qu earned his B.S. in Environmental Engineering and Ph.D. in Environmental Science from Shandong University, China, followed by postdoctoral research in Physics at Nanyang Technological University, Singapore, in Chemistry at Universidade Nova de Lisboa, Portugal, and in Materials at Lawrence Berkeley National Laboratory.

Location: IACS Seminar Room


Event Details & Calendar Link (includes zoom info): https://calendar.stonybrook.edu/site/iacs/event/iacs-seminar-speaker--xiaohui-qu-brookhaven-national-lab/