Abstract: Generative visual models like Stable Diffusion and Sora generate photorealistic images and videos that are nearly indistinguishable from real ones to a naive observer. However, their grasp of the physical world remains an open question: Do they understand 3D geometry, light, and object interactions, or are they mere pixel parrots of their training data? Through systematic probing, I will demonstrate that these models surprisingly learn fundamental scene properties--intrinsic images such as surface normals, depth, albedo, and shading (à la Barrow & Tenenbaum, 1978)--without explicit supervision, which enables applications like image relighting. But I will also show that this knowledge is insufficient. Careful analysis reveals unexpected failures: inconsistent shadows, multiple vanishing points, and scenes that defy basic physics. All these findings suggest these models excel at local texture synthesis but struggle with global reasoning: a crucial gap between imitation and true understanding. I will then conclude by outlining a path toward generative world models that emulate global and counterfactual reasoning, causality, and physics.
Bio: Anand Bhattad is a Research Assistant Professor at the Toyota Technological Institute at Chicago. He earned his PhD from the University of Illinois Urbana-Champaign in 2024 under the mentorship of David Forsyth. His research interests lie at the intersection of computer vision and computer graphics, with a current focus on understanding the knowledge encoded in generative models. Anand has received Outstanding Reviewer honors at ICCV 2023 and CVPR 2021, and his CVPR 2022 paper was nominated for a Best Paper Award. He actively contributes to the research community by leading workshops at CVPR and ECCV, including Scholars and Big Models: How Can Academics Adapt? (CVPR 2023), CV 20/20: A Retrospective Vision (CVPR 2024), Knowledge in Generative Models (ECCV 2024), and How to Stand Out in the Crowd? (CVPR 2025). For more details, visit https://anandbhattad.github.
IACS Student Seminar Speaker: Xiangyan Yang, Dept. of Applied Math & Statistics
Location: IACS Seminar Room or Zoom
Join Zoom Meeting: https://stonybrook.zoom.us/j/91650247483?pwd=fvAGEwadplJh7jFC5RWcdvZ5NWPJth.1
Meeting ID: 916 5024 7483
Passcode: 631055
Capturing the spatio-temporal (4D) dynamics of humans has been a long standing research problem in computer vision and graphics. Synthesizing photorealistic human avatars has broad applications, ranging from immersive telepresence in AR/VR and the movie industry, to enriching the education and healthcare systems. Earlier approaches relied on hand-engineered models that use a small amount of data from one or more subjects. With the advent of neural networks, training on large datasets enhanced the output visual quality. Currently, the combination of neural networks with graphics techniques has achieved natural-looking human animation. However, most approaches are identity-specific, trained only on a single identity, and use only one modality.
In this thesis, we address the problem of learning neural representations of humans in a holistic way. Given that the video data in the real world include multiple modalities (audio and video) and multiple identities, we develop multi-modal and multi-identity representations. First, we propose to reconstruct the 4D face geometry of humans by leveraging both audio and video information. In this way, the network produces accurate lip shapes and is robust to cases when either modality is insufficient. Next, we introduce a NeRF-based representation for audio-driven human face animation that achieves high-quality lip synchronization for cinematic content. Since humans communicate with their full body, combining body pose, hand gestures, and facial expressions, we extend our network to capture the full-body human motion for multiple identities simultaneously. In order to better disentangle identity and non-identity specific information, we subsequently study non-linear interactions between latent factors of variation, and propose a specific multiplicative module. In this way, we learn a multi-identity NeRF that robustly animates human faces under novel expressions and achieves a significant decrease in the total training time. Similarly, we propose a multi-identity gaussian splatting representation for human bodies, by constructing a high-order tensor. Assuming a low-rank structure, we learn a tensor decomposition that leads to a significant decrease in the total number of learnable parameters, as well as to a robust animation under novel poses. In the future, we propose to jointly synthesize audio and visual outputs from just text input. Given the recent rise of large language models, coupling text with natural-looking avatars can enhance the overall interaction between a human and an AI system.
Speaker: Aggelina Chatziagapi
Where: NCS, Room 220
Zoom link: https://stonybrook.zoom.
ID: 98775312249
Passcode: 505777
In this talk I explore an intermediate regime, which is hybrid reduced order models: fast simplified physics approximations where some of the unknown or approximated equations are replaced with data-driven machine learning components. Examples include coarse-grained models where the full macroscopic equations cannot be derived from first-principles microscopic equations, multiscale models with unknown closure terms or sub-grid parameterization schemes, and low-order or latent dynamical systems that learn governing equations on a low-dimensional reduced state space. I discuss how such reduced systems can be identified from very limited data, much less than is often needed in traditional machine learning but at much lower time-to-solution than traditional numerical modeling. This facilitates not only system design and control but also uncertainty quantification approaches that search the space of possible equations for predictive models that can explain the data. I will focus on an example from materials science concerning the design of self-assembling block copolymer nanomaterials.
Speaker: Dr. Nathan Urban, Applied Mathematics Department, Brookhaven National Laboratory
Location: Laufer 101
Zoom: https://stonybrook.zoom.us/j/96090260834?pwd=mw8QTHbMOw9oeU9hazZeoq8bN4VIfH.1
Meeting ID: 960 9026 0834 Passcode: 374969
The University's Main Commencement Ceremony will take place on Friday, May 23, 2025 at 11 am at Kenneth P. LaValle Stadium. Gates open at 10 am.
All guests need a valid ticket to enter LaValle Stadium - no exceptions. Children age 1 and older require a ticket. Seating is first-come, first-served.
Register here.
Abstract: Recent progress in Large Language Models (LLMs) has transformed text and code generation, yet models still falter on scientific reasoning where correctness, constraints, and physical consequences are critical. This talk explores how formal LLM reasoning can advance symbolic scientific modeling. First, our PDE-Controller formalizes informal PDEs (Partial Differential Equations), synthesizes solver-ready code, and plans subgoals to tackle nonconvex control via interactions with external solvers. Second, our Lean Finder accelerates scientific formalization via a semantics-aware search engine for Lean/Mathlib that retrieves relevant theorems, outperforming GPT models and gaining significant traction in the AI-for-math community. Through these efforts, we aim to design a semantics-first LLM that autoformalizes informal scientific problems into machine-checked specifications and synthesizes solver-ready code. This closes the loop between formal analysis and LLM reasoning, ultimately surpassing human heuristics for scientific discovery.
Bio: Dr. Wuyang Chen is a tenure-track Assistant Professor in Computing Science at Simon Fraser University. He is also a visiting research scientist at Microsoft. Previously, he was a postdoctoral researcher in Statistics at the University of California, Berkeley, advised by Professor Michael Mahoney. He obtained his Ph.D. in Electrical and Computer Engineering from the University of Texas at Austin in 2023, advised by Professor Atlas Wang. Dr. Chen's research focuses on integrating AI methods with physical knowledge, scientific machine learning, and theoretical understanding of deep networks. Dr. Chen has published papers at CVPR, ECCV, ICLR, ICML, NeurIPS, and other top conferences. Dr. Chen's research has been recognized by the US NSF newsletter, two Doctoral Dissertation Awards from INNS and iSchools, AAAI New Faculty Highlights, and NVIDIA Academic Grant Award. Dr. Chen also hosted and co-organized many conference workshops at NeurIPS, ICLR, CVPR.
Location: NCS 120
📍 Location: Frey 102
📅 Date: Monday, Nov 11
⏰ Time: 12 PM - 1:50 PM
Scan the QR code or register in the link.
Abstract: This talk shows how machine learning can address challenges in Astrophysics. We specifically focus on black hole simulations and supernova observations. First, we present a super-resolution technique for black hole simulations that avoids the need for high-resolution labels by leveraging the Hamiltonian and momentum constraints from general relativity. This method reduces constraint violations by one to two orders of magnitude. Next, we introduce Maven, a multimodal foundation model for supernova science. Using contrastive learning to align photometric and spectroscopic data, Maven achieves state-of-the-art results in classification and redshift estimation by pre-training on synthetic data and fine-tuning on real observations.
Bio: Thomas Helfer is a computational physicist specializing in deep learning and physics. Currently based at the Institute for Advanced Computational Science at Stony Brook University, Thomas was previously a postdoctoral fellow at Johns Hopkins and did his PhD with Eugene Lim at King's College in London. In his work, he looks to bridge topics; in his PhD, he bridged theoretical particle physics and gravitational waves. Now, in his postdoctoral work, he aims to find novel applications of deep learning in astrophysics.
*please note: this seminar will be held in a hybrid format*
Location: IACS Seminar Room OR Join Zoom Meeting
https://stonybrook.zoom.us/j/
Meeting ID: 986 1763 0652
Passcode: 882994