Abstract: Capturing the spatio-temporal (4D) dynamics of humans has been a long standing research problem in computer vision and graphics. Synthesizing photorealistic human avatars has broad applications, ranging from immersive telepresence in AR/VR and the movie industry, to enriching the education and healthcare systems. Earlier approaches relied on hand-engineered models that use a small amount of data from one or more subjects. With the advent of neural networks, training on large datasets enhanced the output visual quality. Currently, the combination of neural networks with graphics techniques has achieved natural-looking human animation. However, most approaches are identity-specific, trained only on a single identity, and use only one modality.

In this dissertation, we address the problem of learning neural representations of humans in a holistic way. Given that the video data in the real world include multiple modalities (e.g., audio and video) and multiple identities, we develop multi-modal and multi-identity representations. First, we propose to reconstruct the 4D face geometry of humans by leveraging both audio and video information. In this way, the network produces accurate lip shapes and is robust to cases when either modality is insufficient. Next, we introduce a NeRF-based representation for audio-driven human face animation that achieves high-quality lip synchronization for cinematic content. Since humans communicate with their full body, combining body pose, hand gestures, and facial expressions, we extend the network to capture full-body human motion for multiple identities simultaneously. In order to better disentangle identity and non-identity specific information, we subsequently study non-linear interactions between latent factors of variation, and propose a specific multiplicative module. In this way, we learn a multi-identity NeRF that robustly animates human faces under novel expressions and achieves a significant decrease in the total training time. Similarly, we propose a multi-identity Gaussian splatting representation for human bodies, by constructing a high-order tensor. Assuming a low-rank structure, we learn a tensor decomposition that leads to a significant decrease in the total number of learnable parameters, as well as to a robust animation under novel poses. Last but not least, we propose to jointly synthesize audio and visual outputs from just text input. Given the recent rise of large language models, coupling text with natural-looking avatars can enhance the overall interaction between a human and an AI system.

Location: NCS 220 or Zoom

Simons Laufer Mathematical Sciences Institute presents...

In 2023, Tudor Achim co-founded Harmonic with Vlad Tenev to build the world's most advanced reasoning engine. Combining formal verification with informal reasoning, Harmonic's formal reasoning model, Aristotle, achieved gold-medal-equivalent performance on the 2025 International Mathematical Olympiad problems. Aristotle integrates three main components: a Lean proof search system, an informal reasoning system that generates and formalizes lemmas, and a dedicated geometry solver.

Achim is also the Co-Founder and former CTO of Helm.ai. He holds a B.S. in Computer Science from Carnegie Mellon University and was a PhD Candidate in Computer Science at Stanford University.

Register here: https://slmath.us10.list-manage.com/track/click?u=d58ee2e82c69809ff037f56b2&id=f07a675f6f&e=f1b6ba91e6

CSE 600 Seminar Series | Fall 2025


Abstract: The first part of the presentation focuses on the fundamental role that failures play in the Ph.D. journey, highlighting how they offer invaluable learning experiences to build resilience, critical thinking, and adaptability. Instead of viewing failures as signs of inadequacy, they should be recognized as opportunities to learn, re-evaluate, and develop the persistence needed for success in a high-stakes research environment. In the second part of the presentation, we take a quick look at the evolution of distributed databases research at Stony Brook and then focus on different challenges associated with distributed transaction processing systems functioning in untrustworthy environments. Byzantine Fault-Tolerant (BFT) protocols have recently been extensively used by distributed transaction processing systems to establish consensus on the order of transactions. However, the proliferation of different BFT protocols has made it difficult to navigate the BFT landscape, let alone determine the protocol that best meets application needs. Moreover, as novel applications, modern hardware, and new cloud platforms arise, distributed transaction processing systems need to be designed with full-stack adaptivity in mind. This presentation discusses our vision for a reinforcement learning (RL)-based distributed transaction processing system that adjusts effectively in real time to dynamic fault scenarios and evolving workloads.

Bio: Mohammad Javad Amiri is an Assistant Professor in the Department of Computer Science at Stony Brook University. Before joining Stony Brook, he was a postdoctoral researcher in the Computer and Information Science Department at the University of Pennsylvania. He received his Ph.D. in Computer Science from the University of California, Santa Barbara. His research mainly lies at the intersection of data management and distributed systems, focusing on distributed transaction processing, consensus protocols, and blockchains.




https://stonybrook.zoom.us/j/91775729097pwd=Qlc5Nks0NmlyKzJwMjR0S0hrdVZ3QT09

Meeting ID: 917 7572 9097
Passcode: 555459


Abstract: As the saying goes, there are many ways to skin a cat.
While we don't want to go around skinning cats, the world of
optimization is rich with different problems, problem formulations,
and methods and approaches, each with different guarantees and
computational benefits. In this talk we will take a tour down the
problem of structured sparsity in sensing to see how one simple
problem can inspire a wide range of analysis and tools. First, I will
present the optimality conditions for a generalized structured sparse
problem, which can be geometrically visualized as alignment of vectors
and matrices. Then I will introduce three approximation methods for
the problem of phase retrieval, which are a twist on stochastic
gradient and coordinate descent methods. These methods leverage
fundamental numerical linear algebra concepts to give fast approximate
solutions to large-scale problems, which then after postprocessing can
produce more reliable sensing results.

Bio: Yifan Sun received her PhD in Electrical Engineering from the
University of California Los Angeles in 2015, with research focusing
on convex optimization and semidefinite programming. She was then
Technicolor Research and Innovation, focusing on machine learning and
data science applications. More recently, she completed two postdocs,
at the University of British Columbia in Vancouver, Canada and
L'Institut National de Recherche en Informatique et Automatique
(INRIA) in Paris, France.

Please join us this Friday, February 13th for the CSE 600 seminar given by Associate Professor Debswapna Bhattacharya, from the Department of Computer Science at Virginia Tech.

Abstract: Building a model of a biological system that can provide actionable hypotheses to form a solid foundation for experimental and theoretical analyses is one of the key challenges in biology and medicine. In this talk, I will present my group's ongoing work in developing, evaluating, and disseminating a new generation of computational methods for biomolecular modeling powered by artificial intelligence (AI) and machine learning (ML). First, I will introduce a new generation of AI/ML methods for improved modeling and characterization of protein-nucleic acid assemblies by deep graph learning using embeddings from biological large language models (LLMs) as well as geometric attention-enabled pairing of heterogeneous biological LLMs, a previously unexplored avenue. Then, I will present a novel generative deep learning model based on equivariant flow matching for end-to-end generation of all-atom RNA 3D structural ensemble. Finally, I will outline my future research directions on attaining atomic-level accuracy in computational modeling of biomolecules and their assemblies at scale.

Speaker: Debswapna Bhattacharya is an Associate Professor in the Department of Computer Science at Virginia Tech. He received his Ph.D. in Computer Science from the University of Missouri-Columbia in 2016. Before joining Virginia Tech in 2022, he was an Assistant Professor at Auburn University from 2017 to 2021. His research interests lie at the intersection of computational biology and machine learning, with a particular focus on artificial intelligence for computational structural biology, specifically in modeling and characterization of biomolecular structures and interactions. His research group has been developing novel computational and data-driven methods, software, and information systems for diverse biomolecular modeling problems, ranking among the best methods in community-wide blind assessments and serving the worldwide community of biomedical users. He received various research awards (NSF CAREER Award, NIH Maximizing Investigators' Research Award, NSF National AI Research Resource Award) and numerous institutional honors (National Distinction and Outstanding Contributor at Virginia Tech, Ginn Faculty Fellowship at Auburn University, Outstanding Engineering Faculty Award at Auburn University).
Location: NCS 120
The INS (International Neuroethics Society) AI and Consciousness Affinity Group is hosting a talk titled Bringing Trustworthiness in Generative AI and Agentic AI Using Thought Knowledge Graphs featuring speaker Manas Gaur, a computer scientist at UMBC.
The talk will examine the interplay between Thought Knowledge Graphs (TKGs) and how they can form more trustworthy and reasoning-based responses in AI. They will also discuss introducing novel methods on implementing TKGs and their overall impact on creating more trustworthy AI systems.
The talk will be held online via Zoom on Monday, December 2 at 1:00pm (EST).
Register to attend.
Looking to learn about a new topic or skill? Look no further! Gemini's Guided Learning feature acts as your own personal tutor, teaching you about a particular subject through an engaging back and forth conversation. This AI tool helps users develop their knowledge and skills on a wide variety of topics, acting as a patient mentor, breaking down complex topics step-by-step. This session will take place on 2/24 at 11 AM. Please register using the link below!
https://stonybrookuniversity.co1.qualtrics.com/jfe/form/SV_a9PVlBw0E1Bal1A?