In this dissertation, we address the problem of learning neural representations of humans in a holistic way. Given that the video data in the real world include multiple modalities (e.g., audio and video) and multiple identities, we develop multi-modal and multi-identity representations. First, we propose to reconstruct the 4D face geometry of humans by leveraging both audio and video information. In this way, the network produces accurate lip shapes and is robust to cases when either modality is insufficient. Next, we introduce a NeRF-based representation for audio-driven human face animation that achieves high-quality lip synchronization for cinematic content. Since humans communicate with their full body, combining body pose, hand gestures, and facial expressions, we extend the network to capture full-body human motion for multiple identities simultaneously. In order to better disentangle identity and non-identity specific information, we subsequently study non-linear interactions between latent factors of variation, and propose a specific multiplicative module. In this way, we learn a multi-identity NeRF that robustly animates human faces under novel expressions and achieves a significant decrease in the total training time. Similarly, we propose a multi-identity Gaussian splatting representation for human bodies, by constructing a high-order tensor. Assuming a low-rank structure, we learn a tensor decomposition that leads to a significant decrease in the total number of learnable parameters, as well as to a robust animation under novel poses. Last but not least, we propose to jointly synthesize audio and visual outputs from just text input. Given the recent rise of large language models, coupling text with natural-looking avatars can enhance the overall interaction between a human and an AI system.
Location: NCS 220 or Zoom
passcode: 045476
The University's Main Commencement Ceremony will take place on Friday, May 23, 2025 at 11 am at Kenneth P. LaValle Stadium. Gates open at 10 am.
All guests need a valid ticket to enter LaValle Stadium - no exceptions. Children age 1 and older require a ticket. Seating is first-come, first-served.
Register here.
Abstract: Materials used in extreme environments, such as high temperatures, irradiation, and stress, often fail due to rapid defect generation and microstructural evolution, and traditional approaches cannot explore the vast design space needed for next-generation alloys. I will present a machine learning framework powered by massive computing that links individual atomic motion to microstructural evolution. Neural network kinetics models trained on first-principles data map vacancy barrier spectra and capture correlated diffusion in multicomponent alloys, revealing design strategies to suppress radiation damage. At larger scales, simulations uncover dislocation patterning and distinguish between confined and extended slip bands, offering new insight into collective dislocation motion and deformation instabilities. By integrating AI-driven modeling, large-scale computing, and experimental validation, my research goal is to accelerate the discovery of damage-tolerant materials and advance fundamental understanding of defect physics in extreme environments.
Speaker Bio: Penghui Cao is an Associate Professor in Mechanical and Aerospace Engineering at the University of California, Irvine, with a joint appointment in Materials Science and Engineering. He received his PhD in mechanical engineering from Boston University and subsequently worked as a Postdoctoral Associate in the Department of Nuclear Science and Engineering at the Massachusetts Institute of Technology from 2014 to 2018. Dr. Cao's research focuses on understanding the fundamental mechanisms that govern radiation responses and microstructure evolution in materials, and on developing advanced alloys for high-performance nuclear energy systems. His lab advances computational and modeling algorithms, integrates advanced manufacturing techniques to tailor microstructures, and leverages state-of-the-art electron microscopy to characterize and assess underlying mechanisms. He is the recipient of the DOE Early Career Research Program Award and the UCI Samueli School's Mid-Career Award for Faculty Excellence in Research.
Location: Institute for Advanced Computational Science, Seminar Room
*This seminar will be held in-person and online. Zoom link below*
Join Zoom Meeting: https://stonybrook.zoom.us/j/96410717491?pwd=3WGMwbLYNMSbI2IF160VXkvv2JmCQ1.1
Meeting ID: 964 1071 7491
Passcode: 399333
October 19 - 20: Workshops
October 21 - 23: Main Conference
More information can be found here.
Description:
As artificial intelligence and data science reshape the global information landscape, libraries are emerging as key players in both technological innovation and ethical stewardship. This international Zoom discussion brings together library professionals and educators from the U.S., Philippines, and Hong Kong to explore how institutions are integrating AI and data into their pedagogy and services.
Panelists will share concrete examples from their own libraries--ranging from data literacy initiatives to increasing discoverability. The conversation will also examine regional trends in librarianship, spotlighting how institutions in Asia are navigating the evolving role of data and AI.
Join us for a global conversation that highlights the transformative potential of libraries as hubs for innovation and critical inquiry in the age of AI.
Register for this free Zoom panel.
Panelists:
Ahmad Pratama is a Faculty Member and Associate Librarian at Stony Brook University Libraries, where he is working to build a comprehensive, campus-wide data literacy program within the Libraries. As the Data Literacies Lead, his work focuses on empowering students, faculty, and staff to critically and ethically engage with data and AI, including the development of a credit-bearing course in Critical Data & AI Literacies supported by an EDGE Fund Award from the Provost's Office. Previously, Dr. Pratama served as an Associate Professor of Information Technology, and his research and teaching explore the intersections of technology, policy, and society with a focus on data, AI, and innovation in higher education.
Dan Anthony Dorado is a full-time faculty member at the U.P. School of Library and Information Studies, where he teaches information technology, management and marketing, research methodology, and quantitative research. He was also the director of the Diliman Learning Resource Center under the Office of the Vice Chancellor for Student Affairs. Before that, he was an Information Specialist at the College of Engineering Library, in charge of the System and Network Administration and The Learning Commons. He completed his master's degree at the Technology Management Center in U.P. Diliman and is currently pursuing his PhD in Data Science. As a member of Sync.Bio.Optics laboratory and the Publics, Archives, and Data (PANDA) Lab, his research specialization covers Computational Methods, Open Education, Critical Data Studies, and Radical Statistics.
Ryun LEE is Associate University Librarian at The Chinese University of Hong Kong Library, leading Digital Initiatives and Library IT and Systems. He drives digital innovation through emerging technologies, particularly artificial intelligence to enhance services, streamline operations, and support CUHK's mission in research, education, and knowledge advancement. With a background in cataloging and digital repository development, Ryun leads projects in digitization, OCR, data visualization, text and network analysis, GIS, and digital scholarship. He actively promotes knowledge graph applications in Hong Kong studies and oversees efforts to digitize and preserve resources related to Hong Kong and Southern China. His recent work focuses on creating seamless digital experiences and developing data-driven infrastructure. He is currently exploring AI-driven approaches to digitization workflows and entity extraction, aiming to improve access, discovery, and long-term preservation of library materials.
Abstract: Yes, scalable quantum computing should actually work! Sooner than many expect, which will create a huge headache when it breaks the encryption currently used to protect the Internet. But no, we don't think quantum computing can do most of what the popular articles promise in AI and optimization and so forth. Come to this talk to learn about why!
Speaker: Scott Aaronson is Schlumberger Chair of Computer Science at the University of Texas at Austin, and founding director of its Quantum Information Center. He received his bachelor's from Cornell University and his PhD from UC Berkeley. Aaronson's research has focused mainly on the capabilities and limits of quantum computers. His first book, Quantum Computing Since Democritus, was published in 2013 by Cambridge University Press. He received the National Science Foundation's Alan T. Waterman Award, the United States PECASE Award, the Tomassoni-Chisesi Prize in Physics, and the ACM Prize in Computing, and is a Fellow of the ACM and the AAAS and a member of the National Academy of Sciences. He blogs at Shtetl-Optimized, https://www.scottaaronson.com/blog.
Location: Della Pietra Family Auditorium (SCGP 103)
Zoom link to follow.
Abstract: Document classifiers have been widely applied in solving health-related issues, such as suicide prevention, flu vaccination surveillance and disease diagnosis. However, document metadata including time, gender, age and location has an enormous impact on robustness of
document classifiers. Language varies across the metadata bringing both challenges and opportunities to build reliable document classifiers. For example, online written language changes over time, and males and females express opinions differently. This talk describes how to use domain adaptation to integrate temporal and user demographic factors into document classifiers. By adapting knowledge of how language varies across the metadata, models can learn generalized representations of language through the metadata-invariant embeddings.
This approach will lead to metadata-adapted document classifiers and can also extend to personalize classification models by user embedding.
Bio: Xiaolei Huang is a 4th-year PhD candidate in Information Science at the University of Colorado, Boulder. He is currently a visiting scholar at the Johns Hopkins University. His research interests are in Natural Language Processing, Machine Learning and Public Health. Particularly, he focuses on domain adaptation, cross-lingual transfer learning, user modeling and fairness.