Abstract: At XTX Markets, we view algorithmic trading as one of the most compelling real-world frontiers for deep learning and foundation models. Every day, our systems generate forecasts for tens of thousands of financial instruments and execute over $300B in global trading volume: fully automated, with no discretionary human intervention. This domain combines massive data scale with high noise, adversarial dynamics, and frequent regime shifts, making it both scientifically challenging and commercially impactful. For machine learning researchers, it serves as a rigorous proving ground where advances in time-series modeling, large-scale optimization, representation learning, and foundation models can translate directly into measurable real-world outcomes. This talk will provide a high-level overview of our research agenda, infrastructure, and key open challenges at the intersection of large-scale AI and quantitative finance.

Speaker: Dr. Zhangyang Atlas Wang is the Research Director at XTX Markets, one of the world's leading high-frequency trading firms. He founded and leads the firm's AI Lab in New York City, focused on developing large-scale foundation models for financial time series and market data, powered by XTX's proprietary AI infrastructure. He is currently on leave from his position as the Temple Foundation Endowed Associate Professor at The University of Texas at Austin. His academic research has received numerous awards, and he has mentored a broad network of Ph.D. students and postdoctoral researchers. Many of his alumni now hold tenure-track faculty positions (eight to date) or senior research roles in industry (nineteen and counting). For more information about his group and alumni, please visit: https://www.vita-group.space/team.

Location: NCS 120

Refreshments will be served after the seminar in the first-floor atrium.



Join us as we celebrate this year's Brook & Beyond Challenge finalists.
The Office for Research and Innovation invites you to hear about the two-month journey in which the Brook & Beyond team supported eight cohorts in bringing their bold ideas from the lab to the marketplace. It's an energizing evening that highlights the collaboration, creativity, and entrepreneurial spirit driving discovery across the University.
Meet this year's award recipients, hear pitches from the emerging founders, and applaud their achievements.
Connect, celebrate, and be part of the momentum shaping the future of innovation at
Stony Brook University.
Refreshments will be served. Registration is required.
Register Here.
Virtual Talk: Metadata Matters: Robust Document Classification via Adaptation Methods for Text-driven Public Health by Xiaolei Huang

Zoom link to follow.

Abstract: Document classifiers have been widely applied in solving health-related issues, such as suicide prevention, flu vaccination surveillance and disease diagnosis. However, document metadata including time, gender, age and location has an enormous impact on robustness of 
document classifiers. Language varies across the metadata bringing both challenges and opportunities to build reliable document classifiers. For example, online written language changes over time, and males and females express opinions differently. This talk describes how to use domain adaptation to integrate temporal and user demographic factors into document classifiers. By adapting knowledge of how language varies across the metadata, models can learn generalized representations of language through the metadata-invariant embeddings. 
This approach will lead to metadata-adapted document classifiers and can also extend to personalize classification models by user embedding. 

Bio: Xiaolei Huang is a 4th-year PhD candidate in Information Science at the University of Colorado, Boulder. He is currently a visiting scholar at the Johns Hopkins University. His research interests are in Natural Language Processing, Machine Learning and Public Health. Particularly, he focuses on domain adaptation, cross-lingual transfer learning, user modeling and fairness.
Abstract:

Capturing the spatio-temporal (4D) dynamics of humans has been a long standing research problem in computer vision and graphics. Synthesizing photorealistic human avatars has broad applications, ranging from immersive telepresence in AR/VR and the movie industry, to enriching the education and healthcare systems. Earlier approaches relied on hand-engineered models that use a small amount of data from one or more subjects. With the advent of neural networks, training on large datasets enhanced the output visual quality. Currently, the combination of neural networks with graphics techniques has achieved natural-looking human animation. However, most approaches are identity-specific, trained only on a single identity, and use only one modality.

In this thesis, we address the problem of learning neural representations of humans in a holistic way. Given that the video data in the real world include multiple modalities (audio and video) and multiple identities, we develop multi-modal and multi-identity representations. First, we propose to reconstruct the 4D face geometry of humans by leveraging both audio and video information. In this way, the network produces accurate lip shapes and is robust to cases when either modality is insufficient. Next, we introduce a NeRF-based representation for audio-driven human face animation that achieves high-quality lip synchronization for cinematic content. Since humans communicate with their full body, combining body pose, hand gestures, and facial expressions, we extend our network to capture the full-body human motion for multiple identities simultaneously. In order to better disentangle identity and non-identity specific information, we subsequently study non-linear interactions between latent factors of variation, and propose a specific multiplicative module. In this way, we learn a multi-identity NeRF that robustly animates human faces under novel expressions and achieves a significant decrease in the total training time. Similarly, we propose a multi-identity gaussian splatting representation for human bodies, by constructing a high-order tensor. Assuming a low-rank structure, we learn a tensor decomposition that leads to a significant decrease in the total number of learnable parameters, as well as to a robust animation under novel poses. In the future, we propose to jointly synthesize audio and visual outputs from just text input. Given the recent rise of large language models, coupling text with natural-looking avatars can enhance the overall interaction between a human and an AI system.

Speaker: Aggelina Chatziagapi

Where: NCS, Room 220

Zoom link: https://stonybrook.zoom.us/j/98775312249?pwd=uORNAnSdcssrPZdqOsqaMAF5aLcRD9.1
ID: 98775312249
Passcode: 505777
Hidden Biases. Ethical Issues in NLP, and What to Do about Them presented by Dirk Hovy of Bocconi University

ABSTRACT: Through language, we fundamentally express who we are as humans. This property makes text a fantastic resource for research into the complexity of the human mind, from social sciences to humanities. However, it is exactly that property that also creates some ethical problems. Texts reflect the authors' biases, which get magnified by statistical models. This has unintended consequences for our analysis: If our data is not reflective of the population as a whole, if we do not pay attention to the biases contained, we can easily draw the wrong conclusions, and create disadvantages for our users.

In this talk, I will discuss several types of biases that affect NLP models, their sources, and potential counter measures: (1) Bias stemming from data, i.e., selection bias (if our texts do not adequately reflect the population we want to study), label bias (if the labels we use are skewed) and semantic bias (the latent stereotypes encoded in embeddings); (2) Biases deriving from the models themselves, i.e., their tendency to amplify any imbalances that are present in the data; (3) Design bias, i.e., the biases arising from our (the researchers) decisions which topics to analyze, which data sets to use, and what to do with them. For each bias, I will provide examples and discuss the possible ramifications for a wide range of applications, and various ways to address and counteract these biases, ranging from simple labeling considerations to new types of models.

BIO: Dirk Hovey is an associate professor of Computer Science in the department of marketing at Bocconi University. He received his PhD from the University of Southern California in Los Angeles, where he worked as a research assistant at the Information Sciences Institute. 

He works in Natural Language Processing (NLP), a subfield of artificial intelligence. His research focuses on computational social science. His interests include integrating sociolinguistic knowledge into NLP models, using large-scale statistics to model the interaction between people's socio-demographic profile and their language use, and ethics for data science and algorithmic fairness.

Please join us this Friday, February 13th for the CSE 600 seminar given by Associate Professor Debswapna Bhattacharya, from the Department of Computer Science at Virginia Tech.

Abstract: Building a model of a biological system that can provide actionable hypotheses to form a solid foundation for experimental and theoretical analyses is one of the key challenges in biology and medicine. In this talk, I will present my group's ongoing work in developing, evaluating, and disseminating a new generation of computational methods for biomolecular modeling powered by artificial intelligence (AI) and machine learning (ML). First, I will introduce a new generation of AI/ML methods for improved modeling and characterization of protein-nucleic acid assemblies by deep graph learning using embeddings from biological large language models (LLMs) as well as geometric attention-enabled pairing of heterogeneous biological LLMs, a previously unexplored avenue. Then, I will present a novel generative deep learning model based on equivariant flow matching for end-to-end generation of all-atom RNA 3D structural ensemble. Finally, I will outline my future research directions on attaining atomic-level accuracy in computational modeling of biomolecules and their assemblies at scale.

Speaker: Debswapna Bhattacharya is an Associate Professor in the Department of Computer Science at Virginia Tech. He received his Ph.D. in Computer Science from the University of Missouri-Columbia in 2016. Before joining Virginia Tech in 2022, he was an Assistant Professor at Auburn University from 2017 to 2021. His research interests lie at the intersection of computational biology and machine learning, with a particular focus on artificial intelligence for computational structural biology, specifically in modeling and characterization of biomolecular structures and interactions. His research group has been developing novel computational and data-driven methods, software, and information systems for diverse biomolecular modeling problems, ranking among the best methods in community-wide blind assessments and serving the worldwide community of biomedical users. He received various research awards (NSF CAREER Award, NIH Maximizing Investigators' Research Award, NSF National AI Research Resource Award) and numerous institutional honors (National Distinction and Outstanding Contributor at Virginia Tech, Ginn Faculty Fellowship at Auburn University, Outstanding Engineering Faculty Award at Auburn University).
Location: NCS 120