Tuesday, January 28, 2:00 PM to 3:00 PM
In-person: New Computer Science, Seminar Room 120
Zoom link: https://stonybrook.zoom.us/j/91265872116?pwd=gaIUbmJavuafURujQjEa6AdVMs4d54.1
Meeting ID: 912 6587 2116
Passcode: 189681
Zoom link to follow.
Abstract: Document classifiers have been widely applied in solving health-related issues, such as suicide prevention, flu vaccination surveillance and disease diagnosis. However, document metadata including time, gender, age and location has an enormous impact on robustness of
document classifiers. Language varies across the metadata bringing both challenges and opportunities to build reliable document classifiers. For example, online written language changes over time, and males and females express opinions differently. This talk describes how to use domain adaptation to integrate temporal and user demographic factors into document classifiers. By adapting knowledge of how language varies across the metadata, models can learn generalized representations of language through the metadata-invariant embeddings.
This approach will lead to metadata-adapted document classifiers and can also extend to personalize classification models by user embedding.
Bio: Xiaolei Huang is a 4th-year PhD candidate in Information Science at the University of Colorado, Boulder. He is currently a visiting scholar at the Johns Hopkins University. His research interests are in Natural Language Processing, Machine Learning and Public Health. Particularly, he focuses on domain adaptation, cross-lingual transfer learning, user modeling and fairness.
Abstract: Large Language Models (LLMs) have revolutionized how people interact with knowledge, offering unprecedented opportunities to accelerate the pace of scientific discovery. In this talk, I will discuss my research on the synergy between LLMs and scientific knowledge--specifically how these models extract, induce, and verify knowledge to automate the research lifecycle. First, I will cover our work on improving knowledge extraction from vast scientific literature, focusing on enabling models to comprehend long documents in a cost-efficient and comprehensive manner. I will describe a novel paradigm for representing document-level structured information as question-answer pairs and how we address the challenges of long-context understanding by leveraging global context through retrieval-augmented modeling. Next, I present our pioneering work on using LLMs for new scientific hypothesis generation. We introduce a framework employing reinforcement learning with fine-grained reward modeling and adaptive controllers.
This approach balances novelty, feasibility, and effectiveness to generate inspiring and actionable research hypotheses. Finally, I will discuss work on the first LLM Scientist for machine learning research. I will demonstrate how LLMs can move beyond hypothesis generation to participate in the execution and validation of scientific hypotheses, ensuring that the discovered knowledge is not only innovative but also grounded and verified.
Bio: Xinya Du is a tenure-track assistant professor at UT Dallas Computer Science Department. He earned a Ph.D. degree from Cornell University and was a Postdoctoral Research Associate at the University of Illinois (UIUC). He has also worked at Microsoft Research, Google Research, and Allen Institute AI. His research is on large language models, deep learning, and their applications in science.His work has been published in leading NLP and ML conferences (ACL, ICLR, NeurIPS). His research has received multiple recognitions, including a Best Paper Award at AAAI AI for Research and a Best Poster Award at ICML AI for Science workshop. His work was included in the list of Most Influential ACL Papers and has been covered by major media like New Scientist. He was named a Spotlight Rising Star in Data Science by the University of Chicago and is the recipient of several prestigious awards, including the Amazon Research Award, Cisco Research Award, Open Philanthropy Award, and the NSF CAREER Award.
Location: NCS 120
The goal of this workshop is to offer a comprehensive overview of AI agents, bring ML, Systems, and HCI research communities together to share progress, discuss common problems and evaluation setups, and identify opportunities for collaboration. We aim to bring together attendees from diverse disciplines to foster interdisciplinary collaboration and discuss open research questions.
Location: Jane Street Offices, New York
Register here.
The Office for Research and Innovation invites you to hear about the two-month journey in which the Brook & Beyond team supported eight cohorts in bringing their bold ideas from the lab to the marketplace. It's an energizing evening that highlights the collaboration, creativity, and entrepreneurial spirit driving discovery across the University.
Meet this year's award recipients, hear pitches from the emerging founders, and applaud their achievements.
Connect, celebrate, and be part of the momentum shaping the future of innovation at
Stony Brook University.
Refreshments will be served. Registration is required.
Register Here.
As AI drives rapid change across professional fields, how do you bring these developments into your classroom? The CELT AI Panel Discussion will gather academic thought leaders to explore how generative AI is reshaping teaching, learning, and the knowledge students need for today's world. Our panelists will share practical strategies for integrating AI-related advancements into course content, highlight both opportunities and challenges, and discuss how educators can help students build critical thinking, ethical awareness, and hands-on experience with emerging AI technologies. Join us to examine how teaching can evolve alongside an AI-transformed society.
Register here.