Abstract: Large Language Models (LLMs) have revolutionized how people interact with knowledge, offering unprecedented opportunities to accelerate the pace of scientific discovery. In this talk, I will discuss my research on the synergy between LLMs and scientific knowledge--specifically how these models extract, induce, and verify knowledge to automate the research lifecycle. First, I will cover our work on improving knowledge extraction from vast scientific literature, focusing on enabling models to comprehend long documents in a cost-efficient and comprehensive manner. I will describe a novel paradigm for representing document-level structured information as question-answer pairs and how we address the challenges of long-context understanding by leveraging global context through retrieval-augmented modeling. Next, I present our pioneering work on using LLMs for new scientific hypothesis generation. We introduce a framework employing reinforcement learning with fine-grained reward modeling and adaptive controllers.
This approach balances novelty, feasibility, and effectiveness to generate inspiring and actionable research hypotheses. Finally, I will discuss work on the first LLM Scientist for machine learning research. I will demonstrate how LLMs can move beyond hypothesis generation to participate in the execution and validation of scientific hypotheses, ensuring that the discovered knowledge is not only innovative but also grounded and verified.

Bio: Xinya Du is a tenure-track assistant professor at UT Dallas Computer Science Department. He earned a Ph.D. degree from Cornell University and was a Postdoctoral Research Associate at the University of Illinois (UIUC). He has also worked at Microsoft Research, Google Research, and Allen Institute AI. His research is on large language models, deep learning, and their applications in science.His work has been published in leading NLP and ML conferences (ACL, ICLR, NeurIPS). His research has received multiple recognitions, including a Best Paper Award at AAAI AI for Research and a Best Poster Award at ICML AI for Science workshop. His work was included in the list of Most Influential ACL Papers and has been covered by major media like New Scientist. He was named a Spotlight Rising Star in Data Science by the University of Chicago and is the recipient of several prestigious awards, including the Amazon Research Award, Cisco Research Award, Open Philanthropy Award, and the NSF CAREER Award.

Location: NCS 120

This virtual presentation series is designed to inform the Stony Brook University research community about the Research Funding Landscape of key topic areas. Our Strategic Research Initiatives team will provide insight into the rapidly shifting funding environment using policy briefs, budgetary priorities, and relevant legislation. We will highlight federal and state priorities in the current and upcoming years to help Stony Brook researchers develop strategies for pursuing funding in a rapidly shifting environment. This series is moderated by Mónica Bugallo, Interim Vice President for Research & Innovation.

Join us for the third in the series, focused on the artificial intelligence landscape:


Translating the Funding Landscape for Stony Brook Researchers: Artificial Intelligence
Presented by Catherine Chen, Ph.D., Research Development Associate
Faculty Respondent: Assistant Professor Nav Nidhi Rajput, Department of Materials Science and Chemical Engineering
Wednesday, April 22, 2026 at 2 pm to 3 pm

Registration is Required

AI is everywhere and so are the privacy concerns that come with it. At its core, the most common forms of AI we use today are online digital services and thus inherit the usual privacy risks. We'll take a look at indirect prompt injection- a technique that can trick AI tools into revealing or extracting private information as well as techniques being used in academic contexts to manipulate systems and even mislead researchers.

Register here for the online session.

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Abstract: The increasing complexity and volume of data from electron microscopy necessitates advanced computational tools for timely and accurate analysis. In this talk, I will present several machine learning (ML) models developed to interpret diverse datasets from transmission electron microscopy (TEM). First, I demonstrate segmentation models for labelling regions of interest from in situ TEM images, such as atomic column positions or reaction sites that allow atomic-level quantitative analysis of data. Second, I introduce a self-supervised CNN model for denoising of low-dose HRTEM images, enabling clearer visualization of atomic features without sacrificing temporal resolution. Finally, a transformer-based model trained to predict copper oxidation states directly from their electron energy loss spectroscopy spectra will be introduced. Together, these projects showcase the power of tailored ML solutions to extract quantitative insights from complex microscopy data.

Biography: Brian Lee is a research associate working for the Electron Microscopy group and Theory and Computation group at the Center for Functional Nanomaterials. Previously, he has received PhD in Mechanical Engineering from Duke University and worked as a postdoc at Purdue University. His research focuses on applying machine learning and simulation techniques for materials science.

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1604383624?pwd=ffQ5cUPNxTI7nzClKQO6cnsNbhF9Vf.1

Meeting ID: 160 438 3624
Passcode: 558449

The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 10 minutes) by multiple people. Students can register for 1 credit for CSE 656. Registered students must attend and present a minimum of 2 or 3 talks. Everyone else is welcome to attend. Fill in https://forms.gle/pCVXovgfMfQwGqG38 to subscribe to our mailing list for further announcement.

The event will take place on Zoom and will feature two distinguished guest speakers: SBU alumnus, Velchamy Sankarlingam, president of Product and Engineering at Zoom, and Simeon Ananou, vice president for Information Technology and CIO at Stony Brook University. The discussion will be moderated by Haresh Gurnani, dean of the College of Business at Stony Brook University.

Exploring AI's Impact on Communication and Connection

Artificial Intelligence (AI) has rapidly evolved, becoming an integral part of various industries, including education and business. This event aims to delve into how AI is reshaping the way we learn and work, particularly in enhancing communication and fostering human connections. Velchamy Sankarlingam, an SBU alumnus and a key figure at Zoom, will share his insights on how AI-driven tools are revolutionizing virtual communication platforms, making interactions more seamless and effective.

Simeon Ananou, with his extensive experience in information technology, will provide a perspective on how AI is being integrated into educational institutions to improve learning outcomes and administrative efficiency. His role at Stony Brook University places him at the forefront of implementing innovative technologies that benefit both students and staff.

A Conversation Led by Expertise

Dean Haresh Gurnani, known for his leadership and expertise in business education, will guide the conversation, ensuring that the discussion remains focused on the practical implications of AI. He will explore how AI is not only boosting productivity but also enriching overall experiences in the workplace and educational settings. The event will include an interactive Q&A session, allowing attendees to engage directly with the speakers and gain deeper insights into the topics discussed.

As AI continues to develop, events like this are crucial for understanding its impact and potential. Stony Brook University's College of Business is committed to providing platforms for such important discussions, fostering an environment where innovation and education intersect.

This event is open to all. Please visit https://www.givecampus.com/schools/StonyBrookUniversity/events/artificial-intelligence-reshaping-learning-and-work to register.

Towards Saving Lives with Natural Language Processing Andrew Schwartz Dept. of Computer Science Stony Brook Analyzing language use patterns is proving to be a valuable and unique approach to understanding the psychological, social, and health factors of people. On the individual level, Facebook and Twitter have been found predictive of mental health, personality, demographics, and occupational class (among others). At the community or county-level, Twitter has been found predictive of flu and allergy outbreaks, life satisfaction, atherosclerotic heart disease mortality, health behavioral risk factors, excessive drinking, and HIV prevalence. While these techniques have shown robust links over a plethora of important aspects of human life, it is not clear whether any lives have been saved, at least directly, by such work. At their core, some barriers to improving health care and saving lives are likely not NLP or even AI problems, but others are perhaps technical in nature and suggest changing the way we model data. This seminar will have two parts: a presentation and a discussion. I will start by going over recent and on-going work toward predicting mental health outcomes --- depression, addiction relapse, future psychological distress --- from human language use patterns. Then, I will present an imperfect vision of a future where NLP helps to save lives and open the floor for discussion of technical barriers and whether such a vision is practical. Biography: Andrew Schwartz received his PhD in Computer Science from the University of Central Florida in 2011 with research on acquiring lexical semantic knowledge from the Web. He then joined the University of Pennsylvania where he was a Postdoctoral Research Fellow and later Visiting Assistant Professor in Computer & Information Science. He is Lead Research Scientist for the World Well-Being Project, a multidisciplinary group of Computer Scientists and Psychologists studying physical and psychological well-being based on language in social media.