Title: Cultural Biases, World Languages, and User Privacy in Large Language Models
Abstract: In this talk, I will highlight three key aspects of large language models: (1) cultural bias in LLMs and pre-training data, (2) decoding algorithm for low-resource languages, and (3) human-centered design for real-world applications.

The first part focuses on systematically assessing LLMs' favoritism towards Western culture. We take an entity-centric approach to measure the cultural biases among LLMs (e.g., GPT-4, Aya, and mT5) through natural prompts, story generation, sentiment analysis, and named entity tasks. One interesting finding is that a potential cause of cultural biases in LLMs is the extensive use and upsampling of Wikipedia data during the pre-training of almost all LLMs. The second part will introduce a constrained decoding algorithm that can facilitate the generation of high-quality synthetic training data for fine-grained prediction tasks (e.g., named entity recognition, event extraction). This approach outperforms GPT-4 on many non-English languages, particularly low-resource African languages. Lastly, I will showcase an LLM-powered privacy preservation tool designed to safeguard users against the disclosure of personal information. I will share findings from an HCI user study that involves real Reddit users utilizing our tool, which in turn informs our ongoing efforts to improve the design of AI models.
Bio:

Wei Xu is an Associate Professor in the College of Computing and Machine Learning Center at the Georgia Institute of Technology, where she is the director of the NLP X Lab. Her research interests are in natural language processing and machine learning, with a focus on Generative AI, robustness and fairness of large language models, multilingual LLMs, as well as AI for science, education, accessibility, and privacy research. She is a recipient of the NSF CAREER Award, Google Academic Research Award, CrowdFlower AI for Everyone Award, Best Paper Awards and Honorable Mentions at COLING'18, ACL'23, ACL'24. She also received research funds from DARPA and IARPA. She is currently an executive board member of NAACL. Join Zoom Meeting https://stonybrook.zoom.us/j/98855994362?pwd=F2qnpwL85fhCBHAEW9ZBpXihfwGHsj.1 (ID: 98855994362, passcode: 172797) Join by phone (US) +1 646-876-9923 (passcode: 172797) Joining instructions: https://www.google.com/url?q=https://applications.zoom.us/addon/invitation/detail?meetingUuid%3DuDJcUTvyQueZkCaUSAwFlg%253D%253D%26signature%3Da3d49e0f7f2e74e7130f7308c74bd85ba7b99587b98ba2e34238bb657ca51a09%26v%3D1&sa=D&source=calendar&usg=AOvVaw2jTn5cjfRG8vXU8KHHlU2Y Meeting host: H.Andrew.Schwartz@stonybrook.edu

Join Zoom Meeting:
https://stonybrook.zoom.us/j/98855994362?pwd=F2qnpwL85fhCBHAEW9ZBpXihfwGHsj.1
Synthetic Dreams and Ghost Machines

Dr. Steven Skiena will join the Cinema for a presentation exploring the rapidly evolving worlds of artificial intelligence and robotics.

Dr. Skiena is Distinguished Teaching Professor of Computer Science and Associate Director of the AI Innovation Institute at Stony Brook University. A leading researcher in data science and algorithms, he is the author of several influential books on AI and computation, including The Algorithm Design Manual and The Data Science Design Manual.

The presentation will be followed by a screening of Ex Machina, Alex Garland's provocative science-fiction thriller exploring the seductive and unsettling boundaries between human consciousness and artificial intelligence.

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Abstract: This presentation will begin by outlining key challenges facing the modern power grid and summarizing our group's research efforts to address them. It will then discuss how AI and machine learning are reshaping the grid modernization. The major focus of the talk will highlight a range of AI/ML applications we have developed in recent years to enhance grid operation, planning, control, and security.

Biography: Meng Yue is currently leading the Grid Modernization and Security Group in the Interdisciplinary Science Department at Brookhaven National Laboratory (BNL). He received his Ph. D. from Michigan State University in electrical engineering. His major research interests include power system modeling, simulation, and control, and applications of AI/ML- and quantum machine learning and quantum computing in operation, planning, and security of the future grid.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1604383624?pwd=ffQ5cUPNxTI7nzClKQO6cnsNbhF9Vf.1

Meeting ID: 160 438 3624
Passcode: 558449

Abstract: In this talk, we will discuss what a CS PhD entails and the traits and habits that are important for success in PhD programs and future careers. While the talk is targeted to first-year PhD students, PhD students at all levels should derive from it.

Bio: Samir Das is a professor in the Department of Computer Science at Stony Brook
University. He is currently serving as the department chair. He is well recognized in the
community for his research in wireless networks and systems.

Location: NCS120

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Abstract: Two-dimensional (2D) materials such as graphene, hBN, and TMDs offer atomically sharp interfaces and unprecedented tunability when vertically assembled into van der Waals heterostructures. These stacks have enabled discoveries ranging from moiré superconductivity and correlated insulators to quantum emitters and next-generation nanoelectronic devices. Yet constructing high-quality heterostructures remains largely artisanal: researchers manually identify exfoliated flakes, align a polymer stamp by eye, and finely adjust temperature and contact geometry through tacit skill. This manual workflow is difficult to reproduce, scales poorly, and prevents systematic exploration of the enormous combinatorial space of materials, twist angles, and interfacial conditions. AutoLab is an autonomous platform that translates this tacit human expertise into programmable, feedback-driven control. Instead of pressing flakes with predefined trajectories, AutoLab uses machine vision to detect polymer-wafer contact, dynamically regulates contact evolution through closed-loop actuation and temperature control, and captures high-quality flakes with the cleanliness and precision of expert manual fabrication. The system integrates perception, decision making, and motion planning into a single robotic framework, enabling reproducible stacking, wafer-level coverage, and accelerated discovery. Beyond 2D materials, AutoLab illustrates a broader paradigm for AI-native scientific automation: codifying human experimental reasoning into algorithms that interrogate data in real time, adaptively adjust instrumentation, and generate scalable, high-fidelity datasets. Such platforms could generalize to diverse research domains--quantum device fabrication, optical alignment, surface science, autonomous microscopy, and other workflows where expert intuition currently limits throughput and reproducibility. By bridging artisanal manipulation and robotic autonomy, AutoLab points toward a future where scientific discovery is accelerated by machines that not only execute instructions, but learn, respond, and collaborate with human scientists.

Biography: Dr. Yutao Li is a research associate from Department of Condensed Matter Physics and Material Science, Brookhaven National Laboratory. He has 8 years of experience in 2D material sample fabrication, and investigation in their electronic transport, optical and mechanical properties.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1604383624?pwd=ffQ5cUPNxTI7nzClKQO6cnsNbhF9Vf.1

Meeting ID: 160 438 3624
Passcode: 558449

Abstract: My presentation will be focused on introducing the use of Screenomics, a passive sensing approach that directly collects time-intensive data from participants' smartphones, to observe and analyze adolescents' digital behaviors across multiple timescales. I will present our completed and ongoing efforts using Screenomics to (1) evaluate the biases of self-reports of screen time and app use, (2) describe how adolescents use their smartphones during school hours and overnight, (3) examine longitudinal associations between adolescents' social media use and mental health, and (4) capture adolescents' communication pattern with parents. I will also introduce the theoretical framework and study plan for a new NIH-funded project that aims to identify adolescents' social media management strategies (SMMS) and how SMMS are related to adolescents' actual social media use and mental health. I will conclude with a discussion of future directions for interventions to promote healthy digital practices among adolescents.

Bio: Xiaoran Sun, Ph.D., is an assistant professor in the Department of Family Social Science, College of Education and Human Development at University of Minnesota (UMN). She is the director of the UMN Technology, Teens, and Families Lab and a core faculty of the Learning Informatics Lab. She is also affiliated with the UMN Data Science Initiative and the Minnesota Population Center. Her research is mainly focused on using innovative approaches, such as passive sensing and machine learning, to examine children's and parents' use of technology and the implications for their wellbeing. Her work is being funded by the U.S. National Institute of Mental Health and the Spencer Foundation.

We live in a new scientific paradigm: the Big Data era, in which a lot of data is available for almost anything. In this new paradigm, the driving force is to use data directly to learn about chemical and physics systems employing artificial intelligence. This paradigm has proven helpful in simulating realistic physical, biological, and chemical models, yielding impressive results. Similarly, the insight gained in these situations can be used to improve our understanding of fundamental processes. In that regard, we want to answer the question: Can a machine learn chemistry? The answer to this question is still debatable, but we will show our ideas and methods to find the answer. We will also discuss our results on predicting atom-diatom reactions and other avenues and work in progress in our group.

Please register for the STEM Speaker Series: Can a Machine learn Chemistry here.
The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 10 minutes) by multiple people. Students can register for 1 credit for CSE 656. Registered students must attend and present a minimum of 2 or 3 talks. Everyone else is welcome to attend. Fill in https://forms.gle/pCVXovgfMfQwGqG38 to subscribe to our mailing list for further announcement.
Title: AI-Driven Target Selection Methods for Touch and Gaze Input

Abstract: Accurately selecting targets is an essential aspect of  Human-Computer Interaction. Erroneous selections can cause tedious undo and redo actions. Additionally, some selection errors are non-reversible and can lead to undesirable consequences. However, high-accuracy target selection remains a challenge on touchscreen devices due to the small target size and imprecise touch inputs, and in gaze interaction because of the gaze tracking noise and no easy-to-use selection action. We first propose ReLM, a Reinforcement Learning-based Method for touchscreen target selection. ReLM can automatically show suggestions and require a second touch if the input is ambiguous, and can directly select a target candidate when the input is certain. Our empirical evaluation shows that ReLM reduces the error rate from 6.92% to 1.63%, and the selection time from 2.23s to 1.59s over Shift, an existing suggestion-based method. Compared to BayesianCommand, a direct selection-based method, our ReLM reduces the error rate from 3.64% to 0.89%, while increasing the selection time by only 200 ms. Secondly, we investigate how to improve target selection performance for gaze interaction. We propose BayesGaze, an eye-gaze based target selection method. It accumulates the signal of each gaze point for selecting a target calculated by Bayes Theorem, and uses a threshold mechanism to determine the target selection. Our investigation shows that BayesGaze improves target selection accuracy and speed over a dwell-based selection method, and the Center of Gravity Mapping method.

All are welcome. Here  is the zoom meeting link:
https://stonybrook.zoom.us/j/93130953411?pwd=Rm5IRlVPQ3M0cHJsTXpCVFljUlFGUT09Meeting ID: 931 3095 3411Passcode: 999413

Abstract: Generating high-fidelity EEG data at scale is essential for overcoming dataset scale limitation and satisfying privacy requirements in computational neuroscience. However, the predominant reliance on discrete denoising formulations fails to adequately capture the continuous temporal evolution and frequency-domain characteristics intrinsic to EEG recordings. Consequently, such approaches often compromise long-range temporal coherence and introduce structural discrepancies in both spectral and temporal domains.
In this work, We argue that faithful EEG synthesis demands generative models that directly characterize the continuous dynamics underlying neural signals. To this end, we propose a conditional flow matching framework that treats EEG as raw waveform sequences traversing continuous-time trajectories. Rather than relying on discretized denoising steps or handcrafted signal representations, our method learns a smooth velocity field mapping noise distributions to the target EEG manifold, thereby naturally preserving temporal continuity and transient neural phenomena. To enforce fidelity to fundamental EEG characteristics, we incorporate principled constraint terms that maintain spectral consistency, temporal stationarity, and signal-level statistical properties. Evaluated on large-scale benchmarks, our approach establishes new state-of-the-art results compared to competitive baselines. Comprehensive analyses further confirm that the proposed framework faithfully recovers core structural attributes of neural dynamics, offering a scalable and theoretically grounded solution for high-fidelity EEG generation.

Speaker: Yifan Wang

Location: https://stonybrook.zoom.us/j/95300279172?pwd=maxkSIma0Go2O8Wzre3CLYWHflnUi3.1&jst=2
Meeting ID: 953 0027 9172
Passcode: 418725