Hidden Biases. Ethical Issues in NLP, and What to Do about Them presented by Dirk Hovy of Bocconi University

ABSTRACT: Through language, we fundamentally express who we are as humans. This property makes text a fantastic resource for research into the complexity of the human mind, from social sciences to humanities. However, it is exactly that property that also creates some ethical problems. Texts reflect the authors' biases, which get magnified by statistical models. This has unintended consequences for our analysis: If our data is not reflective of the population as a whole, if we do not pay attention to the biases contained, we can easily draw the wrong conclusions, and create disadvantages for our users.

In this talk, I will discuss several types of biases that affect NLP models, their sources, and potential counter measures: (1) Bias stemming from data, i.e., selection bias (if our texts do not adequately reflect the population we want to study), label bias (if the labels we use are skewed) and semantic bias (the latent stereotypes encoded in embeddings); (2) Biases deriving from the models themselves, i.e., their tendency to amplify any imbalances that are present in the data; (3) Design bias, i.e., the biases arising from our (the researchers) decisions which topics to analyze, which data sets to use, and what to do with them. For each bias, I will provide examples and discuss the possible ramifications for a wide range of applications, and various ways to address and counteract these biases, ranging from simple labeling considerations to new types of models.

BIO: Dirk Hovey is an associate professor of Computer Science in the department of marketing at Bocconi University. He received his PhD from the University of Southern California in Los Angeles, where he worked as a research assistant at the Information Sciences Institute. 

He works in Natural Language Processing (NLP), a subfield of artificial intelligence. His research focuses on computational social science. His interests include integrating sociolinguistic knowledge into NLP models, using large-scale statistics to model the interaction between people's socio-demographic profile and their language use, and ethics for data science and algorithmic fairness.
Join Zoom Meeting
https://stonybrook.zoom.us/j/91945227869?pwd=emhoZDFWVTV0MVdPWW5uVk43MjQzUT09

Meeting ID: 919 4522 7869
Passcode: 452304
One tap mobile
+16468769923,,91945227869# US (New York)
+13126266799,,91945227869# US (Chicago)

Dial by your location
        +1 646 876 9923 US (New York)
        +1 312 626 6799 US (Chicago)
        +1 301 715 8592 US (Germantown)
        +1 669 900 6833 US (San Jose)
        +1 253 215 8782 US (Tacoma)
        +1 346 248 7799 US (Houston)
        +1 408 638 0968 US (San Jose)
Meeting ID: 919 4522 7869
Find your local number: https://stonybrook.zoom.us/u/aCvAYWkRg
  
The IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025 will be held from June 11th to June 15th, 2025, at the Music City Center, Nashville, TN. The IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR) is the premier annual computer vision event comprising the main conference and several co-located workshops and short courses. With its high quality and low cost, it provides an exceptional value for students, academics and industry researchers. Register here.
Title: Sustainable NLP

Time: Friday 4/29, 2:40 PM

Location: NCS 120

Abstract:


Natural language processing (NLP) technology has supercharged many real-world applications ranging from intelligent personal assistants (like Alexa, Siri, and Google Assistant) to commercial search engines such as Google and Bing. But current NLP applications use extremely large neural models, making them (i) expensive to deploy on servers, requiring large amounts of compute resources and power, and (ii) impossible to run on mobile devices, making on-device, privacy-preserving applications impractical.

In the first part of the talk, I will describe systems optimizations we have developed that significantly reduce the compute and memory requirement of NLP models. The optimizations we developed can be applied broadly and results in over 10x reduction in latency when deployed on mobile devices. In the second part of the talk, I will describe our recent work on predicting energy consumption of NLP models. Existing energy prediction approaches are not accurate, making it difficult for developers and practitioners to reason about their models in terms of power. We use a multi-level regression approach that produces highly accurate and interpretable energy predictions.



Bio:
Aruna Balasubramanian is an Associate Professor at Stony Brook University. She received her Ph.D from the University of Massachusetts Amherst, where her dissertation won the UMass outstanding dissertation award and was the SIGCOMM dissertation award runner up. She works in the area of networked systems. Her current work consists of two threads: (1) significantly improving Quality of Experience of Internet applications, and (2) improving the usability, accessibility, and privacy of mobile systems. She is the recipient of the SIGMobile Rockstar award, a Ubicomp best paper award, a Computing Innovation Fellowship, a VMWare Early Career award, several Google research awards, an

AI is everywhere -- and so are the privacy concerns that come with it. At its core, the most common forms of AI we use today are online digital services -- and thus inherit the usual privacy risks of any internet-based tool. However, AI also introduces a set of unique and evolving risks. We'll take a closer look at one of the newest developments in this area: indirect prompt injection -- a technique that can trick AI tools into revealing or extracting private information. You'll learn how this emerging form of AI manipulation works, why it matters, and how to protect yourself -- as well as how similar techniques are being used in academic contexts to manipulate systems and even mislead researchers.

Register for this Zoom workshop.

Abstract: As we enter the AI era, domain scientists face a critical question: What can we do to harness AI effectively for scientific discovery? AI has demonstrated remarkable capabilities, from accelerating simulations to uncovering hidden patterns in complex datasets. While these advancements offer unprecedented opportunities, they also raise concerns--AI models often function as black boxes, making it difficult to connect their outputs to established scientific principles. This lack of interpretability can undermine trust and limit adoption, particularly in fields like meteorology where physical understanding is critical.
In this talk, I will explore how interpretable AI can bridge this gap, highlighting its potential to generate explicit, physically meaningful equations rather than opaque neural networks. Through four case studies from my lab, I will showcase how interpretable AI can enhance scientific understanding:
  1. Satellite Precipitation Retrieval: Using AI-based approaches to interpret precipitation retrieval algorithms from AMSU data, we identified critical microwave channels (89 and 150 GHz) that directly link to physical processes in the atmosphere.
  2. Quantitative Precipitation Estimation (QPE): By applying symbolic regression models to polarimetric radar data, we derived mathematical expressions that outperform traditional Z-R relationships and existing QPE algorithms, offering new insights into rainfall microphysics.
  3. Tornado Probability Prediction: Leveraging reinforcement learning-based symbolic deep learning models, we developed interpretable equations that outperform the traditional Significant Tornado Parameter (STP) index, providing a clearer understanding of the relationships between key atmospheric variables and tornado risk.
  4. Domain-Aware Symbolic Regression for Scientific Equations: In our latest work, we introduced a symbolic regression framework that incorporates domain-specific symbol priors extracted from thousands of scientific publications. By encoding common mathematical structures--such as the prevalence of trigonometric functions in physics or logarithmic forms in biology--into a tree-structured reinforcement learning model, we improved both the accuracy and interpretability of discovered equations. This approach accelerates convergence, enforces physical plausibility, and reveals new governing relationships in climate and geophysical data.
Through these examples, I hope to spark discussion on the evolving role of domain scientists in the AI era and inspire new ways to integrate AI with physical understanding in atmospheric research.

IACS Seminar Speaker: Yixin Wen, University of Florida

Location: IACS Seminar Room or Zoom

Join Zoom Meeting: https://stonybrook.zoom.us/j/97596399106?pwd=0PBvElFLqov3biO6OlQxSWLWudkIuH.1
Meeting ID: 975 9639 9106
Passcode: 096213
CSE 600 Talk: Squeezing Software Performance via Eliminating Wasteful Operations presented by Xu Liu

ABSTRACT: Inefficiencies abound in complex, layered software. A variety of inefficiencies show up as wasteful memory operations, such as redundant or useless memory loads and stores. Aliasing, limited optimization scopes, and insensitivity to input and execution contexts act as severe deterrents to static program analysis. Microscopic observation of whole executions at instruction- and operand-level granularity breaks down abstractions and helps recognize redundancies that masquerade in complex programs. In this talk, I will describe various wasteful memory operations, which pervasively exist in modern
software packages and expose great potential for optimization. I will discuss the design of a fine-grained instrumentation-based profiling framework that identifies wasteful operations in their contexts, which guides nontrivial performance improvement. Furthermore, I will show our recent improvement to the profiling framework by abandoning
instrumentation, which reduces the runtime overhead from 10x to 3% on average. I will show how our approach works for native binaries and various managed languages such as Java, yielding new performance insights for optimization.

BIO: Xu Liu is an assistant professor in the Department of Computer Science at College of William & Mary. He obtained his PhD from Rice University in 2014 and joined the College of William & Mary in the same year. Prof. Liu works on building performance tools to pinpoint and optimize inefficiencies in HPC code bases. He has developed several open-source profiling tools, which are used worldwide at universities, DOE national laboratories and industrial companies. Prof. Liu has published a number of papers in high-quality venues. His papers received Best Paper Award at SC'15, PPoPP'18, PPoPP'19 and ASPLOS'17 Highlights, as well as Distinguished Paper Award at ICSE'19. His recent ASPLOS'18 paper has been selected as ACM SIGPLAN Research Highlights in 2019 and nominated for CACM Research Highlights. Prof. Liu is the receipt of 2019 IEEE TCHPC Early Career Researchers Award for Excellence in High Performance Computing. Prof. Liu served on the program committee of conferences such as SC, PPoPP, IPDPS, CGO, HPCA and ASPLOS.
CSE 656 Seminars in Computer Vision - Wednesdays 11:30am-12:50pm, Room NCS 120

The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 10 minutes) by multiple people. Students can register for 1 credit for CSE656. Registered students must attend and present a minimum of 2 or 3 talks. Everyone else is welcome to attend. Fill in https://forms.gle/pCVXovgfMfQwGqG38 to subscribe to our mailing list for further announcement.

The AI Community at Stony Brook University is proud to announce Datathon 2026.

Dive into data analysis and AI/ML, and get ready to build something big. In this year's underwater-themed event, enjoy a weekend of data analysis, hacking, networking, fun activities, and minigames.

Whether you're a seasoned developer, data scientist, designer, or completely new to hacking, this event is your chance to collaborate, learn data science, and create something impactful with data and AI/ML.

What is Datathon?

AI Community's Datathon is the premier data science competition at Stony Brook University, bringing together students of all skill levels for a weekend of data exploration, analysis, and innovation. Just like a typical hackathon, you will be using your skills to build your dream project.

Unlike a regular hackathon, Datathon is focused on data science. You will be given a set of data to work with, analyze, and apply to your project. You can also find your own data to use. Your project will be presented to a panel of judges consisting of professors and industry professionals!

Who Can Participate

  • Students of all skill levels and majors are welcome.
  • Come with a team or find one at the event or on Discord.
  • This event is open to SBU and non-SBU students.

* Non-SBU Undergraduate Students are ineligible to receive prizes
* You must be 18+ or older (Excludes minors who are active SBU students)

Location: SAC Ballroom B

Register here.