Hidden Biases. Ethical Issues in NLP, and What to Do about Them presented by Dirk Hovy of Bocconi University

ABSTRACT: Through language, we fundamentally express who we are as humans. This property makes text a fantastic resource for research into the complexity of the human mind, from social sciences to humanities. However, it is exactly that property that also creates some ethical problems. Texts reflect the authors' biases, which get magnified by statistical models. This has unintended consequences for our analysis: If our data is not reflective of the population as a whole, if we do not pay attention to the biases contained, we can easily draw the wrong conclusions, and create disadvantages for our users.

In this talk, I will discuss several types of biases that affect NLP models, their sources, and potential counter measures: (1) Bias stemming from data, i.e., selection bias (if our texts do not adequately reflect the population we want to study), label bias (if the labels we use are skewed) and semantic bias (the latent stereotypes encoded in embeddings); (2) Biases deriving from the models themselves, i.e., their tendency to amplify any imbalances that are present in the data; (3) Design bias, i.e., the biases arising from our (the researchers) decisions which topics to analyze, which data sets to use, and what to do with them. For each bias, I will provide examples and discuss the possible ramifications for a wide range of applications, and various ways to address and counteract these biases, ranging from simple labeling considerations to new types of models.

BIO: Dirk Hovey is an associate professor of Computer Science in the department of marketing at Bocconi University. He received his PhD from the University of Southern California in Los Angeles, where he worked as a research assistant at the Information Sciences Institute. 

He works in Natural Language Processing (NLP), a subfield of artificial intelligence. His research focuses on computational social science. His interests include integrating sociolinguistic knowledge into NLP models, using large-scale statistics to model the interaction between people's socio-demographic profile and their language use, and ethics for data science and algorithmic fairness.


Abstract:
Large language models (LLMs) have transformed the way humans write code, bringing unprecedented automation to software development. In this talk, I will first provide an overview of my research on enhancing LLMs' code intelligence, optimizing each step of the development pipeline towards more complex software engineering tasks. I will then delve into my key contributions, focusing on how to equip LLMs with a deeper, more comprehensive understanding of software programs. Finally, I will discuss the future of AI-driven software engineering, envisioning a new era of automation that is more reliable, intelligent, and cost-efficient.

Bio:
Yangruibo (Robin) Ding is a Ph.D. candidate in the Department of Computer Science at Columbia University. His research is at the intersection of Software Engineering and Machine Learning, focusing on developing large language models (LLMs) for code. He trains LLMs to generate, analyze, and refine software programs and constructs benchmarks to systematically evaluate LLMs in solving software engineering tasks. He also studies how to improve LLMs' reasoning capability to tackle complex programming tasks, such as debugging and patching. His interdisciplinary research has been published in top-tier conferences of software engineering, programming languages, natural language processing, and machine learning. He won an ACM SIGSOFT Distinguished Paper Award, an IEEE TSE Best Paper Runner-up, and received an IBM Ph.D. Fellowship.
Location:
NCS 120

Submit an abstract celebrating research, new discoveries and achievements in medicine and science!

We encourage faculty, nurse practitioners, post-doctoral fellows, fellows, residents, medical students, graduate students and undergraduate students to submit an abstract. Original research, case reports and case series are welcome.

Abstract submission deadline: FEBRUARY 7, 2025

For more details, visit here.

Description:

Curious about what AI image generation tools are out there and how they work? Come down to the library Galleria space (outside the Central Reading Room) to see some demonstrations and learn more about them.

Librarians Chris Kretz and Ahmad Pratama, along with David Ecker of DoIT, will be hosting Explore AI demos from Monday - Wednesday this week on different topics. Whether you're new to AI or an experienced user, stop by and take a look!

Location: Library Galleria

Join the Center of Excellence in Wireless and Information Technology (CEWIT) and their co-host IEEE-USA for a livestream panel discussion on Generative Artificial Intelligence (Gen AI). In this engaging livestream, we will dive into the technologies that continue to transform what is possible and explore the dynamic intersection of innovation, creativity, ethics, and Gen AI.

CEWIT is joined by Stony Brook University experts who will provide their insights and perspectives on this rapidly changing technology.

Meet the Panel

Laura Lindenfeld, PhD

Executive Director
Alan Alda Center for Communicating Science®
Dean
School of Communication & Journalism
BIO

Margaret Schedel, PhD
Associate Professor
Composition and Computer Music
Co-Founder
Lyrai
BIO

Steven Skiena, PhD

Interim Director
AI Innovation Institute
Distinguished Professor
Computer Science
BIO

Vivian Zhang
CTO/School Director
NYC Data Science Academy
Chief Data Officer
GoDental.ai
BIO


Register here.

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes three short talks on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Tuesday, January 7, 2025, 12:00 pm -- CDS, Bldg. 725, Training Room


Speakers

Sanket Jantre
Tao Zhang
Xi Yu


Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1615289117?pwd=Hqkbj9itxWrFnkhZ8rQXHPInO2gxdF.1

Meeting ID: 161 528 9117
Passcode: 991382

​This session brings together the scientists, agencies, and community partners generating environmental data across New York City to confront a shared challenge: critical atmospheric and marine data is being collected across the region, but too often in silos that limit its reach and impact.

​Using Governors Island's environmental sensing efforts as a working case, the program opens into a broader conversation. We will discuss how disparate data streams and objectives across NYC can be coordinated, shared, and activated for the benefit of the broader NYC community. And how that data can support healthier and safer communities, emergency preparedness and resilience, more informed city planning and operations, and better decision-making by businesses and investors.

​This session will include a fireside chat on the state of hyperlocal data in NYC with Assistant Commissioner Carolyn Olson, a panel and discussion on deploying hyperlocal data, and a tour of the Governors Island Environmental Observatory (GIEO). We hope to see you there.

Location: 110 Andes Rd New York, NY

Register to join.

An interactive session to discover how to create ALT text tags from images and create high-impact visuals, from identification to communicating ideas with images.

Discover how to use AI to create ALT text from images as well as identify objects in your environment, and build relatable visuals for high-impact presentations. Images communicate ideas as a way to understand concepts. AI-generated images have helped allow anyone to create these.

In this session, you will

  1. Creating image ALT Tags
  2. Transform ideas into images that are visually appealing
  3. Identify objects from visuals

Register here.
Abstract:

Recent advances in deep learning have significantly enhanced the capabilities of Natural Language Processing (NLP) and Vision-Language Models (VLMs). However, these advancements come with increased vulnerabilities, notably through backdoor attacks that pose severe security threats. This thesis addresses two critical dimensions of Trustworthy AI and Efficient Multimodal Representation Learning: (1) security through analyzing, detecting, and designing backdoor attacks in NLP and VLMs, and (2) efficiency through advanced multimodal representation methods tailored for clinical and medical imaging applications.

In the first dimension, we explore the internal mechanisms exploited by backdoor attacks, identifying the distinctive phenomenon of attention focus drifting in compromised transformer models, where trigger tokens consistently hijack attention. Leveraging these insights, we propose robust detection frameworks, including the attention-based Trojan detector (AttenTD) and a task-agnostic logit-based detection method (TABDet), achieving effective identification of backdoored NLP models across diverse tasks. We further introduce novel backdoor attack methodologies: the Trojan Attention Loss (TAL), enhancing attack efficiency and stealth through direct attention manipulation, and BadCLM, demonstrating critical vulnerabilities in clinical decision-support systems by effectively compromising clinical language models.

Extending our security exploration to multimodal settings, we investigate backdoor attacks on Vision-Language Models (VLMs), particularly in complex image-to-text generation tasks, proposing innovative techniques (TrojVLM, VLOOD) capable of embedding backdoors without direct access to original training data, thus showcasing practical risks in real-world scenarios.

In the second dimension, we address efficiency and interpretability challenges in clinical and pathology applications. We introduce TCP-LLaVA, the first multimodal large language model (MLLM) designed explicitly for Whole Slide Image (WSI) Visual Question Answering (VQA). Utilizing a novel token compression mechanism inspired by transformer-based models, TCP-LLaVA substantially reduces computational resource consumption while maintaining superior VQA performance across multiple tumor subtypes. Additionally, we present a multimodal transformer model integrating structured Electronic Health Records (EHR) with clinical notes, demonstrating enhanced predictive accuracy and interpretability for in-hospital mortality prediction through integrated gradient-based interpretability methods.

Together, these contributions present a comprehensive approach to ensuring AI models are not only secure against malicious manipulation but also efficient and interpretable for critical clinical applications, underscoring the essential need for trustworthy and effective AI systems.

Speaker: Weimin Lyu

Zoom: https://stonybrook.zoom.us/j/2392326575?pwd=SVQ2VkFXTnZZYmJUMXgvTXBuZWM3UT09

Meeting ID: 239 232 6575
Passcode: 436192