Abstract: Language is not just something we generate, but something humans use to interact with the world around them. Indeed, today's conversational AI agents speak fluently, but often treat language as prediction rather than interaction, producing responses that sound correct while failing to recognize when requests are ungrounded, impossible, or misunderstood. My research asks what it would take for multimodal agents to take the next step and use language effectively to take actions in grounded contexts and in this talk, I argue that many challenges in multimodal LLM design, including alignment, hallucination, and adaptation, are due to a lack of pragmatics: an understanding of the implicit context behind the implied actions of the words in the query. From visual understanding, to automatic speech recognition, to hallucination detection, I will demonstrate that incorporating pragmatic/contextual reasoning substantially improves agent behavior, and that pragmatic reasoning will drive a necessary shift in how we build multimodal conversational agents that can see, listen, act, and speak in context.


Bio: David M. Chan, Ph.D., is a postdoctoral scholar at the University of California, Berkeley, specializing in multimodal conversational AI. His research focuses on developing scalable AI systems that move beyond language prediction toward grounded language use, integrating vision, audio, and language to enable pragmatic interaction, improve AI-human collaboration, and reduce hallucinations in generative models. Beyond academia, he has worked with leading organizations including Amazon, Google, and NASA, to build and deploy safe, efficient, and accessible machine learning systems. He is also the developer and maintainer of TSNE-CUDA, an open-source tool for high-dimensional data visualization, used by over 40,000 researchers in fields ranging from biomedical technology to industrial manufacturing. David holds a Ph.D. and M.Sc. in Computer Science from UC Berkeley, where he was a graduate fellow with the Center for Technology, Society & Policy (CTSP), and dual B.Sc. degrees in Computer Science and Mathematics from the University of Denver.

Location: NCS 120
Le Hou Dissertation Defense: Deep Learning for Digital Histopathology across Multiple Scales

ABSTRACT: Histopathology is the study of tissue changes caused by diseases such as cancer. It plays a crucial role in disease diagnosis, survival analysis and development of new treatments. Using computer vision techniques, I focus on multiple tasks for automated analysis in digital histopathology images, which are challenging because histopathology images are heterogeneous and complex, due to the large variation of hundreds of cancer types in gigapixel resolution. In this thesis, I show how histopathology image analysis tasks can be viewed in three scales: Whole Slide Image (WSI)-level, patch-level and cellular-level, and present my contributions in each resolution level.

BIO: WSI-level analysis such as classifying WSIs into cancer types is challenging, because conventional classification methods such as off-the-shelf deep learning models cannot be applied directly on gigapixel WSIs due to computational limitations. I contribute a patch-based deep learning method that classifies gigapixel WSIs into cancer types and subtypes with close-to-human performance. This method is useful for computer-aided diagnosis. At patch-level, I contribute a novel method for histopathology image patch classification. On the task of identifying Tumor Infiltrating Lymphocyte (TIL) regions, the prediction result of this method correlates to the survival rate of patients. At cellular-level, I contribute novel methods for nucleus classification and roundness regression, which are interpretable features for histopathology studies. With this method, I generated a large-scale dataset of segmented nuclei, in WSIs from a large publicly available digital histopathology image dataset, to help advance histopathology research.

Abstract: The advent of ChatGPT has redrawn the boundary of pedagogical discourse, where the dyadic configuration of teacher-student has, for many, become triadic -- one that includes AI as an relevant third party, not to be missed or dismissed. Within applied linguistics, AI-focused research has predominantly targeted the teaching and learning of writing (Fang & Han, 2025). The work on AI and speaking, on the other hand, has largely involved perception studies documenting its positive impact on learners' willingness to communicate (Goh & Aryadoust, 2025). In this talk, I explore the role of AI in the teaching and learning of speaking, and in particular, the development of interactional competence. Based on a corpus of learner-AI interactions, I demonstrate the ways in which ChatGPT excels and fails at acting as a useful conversation partner, with a view towards furthering our ongoing deliberation on the affordances and constraints of AI in language education.

Speaker: Hansun Zhang Waring (Teachers College, Columbia University)

Hansun Zhang Waring is Professor of Linguistics and Education at Columbia University and founder The Language and Social Interaction Working Group (LANSI). As an applied linguist and a conversation analyst, Hansun is interested in all things interaction -- (second language) pedagogical interaction, communication with the public, parent-child interaction, and human-AI interaction (HAI). Her work has appeared in leading journals in applied linguistics and discourse analysis as well as numerous book volumes, some of which she (co-)authored or co-edited. She is on the editorial boards of Chinese Language and Discourse (CLD), Classroom Discourse (CD), and International Review of Applied Linguistics (IRAL).

Location: Wang Center, Lecture Hall #1

If you need special accommodation, please contact chikako.nakamura@stonybrook.edu.

AI for Conservation: AI and Humans Combating Extinction Together by Daniel I. Rubenstein of Princeton University

ABSTRACT: The state of our planet is not good. We have lost more than 60% of the world's wildlife. Stopping the decline remains a challenge, especially since acquiring appropriate knowledge is expensive, time consuming and risky. Visual observations following the fates of a few individuals was the currency of the realm. But GPS technology and now machine learning provide a non-invasive scalable alternative. Photographs, taken by field scientists, tourists, automated cameras and incidental photographers, are the most abundant source of data on wildlife today. Wildbook, a project of tech for conservation coordinated by a non-profit Wild Me, is an autonomous computational system that starts from massive collections of images and, by detecting various species of animals and identifying individuals, combined with sophisticated data management, turns them into high-resolution information databases, enabling scientific inquiry, conservation and citizen science.

BIO: Dan Rubenstein is the Class of 1877 Professor of Zoology. He is currently Director of Princeton's Environmental Studies Program and is former Chair of Princeton University's Department of Ecology and Evolutionary Biology and Director of Princeton's Program in African Studies. He is a behavioral ecologist who studies how environmental variation and individual differences shape social behavior, social structure, sex
roles and the dynamics of populations. He has special interests in all species of wild horses, zebras and asses, and has done field work on them throughout the world identifying rules governing decision-making, the emergence of complex behavioral patterns and how these understandings influence their management
and conservation. In Kenya he also works with pastoral communities to develop and assess impacts of various grazing strategies on rangeland quality, wildlife use and livelihoods. He has also developed a scout program for gathering data on Grevy's zebras and created curricular modules for local schools to raise awareness about the plight of this endangered species. He engages people as 'Citizen Scientists' and has recently extended his work to measuring the effects of environmental change, including issues pertaining to the global commons
and changes wrought by management and by global warming, on behavior.
Abstract: Robot control has evolved from optimization-based controllers---precise but task-specific---through deep reinforcement learning's learned policies, to Vision-Language-Action (VLA) models that leverage pretrained vision-language backbones for language-conditioned manipulation across diverse tasks.
Despite their promise, VLAs exhibit a critical limitation: they function primarily as trajectory learners rather than skill learners. Recent evaluations reveal that VLAs often fail when faced with even minor variations in object initialization or environmental conditions, suggesting they memorize specific trajectories rather than acquiring generalizable manipulation skills. Attempts to address this through 3D spatial representations have shown limited success, indicating that the missing component may be more fundamental than geometric understanding alone.
This work argues that World Models (WMs)---internal representations that predict future states given actions---constitute the missing piece for robust VLA systems. We present one completed contribution and two ongoing investigations.
We developed a dual-layer world model for human-robot interaction that anticipates both physical scene evolution and latent human preferences for assistive tasks. Building on these foundations, we present ongoing work probing VLA internal representations to verify implicit world model existence, and propose a WM-VLA integration approach operating in the native visual domain through embedding prediction and image decoding.
Together, these contributions and investigations establish a foundation for WM-VLA systems, pointing toward robust, generalizable robot policies.
Speaker: Jason Qin
Location: NCS 220

As part of a grant project funded by the AI3 Institute, a group of instructors participated in a faculty development program, Fostering Writing-to-Learn Skills with Critical AI Literacy: A Faculty Development and Student Support Program. This program was developed to support instructors across campus with navigating/integrating AI in their courses specifically around writing intensive/involved assignments. We would like to invite anyone interested to the culmination of this program, a mini-symposium, where the participants will share practical changes they made or are making around writing intensive/involved assignments and AI.

Location: Wang 201

A light lunch will be served. Please register by Friday, November 7th.