This presentation will be hosted both in-person and via Zoom.
Thursday, January 23, 1:00 PM to 2:00 PM
In-person: New Computer Science, Seminar Room 120
Zoom link: https://stonybrook.zoom.us/j/93425835490?pwd=jHwGG9A868eVNXZoK747OanDJk2pCq.1
Meeting ID: 934 2583 5490
Passcode: 251450
Time: Mar 17, 2021 10:00 AM Eastern Time (US and Canada)
Join Zoom Meeting
https://stonybrook.zoom.us/j/
Meeting ID: 936 1464 4178. Passcode: 965936
Natural Language Understanding and Semantic Parsing
(Partly joint work with former colleagues at Elemental Cognition)
Semantic parsing refers to the task of determining the propositional content of language: who did what to whom. It is part of the larger task of natural language understanding (NLU). I will start out by discussing what full NLU means, and argue that we are still far away, as a field, from solving full NLU, or even from knowing how to evaluate it.
In the second part of the talk, I will situate semantic parsing in the context of several other NLU subtasks. Typically, the target representation of semantic parsing uses an ontology (such as PropBank or FrameNet). Semantic parsing includes the subtasks of word sense disambiguation, argument detection, and argument role labeling. I will discuss choices among possible target ontologies. I will justify why we created a new ontology, Hector, based on FrameNet and the lexical resource NOAD, and explain some of its characteristics.
In the third part of the talk, I will present experiments we performed using transformer models. We obtain best results using a two-phase model, in which we first choose the frame, and then, given the frame, choose the arguments. We encode the problem for both tasks using indices in the sentence. While we develop the parser for our new ontology Hector, this approach also beats the state of the art for FrameNet and PropBank parsing.Biography: I am a professor in the Department of Linguistics at Stony Brook University with a joint appointment in IACS.
Until recently, I was a research scientist at Elemental Cognition. Elemental Cognition is working on deep natural language understanding.
I got my PhD with Aravind Joshi at the University of Pennsylvania in 1994. I have worked at CoGenTex, and at AT&T Labs -- Research, and for many years I was a research scientist at Columbia University in the Center for Computational Learning Systems.
This workshop synthesizes the latest research on the impact of AI usage in education so that you could make informed decisions on whether and how to use AI to facilitate your learning. You might have seen conflicting reports on whether the use of AI is good for learning. In this workshop, we are going to tease out, drawing on the latest research, which types of AI usage are beneficial or harmful for different kinds of learning. At the end of the workshop, you should walk away with more clarity on when and how to use AI for your own learning. Join PRODIG+ fellow on critical AI, Zheng Fu, in this informative workshop.
In this talk, I will explore how interpretable AI can bridge this gap, highlighting its potential to generate explicit, physically meaningful equations rather than opaque neural networks. Through four case studies from my lab, I will showcase how interpretable AI can enhance scientific understanding:
- Satellite Precipitation Retrieval: Using AI-based approaches to interpret precipitation retrieval algorithms from AMSU data, we identified critical microwave channels (89 and 150 GHz) that directly link to physical processes in the atmosphere.
- Quantitative Precipitation Estimation (QPE): By applying symbolic regression models to polarimetric radar data, we derived mathematical expressions that outperform traditional Z-R relationships and existing QPE algorithms, offering new insights into rainfall microphysics.
- Tornado Probability Prediction: Leveraging reinforcement learning-based symbolic deep learning models, we developed interpretable equations that outperform the traditional Significant Tornado Parameter (STP) index, providing a clearer understanding of the relationships between key atmospheric variables and tornado risk.
- Domain-Aware Symbolic Regression for Scientific Equations: In our latest work, we introduced a symbolic regression framework that incorporates domain-specific symbol priors extracted from thousands of scientific publications. By encoding common mathematical structures--such as the prevalence of trigonometric functions in physics or logarithmic forms in biology--into a tree-structured reinforcement learning model, we improved both the accuracy and interpretability of discovered equations. This approach accelerates convergence, enforces physical plausibility, and reveals new governing relationships in climate and geophysical data.
IACS Seminar Speaker: Yixin Wen, University of Florida
Location: IACS Seminar Room or Zoom
Join Zoom Meeting: https://stonybrook.zoom.us/j/97596399106?pwd=0PBvElFLqov3biO6OlQxSWLWudkIuH.1
Meeting ID: 975 9639 9106
Passcode: 096213
ABSTRACT: The key success of deep learning is the increasing size of models that can achieve high accuracy. At the same time, it is difficult to train the complex models with large data sets. Therefore, it is crucial to accelerate training with distributed systems and architectures, where communication and heterogeneity are two key challenges. In this talk, I will present two heterogeneity-aware decentralized training protocols without communication bottleneck. Specifically, Hop supports arbitrary iteration gap between workers by novel queue-based synchronization which can tolerate heterogeneity with system techniques. Prague uses randomized communication to tolerate heterogeneity with a new training algorithm based on partial reduce -- an efficient communication primitive. If time permits, I will present the systematic tensor partitioning for training on heterogeneous accelerator arrays (e.g., GPU/TPU). We believe that our principled approaches are crucial for achieving high-performance and efficient distributed training.
BIO: Xuehai Qian is an assistant professor at University of Southern California. His research interests include domain-specific systems and architectures, performance tuning and resource management of cloud systems and parallel computer architectures. He received his PhD from the University of Illinois Urbana Champaign and was a postdoc at UC Berkeley. He is the recipient of W.J Poppelbaum Memorial Award at UIUC, NSF CRII and CAREER Award, and the inaugural ACSIC (American Chinese Scholar In Computing) Rising Star Award.
SBUHacks is back for a second year! Join us for our 24-hour hackathon.
Location: Frank Melville Jr. Memorial Library, 100 Nicolls Rd, Stony Brook, NY 11794, USA
http://sbuhacks.org/ for more info.
Abstract:
In recent years, the landscape of artificial intelligence (AI) has been reshaped by the rapid emergence of Foundation Models (FMs). These versatile models have garnered widespread attention for their remarkable ability to transcend the boundaries of traditional, bespoke AI solutions and to generalize to a large set of downstream tasks. In this presentation we will describe the development of geospatial FMs with earth observation and weather data and discuss initial results of such models. We will also show how such foundation models can be a new and exciting tool for assisting with and accelerating scientific discovery.
Speaker:
Hendrik Hamann
Distinguished Researcher
IBM T.J. Watson Research Center