Abstract:

Many real world complex problems are multi-step reasoning tasks. These range from analytic tasks such as answering questions to automation tasks where agents complete tasks on behalf of users.. Evaluation, datasets, and models for such tasks can be unreliable for multiple reasons. (i) Datasets often have annotation artifacts and biases, allowing models to take reasoning shortcuts. Such shortcuts can allow models to make effective guesses -- or, in a sense, cheat -- to achieve high performance without any multi-step reasoning. This issue is further exacerbated for complex tasks because as the number of the required reasoning steps increases, so do the avenues for bypassing those steps. (ii) Models trained on such dataset/s learn to solve the task by taking reasoning shortcuts instead of proper multi-step reasoning. As a result, these models are not robust (reliable) when evaluated in an out-of-distribution evaluation setting. (iii) Lastly, recent works have shown that language models can solve complex multi-step tasks by producing a step-by-step explanation without any training. However, these methods often hallucinate factually incorrect (i.e., unreliable) explanations when posed with knowledge-intensive tasks.

I address these challenges by carefully characterizing the requirements of robust multi-step reasoning and designing reliable evaluation datasets and training methods that necessitate thorough multi-step reasoning. In DiRe, I first formalize and introduce Disconnected Reasoning, i.e., reasoning that allows models to arrive at the correct answer by bypassing necessary reasoning steps, and use this formalization to measure how much multi-step reasoning a model does on a dataset. In MuSiQue, I built a multi-step reasoning dataset for QA from scratch that avoids cheatability via disconnected reasoning, providing a more reliable evaluation. In TeaBReaC, I developed a synthetically generated multi-step QA pretraining dataset designed to force models to avoid disconnected reasoning and learn reliable multi-step reasoning. In IRCoT, I address the reliability of model-generated multi-step reasoning chains by interleaving models' step-by-step reasoning with a step-by-step retrieval from an external corpus, resulting in more factually correct reasoning. Finally, in AppWorld, I built a multi-step reasoning dataset that requires highly interactive problem-solving in an environment carefully designed to ensure models need thorough reasoning to succeed.
Speaker: Harsh Trivedi

Location: NCS 220 or Zoom

https://stonybrook.zoom.us/j/99096379762?pwd=zYCJZQVxRuZd9BboscO4nlodCwsKBr.1
Abstract: Datalog is a powerful language for expressing recursive computations through rules: Horn clauses in first order logic. Although effective at expressing queries over existential properties, Datalog and many of its popular implementations struggle with queries that involve more complex aggregates, requiring users to apply verbose, non-composable, and/or inefficient workarounds. Recent work on lattice-based datalogs addresses many of these concerns for aggregates that can be encoded as lattices (e.g., min or max), but more general aggregates like count remain problematic. In this talk, I will argue that this is not a fundamental limitation of Datalog, but rather from its model of truth: Both datalog semantics and evaluation rules make heavy use of the fact that insertion is both monotone and idempotent. Once a fact is known to be true, it can not be retracted, nor can further discoveries of the same fact alter its truth. Monotonicity is critical for forward progress under Datalog's ``open world'' model, as it allows us to safely assert the truth of a body. Meanwhile, idempotence makes it easier to reason about evaluation, as we need only guarantee that each head atom will be derived at-least-once. Unfortunately, more general aggregates like sum() are neither idempotent, nor monotone. I will introduce Hedgelog, a strict generalization of Datalog that uses general monoids as a basis for truth. I will show that this generalization remains compatible with Datalog's open world model, how it enables cleaner and more composable datalog programs, and how the underlying monoid relations open the door to interesting datastructure-level optimizations.

Bio: Oliver Kennedy is an associate professor at the University at Buffalo. He earned his PhD from Cornell University in 2011 and now leads the Online Data Interactions (ODIn) lab, which operates at the intersection of databases and programming languages. Oliver is the recipient of an NSF CAREER award, an IEEE Region 1 Technological Innovation Award, UB's Exceptional Scholar Award, and several UB SEAS teaching awards. Oliver is also one of the founding board members of Breadcrumb Analytics. Several of Oliver's papers have been invited to Best of compilations from SIGMOD and VLDB. The ODIn lab is currently exploring (i) how we can leverage database techniques like incremental view maintenance to make compilers faster, (ii) how to make it easier for data scientists to track how sources of uncertainty, ambiguity, and/or bias affect analyses, and (iii) how to streamline the interfaces --- both human and software --- between different tools for data science, like python, sql, and spreadsheets.

Location: NCS 120
The Future Histories Studio welcomes Moontae Lee, LG AI Research.


Generative AI is transforming how we understand, create, and interact with information. Large Language Models (LLMS) comprehend contexts, answer non-trivial questions, and spark creative ideas. This talk introduces the evolution of these models, highlighting the most recent advancements in planning, reasoning, and evaluation. The talk also touches on the criticalconsiderations for both model developers and users, carefully addressing limitations of LLMs as well as ethical and societal implications. Finally, the talk provides ongoing directions in researchand production: from the rise of personalized AI agents to the future frontiers of AI.

Moontae Lee is the Director of the Superintelligence Lab at LG AI Research and an Assistant Professor of Information and Decision Sciences at the University of Illinois Chicago. His journey with Large Language Models began as a visiting scholar at Microsoft Research in 2019, continuously consulting the Deep Learning Group at Redmond until joining LG. He holds a PhD in Computer Science from Cornell, an MS from Stanford, and BS degrees in Computer Science, Mathematics, and Psychology from Sogang University. He has been an area chair for major AI conferences and earned recognition in Operations Research and Computational Social Science, including awards from INFORMS and Amazon.

His research interests include:
● Computational Creativity, Algorithmic Awareness
● Retrieval-Augmented Generation and Evaluation
● Code Generation, Reasoning, Planning
● Fine-grained Alignment from Human/AI Feedback in Generative AI
● Large Time-series Models, Diffusion/Consistency
● Machine Unlearning
● Ranking Monopoly, Voting Fairness
● AI Safety, Ethics, and Market Impacts

Join us in person @ Future Histories Studio Staller Center for the Arts, 4222
The annual conference on Neural Information Processing Systems is a multi-track interdisciplinary annual meeting that includes invited talks, demonstrations, symposia, and oral and poster presentations of refereed papers. Along with the conference is a professional exposition focusing on machine learning in practice, a series of tutorials, and topical workshops that provide a less formal setting for the exchange of ideas.

For more information and registration, visit the official website.
What AI tools are available to help with the scholarly research process? Are they helpful? What do they do and is it worth the time and energy to try them out? Join librarian Christine Fena to explore and compare established and emerging AI research tools such as Elicit, Scite, Consensus, and Undermind. The online workshop will provide a starting point to understanding what these tools are, the basics of how they work, and how AI research assistants might bring changes to your search process in the future. All are welcome!



Register here via Zoom.
The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 10 minutes) by multiple people. Students can register for 1 credit for CSE 656. Registered students must attend and present a minimum of 2 or 3 talks. Everyone else is welcome to attend. Fill in https://forms.gle/pCVXovgfMfQwGqG38 to subscribe to our mailing list for further announcement.

What AI tools are available to help with the scholarly research process? Are they helpful? What do they do and is it worth the time and energy to try them out? Join librarian Christine Fena to explore and compare established and emerging AI research tools such as Elicit, Scite, Consensus, and Undermind. The workshop will not offer a lengthy tutorial on how to use any of these tools, but will provide a starting point to understanding what they are, what new ones are emerging, and how AI research assistants might bring changes to your search process. All are welcome!

Register for this Zoom workshop.

Title: Cyberinfrastructure for forward prediction and inversion estimation with uncertainty quantification

Seminar Speaker: Dr. Mengyang Gu, Assistant Professor, Department of Statistics and Applied Probability, University of California, Santa Barbara

Abstract: In this talk, we introduce four useful tools for forward prediction and inversion estimation. The first tool is the parallel partial Gaussian process surrogate model for emulating expensive computer simulations with massive coordinates. The tool is implemented in the RobustGaSP package available in R, MATLAB, and Python, for predicting both scalar- and vector-valued outputs with uncertainty assessment. The second tool is implemented in the RobustCalibration package, which handles Bayesian data inversion or model calibration by one or multiple types of experimental observations. A unique feature of the package is the inclusion of fast surrogate models of both scalar- and vector-valued computer simulations that bypass the expensive simulation in one line of code. The third tool is implemented in the AIUQ package, available in both R and MATLAB. In this approach, we show that differential dynamic microscopy, a scattering-based analysis tool that extracts dynamical information from microscopy videos, is equivalent to fitting the temporal auto-covariance in Fourier space, based on a latent factor model we construct. We develop a more efficient estimator and reduce the computational cost to pseudolinear order with respect to the number of observations without approximation, by utilizing the generalized Schur algorithm for the Toeplitz covariance. In the last tool, we developed a new method called the inverse Kalman filter, which enables fast matrix-vector multiplication between a covariance matrix from a dynamic linear model and any real-valued vector with a linear computational cost. These new approaches outline a wide range of applications that include emulating expensive simulation at molecular-, meso- and macro-scales, active learning with error control, nonparametric estimation of particle interaction functions, and data inversion from microscopy and velocity fields.

Join Zoom Meeting: https://bnl.zoomgov.com/j/1606285496?pwd=2yJYSG6lx8gMPiibzgAIBQtKHIjuHV.1
Meeting ID: 160 628 5496
Passcode: 472506