Abstract:

Many real world complex problems are multi-step reasoning tasks. These range from analytic tasks such as answering questions to automation tasks where agents complete tasks on behalf of users.. Evaluation, datasets, and models for such tasks can be unreliable for multiple reasons. (i) Datasets often have annotation artifacts and biases, allowing models to take reasoning shortcuts. Such shortcuts can allow models to make effective guesses -- or, in a sense, cheat -- to achieve high performance without any multi-step reasoning. This issue is further exacerbated for complex tasks because as the number of the required reasoning steps increases, so do the avenues for bypassing those steps. (ii) Models trained on such dataset/s learn to solve the task by taking reasoning shortcuts instead of proper multi-step reasoning. As a result, these models are not robust (reliable) when evaluated in an out-of-distribution evaluation setting. (iii) Lastly, recent works have shown that language models can solve complex multi-step tasks by producing a step-by-step explanation without any training. However, these methods often hallucinate factually incorrect (i.e., unreliable) explanations when posed with knowledge-intensive tasks.

I address these challenges by carefully characterizing the requirements of robust multi-step reasoning and designing reliable evaluation datasets and training methods that necessitate thorough multi-step reasoning. In DiRe, I first formalize and introduce Disconnected Reasoning, i.e., reasoning that allows models to arrive at the correct answer by bypassing necessary reasoning steps, and use this formalization to measure how much multi-step reasoning a model does on a dataset. In MuSiQue, I built a multi-step reasoning dataset for QA from scratch that avoids cheatability via disconnected reasoning, providing a more reliable evaluation. In TeaBReaC, I developed a synthetically generated multi-step QA pretraining dataset designed to force models to avoid disconnected reasoning and learn reliable multi-step reasoning. In IRCoT, I address the reliability of model-generated multi-step reasoning chains by interleaving models' step-by-step reasoning with a step-by-step retrieval from an external corpus, resulting in more factually correct reasoning. Finally, in AppWorld, I built a multi-step reasoning dataset that requires highly interactive problem-solving in an environment carefully designed to ensure models need thorough reasoning to succeed.
Speaker: Harsh Trivedi

Location: NCS 220 or Zoom

https://stonybrook.zoom.us/j/99096379762?pwd=zYCJZQVxRuZd9BboscO4nlodCwsKBr.1

Description:

As artificial intelligence and data science reshape the global information landscape, libraries are emerging as key players in both technological innovation and ethical stewardship. This international Zoom discussion brings together library professionals and educators from the U.S., Philippines, and Hong Kong to explore how institutions are integrating AI and data into their pedagogy and services.

Panelists will share concrete examples from their own libraries--ranging from data literacy initiatives to increasing discoverability. The conversation will also examine regional trends in librarianship, spotlighting how institutions in Asia are navigating the evolving role of data and AI.

Join us for a global conversation that highlights the transformative potential of libraries as hubs for innovation and critical inquiry in the age of AI.

Register for this free Zoom panel.

Panelists:

Ahmad Pratama is a Faculty Member and Associate Librarian at Stony Brook University Libraries, where he is working to build a comprehensive, campus-wide data literacy program within the Libraries. As the Data Literacies Lead, his work focuses on empowering students, faculty, and staff to critically and ethically engage with data and AI, including the development of a credit-bearing course in Critical Data & AI Literacies supported by an EDGE Fund Award from the Provost's Office. Previously, Dr. Pratama served as an Associate Professor of Information Technology, and his research and teaching explore the intersections of technology, policy, and society with a focus on data, AI, and innovation in higher education.

Dan Anthony Dorado is a full-time faculty member at the U.P. School of Library and Information Studies, where he teaches information technology, management and marketing, research methodology, and quantitative research. He was also the director of the Diliman Learning Resource Center under the Office of the Vice Chancellor for Student Affairs. Before that, he was an Information Specialist at the College of Engineering Library, in charge of the System and Network Administration and The Learning Commons. He completed his master's degree at the Technology Management Center in U.P. Diliman and is currently pursuing his PhD in Data Science. As a member of Sync.Bio.Optics laboratory and the Publics, Archives, and Data (PANDA) Lab, his research specialization covers Computational Methods, Open Education, Critical Data Studies, and Radical Statistics.

Ryun LEE is Associate University Librarian at The Chinese University of Hong Kong Library, leading Digital Initiatives and Library IT and Systems. He drives digital innovation through emerging technologies, particularly artificial intelligence to enhance services, streamline operations, and support CUHK's mission in research, education, and knowledge advancement. With a background in cataloging and digital repository development, Ryun leads projects in digitization, OCR, data visualization, text and network analysis, GIS, and digital scholarship. He actively promotes knowledge graph applications in Hong Kong studies and oversees efforts to digitize and preserve resources related to Hong Kong and Southern China. His recent work focuses on creating seamless digital experiences and developing data-driven infrastructure. He is currently exploring AI-driven approaches to digitization workflows and entity extraction, aiming to improve access, discovery, and long-term preservation of library materials.

CSE 600 Talk: Squeezing Software Performance via Eliminating Wasteful Operations presented by Xu Liu

ABSTRACT: Inefficiencies abound in complex, layered software. A variety of inefficiencies show up as wasteful memory operations, such as redundant or useless memory loads and stores. Aliasing, limited optimization scopes, and insensitivity to input and execution contexts act as severe deterrents to static program analysis. Microscopic observation of whole executions at instruction- and operand-level granularity breaks down abstractions and helps recognize redundancies that masquerade in complex programs. In this talk, I will describe various wasteful memory operations, which pervasively exist in modern
software packages and expose great potential for optimization. I will discuss the design of a fine-grained instrumentation-based profiling framework that identifies wasteful operations in their contexts, which guides nontrivial performance improvement. Furthermore, I will show our recent improvement to the profiling framework by abandoning
instrumentation, which reduces the runtime overhead from 10x to 3% on average. I will show how our approach works for native binaries and various managed languages such as Java, yielding new performance insights for optimization.

BIO: Xu Liu is an assistant professor in the Department of Computer Science at College of William & Mary. He obtained his PhD from Rice University in 2014 and joined the College of William & Mary in the same year. Prof. Liu works on building performance tools to pinpoint and optimize inefficiencies in HPC code bases. He has developed several open-source profiling tools, which are used worldwide at universities, DOE national laboratories and industrial companies. Prof. Liu has published a number of papers in high-quality venues. His papers received Best Paper Award at SC'15, PPoPP'18, PPoPP'19 and ASPLOS'17 Highlights, as well as Distinguished Paper Award at ICSE'19. His recent ASPLOS'18 paper has been selected as ACM SIGPLAN Research Highlights in 2019 and nominated for CACM Research Highlights. Prof. Liu is the receipt of 2019 IEEE TCHPC Early Career Researchers Award for Excellence in High Performance Computing. Prof. Liu served on the program committee of conferences such as SC, PPoPP, IPDPS, CGO, HPCA and ASPLOS.
TITLE: Sampling Using Langevin Diffusions Beyond the Worst-Case by Andrej Risteski of CMU


ABSTRACT: Many tasks involving generative models involve being able to sample from distributions parametrized as p(x) = e^{-f(x)}/Z where Z is the normalizing constant, for some function f whose values and gradients we can query. This mode of access to f is natural -- for instance sampling from posteriors in latent-variable models. Classical results show that a natural random walk, Langevin diffusion, mixes rapidly when f is convex. Unfortunately, even in simple examples, the applications listed above will entail working with functions f that are nonconvex.

We exhibit instances where Langevin diffusion (combined with other tools) can provably be shown to mix rapidly in instances of relevance in practice: distributions p that are multimodal, as well as distributions p that have a natural manifold structure on their level sets. 

Zoom Link: https://stonybrook.zoom.us/j/98533029054?pwd=5FXO6lWGTJssCADEYkYbA7sjaacPRX.1

Meeting ID: 985 3302 9054

Passcode: 436997

Abstract:

Semantic segmentation, the task of assigning a semantic label to each pixel in an image, is a fundamental problem in the field of Computer Vision. with crucial applications in domains like autonomous driving, drone imagery and medical image analysis. Despite advancements in deep learning architectures, state-of-the-art models still heavily depend on large-scale pixel-level annotations, which are costly and time-consuming to acquire. To address this issue, Semi-Supervised Segmentation (SSS) has emerged as a promising solution, leveraging a small set of labeled images alongside a larger corpus of unlabeled data to reduce the annotation burden. In this proposal, I aim to investigate the challenges of SSS and propose approaches to address them. Existing SSS methods rely on a teacher-student framework to generate pseudo-labels for unlabeled images, which are then used for model training. However, this approach presents two major challenges. Pixel-level consistency fails to effectively capture contextual information, and pseudo-labels are noisy, especially in the early stages of training. To address the challenge of noisy pseudo-labels, existing methods rely on confidence-based thresholding to identify reliable pseudo-labels. However, during early training phases, when the model is poorly calibrated, this approach can select high-confidence but noisy pseudo-labels. To address this, we propose a novel approach that reduces reliance on model confidence to select reliable pseudo-labels. Our method employs an ensemble of a segmentation model and an object detection model to select more reliable pseudo-labels, which are then used to weight pseudo-labels using rank statistics, reducing the influence of noisy labels in training. Next, to address both the challenge of capturing contextual information and noisy pseudo-labels I introduce a novel Multi-scale Patch-based Multi-label Classifier (MPMC), which incorporates patch-level contextual information and reduces the impact of noisy pixel pseudo-labels by using the predictions of the patch-level Multi-label classifier to detect noisy labels, enhancing overall segmentation performance. While my work so far has focused on effectively utilizing unlabeled data to improve segmentation performance, as part of our future work, I will explore the use of textual information, such as category descriptions, for segmentation tasks. In limited labeled data scenarios it is more challenging to align visual features with textual features from large language models (LLMs).

An interactive session to discover how to create ALT text tags from images and create high-impact visuals, from identification to communicating ideas with images.

Discover how to use AI to create ALT text from images as well as identify objects in your environment, and build relatable visuals for high-impact presentations. Images communicate ideas as a way to understand concepts. AI-generated images have helped allow anyone to create these.

In this session, you will

  1. Creating image ALT Tags
  2. Transform ideas into images that are visually appealing
  3. Identify objects from visuals

Register here.

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

HPCortex - a new, general-purpose machine learning library for HPC

Abstract: I will introduce HPCortex, a lightweight, C++, MPI-native machine-learning library for heterogeneous HPC systems. It implements many common architecture patterns including transformers, graph neural networks, and convolutional networks, and delivers performance portability across NVIDIA, AMD, and Intel GPUs while depending only on MPI and standard compiler/BLAS stacks. I will illustrate its capabilities via a surrogate model for the RHIC AGS Booster digital twin, a simple GNN for a coupled spring system, and a compact language model, then outline the roadmap.

Biography: Christopher is a research scientist and head of the Scientific Computing Applications Group in the Computational Science Department at Brookhaven National Laboratory. Previously he was an assistant staff scientist in the Physics Dept. at Columbia University, and held physics postdoctoral research positions at both Brookhaven and Columbia. He earned his Ph.D in Theoretical Physics from the University of Edinburgh, UK.
His scientific background is in lattice QCD and high performance computing, but since joining Brookhaven in 2020 his research interests have expanded to include machine learning, applied mathematics and performance analysis, with a particular emphasis on building tools to support scientific research on HPC systems.

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1604143373?pwd=hHT2yaIjahBIQ6tieURFqs8Pwex9gU.1

Meeting ID: 160 414 3373
Passcode: 277410

Synthetic Dreams and Ghost Machines

Dr. Steven Skiena will join the Cinema for a presentation exploring the rapidly evolving worlds of artificial intelligence and robotics.

Dr. Skiena is Distinguished Teaching Professor of Computer Science and Associate Director of the AI Innovation Institute at Stony Brook University. A leading researcher in data science and algorithms, he is the author of several influential books on AI and computation, including The Algorithm Design Manual and The Data Science Design Manual.

The presentation will be followed by a screening of Ex Machina, Alex Garland's provocative science-fiction thriller exploring the seductive and unsettling boundaries between human consciousness and artificial intelligence.

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Uncertainty-Aware Adaptation of LLMs for Protein-Protein Interaction Analysis

Abstract: Identification of protein-protein interactions (PPIs) helps derive cellular mechanistic understanding, particularly in the context of complex conditions such as neurodegenerative disorders, metabolic syndromes, and cancer. Large Language Models (LLMs) have demonstrated remarkable potential in predicting protein structures and interactions via automated mining of vast biomedical literature; yet their inherent uncertainty remains a key challenge for deriving reproducible findings, critical for biomedical applications. In this study, we present an uncertainty-aware adaptation of LLMs for PPI analysis, leveraging fine-tuned LLaMA-3 and BioMedGPT models. To enhance prediction reliability, we integrate LoRA ensembles and Bayesian LoRA models for uncertainty quantification (UQ), ensuring confidence- calibrated insights into protein behavior. Our approach achieves competitive performance in PPI identification across diverse disease contexts while addressing model uncertainty, thereby enhancing trustworthiness and reproducibility in computational biology. These findings underscore the potential of uncertainty-aware LLM adaptation for advancing precision medicine and biomedical research.

Biography: Sanket is a research staff member in the Applied Mathematics department within the Computing and Data Sciences Directorate at Brookhaven National Laboratory. Previously, he was the Amalie Emmy Noether Postdoctoral Fellow in the same department. He earned his Ph.D. in Statistics from Michigan State University.. Sanket's research interests span Bayesian statistics, uncertainty quantification (UQ), deep learning, Markov Chain Monte Carlo (MCMC), variational inference, and sparsity methods. He also focuses on dimensionality reduction, surrogate modeling, hybrid physical-data driven models, and active learning, with applications across climate science, materials science, and life sciences.

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1605691898?pwd=xC7GebG7Kvzxa4AjPSIxJw7e9IZtoY.1

Meeting ID: 160 569 1898
Passcode: 303888