The Natural Language Processing Reading Group at Stony Brook University meets weekly to discuss recent research papers in NLP and related fields.
Join the Google Group here.

Title: Don't Just Fix it in Post: A Science of AI Must Study Training Dynamics

Abstract: What would it mean to have a scientific understanding of AI? Models are not static objects: they are snapshots of time-evolving processes shaped by data, objectives, architectures, and optimization dynamics. Yet much of AI research treats models as fixed artifacts, analyzing behaviors after training rather than asking why they emerge. This position paper argues that a science of AI must move beyond post hoc fixes and study the training dynamics that produce model behavior. Such a science should support progressively stronger forms of understanding: predicting outcomes from early training signals, intervening when trajectories go wrong, and ultimately designing training procedures that more reliably produce desired properties. Scaling laws have made prediction routine for loss; the challenge is extending this success to capabilities, biases, robustness, and safety-relevant behaviors. We articulate requirements for such theories grounded in the history and philosophy of science, examine progress in mechanistic interpretability, fairness, memorization, and simplicity bias, and identify concrete open problems.

Location: NCS 220


Zoom Link: https://stonybrook.zoom.us/j/92942833294?pwd=927683
Abstract: Current Chain-of-Thought (CoT) verification methods predict reasoning correctness based on outputs (black-box) or activations (gray-box), but offer limited insight into \textit{why} a computation fails. We introduce a white-box method: \textbf{Circuit-based Reasoning Verification (CRV)}. We hypothesize that attribution graphs of correct CoT steps, viewed as \textit{execution traces} of the model's latent reasoning circuits, possess distinct structural fingerprints from those of incorrect steps. By training a classifier on structural features of these graphs, we show that these traces contain a powerful signal of reasoning errors. Our white-box approach yields novel scientific insights unattainable by other methods. (1) We demonstrate that structural signatures of error are highly predictive, establishing the viability of verifying reasoning directly via its computational graph. (2) We find these signatures to be highly domain-specific, revealing that failures in different reasoning tasks manifest as distinct computational patterns. (3) We provide evidence that these signatures are not merely correlational; by using our analysis to guide targeted interventions on individual transcoder features, we successfully correct the model's faulty reasoning. Our work shows that, by scrutinizing a model's computational process, we can move from simple error detection to a deeper, causal understanding of LLM reasoning.

Speaker: Xianjun Yang

Location: Old Computer Science Building - CS2311

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Abstract: This presentation will begin by outlining key challenges facing the modern power grid and summarizing our group's research efforts to address them. It will then discuss how AI and machine learning are reshaping the grid modernization. The major focus of the talk will highlight a range of AI/ML applications we have developed in recent years to enhance grid operation, planning, control, and security.

Biography: Meng Yue is currently leading the Grid Modernization and Security Group in the Interdisciplinary Science Department at Brookhaven National Laboratory (BNL). He received his Ph. D. from Michigan State University in electrical engineering. His major research interests include power system modeling, simulation, and control, and applications of AI/ML- and quantum machine learning and quantum computing in operation, planning, and security of the future grid.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1604383624?pwd=ffQ5cUPNxTI7nzClKQO6cnsNbhF9Vf.1

Meeting ID: 160 438 3624
Passcode: 558449

This is Stony Brook's quantum moment. Join us for a spotlight on the core achievements and research excellence of faculty across the Colleges of Arts and Sciences (CAS), and Engineering and Applied Sciences (CEAS) - and their collaborative advancements in quantum science and technology. Learn about the real world impact of their enduring work, their leadership in translating foundational science into entrepreneurial opportunities, and their impetus for making connections to next generation innovation.

Presented by: Catherine Chen, Ph.D., Research Development Associate

Welcome remarks: President Andrea Goldsmith

Panel moderators: Dean David Wrobel, CAS, and Dean Andrew Singer, CEAS

Presentations and panel featuring our faculty:

  • Jennifer Cano, CAS, Physics and Astronomy

  • P. Scott Carney, CEAS, Mechanical Engineering

  • Hyeongrak Chuck Choi, CEAS, Electrical and Computer Engineering

  • Eden Figueroa, CAS, Physics and Astronomy

  • Humanshu Gupta, CEAS, Computer Science

  • Angela Kelly, CAS, Physics and Astronomy

Location: Theatre at the Charles B. Wang Center, Stony Brook University

Reserve your tickets by March 26!



Place:  https://stonybrook.zoom.us/j/99167126152?pwd=TFpEYzM0aFhiOFJxSFJEb1JSS3YyQT09  

Time: 3 PM EST - Dec, 16th, 2020 

Abstract: 

Shadows provide useful cues to analyze visual scenes but also hamper many computer vision algorithms such as image segmentation, object detection, or tracking. For those reasons, shadow detection and shadow removal have been well-studied in computer vision.

Early work on shadow detection and removal focused on physical illumination models of shadows. These methods can express, identify, and remove shadows in a physically plausible manner. However, these models are often hard to optimize and are slow during inference due to their reliance on hand-designed image features. Recently, deep-learning approaches have achieved breakthroughs in performance for both shadow detection and removal. They learn to extract useful features through training while being extremely efficient during inference. However, these models are data-dependent, opaque, and ignore the physical aspects of shadows. Thus they often lack generalization and produce inconsistent results.

We propose incorporating physical illumination constraints of shadows into deep-learning models. These constraints force the networks to more closely follow the physics of shadows, enabling them to systematically and realistically modify shadows in images. For shadow detection, we present a novel Generative Adversarial Network (GAN) based model where the generator learns to generate images with realistic attenuated shadows that can be used to train a shadow detector. For shadow removal, we propose a method that uses deep-networks to estimate the unknown parameters of a shadow image formation model that removes shadows. The system outputs high-quality shadow-free images with little or no image artifacts and achieves state-of-the-art performance in shadow removal when trained on a fully-supervised setting. Moreover, the system is easy to train and constrain since the shadow removal mapping is strictly defined by the simplified illumination model with interpretable parameters. Thus, it can be trained even with a much weaker form of supervision signal. In particular, we show that we can use two sets of patches, shadow and shadow-free, to train our shadow decomposition framework via an adversarial system. These patches are cropped from the shadow images themselves.
Therefore, this is the first deep-learning method for shadow removal that can be trained without any shadow-free images, providing an alternative solution to the paired data dependency issue. The advantage of this training scheme is even more pronounced when tested on a novel domain such as video shadow removal where the method can be fine-tuned on a testing video with only the shadow masks generated by a pre-trained shadow detector and further improves shadow removal results.
Abstract: Neurosymbolic AI, as the combination of deep learning and knowledge representation/reasoning methods, is currently receiving significant attention due to its promise to both overcome apparent limitations of pure deep learning systems, and to address the knowledge acquisition bottleneck. This presentation will focus on some recent reseach advances made at Kansas State University related to Semantic Web / knowledge graphs and neurosymbolic AI, with particular emphasis on the use of symbolic AI methods for understanding hidden neuron activations, and on using large language models for knowledge graph and ontology engineering.

Short bio: Pascal Hitzler is University Distinguished Professor and endowed Lloyd T. Smith Creativity in Engineering Chair at the Department of Computer Science at Kansas State University, one of the Directors of the Institute for Digital Agriculture and Advanced Analytics (ID3A), and Director of the Center for Artificial Intelligence and Data Science (CAIDS). His research record lists over 400 publications in such diverse areas as neurosymbolic artificial intelligence, semantic web, knowledge graphs, knowledge representation and reasoning, denotational semantics, and set-theoretic topology. He was founding Editor-in-chief of the Semantic Web journal, the leading journal in the field, and is founding Editor-in-chief of the new Neurosymbolic Artificial Intelligence journal. He is co-author of the W3C Recommendation OWL 2 Primer, and of the book Foundations of Semantic Web Technologies by CRC Press, 2010, which was named as one out of seven Outstanding Academic Titles 2010 in Information and Computer Science by the American Library Association's Choice Magazine, and has translations into German and Chinese. He is a founding steering committee member of the Neural-Symbolic Learning and Reasoning Association that runs the annual International Conference on Neurosymbolic Learning and Reasoning (NeSy). For more information about him, see http://www.pascal-hitzler.de.

Location: NCS 120