University Libraries Present: Qualitative data can be challenging to analyze and interpret effectively. In this workshop, SBU Libraries' Data Literacies Lead, Ahmad Pratama will show you how to extract meaningful insights from textual data, including understanding sentiment trends. Learn to explore qualitative data with Python using word clouds, basic natural language processing (NLP) techniques, and lexicon-based sentiment analysis with VADER.
https://stonybrook.zoom.us/meeting/register/k0r6mPYCRayk2AOGmyd0qw#/registration
Abstract: Many foundation models for digital pathology have been released recently. Benchmarking available methods then becomes paramount to get a clearer view of the research landscape. For this reason, we introduce THUNDER, a tile-level benchmark for digital pathology foundation models, allowing for efficient comparison of many models on diverse datasets with a series of downstream tasks, studying their feature spaces and assessing the robustness and uncertainty of predictions informed by their embeddings. Such foundation models are often used as feature extractors and combined with Multiple Instance Learning (MIL) aggregators at downstream time. Such aggregation must be efficient and reliable. We will focus on two specific examples of this: (I) HistAug, a fast and efficient generative model for controllable augmentations in the latent space of foundation models to perform data augmentation for MIL, and (ii) CAR-MIL, a method based on counterfactual attention regularisation to improve the reliability of attention maps of MIL methods.

Short-bio: Pierre Marza is a Postdoctoral Researcher at CentraleSupelec in the Biomathematics team of the MICS lab, studying Computer Vision and Deep Learning for Medical Imaging, with a focus on Digital Pathology. Prior to this, he was a PhD student at INSA Lyon, in the LIRIS and CITI labs, advised by Christian Wolf, and co-advised by Laetita Matignon and Olivier Simonin. He studied Visual Navigation, Embodied AI, Spatial Reasoning, more specifically how to learn to represent 3D space, generalize to new environments and master diverse tasks from light supervision.

Location: NCS 220

Zoom: https://stonybrook.zoom.us/j/94798224254?pwd=CFraer25qnpORbJ14aAVHRwaSJOjJM.1
Abstract: Current Chain-of-Thought (CoT) verification methods predict reasoning correctness based on outputs (black-box) or activations (gray-box), but offer limited insight into \textit{why} a computation fails. We introduce a white-box method: \textbf{Circuit-based Reasoning Verification (CRV)}. We hypothesize that attribution graphs of correct CoT steps, viewed as \textit{execution traces} of the model's latent reasoning circuits, possess distinct structural fingerprints from those of incorrect steps. By training a classifier on structural features of these graphs, we show that these traces contain a powerful signal of reasoning errors. Our white-box approach yields novel scientific insights unattainable by other methods. (1) We demonstrate that structural signatures of error are highly predictive, establishing the viability of verifying reasoning directly via its computational graph. (2) We find these signatures to be highly domain-specific, revealing that failures in different reasoning tasks manifest as distinct computational patterns. (3) We provide evidence that these signatures are not merely correlational; by using our analysis to guide targeted interventions on individual transcoder features, we successfully correct the model's faulty reasoning. Our work shows that, by scrutinizing a model's computational process, we can move from simple error detection to a deeper, causal understanding of LLM reasoning.

Speaker: Xianjun Yang

Location: Old Computer Science Building - CS2311
CSE 600 Seminar Series | Fall 2025

Abstract: Imagine machines that can see the invisible: drones locating wildfire survivors, cameras predicting building failures, and smartphones detecting skin tumors. These applications lie beyond today's vision systems, which focus only on human-visible information. In this talk, I argue that a wealth of scene information is hidden in light properties invisible to the human eye, such as the travel time of photons and polarization of light waves. I will present how co- designing camera hardware, graphics models, and learning algorithms unlocks these invisible properties to create superhuman vision systems. I will present three superhuman vision capabilities: seeing around blind corners, turning objects into cameras, and extracting internal stress fields. By analyzing faint light reflections on diffuse walls and shiny objects, we create virtual cameras that reveal scenes hidden from the line of sight - enabling autonomous systems to navigate safely. Using the polarization of light, we recover mechanical stress fields hidden inside objects - opening new possibilities for non-destructive material characterization. These capabilities point toward a future where machines can see the invisible: around us, beneath our bodies, and beyond our scientific understanding.

Bio:
Akshat Dave is an Assistant Professor in the Department of Computer Science at Stony Brook
University, USA. His research lies at the intersection of applied optics, computer vision, and
machine learning. His work has been recognized by Rice University's Best Thesis Award, Optica Best Paper Prize, SIGGRAPH Asia Doctoral Consortium, and fellowships by Qualcomm, Texas Instruments, and INK Global Foundation. Prior to Stony Brook, he was a Postdoctoral Associate at MIT Media Lab. He holds a Ph.D. from Rice University and a Masters and a Bachelors from Indian Institute of Technology Madras.
Abstract: Retrieval-augmented generation (RAG) systems empower large language models (LLMs) to access external knowledge during inference. Recent advances have enabled LLMs to act as search agents via reinforcement learning (RL), improving information acquisition through multi-turn interactions with retrieval engines. However, existing approaches either optimize retrieval using search-only metrics (e.g., NDCG) that ignore downstream utility or fine-tune the entire LLM to jointly reason and retrieve--entangling retrieval with generation and limiting the real search utility and compatibility with frozen or proprietary models. In this work, we propose s3, a lightweight, model-agnostic framework that decouples the searcher from the generator and trains the searcher using a Gain Beyond RAG reward: the improvement in generation accuracy over naïve RAG. s3 requires only 2.4k training samples to outperform baselines trained on over 70 × more data, consistently delivering stronger downstream performance across six general QA and five medical QA benchmarks.

Speaker: Peter Zeng

Location: CS2311

Abstract:

Conventional approaches to scientific discovery often prioritize building larger sensors, gathering more data, and scaling up computational power. In this talk, I will present a complementary perspective: extracting insights hidden in the data we already have. The key lies in using AI not as a black-box predictor, but as a tool for interpreting data through its underlying physical process.

I will demonstrate how AI, when integrated with the physics of light propagation, can serve as a computational lens to overcome fundamental limitations in fields ranging from biomedicine to astrophysics. Specifically, I will showcase two compelling applications: non-invasive imaging through scattering biological tissues, and detecting faint exoplanets against the overwhelming brightness of their host stars.

These methods represent a departure from traditional learning-based approaches that rely on fitting models to training labels and hoping for generalization. Instead, with physics-informed strategies that decode how light propagates, we can transform raw measurements into scientifically meaningful insights--without requiring costly hardware upgrades or human-annotated datasets. Finally, I will outline future directions for combining AI with physical principles, enabling us to unlock more phenomena once considered hidden and accelerating discoveries in healthcare, astronomy, and beyond.

Short Bio:

Brandon Y. Feng is a Postdoctoral Associate at MIT Computer Science and Artificial Intelligence Laboratory (CSAIL) and a Visiting Scientist at the Harvard-Smithsonian Center for Astrophysics. His research bridges artificial intelligence and physics to expand the limits of human and machine vision. He develops AI-driven methods that reveal hidden patterns in complex visual data, driving breakthroughs in areas such as exoplanet detection and imaging through scattering tissues. His work has been published in top venues, including Science Advances, CVPR, ICCV, ECCV, and NeurIPS, and has been featured in Science.org, New Scientist, and Phys.org. He holds a Ph.D. in Computer Science from the University of Maryland, along with a B.A. in Computer Science and Statistics and an M.S. in Statistics from the University of Virginia.

Location: NCS 220

Over the past decade, researchers in neuroscience, psychology and artificial intelligence have come together to build advanced computer models that mimic how our brain processes what we see. These models are designed to closely copy the brain's visual system, all the way to a key area called the inferior temporal cortex, which plays an important role in recognizing objects.

Because these computer models can be fully observed, scientists can use them to make detailed predictions about how the brain works -- something older, more theoretical models could not do.

Dr. James DiCarlo's work explores whether these computer digital twin models of the brain could help guide safe, non- invasive ways to infl uence brain activity. In his talk, he explains how such a model could be used to design specific patterns of light. When this carefully designed light is added to what the eye naturally sees, it can precisely influence activity in groups of neurons in the inferior temporal cortex.

Since neural activity in this visual brain area may be connected to emotional states like anxiety, this research could eventually open the door to non-invasive approaches that may benefit mental well-being in the future.

Speaker: James J. DiCarlo, MD, PhD, Peter de Florez Professor, MIT Brain and Cognitive Sciences, and Director, MIT Siegel Family Quest for Intelligence

Location: Staller Center Main Stage

The event will be livestreamed at stonybrook.edu/live



New York Scientific Data Summit (NYSDS) is a premier annual conference that brings together researchers and thought leaders from academia, national labs and industry to exchange ideas and foster collaboration focused on data-driven science and technology. Co-hosted by Brookhaven National Laboratory and the Institute for Advanced Computational Science (IACS) at Stony Brook University, NYSDS 2025 will take place on September 11-12, 2025, in the SUNY Global Center in New York City.

NYSDS 2025 will spotlight artificial intelligence (AI), machine learning (ML) and robotics - fields currently at a pivotal point with transformative impacts on science and technology. From accelerating computationally demanding simulations to discerning signals from noisy data, AI/ML has become an integral part of the scientific workflows. Despite many advances, challenges remain to ensure that AI/ML applications are reliable, explainable and trustworthy.

Robotics, a growing field that couples AI with physically actuated mechanical bodies, has seen increased interest in areas spanning science, technology and manufacturing. The need for real-time decision-making and control, along with the intricate morphology of robots, makes robotics an intriguing application of AI, advanced computing and optimization.


This NYSDS 2025 is open to the public. To be eligible to attend, all participants must register online by August 30, 2025. For questions or assistance with registering, please contact the Summit Coordinator.

Register here.

Ready for Round Two? Dr. Zach Justus Returns! Join us on October 30, 2025, in the SBU Hilton Garden Inn. Buckle up your curiosity for a high-energy morning session with the engaging Dr. Zach Justus as we navigate how GenAI is reshaping not just how we teach, but what we teach. With real talk and questions that hit hard like Are students learning what we think we're teaching? This is your chance to rethink your program's true destination. Whether you're looking to pick up a few takeaways or chart a new direction entirely, this symposium is your space to explore, reflect, and act.

Check-in and breakfast will begin at 8:30 a.m. in order to begin our program promptly at 9:00 a.m.

Registration will remain open until October 15 or until the event reaches capacity. If closed, please contact educationaleffectiveness@stonybrook.edu to request a spot on the waitlist.

Event Website: bnl.gov/nysds
Dates: September 28-29, 2026
Location: SUNY Global Center, New York, NY
Co-hosts: Brookhaven National Laboratory, the Institute for Advanced Computational Science (IACS), and the AI Innovation Institute at Stony Brook University.

Join us for this premier annual conference that brings together researchers and thought leaders from academia, national labs, and industry to exchange ideas and foster cross-disciplinary collaboration centered on data-driven science and technology.

The theme of NYSDS 2026 is Transformational AI from Science to Society. This year's conference will focus on the ways in which artificial intelligence (AI), machine learning (ML) and robotics are impacting everything from our everyday lives to scientific endeavors. NYSDS2026 will feature the following main tracks:

  • Robotics and Embodied AI: advances in perception, control, learning and interaction for autonomous physical agents in real-world and scientific environments.

  • AI for Science: innovative applications ranging from the physical sciences to biology and medicine, and discussion of important ways in which AI is changing the nature of science.

  • AI for Energy: how energy is discovered and produced, how hazards are mitigated to ensure reliability of the power grid, and our search for new sources of critical minerals and materials.

  • AI for Risk Assessment: inventive uses of large amounts of data and the AI tools to assess and mitigate risks in national security, climate, and financial markets.

  • AI for Education: the ways in which education is being reimagined in the age of AI and the roadblocks to training the next generation of students.

Each track will include invited presentations, contributed talks, posters, and panel discussions.