Abstract: In high-dimensional data spaces, vast empty regions often exist where no known data points are present. These empty spaces are not merely gaps but hold untapped potential for discovering novel configurations, optimizing parameters, and improving decision-making processes. However, traditional exploration techniques struggle to identify and leverage these regions due to the curse of dimensionality. To address this, we introduce the Empty Space Search Algorithm (ESA), a scalable, physics-inspired method that systematically identifies and explores these uncharted voids. ESA operates by modeling the data space as a dynamic system, using a repulsion-attraction mechanism to locate optimal empty space configurations (ESCs) without requiring exhaustive search. Building upon ESA, we present GapMiner, a visual analytics system that integrates human-in-the-loop AI to iteratively refine and validate ESCs. GapMiner combines parallel coordinate visualization, interactive optimization, and deep learning-based predictive modeling to enhance the efficiency of empty space exploration. This methodology has broad applications, including accelerating convergence in evolutionary algorithms through a more diverse initial population, optimizing adversarial learning strategies, and discovering novel parameter configurations in reinforcement learning. Our approach demonstrates that empty space is not just an absence of data but a frontier for new possibilities in high-dimensional problem-solving.
Bio: Xinyu Zhang received his B.E. in Computer Science from Shandong University, Taishan College, in 2019. He is currently a final-year Ph.D. candidate in the Department of Computer Science at Stony Brook University, advised by Prof. Klaus Mueller. His research focuses on multivariate data analysis, scientific visualization, and reinforcement learning. He has published multiple papers in top-tier journals and conferences, including IEEE TVCG and NeurIPS.
*this seminar will be held in person (food provided on a first come, first serve basis), and online (zoom link below)!
Topic: IACS Student Seminar Speaker: Xinyu Zhang
Time: Feb 26, 2025 12:00 PM Eastern Time (US and Canada)
Join Zoom Meeting
https://stonybrook.zoom.us/j/91848218975?pwd=lfITFa61GaXZ2Wsa1B1OnbLQMmXvOE.1

Meeting ID: 918 4821 8975
Passcode: 027337
Date: March 11, 2022
Time: 2:40PM EST

Title: Towards Scalable and Efficient Machine Learning as a Service (MLaaS)

Abstract:
Driven by the explosive growth of big data, the sustained advances of
Machine Learning (ML), and the fast evolving of computer system
techniques, the past few years have witnessed a surging demand for
Machine Learning as a Service (MLaaS). MLaaS is an emerging computing
paradigm that facilitates ML model design, training, inference serving
and provides optimized executions of ML tasks in an automated,
scalable, and efficient manner. In this talk, I will demonstrate how
to integrate ML algorithm research and system research in synergy to
address the pressing challenges in MLaaS. I will first share a story
about how our system experience led to a novel large batching
algorithm design that revolutionizes large-scale training. Then I will
tell another story about how our gradient compression algorithm
research helped us to discover overlooked critical features of modern
ML systems and thereby build a compression-aware distributed ML
system. I will also briefly discuss a promising future of harnessing
serverless computing for MLaaS model inference serving. I will
conclude my talk with a discussion of interdisciplinary research and
future plans.

Bio:
Dr. Feng Yan is an Assistant Professor of Computer Science and
Engineering at University of Nevada, Reno (UNR) and director of the
Intelligent Data and Systems Lab (IDS Lab). Dr. Yan received M.S. and
Ph.D. degrees in Computer Science from the College of William and Mary
and worked at Microsoft Research and HP Labs. Dr. Yan's research
bridges the fields of big data, machine learning, and systems. The
focus of his research is on developing methodologies and building
systems that are automated, high-performing, efficient, robust, and
user-centric. Some of his recent research topics include large-scale
distributed deep learning, machine learning as a service (MLaaS),
federated learning, AutoML, serverless computing, and broad topics in
cloud and HPC. Dr. Yan is also dedicated to interdisciplinary research
and has established fruitful collaborations with domain experts in
areas such as health, physics, geography, material science, mechanical
engineering, civil engineering, and innovated big data and AI-driven
approaches for these domains. Dr. Yan and his team are actively
publishing at the most prestigious venues in computer system area
(such as SOSP, SC, HPDC, USENIX ATC, EuroSys, FAST, VLDB, etc.) and
machine learning area (such as NIPS/NeurIPS, KDD, AAAI, etc.). Dr. Yan
and his students are the recipients of the Best Student Paper Award of
IEEE CLOUD 2018, the Best Paper Award of CLOUD 2019, and the Best
Student Paper Award of ITNG 2021. Dr. Yan is the recipient of the NSF
CAREER Award, the NSF CRII Award, the CSE Best Researcher Award, and
has been nominated for the Regents' Rising Researcher Award. Dr. Yan
serves as Social Media Chair of ACM SIGMETRICS. To learn more
information, please visit Dr. Yan's homepage:
https://www.cse.unr.edu/~fyan.
Optimization and Machine Learning - presented by Yifan Sun

Abstract: Optimization is a growing topic of interest in the machine learning community. It starts out as an option to check in Tensorflow (SGD? Adam? Adagrad?), but as we get more into the how and why of these options, we uncover many fundamental principles relating to operations research, control theory, and dynamical systems, dating back as far as the Cold World era. 

In this talk I will give a broad overview of some of the important optimization themes in machine learning. I will try to give connections between tools we are used to seeing in popular packages 
and fundamental optimization concepts like duality, convexity, contractive operators, etc. While we cannot hope to completely cover this diverse research area, I hope to provide a glimpse of this exciting research area that is permeating more and more into the machine learning world. 

Bio: Yifan Sun received her PhD in Electrical Engineering from the University of California Los Angeles in 2015, with research focusing on convex optimization and semidefinite programming. She was then Technicolor Research and Innovation, focusing on machine learning and 
data science applications. More recently, she completed two postdocs focusing on optimization, at the University of British Columbia in Vancouver, Canada and INRIA, in Paris, France.

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

#1 How to train your Scientific Chatbot by Alexandr Prozorov, Post-Doctoral Research Associate


Abstract: RHIC is closing its 25-year run with ~1 EB of data and decades of hard-won know-how that risk drifting into obscurity. The RHIC Data & Analysis Preservation Plan (DAPP) pilots an AI assistant that lets physicists talk to RHIC in natural language--searching internal notes, code, workflows, and docs, and pointing to runnable, containerized analyses. Built on Retrieval-Augmented Generation(RAG) with a Model Context Protocol orchestration layer, the system indexes heterogeneous, experiment-specific content and enforces role-aware access
for public vs. collaboration-restricted materials. Takeaway: domain-adapted AI can turn a legacy exabyte into reproducible answers, training assets, and new discovery paths.

Biography: Alexandr Prozorov is a postdoc from Czech Technical University in Prague working in STAR experiment. Fascinated by AI

#2 Quantum AI: Atoms, Cavities and Learning by Raman Kumar, Post-Doctoral Research Associate, Instrumentation Department

Abstract: The Instrumentation Department (IO) in the Discovery Technologies directorate at BNL is engaged in exploring various aspects of quantum systems research. One of the main goals of our group's effort is in developing neutral atom-cavity array platforms for remote entanglement generation and distributed quantum processing. This platform promises to herald truly scalable quantum computing systems and open new paradigms for networking and sensing. In this talk, I will explain our group's research and the role AI is playing in unlocking new insights with two examples. The first application of AI is in fabrication process prediction of micro-cavity structures. The second application revolves around role of AI in quantum error detection and correction in modern quantum computing systems.

Biography: Dr. Raman Kumar is a postdoctoral research associate in the IO department at BNL working with Dr./Prof. Sebastian Will (Columbia U.). Kumar obtained his Ph.D. degree in Electrical and Computer Engineering from the University of Illinois Urbana-Champaign. Prior to joining BNL in Nov 2024, Kumar worked as a postdoc at the City College in New York working on topological photonic quantum sensing using NV centers in diamond. Kumar and Will combined have an extremely wide moat and expertise in a variety of different areas which include Ultra cold atoms and molecules, quantum optics, quantum condensed matter, nanofabrication, semiconductor devices and advanced electromagnetics. Their areas of research interest include scalable quantum computing, communications and sensing, all enabled by AI.

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting https://bnl.zoomgov.com/j/1607892208?pwd=MSjxN5btSeToZsQMwEQzCCbBo5h58V.1

Meeting ID: 160 789 2208
Passcode: 753871

Do Natural Language Understanding Systems Learn to Understand or to
Find Shortcuts? (Naoya Inoue, http://naoya-i.github.io/)

ABSTRACT: Recent studies have suggested that natural language understanding (NLU) systems learn to exploit superficial, task-unrelated cues (a.k.a. annotation artifacts) in current datasets. This prevents the community from reliably measuring the progress of NLU systems. In this talk, I will discuss two latest studies from our research team: (i) analysis of annotation artifacts in commonsense causal reasoning and (ii) creation of benchmark for evaluating NLU systems' internal reasoning.
---------------------------------------------------------------------------------------------------------------------------------------------
---------------------------------------------------------------------------------------------------------------------------------------------
Learning graph-structured sparse models (Baojian Zhou, https://baojianzhou.github.io/

ABSTRACT: Learning graph-structured sparse models has recently received significant attention thanks to their broad applicability to many important real-world problems. However, such models, of more effective and stronger interpretability compared with their counterparts, are difficult to learn due to optimization challenges. In this talk, we will discuss how to learn graph-structured sparse models under stochastic and online learning settings. Some interesting related problems will also be discussed.
Le Hou Dissertation Defense: Deep Learning for Digital Histopathology across Multiple Scales

ABSTRACT: Histopathology is the study of tissue changes caused by diseases such as cancer. It plays a crucial role in disease diagnosis, survival analysis and development of new treatments. Using computer vision techniques, I focus on multiple tasks for automated analysis in digital histopathology images, which are challenging because histopathology images are heterogeneous and complex, due to the large variation of hundreds of cancer types in gigapixel resolution. In this thesis, I show how histopathology image analysis tasks can be viewed in three scales: Whole Slide Image (WSI)-level, patch-level and cellular-level, and present my contributions in each resolution level.

BIO: WSI-level analysis such as classifying WSIs into cancer types is challenging, because conventional classification methods such as off-the-shelf deep learning models cannot be applied directly on gigapixel WSIs due to computational limitations. I contribute a patch-based deep learning method that classifies gigapixel WSIs into cancer types and subtypes with close-to-human performance. This method is useful for computer-aided diagnosis. At patch-level, I contribute a novel method for histopathology image patch classification. On the task of identifying Tumor Infiltrating Lymphocyte (TIL) regions, the prediction result of this method correlates to the survival rate of patients. At cellular-level, I contribute novel methods for nucleus classification and roundness regression, which are interpretable features for histopathology studies. With this method, I generated a large-scale dataset of segmented nuclei, in WSIs from a large publicly available digital histopathology image dataset, to help advance histopathology research.
The 20th International Conference on Emerging Technologies for a Smarter World (CEWIT 2025)

The Innovation Edge: Harnessing AI for the Future
Exploring Generative AI, Agentic AI, and Frontier Technologies Revolutionizing Healthcare, Defense, Energy, FinTech, and Beyond

Organized by the New York State Center of Excellence in Wireless and Information Technology (CEWIT) at Stony Brook University, our international conference is a destination for researchers, innovators and entrepreneurs, across borders and disciplines. CEWIT2023 conference attracted over 150 industry and academic participants worldwide. Over twenty-three presenters took the podium in breakout sessions and engaging panel discussions.

Continuing the tradition since the inception of our conference in 2003, CEWIT2025 will be a premier forum for presentations of cutting-edge research as well as the exchange and transfer of emerging technologies and innovative applications. We are expecting renowned speakers, presenters and panelists from industry, academia and government, beginning with a series of plenary presentations & a keynote, and followed by several conversational panels - all for an audience ready to network!


Location: The Center of Excellence in Wireless and Information Technology (CEWIT), Stony Brook University

Event Details: Visit CEWIT2025 site to learn more about the event

Questions/Concerns: CEWIT Conference Team at 631-216-7114 or info@cewit.org