Abstract: Visual generation is a fundamental problem in computer vision and graphics, with applications ranging from 3D capture to image and video synthesis. Despite rapid progress in neural rendering and generative models, efficiency remains a key bottleneck: high-quality 3D reconstruction often relies on dense multi-view supervision; high-fidelity 3D synthesis requires costly optimization, training, and rendering; and modern image and video generators require substantial computation as the number of tokens grows rapidly for high-resolution generation. This dissertation focuses on efficient visual generation by improving sample efficiency in 3D reconstruction, representation efficiency in 3D generation, and computational efficiency in image and video synthesis. First, we improve the sample efficiency of neural implicit surface reconstruction. We integrate multi-view stereo probability volumes as a geometric regularizer, enabling high-quality sparse-view reconstruction from as few as three input images. Next, we introduce an explicit 3D representation for 3D generation, built from multi-view depth and RGB images with 3D Gaussian features. This design allows the model to use 2D generative priors while enforcing multi-view consistency through epipolar attention. We then address the computational bottleneck in image and video synthesis with importance-based token merging. Our method uses importance signals available during generation to preserve critical information while merging redundant tokens. Finally, we enable efficient mixed-resolution diffusion transformers via phase-aligned attention. This approach stabilizes attention under mixed-resolution token grids and unlocks high-fidelity image and video generation at reduced cost. Taken together, these contributions reduce the data requirements, representational overhead, and computational demands, thereby providing a foundation for high-quality, scalable, and efficient visual generation.

Speaker: Haoyu Wu

Location: NCS 120

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes one short talk on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Abstract: Two-dimensional (2D) materials such as graphene, hBN, and TMDs offer atomically sharp interfaces and unprecedented tunability when vertically assembled into van der Waals heterostructures. These stacks have enabled discoveries ranging from moiré superconductivity and correlated insulators to quantum emitters and next-generation nanoelectronic devices. Yet constructing high-quality heterostructures remains largely artisanal: researchers manually identify exfoliated flakes, align a polymer stamp by eye, and finely adjust temperature and contact geometry through tacit skill. This manual workflow is difficult to reproduce, scales poorly, and prevents systematic exploration of the enormous combinatorial space of materials, twist angles, and interfacial conditions. AutoLab is an autonomous platform that translates this tacit human expertise into programmable, feedback-driven control. Instead of pressing flakes with predefined trajectories, AutoLab uses machine vision to detect polymer-wafer contact, dynamically regulates contact evolution through closed-loop actuation and temperature control, and captures high-quality flakes with the cleanliness and precision of expert manual fabrication. The system integrates perception, decision making, and motion planning into a single robotic framework, enabling reproducible stacking, wafer-level coverage, and accelerated discovery. Beyond 2D materials, AutoLab illustrates a broader paradigm for AI-native scientific automation: codifying human experimental reasoning into algorithms that interrogate data in real time, adaptively adjust instrumentation, and generate scalable, high-fidelity datasets. Such platforms could generalize to diverse research domains--quantum device fabrication, optical alignment, surface science, autonomous microscopy, and other workflows where expert intuition currently limits throughput and reproducibility. By bridging artisanal manipulation and robotic autonomy, AutoLab points toward a future where scientific discovery is accelerated by machines that not only execute instructions, but learn, respond, and collaborate with human scientists.

Biography: Dr. Yutao Li is a research associate from Department of Condensed Matter Physics and Material Science, Brookhaven National Laboratory. He has 8 years of experience in 2D material sample fabrication, and investigation in their electronic transport, optical and mechanical properties.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Location: CDS, Bldg. 725, Training Room

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1604383624?pwd=ffQ5cUPNxTI7nzClKQO6cnsNbhF9Vf.1

Meeting ID: 160 438 3624
Passcode: 558449

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes three short talks on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Tuesday, January 7, 2025, 12:00 pm -- CDS, Bldg. 725, Training Room

Speakers

Maria Zawadowicz, EBNN--ML for Atmospheric Aerosol Research

Mohammad Atif, CDS--An Extensible Digital Twin Framework

Guang Zhao, CDS--Pareto Prompt Optimization

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1615289117?pwd=Hqkbj9itxWrFnkhZ8rQXHPInO2gxdF.1

Meeting ID: 161 528 9117
Passcode: 991382

Visual Analytics and Machine Learning for Biomedical Imaging Diagnosis

 

Arie Kaufman

 

We present an integrated approach using visual analytics and machine learning (ML) to diagnose abnormalities in 3D radiological imaging and biological microscopes. The primary example will involve 3D virtual pancreatography (VP), a novel visualization-ML procedure and application for non-invasive diagnosis and classification of pancreatic lesions, the precursors of pancreatic cancer. Currently, non-invasive screening of patients is performed through visual inspection of 2D axis-aligned CT images, though the relevant features are often not clearly visible nor automatically detected. VP is an end-to-end visual diagnosis system that includes an ML-based automatic segmentation of the pancreatic gland and the lesions, a semi-automatic approach to extract the primary pancreatic duct, an ML-based automatic classification of lesions into four prominent types, and specialized 3D and 2D exploratory visualizations of the pancreas, lesions and surrounding anatomy. We combine volume rendering with pancreas- and lesion-centric visualizations and measurements for effective diagnosis. We designed VP through close collaboration and feedback from expert radiologists, and evaluated it on multiple real-world CT datasets with various pancreatic lesions and case studies examined by the expert radiologists. Other applications include virtual colonoscopy, COVID-19, pathology, brain neurites, etc.


Biography: Arie Kaufman is Distinguished Professor and formerChair of the Department of Computer Science at Stony Brook University, where he is also Director of the Center for Visual Computing (CVC), and Chief Scientist at the Center of Excellence in Wireless and Information Technology (CEWIT). 

He received his PhD in Computer Science at Ben-Gurion University of the Negev in 1977.   He is known for his work in visualization, graphics, virtual reality, user interfaces, multimedia, and their applications, especially in bio-medicine. He is especially well known for his work on the 3-dimensional virtual colonoscopy, a revolutionary low-risk technique for colon cancer screening, and for pioneering the use of Graphics Processing Units (GPUs) and GPU-clusters. In 2012, he presided over the development and opening of the Reality Deck, the largest virtual reality display in the world, at Stony Brook University.

Kaufman was the founding Editor in Chief of IEEE Transactions on Visualization and Computer Graphics (TVCG), co-founded the IEEE Visualization Conference and Volume Graphics series, and is currently the director of IEEE Computer Society Technical Committee on Visualization and Graphics. He is an IEEE Fellow, ACM Fellow, winner of many awards, including the IEEE Visualization Career Award, and member of the European Academy of Sciences.



Steven Skiena is inviting you to a scheduled Zoom meeting.

Topic: AI Seminar: Arie Kaufman
Time: Apr 21, 2021 10:00 AM Eastern Time (US and Canada)

Join Zoom Meeting
https://stonybrook.zoom.us/j/96017498640?pwd=SE0rdHB6ZVlCM2ZpY2RnRUxyVnR3Zz09

This workshop synthesizes the latest research on the impact of AI usage in education so that you could make informed decisions on whether and how to use AI to facilitate your learning. You might have seen conflicting reports on whether the use of AI is good for learning. In this workshop, we are going to tease out, drawing on the latest research, which types of AI usage are beneficial or harmful for different kinds of learning. At the end of the workshop, you should walk away with more clarity on when and how to use AI for your own learning. Join PRODIG+ fellow on critical AI, Zheng Fu, in this informative workshop.

Register for this Zoom workshop.

Abstract: Recent advances in Spatial Transcriptomics (ST) pair histology images with spatially resolved gene expression profiles, enabling predictions of gene expression across different tissue locations based on image patches. This opens up new possibilities for enhancing whole slide image (WSI) prediction tasks with localized gene expression. However, existing methods do not fully leverage the interactions between different tissue locations, which are crucial for accurate joint prediction. To address this, we introduce MERGE (Multi-faceted hiErarchical gRaph for Gene Expressions), which combines a multi-faceted hierarchical graph construction strategy with graph neural networks (GNN) to improve gene expression predictions from WSIs. By clustering tissue image patches based on both spatial and morphological features, and incorporating intra- and inter-cluster edges, our approach fosters interactions between distant tissue locations during GNN learning. As an additional contribution, we evaluate different data smoothing techniques that are necessary to mitigate artifacts in ST data, often caused by technical imperfections. We advocate for adopting gene-aware smoothing methods that are more biologically justified. Experimental results on gene expression prediction show that our GNN method outperforms state-of-the-art techniques on multiple metrics such as mean squared error (MSE), mean absolute error (MAE), and pearson correlation coefficient (PCC). Qualitative analysis establishes the effectiveness of MERGE in capturing cancer marker genes, thus consolidating its utility in diagnostics. As an extension of this work, we use MERGE in a setting with an uncertainty calibration branch to perform robust gene expression smoothing. We show that using patch-wise uncertainty from an uncertainty calibration model and the gene expression predictions from MERGE to enrich the ground truth gene expression matrix, results in better alignment with pathologist annotations, thus establishing that the smoothing is biologically informed.

Speaker: Aniruddha Ganguly

Location: Virtual Zoom Meeting


https://stonybrook.zoom.us/j/5474847973?pwd=Sng0Q2h1c1d3cm9sbFBmYUczMHZNdz09
Meeting ID: 547 484 7973
Passcode: 206739
The Future of Learning: Rethinking Practice in a Changing World

Thursday, March 26, 2026 (Workshops)
Friday, March 27, 2026 (Symposium)

Open to Stony Brook University Faculty, Staff, and Graduate Students. Hosted by the Center for Excellence in Learning and Teaching, Office of the Provost.

Thursday, March 26, 2026
Workshop: AI Tools and Techniques
  • Open to all faculty & staff
  • Hands-on, exploratory
  • Registration only limited to the size of the room
  • Location: In-person, TBD
  • Time: 10 AM - 12 PM
  • Registration required

Friday, March 27, 2026
Keynote: Teaching and Thinking with AI
  • Faculty, TAs, postdocs, and academic staff
  • In-person on-campus conference venue
  • Location: SAC Balroom
  • Time: 9 AM - 3 PM
  • Registration required

Keynote Speaker: José Antonio Bowen

José Antonio Bowen has been leading innovation and change for over 40 years at Stanford, Georgetown and the University of Southampton (UK), as a dean at Miami University and SMU and as President of Goucher College. Bowen has worked as a musician with Stan Getz, Dave Brubeck, and many others and his symphony was nominated for the Pulitzer Prize in Music (1985).
Bowen holds four degrees from Stanford and has written over 100 scholarly articles and books, including the Cambridge Companion to Conducting (2003), Teaching Naked (2012 and the winner of the Ness Award for Best Book on Higher Education), Teaching Naked Techniques with C. Edward Watson (2017) and Teaching Change: How to Develop Independent Thinkers using Relationships, Resilience and Reflection (Johns Hopkins University Press, 2021).
Bowen has appeared in The New York Times, Forbes, The Wall Street Journal, and has three TED talks. Stanford honored him as a Distinguished Alumni Scholar (2010) and he has presented keynotes and workshops at more than 300 campuses and conferences 46 states and 17 countries around the world. In 2018, he was awarded the Ernest L. Boyer Award (for significant contributions to American higher education). He is a senior fellow for the American Association of Colleges and Universities.

Register here.
All are welcome to attend BMI grand rounds talk by Dr. Le Lu on 04/14. 

Le Lu, Ph.D 
Executive Director, PAII Inc 
Johns Hopkins University
IEEE Fellow, MICCAI Board Member


Time: Wednesday, April 14, 2021 3:00 pm - 4:00 pm 

Zoom Meeting 
https://stonybrook.zoom.us/j/95617197636?pwd=KytzZ2pVRG9SZGpKZUtpNXJISjNjZz09 
Meeting ID: 956 1719 7636 Passcode: 924293

Title: 
In Search of Effective and Reproducible Clinical Imaging Biomarkers for Population Health and Oncology Applications of Screening, Diagnosis and Prognosis

Bio: 
Le Lu received a PhD in 2007 from Johns Hopkins University. During his first six years at Siemens, he made significant contributions to the company's CT colonography and Lung CAD product lines. From 2013 to 2017, Dr. Lu served as a staff scientist in the Radiology and Imaging Sciences department of the National Institutes of Health Clinical Center. He then went on to found Nvidia's medical image analysis group and he held the position of senior research manager until June 2018. Since then, he has been the Executive Director at PAII Inc., Bethesda Research lab, Maryland, USA which has become one of the leading industrial research labs in medical imaging. He was the main technical leader for two of the most-impactful public radiology image dataset releases (NIH ChestXray14, NIH DeepLesion 2018). He won NIH Clinical Center Director Award in 2017, NIH Mentor of the year award in 2015, and won numerous best paper awards in MICCAI and RSNA from 2016 to 2020 (over 10000 citations). In 2021, He was elected into IEEE Fellow class cited for his contribution to machine learning for cancer detection and diagnosis, and MICCAI society board member (MICCAI-Industry Workgroup Chair). He is currently an Associate Editor for IEEE Trans. Pattern Analysis and Machine Intelligence and IEEE Signal Processing Letters. He has served as an Area Chair for recent MICCAI, AAAI, CVPR, WACV, ICIP and ICHI conferences for 14 times.

Abstract: 
This talk will first give an overall on the work of employing deep learning to permit novel clinical workflows in two population health tasks, namely using conventional ultrasound for liver steatosis screening and quantitative reporting; osteoporosis screening via conventional X-ray imaging and AI readers. These two tasks were generally considered as infeasible tasks for human readers, but as proved by our scientific and clinical studies and peer-reviewed publications, they are suitable for AI readers. AI can be a supplementary and useful tool to assist physicians for cheaper and more convenient/precision patient management. Next, the main part of this talk describes a roadmap on three key problems in pancreatic cancer imaging solution: early screening, precision differential diagnosis, and deep prognosis on patient survival prediction. (1) Based on a new self- learning framework, we train the pancreatic ductal adenocarcinoma (PDAC) segmentation model using a larger quantity of patients (≈1,000, four institutions), with a mix of annotated/unannotated venous or multi-phase CT images. Pseudo annotations are generated by combining two teacher models with different PDAC segmentation specialties on unannotated images, and can be further refined by a teaching assistant model that identifies associated vessels around the pancreas. Our approach makes it technically feasible for robust large-scale PDAC screening from multi-institutional multi-phase partially-annotated CT scans. (2) We propose a holistic segmentation-mesh classification network (SMCN) to provide patient-level diagnosis, by fully utilizing the geometry and location information. SMCN learns the pancreas and mass segmentation task and builds an anatomical correspondence-aware organ mesh model by progressively deforming a pancreas prototype on the raw segmentation mask. Our results are comparable to a multimodality clinical test that combines clinical, imaging, and molecular testing for clinical management of patients with cysts. (3) Accurate preoperative prognosis of resectable PDACs for personalized treatment is highly desired in clinical practice. We present a novel deep neural network for the survival prediction of resectable PDAC patients, 3D Contrast-Enhanced Convolutional Long Short-Term Memory network (CE- ConvLSTM), to derive the tumor attenuation signatures from CE-CT imaging studies. Our framework can significantly improve the prediction performances upon existing state-of-the-art survival analysis methods. This deep tumor signature has evidently added values (as a predictive biomarker) to be combined with the existing clinical staging system.

More information can be found at:
https://bmi.stonybrookmedicine.edu/sites/default/files/Lu_le_04_14.pdf

Abstract: Generative visual models like Stable Diffusion and Sora generate photorealistic images and videos that are nearly indistinguishable from real ones to a naive observer. However, their grasp of the physical world remains an open question: Do they understand 3D geometry, light, and object interactions, or are they mere pixel parrots of their training data? Through systematic probing, I will demonstrate that these models surprisingly learn fundamental scene properties--intrinsic images such as surface normals, depth, albedo, and shading (à la Barrow & Tenenbaum, 1978)--without explicit supervision, which enables applications like image relighting. But I will also show that this knowledge is insufficient. Careful analysis reveals unexpected failures: inconsistent shadows, multiple vanishing points, and scenes that defy basic physics. All these findings suggest these models excel at local texture synthesis but struggle with global reasoning: a crucial gap between imitation and true understanding. I will then conclude by outlining a path toward generative world models that emulate global and counterfactual reasoning, causality, and physics.

Bio: Anand Bhattad is a Research Assistant Professor at the Toyota Technological Institute at Chicago. He earned his PhD from the University of Illinois Urbana-Champaign in 2024 under the mentorship of David Forsyth. His research interests lie at the intersection of computer vision and computer graphics, with a current focus on understanding the knowledge encoded in generative models. Anand has received Outstanding Reviewer honors at ICCV 2023 and CVPR 2021, and his CVPR 2022 paper was nominated for a Best Paper Award. He actively contributes to the research community by leading workshops at CVPR and ECCV, including Scholars and Big Models: How Can Academics Adapt? (CVPR 2023), CV 20/20: A Retrospective Vision (CVPR 2024), Knowledge in Generative Models (ECCV 2024), and How to Stand Out in the Crowd? (CVPR 2025). For more details, visit https://anandbhattad.github.io/




Abstract:
Large language models (LLMs) have transformed the way humans write code, bringing unprecedented automation to software development. In this talk, I will first provide an overview of my research on enhancing LLMs' code intelligence, optimizing each step of the development pipeline towards more complex software engineering tasks. I will then delve into my key contributions, focusing on how to equip LLMs with a deeper, more comprehensive understanding of software programs. Finally, I will discuss the future of AI-driven software engineering, envisioning a new era of automation that is more reliable, intelligent, and cost-efficient.

Bio:
Yangruibo (Robin) Ding is a Ph.D. candidate in the Department of Computer Science at Columbia University. His research is at the intersection of Software Engineering and Machine Learning, focusing on developing large language models (LLMs) for code. He trains LLMs to generate, analyze, and refine software programs and constructs benchmarks to systematically evaluate LLMs in solving software engineering tasks. He also studies how to improve LLMs' reasoning capability to tackle complex programming tasks, such as debugging and patching. His interdisciplinary research has been published in top-tier conferences of software engineering, programming languages, natural language processing, and machine learning. He won an ACM SIGSOFT Distinguished Paper Award, an IEEE TSE Best Paper Runner-up, and received an IBM Ph.D. Fellowship.
Location:
NCS 120