Fantastic Futures is an international conference series organized by AI4LAM as one of the most important global events at the intersection of AI and cultural heritage, shaping how memory institutions adapt to rapidly evolving technologies while maintaining public trust.

The conference brings together professionals, researchers, technologists, and cultural‑heritage institutions to explore how AI can support organizational workflows, users services and data preservation, access, discovery, as well as innovation in the cultural‑heritage sector:

  • Librarians, archivists, museum professionals and all other GLAM enthusiasts
  • Technologists working on cultural‑heritage applications, AI researchers and developers
  • Digital humanists, researchers of cultural heritage
  • University professors, students, and librarians
  • Policy and ethics experts, governmental employees
  • Strategic leaders and innovation managers

To register, visit AI4LAM's official website.

Time: Jan 26, 2021 03:00 PM Eastern Time (US and Canada)

All are welcome!

Zoom Meeting:
https://stonybrook.zoom.us/j/93818552212?pwd=ajZkT2x4a2tiaDJUL1h3VFhLZEgwQT09

Meeting ID: 938 1855 2212
Passcode: 802722

Title: Data-Driven Document Unwarping

Abstract: Capturing document images is a common way to digitize and record physical documents due to the ubiquitousness of mobile cameras. To make text recognition easier, it is often desirable to digitally flatten a document image when the physical document sheet is folded or curved. However, unwarping a document from a single image in natural scenes is very challenging due to the complexity of document sheet deformation, document texture, and environmental conditions. Previous model-driven approaches struggle with inefficiency and limited generalizability. In this thesis, I investigate several data-driven approaches to tackle the document unwarping problem.

Data acquisition is the central challenge in data-driven methods. I first design an efficient data synthesis pipeline based on 2D image warping and train DocUNet, the pioneering data-driven document unwarping model, on the synthetic data. A benchmark dataset is also created to facilitate comprehensive evaluation and comparison. To improve the unwarping performance by training on more realistic data, I introduce the Doc3D dataset and DewarpNet. Supervised by 3D shape ground truth in Doc3D, DewarpNet is significantly better than DocUNet. DocUNet and DewarpNet depend on the synthetic data for the ground truth deformation annotation. To exploit the real-world images, I propose PaperEdge, a weakly supervised model trained with in-the-wild document images with easy-to-obtain boundary information. PaperEdge surpasses DewarpNet by utilizing both the synthetic data and weakly annotated real data in the Document In the Wild (DIW) dataset. Finally, I propose to incorporate the 3D physical constraints in training DewarpNet and PaperEdge. The constraints regulate the possible deformations on document papers. I also propose to augment the Doc3D and DIW dataset by introducing an online document segmentation model and better hardware.
CSE 600 Talk: Haibin Ling - Computer Vision Research and Applications


Abstract: Having been intensively studied over half a decade, computer vision has evolved as a broad research area and become mature in many applications. In this talk, we will summarize our work in computer vision in both core vision topics and application-oriented ones. In particular, for core vision problems, we will report studies on visual tracking, visual matching and visual detection; for applications, we will describe our work on medical image analysis, intelligent transportation, smart projector systems and preliminary work on material property prediction.

Bio: Haibin Ling received the BS and MS degrees from Peking University in 1997 and 2000, respectively, and the PhD degree from the University of Maryland, College Park, in 2006. From 2000 to 2001, he was an assistant researcher at Microsoft Research Asia. From 2006 to 2007, he worked as a postdoctoral scientist at the University of California Los Angeles. In 2007, he joined Siemens Corporate Research as a research scientist. From 2008 to 2019, he worked as a faculty member of Temple University. In fall 2019, he joined the Department of Computer Science of Stony Brook University where he is currently a SUNY Empire Innovation Professor. His research interests include computer vision, augmented reality, medical image analysis, and human computer interaction. He received the Best Student Paper Award at the ACM UIST in 2003, and the NSF CAREER Award in 2014. He serves as Associate Editor for several journals including IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Pattern Recognition (PR), and Computer Vision and Image Understanding (CVIU). He has served or will serve as Area Chair for CVPR 2014, 2016, 2019 and 2020.

The Pittsburgh Supercomputing Center is pleased to present a Machine Learning and Big Data workshop.

This workshop will focus on topics including big data analytics and machine learning with Spark, as well as deep learning.

This will be an IN PERSON event hosted by various satellite sites, there WILL NOT be a direct to desktop option for this event. SBU's Institute for Advanced Computational Science (IACS) is one of those satellite sites!

Location: IACS Conference Room #2

Interested applicants must first have an ACCESS ID. If you don't have the ID, please visit this page to create one: ACCESS USER REGISTRATION.


Once you have an ACCESS ID, please login (see top right here) then register here.
Abstract: Artificial intelligence (AI) is rapidly transforming scientific discovery, enabling breakthroughs in areas ranging from drug discovery to modeling complex physical systems. In the life sciences, AI has traditionally been applied to prediction tasks such as classifying molecules as toxic or non-toxic, estimating drug properties, or solving partial differential equations. These discriminative models have proven powerful, but they are inherently limited to mapping existing inputs to deterministic outputs. A new wave of methods is shifting the paradigm from discrimination to generation: creating new possibilities, such as generating novel molecules or designing new drugs. By reframing AI as both a predictive and generative engine, this shift offers new pathways for accelerating discovery and innovation in life sciences at an unprecedented scale. This talk will cover several aspects of AI for Science (AI4Sci), beginning with advances in discriminative models for molecular systems and solving PDEs, and then turning to generative approaches, including diffusion models for 3D molecular generation and large language models for drug editing. Together, these developments illustrate how moving from prediction to creation is redefining what AI can contribute to science.

Bio: Wenhan Gao is a fourth-year Ph.D. student in Applied Mathematics under the supervision of Professor Yi Liu. He was also a Staff Research Scientist Intern at VISA Research, where he worked on large language models (LLMs) and multi-agent systems for commerce. Wenhan's research focuses on AI for Science (AI4Sci), with a particular emphasis on generative AI. His work looks deep into the fundamental mechanisms of AI models when applied to scientific tasks, and he strives to incorporate established scientific priors, such as symmetry, into model design. He has published papers as a first or corresponding author in leading AI and computational venues, including ICLR, ICML, NeurIPS, TMLR, ACL, and the Journal of Computational Physics. In addition to his research, Wenhan has served as a reviewer and oral session chair for top AI conferences and as a lecturer for both undergraduate and graduate courses at Stony Brook University.

Location: IACS Seminar Room or Zoom

This seminar will take place in person and online*

Join Zoom Meeting: https://stonybrook.zoom.us/j/91670093552?pwd=2EcniXqPZLTpa4ZBKRs1zAjYqs1LS0.1

Meeting ID: 916 7009 3552
Passcode: 434045

Abstract: Computational pathology has revolutionized cancer diagnosis and research through the analysis of digitized whole slide images (WSIs). However, the giga-pixel size of these images presents profound technical challenges, creating two intertwined bottlenecks: computational inefficiency and label inefficiency. The immense data scale makes standard end-to-end (E2E) training of deep neural networks infeasible due to prohibitive GPU memory requirements, while the reliance on expert pathologists for annotations makes obtaining high-quality labeled data a tedious and expensive process. This proposal confronts these dual challenges by developing a series of novel model architectures, training paradigms, and self-supervised learning methods designed to create a more efficient and effective framework for WSI analysis.

To improve computational efficiency, this proposal first introduces a locally supervised learning paradigm that enables E2E training on entire WSIs by partitioning a network into gradient-isolated modules, circumventing the memory bottleneck of backpropagation. Second, it presents Prompt-MIL, a parameter-efficient fine-tuning framework that reduces the number of trainable parameters, memory consumption, and training time by fine-tuning only few prompts to guide large pre-trained models. Third, this work advances the efficient architecture on WSIs by developing novel State-Space Models (SSMs). It proposes 2DMamba, the first intrinsic Mamba architecture that preserves the crucial 2D spatial structure of images, overcoming the spatial discrepancy inherent in 1D models. Fourth, to address the inefficiency of multi-directional scans in Mamba models, including 2DMamba, it presents Locally Bi-directional Mamba (LBMamba), which introduces a novel, hardware-aware local backward scan that integrates bi-directional scan into a single forward pass, significantly improving throughput performance trade-off. Lastly, it proposes an extension to the LBMamba, warp-level Bi-directional Mamba (WLBMamba) that extends the thread-level bidirectional scan to warp-level bidirectional scan that further improves the throughput performance trade-off.

To improve label efficiency, this proposal proposes a Precise Location-based Matching strategy for self-supervised dense contrastive learning. By allowing a local patch in one augmented view to match multiple overlapping patches in another, creates a more accurate correspondence, leading to superior feature representations for dense prediction tasks like segmentation and detection.

In summary, this proposal presents a holistic investigation into the efficiency bottlenecks in computational pathology. Through these combined contributions in model architecture, training paradigms, and self-supervised learning, this work establishes a more scalable, efficient, and powerful computational framework for analyzing giga-pixel pathology images.

Speaker: Jingwei Zhang

Location: Old Computer Science Room 2114

Zoom: https://stonybrook.zoom.us/j/95187903649?pwd=tV0CNxLu1QKqw7hGmcE1h0rJ2C6n1b.1
Meeting ID: 951 8790 3649 | Passcode: 488916
Speaker Petar Djuric Refreshments will be provided Deep Gaussian processes: Theory and applications Petar M. Djurić Department of Electrical and Computer Engineering Stony Brook University Abstract: Gaussian processes are an infinite-dimensional generalization of multivariate normal distributions. They provide a principled approach to learning with kernel machines and they have found wide applications in many fields. More recently, with the advance of deep learning, the concept of deep Gaussian processes has emerged. Deep Gaussian processes can be viewed as multilayer hierarchical organizations of Gaussian processes that are equivalent to infinitely wide multiple layer neural networks. Deep Gaussian processes have improved capacity for prediction and classification over standard Gaussian processes, while models based on them continue to allow for full Bayesian treatment and for applications when the amount of available data is limited. The theory of recent progress in deep Gaussian processes will be presented and some applications will be provided. Biosketch: Petar M. Djurić received the B.S. and M.S. degrees in electrical engineering from the University of Belgrade, Belgrade, Yugoslavia, respectively, and the Ph.D. degree in electrical engineering from the University of Rhode Island, Kingston, RI, USA. He is a SUNY Distinguished Professor and currently, he is a Chair of the Department of Electrical and Computer Engineering, Stony Brook University, Stony Brook, NY, USA. Djurić was a recipient of the IEEE Signal Processing Magazine Best Paper Award in 2007 and the EURASIP Technical Achievement Award in 2012. From 2008 to 2009, he was a Distinguished Lecturer of the IEEE Signal Processing Society. He was the Editor-in-Chief of the IEEE Transactions on Signal and Information Processing over Networks (2015-2018). Djurić is a Fellow of IEEE and EURASIP

Submit an abstract celebrating research, new discoveries and achievements in medicine and science!

We encourage faculty, nurse practitioners, post-doctoral fellows, fellows, residents, medical students, graduate students and undergraduate students to submit an abstract. Original research, case reports and case series are welcome.

Abstract submission deadline: FEBRUARY 7, 2025

For more details, visit here.

Abstract: Language offers a uniquely powerful lens for understanding the mind: one that can access latent psychological realities often missed by traditional measurement tools. However, as language models expand their ability to capture semantics through context length, expansion into deeper levels of semantics is less explored, especially with respect to understanding cognitive patterns of authors. This dissertation proposes that we can uncover deeper cognitive and affective patterns that reflect more accurate underlying mental states by analyzing language at higher levels of discourse semantics and by modeling latent states.


First, the dissertation focuses on uncovering cognitive styles or thinking patterns manifesting in language. We demonstrate that modeling language at deeper semantic levels such as discourse relations, can unveil latent psychological states and traits, including cognitive styles that influence both mental health and behavior. Introducing a novel blend of transfer and active learning, we efficiently curated a new set of linguistic data on cognitive styles like dissonance. This approach allows for more precise measurement when dealing with rare-classes and low-resource tasks. As a second contribution, effective validation methods are introduced to language-based assessments of the underlying cognitive styles. Controlled behavioral experiments and online studies show that cognitive styles detected through linguistic signals reliably predict real-world behaviors such as decision-making and engagement with extremist communities, both at the individual and community levels, sometimes months in advance

The research further moves beyond traditional measurement tools like questionnaires and expert judgments, which rely on Classical Test Theory, by establishing that language-based assessments more closely approximate true psychological states. The mechanisms by which these assessments outperform standard tools are explained, highlighting their predictive power for behaviors linked to underlying traits. Finally, a more sophisticated approach is explored by modeling psychological outcomes with Item Response Theory (IRT), an improvement over Classical Test Theory. Adaptive language-based assessments are introduced, showing that targeted, adaptive testing based on latent IRT scores can efficiently and accurately capture multiple psychological dimensions.

Taken together, these contributions argue for a shift towards language-based psychological assessments. By integrating deeper discourse-level semantics with measurement theory, this dissertation charts a path towards truer scores of mental states: ones that are more precise, and reflective of the complexity of human cognition and emotions.

Speaker: Vasudha Varadarajan

https://stonybrook.zoom.us/j/99180374682?pwd=w2zZTkQsfunrBZhHgEweR54NjKabZ2.1&jst=2