The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 10 minutes) by multiple people. Students can register for 1 credit for CSE 656. Registered students must attend and present a minimum of 2 or 3 talks. Everyone else is welcome to attend. Fill in https://forms.gle/pCVXovgfMfQwGqG38 to subscribe to our mailing list for further announcement.
Abstract: Retrieval-augmented generation (RAG) systems empower large language models (LLMs) to access external knowledge during inference. Recent advances have enabled LLMs to act as search agents via reinforcement learning (RL), improving information acquisition through multi-turn interactions with retrieval engines. However, existing approaches either optimize retrieval using search-only metrics (e.g., NDCG) that ignore downstream utility or fine-tune the entire LLM to jointly reason and retrieve--entangling retrieval with generation and limiting the real search utility and compatibility with frozen or proprietary models. In this work, we propose s3, a lightweight, model-agnostic framework that decouples the searcher from the generator and trains the searcher using a Gain Beyond RAG reward: the improvement in generation accuracy over naïve RAG. s3 requires only 2.4k training samples to outperform baselines trained on over 70 × more data, consistently delivering stronger downstream performance across six general QA and five medical QA benchmarks.

Speaker: Peter Zeng

Location: CS2311
https://stonybrook.zoom.us/j/99820812332?pwd=c05BSTVLNmw3L04yZjdEcG5pem1OZz09 Speaker: Alexei Koulakov of Cold Spring Harbor Laboratory Brain evolution as a machine learning problem We have entered a golden age of artificial intelligence research, driven mainly by the advances in ANNs over the last decade or so. Applications of these techniques--to machine vision, speech recognition, autonomous vehicles, machine translation and many other domains--are coming so quickly that many observers predict that the long-elusive goal of Artificial General Intelligence (AGI) is within our grasp. However, we still cannot build a machine capable of building a nest, stalking prey, or loading a dishwasher. I will describe several projects, ranging from theories of evolution of neural development to the perception of smells, in which we are attempting to understand the algorithms that the nervous system is using to solve some of these challenging problems.
Objectives:
1. Explain the clinical radiology workflow, and highlight how AI is currently in use to impact each step
2. Describe how radiologists interact with the currently available tools, highlighting both positive andnegative examples
3. Offer a brief description of how these tools are approved, validated, and reimbursed
4. Explore the utility of cutting edge AI techniques in diagnostic radiology

Speaker:
Dr. David Payne, MD Neuroradiologist and Assistant Professor, Rush University Medical Centre

Remote Access:
Zoom: https://stonybrook.zoom.us/j/95617197636?pwd=KytzZ2pVRG9SZGpKZUtpNXJISjNjZz09
Meeting ID: 95617197636
Passcode: 924293

Time: Jan 26, 2021 03:00 PM Eastern Time (US and Canada)

All are welcome!

Zoom Meeting:
https://stonybrook.zoom.us/j/93818552212?pwd=ajZkT2x4a2tiaDJUL1h3VFhLZEgwQT09

Meeting ID: 938 1855 2212
Passcode: 802722

Title: Data-Driven Document Unwarping

Abstract: Capturing document images is a common way to digitize and record physical documents due to the ubiquitousness of mobile cameras. To make text recognition easier, it is often desirable to digitally flatten a document image when the physical document sheet is folded or curved. However, unwarping a document from a single image in natural scenes is very challenging due to the complexity of document sheet deformation, document texture, and environmental conditions. Previous model-driven approaches struggle with inefficiency and limited generalizability. In this thesis, I investigate several data-driven approaches to tackle the document unwarping problem.

Data acquisition is the central challenge in data-driven methods. I first design an efficient data synthesis pipeline based on 2D image warping and train DocUNet, the pioneering data-driven document unwarping model, on the synthetic data. A benchmark dataset is also created to facilitate comprehensive evaluation and comparison. To improve the unwarping performance by training on more realistic data, I introduce the Doc3D dataset and DewarpNet. Supervised by 3D shape ground truth in Doc3D, DewarpNet is significantly better than DocUNet. DocUNet and DewarpNet depend on the synthetic data for the ground truth deformation annotation. To exploit the real-world images, I propose PaperEdge, a weakly supervised model trained with in-the-wild document images with easy-to-obtain boundary information. PaperEdge surpasses DewarpNet by utilizing both the synthetic data and weakly annotated real data in the Document In the Wild (DIW) dataset. Finally, I propose to incorporate the 3D physical constraints in training DewarpNet and PaperEdge. The constraints regulate the possible deformations on document papers. I also propose to augment the Doc3D and DIW dataset by introducing an online document segmentation model and better hardware.
The Future Histories Studio welcomes Moontae Lee, LG AI Research.


Generative AI is transforming how we understand, create, and interact with information. Large Language Models (LLMS) comprehend contexts, answer non-trivial questions, and spark creative ideas. This talk introduces the evolution of these models, highlighting the most recent advancements in planning, reasoning, and evaluation. The talk also touches on the criticalconsiderations for both model developers and users, carefully addressing limitations of LLMs as well as ethical and societal implications. Finally, the talk provides ongoing directions in researchand production: from the rise of personalized AI agents to the future frontiers of AI.

Moontae Lee is the Director of the Superintelligence Lab at LG AI Research and an Assistant Professor of Information and Decision Sciences at the University of Illinois Chicago. His journey with Large Language Models began as a visiting scholar at Microsoft Research in 2019, continuously consulting the Deep Learning Group at Redmond until joining LG. He holds a PhD in Computer Science from Cornell, an MS from Stanford, and BS degrees in Computer Science, Mathematics, and Psychology from Sogang University. He has been an area chair for major AI conferences and earned recognition in Operations Research and Computational Social Science, including awards from INFORMS and Amazon.

His research interests include:
● Computational Creativity, Algorithmic Awareness
● Retrieval-Augmented Generation and Evaluation
● Code Generation, Reasoning, Planning
● Fine-grained Alignment from Human/AI Feedback in Generative AI
● Large Time-series Models, Diffusion/Consistency
● Machine Unlearning
● Ranking Monopoly, Voting Fairness
● AI Safety, Ethics, and Market Impacts

Join us in person @ Future Histories Studio Staller Center for the Arts, 4222
We live in a new scientific paradigm: the Big Data era, in which a lot of data is available for almost anything. In this new paradigm, the driving force is to use data directly to learn about chemical and physics systems employing artificial intelligence. This paradigm has proven helpful in simulating realistic physical, biological, and chemical models, yielding impressive results. Similarly, the insight gained in these situations can be used to improve our understanding of fundamental processes. In that regard, we want to answer the question: Can a machine learn chemistry? The answer to this question is still debatable, but we will show our ideas and methods to find the answer. We will also discuss our results on predicting atom-diatom reactions and other avenues and work in progress in our group.

Please register for the STEM Speaker Series: Can a Machine learn Chemistry here.


Time:
Sep 7, Tue, 11:00am EDT

Place:
NCS 220 or on Zoom (info below)

Title: Data-Driven Document Unwarping


Abstract:
Capturing document images is a common way to digitize and record physical documents due to the ubiquitousness of mobile cameras. To make text recognition easier, it is often desirable to digitally flatten a document image when the physical document sheet is folded or curved. However, unwarping a document from a single image in natural scenes is very challenging due to the complexity of document sheet deformation, document texture, and environmental conditions. Previous model-driven approaches struggle with inefficiency and limited generalizability. In this thesis, I investigate several data-driven approaches to tackle the document unwarping problem.

Data acquisition is the central challenge in data-driven methods. I first design an efficient data synthesis pipeline based on 2D image warping and train DocUNet, the pioneering data-driven document unwarping model, on the synthetic data. A benchmark dataset is also created to facilitate comprehensive evaluation and comparison. To improve the unwarping performance by training on more realistic data, I introduce the Doc3D dataset and DewarpNet. Supervised by 3D shape ground truth in Doc3D, DewarpNet is significantly better than DocUNet. DocUNet and DewarpNet depend on the synthetic data for the ground truth deformation annotation. To exploit the real-world images, I propose PaperEdge, a weakly supervised model trained with in-the-wild document images with easy-to-obtain boundary information. PaperEdge surpasses DewarpNet by utilizing both the synthetic data and weakly annotated real data in the Document In the Wild (DIW) dataset. Finally, I propose directly predicting the $uv$ parameterized 3D mesh of the document with 3D constraints and using the accessible 3D presentations like depth maps as training targets. Predicting the 3D mesh of the document solves the unwarping task and also benefits VR/AR applications.

Join Zoom Meeting
https://stonybrook.zoom.us/j/96440592912?pwd=ZU5waTdyUzRFNW5SRHM5ME84TWdFQT09

Meeting ID: 964 4059 2912
Passcode: 793149
One tap mobile
+16468769923,,96440592912# US (New York)
+13017158592,,96440592912# US (Washington DC)

Dial by your location
        +1 646 876 9923 US (New York)
        +1 301 715 8592 US (Washington DC)
        +1 312 626 6799 US (Chicago)
        +1 253 215 8782 US (Tacoma)
        +1 346 248 7799 US (Houston)
        +1 408 638 0968 US (San Jose)
        +1 669 900 6833 US (San Jose)
Meeting ID: 964 4059 2912
Find your local number: https://stonybrook.zoom.us/u/adxTt9ZbuJ
West Campus - SAC- Student Activities Center - Ballrooms A & B 100 Nicolls Road Stony Brook NY 11794 Job Fair.jpg The Career Center invites Alumni Employers and Job Seekers to the IT/Computer Science Job and Internship Fair this spring. Job Seekers: A job fair is an opportunity for you to present yourself professionally in person to a potential employer, while showcasing your communication skills. Get more information Alumni Employers: Held in both the fall and spring semesters, this event is ideal for employers looking to fill internship, co-op, part-time and full-time opportunities in the field of information technology (i.e. Software Engineering, Network Administration, Web Development, etc.). Register here to recruit top SBU talent.

Abstract: Computational pathology has revolutionized cancer diagnosis and research through the analysis of digitized whole slide images (WSIs). However, the giga-pixel size of these images presents profound technical challenges, creating two intertwined bottlenecks: computational inefficiency and label inefficiency. The immense data scale makes standard end-to-end (E2E) training of deep neural networks infeasible due to prohibitive GPU memory requirements, while the reliance on expert pathologists for annotations makes obtaining high-quality labeled data a tedious and expensive process. This proposal confronts these dual challenges by developing a series of novel model architectures, training paradigms, and self-supervised learning methods designed to create a more efficient and effective framework for WSI analysis.

To improve computational efficiency, this proposal first introduces a locally supervised learning paradigm that enables E2E training on entire WSIs by partitioning a network into gradient-isolated modules, circumventing the memory bottleneck of backpropagation. Second, it presents Prompt-MIL, a parameter-efficient fine-tuning framework that reduces the number of trainable parameters, memory consumption, and training time by fine-tuning only few prompts to guide large pre-trained models. Third, this work advances the efficient architecture on WSIs by developing novel State-Space Models (SSMs). It proposes 2DMamba, the first intrinsic Mamba architecture that preserves the crucial 2D spatial structure of images, overcoming the spatial discrepancy inherent in 1D models. Fourth, to address the inefficiency of multi-directional scans in Mamba models, including 2DMamba, it presents Locally Bi-directional Mamba (LBMamba), which introduces a novel, hardware-aware local backward scan that integrates bi-directional scan into a single forward pass, significantly improving throughput performance trade-off. Lastly, it proposes an extension to the LBMamba, warp-level Bi-directional Mamba (WLBMamba) that extends the thread-level bidirectional scan to warp-level bidirectional scan that further improves the throughput performance trade-off.

To improve label efficiency, this proposal proposes a Precise Location-based Matching strategy for self-supervised dense contrastive learning. By allowing a local patch in one augmented view to match multiple overlapping patches in another, creates a more accurate correspondence, leading to superior feature representations for dense prediction tasks like segmentation and detection.

In summary, this proposal presents a holistic investigation into the efficiency bottlenecks in computational pathology. Through these combined contributions in model architecture, training paradigms, and self-supervised learning, this work establishes a more scalable, efficient, and powerful computational framework for analyzing giga-pixel pathology images.

Speaker: Jingwei Zhang

Location: Old Computer Science Room 2114

Zoom: https://stonybrook.zoom.us/j/95187903649?pwd=tV0CNxLu1QKqw7hGmcE1h0rJ2C6n1b.1
Meeting ID: 951 8790 3649 | Passcode: 488916