CSE 656 Seminars in Computer Vision - Wednesdays 11:30am-12:50pm, Room NCS 120

The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 10 minutes) by multiple people. Students can register for 1 credit for CSE656. Registered students must attend and present a minimum of 2 or 3 talks. Everyone else is welcome to attend. Fill in https://forms.gle/pCVXovgfMfQwGqG38 to subscribe to our mailing list for further announcement.

The first meeting will be Wed Jan 29 at 11.30am, room 120 New CS. The meeting will deal with organizational matters and we will start right away with some presentations. Send David Paredes Merino <dparedesmeri@cs.stonybrook.edu> an email if you are interested but cannot attend the first meeting. Please forward to people outside the CS department that you think might be interested.
Title:Deep Contextual Modeling for Natural Language Understanding, Generation, and Grounding Zoom instructions: Join Zoom Meeting https://stonybrook.zoom.us/j/645050299?pwd=TVJVRkc3dlhxdDF5d00xWGlDQkovZz09 Meeting ID: 645 050 299 Password: 810247 One tap mobile +16468769923,,645050299#,,#,810247# US (New York) +13126266799,,645050299#,,#,810247# US (Chicago) Dial by your location +1 646 876 9923 US (New York) +1 312 626 6799 US (Chicago) +1 301 715 8592 US +1 346 248 7799 US (Houston) +1 408 638 0968 US (San Jose) +1 669 900 6833 US (San Jose) +1 253 215 8782 US Meeting ID: 645 050 299 Password: 810247 Find your local number: https://stonybrook.zoom.us/u/aemTiJMXu6 Abstract: Natural language is a fundamental form of information and communication. In both human-human and human-computer communication, people reason about the context of text and world state to understand language and produce language response. In this talk, I present several deep neural network based systems that first understand the meaning of language grounded in various contexts where the language is used, and then generate effective language responses in different forms for information access and human-computer communication. First, I will introduce Speaker Interaction RNNs for addressee and response selection in multi-party conversations based on explicit representations for different discourse participants. Then, I will present a text summarization approach for generating email subject lines by optimizing quality scores in a reinforcement learning framework. Finally, I will show an editing-based multi-turn SQL query generation system towards intelligent natural language interfaces to databases. Bio:Rui Zhang is a final year Ph.D. student at Yale University advised by Professor Dragomir Radev. His research interest lies in various natural language processing problems in understanding, generation, and grounding. He has been working on (1) End-to-End Neural Modeling for Entities, Sentences, Documents, and Multi-party Multi-turn Dialogues, (2) Text Summarization for Emails, News, and Scientific Articles, (3) Cross-lingual Information Retrieval for Low-Resource Languages, (4) Context-Dependent Text-to-SQL Semantic Parsing in Human-Computer Interaction. Rui Zhang has published papers and served as Program Committee members at top-tier NLP and AI conferences including ACL, NAACL, EMNLP, AAAI, CoNLL. During his Ph.D., He has done research internships at IBM Thomas J. Watson Research Center, Grammarly Research, and Google AI. He was a graduate student at the University of Michigan and got his bachelor's degrees at both the University of Michigan and Shanghai Jiao Tong University from the UM-SJTU Joint Institute.

Abstract: Much like other AI for Science domains, polymer design poses significant challenges. It requires grounding in empirical data and physical laws, precise handling of domain-specific structured representations, and compositional reasoning over multiple interacting constraints--all while working with limited data.

To address these limitations, we introduce PolyBench, a large-scale benchmark comprising over 125K polymer design and analysis tasks grounded in verified experimental and synthetic data. PolyBench includes tasks created from a wide range of data sources and presents diverse structural, property-driven, and synthesis-oriented reasoning problems. Tasks in PolyBench are organized from simple to complex analytical reasoning problems, enabling generalization tests and includes diagnostic probes to evaluate model capabilities. Additionally, to support effective domain alignment, we propose a knowledge-augmented reasoning distillation framework that enriches the dataset with structured chain-of-thought supervision derived from expert-informed reasoning strategies.

Small language models (7B-14B parameters) trained on PolyBench substantially outperform comparably sized baselines and, in many cases, exceed the performance of larger closed-source frontier models on polymer reasoning tasks, while also demonstrating improved transfer to external polymer benchmarks. Last, we conduct a diagnostic study that reveals a compositionality gap: despite strong performance on decomposed sub-questions, models struggle to integrate multiple interacting constraints and intermediate reasoning steps, highlighting fundamental limitations in current scientific language models.

Speaker: Dikshya Mohanty

Location: NCS 115/Online

Zoom: https://stonybrook.zoom.us/j/94746001760?pwd=BCAd8gu7cXLn3PXM6kkbh11V6r0Mr7.1
Meeting ID: 947 4600 1760 Passcode: 987917

Abstract: Computational pathology has revolutionized cancer diagnosis and research through the analysis of digitized whole slide images (WSIs). However, their giga-pixel size creates two intertwined bottlenecks: computational inefficiency, as prohibitive GPU memory makes standard end-to-end (E2E) training infeasible, and label inefficiency, as expert annotation is tedious and expensive. This dissertation confronts both challenges through novel architectures, training paradigms, and self-supervised learning methods for efficient WSI analysis.
To improve computational efficiency, this dissertation first introduces a locally supervised learning paradigm that enables E2E training on entire WSIs by partitioning a network into gradient-isolated modules, circumventing the memory bottleneck of backpropagation. Second, it presents Prompt-MIL, a parameter-efficient fine-tuning framework training only a few prompts to guide large pre-trained models, reducing trainable parameters, memory, and training time. Third, this work proposes 2DMamba, the first intrinsic Mamba architecture that preserves the crucial 2D spatial structure of images, overcoming the spatial discrepancy in 1D models. Fourth, it presents Locally Bi-directional Mamba (LBMamba), whose hardware-aware local backward scan integrates bi-directional scanning into a single forward pass, improving the throughput-performance trade-off of Mamba models.
To improve label efficiency, this dissertation proposes a precise location based matching strategy for self-supervised dense contrastive learning, which allows a local patch in one augmented view to match multiple overlapping patches in another, producing more accurate correspondences and superior features for dense prediction tasks like segmentation and detection. Additionally, to better scale multi-channel cell imaging modalities, this dissertation introduces ChannelSFormer, a channel-agnostic vision transformer that disentangles spatial and channel-wise reasoning through divided attention and channel class token, enabling effective representation learning across variable channel configurations in both self-supervised and supervised settings.
In summary, this dissertation presents a holistic investigation into the efficiency bottlenecks in computational pathology. Through these combined contributions in model architecture, training paradigms, and self-supervised learning, this work establishes a more scalable, efficient, and powerful computational framework for analyzing giga-pixel pathology images.

Speaker: Jingwei Zhang

Location: NCS 220

Zoom: https://stonybrook.zoom.us/j/93175806292?pwd=xbtxnQyYGoThz5B1DyJxJxPF9lxiJE.1
Meeting ID: 931 7580 6292
Passcode: 314091
CSE 600 Talk: Haibin Ling - Computer Vision Research and Applications


Abstract: Having been intensively studied over half a decade, computer vision has evolved as a broad research area and become mature in many applications. In this talk, we will summarize our work in computer vision in both core vision topics and application-oriented ones. In particular, for core vision problems, we will report studies on visual tracking, visual matching and visual detection; for applications, we will describe our work on medical image analysis, intelligent transportation, smart projector systems and preliminary work on material property prediction.

Bio: Haibin Ling received the BS and MS degrees from Peking University in 1997 and 2000, respectively, and the PhD degree from the University of Maryland, College Park, in 2006. From 2000 to 2001, he was an assistant researcher at Microsoft Research Asia. From 2006 to 2007, he worked as a postdoctoral scientist at the University of California Los Angeles. In 2007, he joined Siemens Corporate Research as a research scientist. From 2008 to 2019, he worked as a faculty member of Temple University. In fall 2019, he joined the Department of Computer Science of Stony Brook University where he is currently a SUNY Empire Innovation Professor. His research interests include computer vision, augmented reality, medical image analysis, and human computer interaction. He received the Best Student Paper Award at the ACM UIST in 2003, and the NSF CAREER Award in 2014. He serves as Associate Editor for several journals including IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Pattern Recognition (PR), and Computer Vision and Image Understanding (CVIU). He has served or will serve as Area Chair for CVPR 2014, 2016, 2019 and 2020.

Ready for Round Two? Dr. Zach Justus Returns! Join us on October 30, 2025, in the SBU Hilton Garden Inn. Buckle up your curiosity for a high-energy morning session with the engaging Dr. Zach Justus as we navigate how GenAI is reshaping not just how we teach, but what we teach. With real talk and questions that hit hard like Are students learning what we think we're teaching? This is your chance to rethink your program's true destination. Whether you're looking to pick up a few takeaways or chart a new direction entirely, this symposium is your space to explore, reflect, and act.

Check-in and breakfast will begin at 8:30 a.m. in order to begin our program promptly at 9:00 a.m.

Registration will remain open until October 15 or until the event reaches capacity. If closed, please contact educationaleffectiveness@stonybrook.edu to request a spot on the waitlist.

Join Zoom Meeting
https://stonybrook.zoom.us/j/91945227869?pwd=emhoZDFWVTV0MVdPWW5uVk43MjQzUT09

Meeting ID: 919 4522 7869
Passcode: 452304
One tap mobile
+16468769923,,91945227869# US (New York)
+13126266799,,91945227869# US (Chicago)

Dial by your location
        +1 646 876 9923 US (New York)
        +1 312 626 6799 US (Chicago)
        +1 301 715 8592 US (Germantown)
        +1 669 900 6833 US (San Jose)
        +1 253 215 8782 US (Tacoma)
        +1 346 248 7799 US (Houston)
        +1 408 638 0968 US (San Jose)
Meeting ID: 919 4522 7869
Find your local number: https://stonybrook.zoom.us/u/aCvAYWkRg
  




https://stonybrook.zoom.us/j/91775729097pwd=Qlc5Nks0NmlyKzJwMjR0S0hrdVZ3QT09

Meeting ID: 917 7572 9097
Passcode: 555459


Abstract: As the saying goes, there are many ways to skin a cat.
While we don't want to go around skinning cats, the world of
optimization is rich with different problems, problem formulations,
and methods and approaches, each with different guarantees and
computational benefits. In this talk we will take a tour down the
problem of structured sparsity in sensing to see how one simple
problem can inspire a wide range of analysis and tools. First, I will
present the optimality conditions for a generalized structured sparse
problem, which can be geometrically visualized as alignment of vectors
and matrices. Then I will introduce three approximation methods for
the problem of phase retrieval, which are a twist on stochastic
gradient and coordinate descent methods. These methods leverage
fundamental numerical linear algebra concepts to give fast approximate
solutions to large-scale problems, which then after postprocessing can
produce more reliable sensing results.

Bio: Yifan Sun received her PhD in Electrical Engineering from the
University of California Los Angeles in 2015, with research focusing
on convex optimization and semidefinite programming. She was then
Technicolor Research and Innovation, focusing on machine learning and
data science applications. More recently, she completed two postdocs,
at the University of British Columbia in Vancouver, Canada and
L'Institut National de Recherche en Informatique et Automatique
(INRIA) in Paris, France.
Abstract: In recent years, we have been developing generative AI methods to design increasingly complex objects. Our goal is to improve performance while ensuring that these objects remain controllable. This requires addressing several challenging problems, including:
  • Modeling composite objects in a way that preserves consistency under deformation.
  • Estimating the uncertainty of the surrogate models used to predict performance during optimization.
  • Co-designing objects for both high performance and ease of control.
In this talk, I will present our approach to these challenges and describe end-to-end design pipelines that have the potential to radically transform computer-assisted engineering.

Speaker: Pascal Fua received an engineering degree from Ecole Polytechnique, Paris, in 1984 and a Ph.D. in Computer Science from the University of Orsay in 1989. He joined EPFL (Swiss Federal Institute of Technology) in 1996, where he is a Professor in the School of Computer and Communication Science and head of the Computer Vision Lab. Before that, he worked at SRI International and at INRIA Sophia-Antipolis as a Computer Scientist.
His research interests include shape modeling and motion recovery from images, analysis of microscopy images, and machine learning. He has (co)authored over 400 publications in refereed journals and conferences. He has received several ERC grants. He is an IEEE Fellow and has been an Associate Editor of the IEEE journal Transactions For Pattern Analysis and Machine Intelligence. He often serves as the program committee member, area chair, and program chair of major vision conferences and has cofounded three spinoff companies.