CSE 600 Talk: Haibin Ling - Computer Vision Research and Applications


Abstract: Having been intensively studied over half a decade, computer vision has evolved as a broad research area and become mature in many applications. In this talk, we will summarize our work in computer vision in both core vision topics and application-oriented ones. In particular, for core vision problems, we will report studies on visual tracking, visual matching and visual detection; for applications, we will describe our work on medical image analysis, intelligent transportation, smart projector systems and preliminary work on material property prediction.

Bio: Haibin Ling received the BS and MS degrees from Peking University in 1997 and 2000, respectively, and the PhD degree from the University of Maryland, College Park, in 2006. From 2000 to 2001, he was an assistant researcher at Microsoft Research Asia. From 2006 to 2007, he worked as a postdoctoral scientist at the University of California Los Angeles. In 2007, he joined Siemens Corporate Research as a research scientist. From 2008 to 2019, he worked as a faculty member of Temple University. In fall 2019, he joined the Department of Computer Science of Stony Brook University where he is currently a SUNY Empire Innovation Professor. His research interests include computer vision, augmented reality, medical image analysis, and human computer interaction. He received the Best Student Paper Award at the ACM UIST in 2003, and the NSF CAREER Award in 2014. He serves as Associate Editor for several journals including IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Pattern Recognition (PR), and Computer Vision and Image Understanding (CVIU). He has served or will serve as Area Chair for CVPR 2014, 2016, 2019 and 2020.
Towards Saving Lives with Natural Language Processing Andrew Schwartz Dept. of Computer Science Stony Brook Analyzing language use patterns is proving to be a valuable and unique approach to understanding the psychological, social, and health factors of people. On the individual level, Facebook and Twitter have been found predictive of mental health, personality, demographics, and occupational class (among others). At the community or county-level, Twitter has been found predictive of flu and allergy outbreaks, life satisfaction, atherosclerotic heart disease mortality, health behavioral risk factors, excessive drinking, and HIV prevalence. While these techniques have shown robust links over a plethora of important aspects of human life, it is not clear whether any lives have been saved, at least directly, by such work. At their core, some barriers to improving health care and saving lives are likely not NLP or even AI problems, but others are perhaps technical in nature and suggest changing the way we model data. This seminar will have two parts: a presentation and a discussion. I will start by going over recent and on-going work toward predicting mental health outcomes --- depression, addiction relapse, future psychological distress --- from human language use patterns. Then, I will present an imperfect vision of a future where NLP helps to save lives and open the floor for discussion of technical barriers and whether such a vision is practical. Biography: Andrew Schwartz received his PhD in Computer Science from the University of Central Florida in 2011 with research on acquiring lexical semantic knowledge from the Web. He then joined the University of Pennsylvania where he was a Postdoctoral Research Fellow and later Visiting Assistant Professor in Computer & Information Science. He is Lead Research Scientist for the World Well-Being Project, a multidisciplinary group of Computer Scientists and Psychologists studying physical and psychological well-being based on language in social media.
Join us for the New York State Innovation Summit on October 28-29, 2024 in Syracuse, NY. This multi-day is event for NYS organizations that want to showcase and discover new and emerging technologies that support innovation and drive business growth. The event serves as an opportunity to foster collaboration; introduce industry to experts that can assist growth, strengthen our statewide innovation ecosystem and showcase promising early stage companies. Whether you're a startup, an economic developer, or an established manufacturer, the NYS Innovation Summit is for you. The 2024 New York State Innovation Summit will showcase companies and researchers at the forefront of emerging technologies and new advancements in production capabilities. This event celebrates New York State leadership in technology-led economic growth with experts in biotechnology, new materials, energy innovation, and artificial intelligence that will explore current technology convergence opportunities, ways to accelerate commercialization, and issues of manufacturing sustainability.
The Future of Learning: Rethinking Practice in a Changing World

Thursday, March 26, 2026 (Workshops)
Friday, March 27, 2026 (Symposium)

Open to Stony Brook University Faculty, Staff, and Graduate Students. Hosted by the Center for Excellence in Learning and Teaching, Office of the Provost.

Thursday, March 26, 2026
Workshop: AI Tools and Techniques
  • Open to all faculty & staff
  • Hands-on, exploratory
  • Registration only limited to the size of the room
  • Location: In-person, TBD
  • Time: 10 AM - 12 PM
  • Registration required

Friday, March 27, 2026
Keynote: Teaching and Thinking with AI
  • Faculty, TAs, postdocs, and academic staff
  • In-person on-campus conference venue
  • Location: SAC Balroom
  • Time: 9 AM - 3 PM
  • Registration required

Keynote Speaker: José Antonio Bowen

José Antonio Bowen has been leading innovation and change for over 40 years at Stanford, Georgetown and the University of Southampton (UK), as a dean at Miami University and SMU and as President of Goucher College. Bowen has worked as a musician with Stan Getz, Dave Brubeck, and many others and his symphony was nominated for the Pulitzer Prize in Music (1985).
Bowen holds four degrees from Stanford and has written over 100 scholarly articles and books, including the Cambridge Companion to Conducting (2003), Teaching Naked (2012 and the winner of the Ness Award for Best Book on Higher Education), Teaching Naked Techniques with C. Edward Watson (2017) and Teaching Change: How to Develop Independent Thinkers using Relationships, Resilience and Reflection (Johns Hopkins University Press, 2021).
Bowen has appeared in The New York Times, Forbes, The Wall Street Journal, and has three TED talks. Stanford honored him as a Distinguished Alumni Scholar (2010) and he has presented keynotes and workshops at more than 300 campuses and conferences 46 states and 17 countries around the world. In 2018, he was awarded the Ernest L. Boyer Award (for significant contributions to American higher education). He is a senior fellow for the American Association of Colleges and Universities.

Register here.

This course is designed to help you approach AI with clarity and intention. Explore how to ask better questions, apply critical judgment to AI-generated outputs, and use these tools to support--not replace--human decision-making. With the right prompts, a thoughtful lens, and a focus on impact, AI can help reduce friction and free you to focus on what matters most: people, purpose, and leadership.

Audience: Any current or aspiring leader who currently uses AI and is interested in bringing their skills to the next level.

Register here.

Abstract:

Photorealistic editing of human facial expressions and head articulations remains a long-standing topic in the computer graphics and computer vision community. Methods enabling such control have great potential in AR/VR applications where a 3D immersive experience is valuable, especially when this control extends to novel views of the scene in which the human subject appears. Traditionally, 3D Morphable Face Models (3DMMs) have been used to control the facial expressions and head pose of a human head. However, the PCA-based shape and expression spaces of 3DMMs lack the expressivity. They cannot model essential elements of the human head such as hair, skin details, and accessories such as glasses that are paramount for realistic reanimation. In this thesis, we present a set of methods that enables facial reanimation, starting from editing expressions in still face images to creating fully controllable neural 3D portraits with control over facial expressions, head pose, and viewing direction of the scene using only casually captured monocular videos from a smartphone to finally achieving studio-like quality from the said monocular captures.
First, we propose a method for editing facial expressions in near-frontal facial images through the unsupervised disentangling of expression-induced deformations and texture changes. Next, we extend facial expression editing to human subjects in 3D scenes. We represent the scene and the subject in it using a semantically guided neural field. This enables control over the subject's facial expressions and the viewing direction of the scene they're in. We then present a method that learns, in an unsupervised manner, to deform static 3D neural fields using facial expression and head-pose dependent deformations, enabling control over facial expressions and head pose of the subject along with the viewing direction of the 3D scene they're in. Next, we propose a method that makes the learning of the aforementioned deformation field robust to strong illumination effects, which adversely impact the registration of the deformation. We then propose an extension of this unsupervised deformation model to 3D Gaussian splatting by constraining it using a 3D morphable model, resulting in a rendering speed of 18 FPS--a 100x speed improvement over prior work. Finally, we propose a method that bridges the quality gap between 3D portraits created using in-the-wild monocular data and multi-view studio capture data. We accomplish this using a two-stage method. First, we train a StyleGAN to relight and inpaint in-the-wild face texture maps (with strong illumination effects and incompletely captured regions). Next, we both reconstruct and generate identity-specific facial details that may be poorly captured in the in-the-wild captures. Once trained, we can generate studio-like complete avatars from monocular phone captures.

Speaker: Shahrukh Athar

Zoom Link:
https://stonybrook.zoom.us/j/94228500743?pwd=RqOBgG6tbJkKaFBlWFwBkYFX0VRovV.1

Meeting ID: 94228500743
Passcode: 661599
Abstract: Humans perceive the world around them by recognizing global patterns and structures such as object parts, branches, their spatial arrangement, and so on. Most deep learning models, however, take a fundamentally local approach. They process images pixel-by-pixel rather than focusing on structures as a whole. While these models indeed perform well on many tasks, the local (pixel-level) versus global (structure-level) disconnect makes them harder to interpret and control.

Topology, in a general sense, is a mathematical language for describing structure. It delineates how different parts of an image relate to one another, capturing both individual structures and their overall layout. Preserving topology enforces structural correctness and, by extension, semantic validity.

In this thesis, we investigate how topological constraints can be used to bridge the gap between local and global understanding. We use topology to inform the design of deep learning models that are explicitly structure-aware. Our thesis focuses on dense prediction tasks, which include image segmentation, uncertainty estimation, and generative modeling. First, we introduce a topological interaction module for semantic segmentation that encodes containment and exclusion constraints directly into the learning process. This preserves anatomical hierarchies and improves multi-class consistency. Next, since segmentation models can never be truly perfect, we address the need for reliable uncertainty estimation to identify error-prone regions. Unlike conventional pixel-wise uncertainty maps, which tend to be noisy and difficult to interpret, we propose reasoning at the level of structural units--branches and connections--which are more visually discernible and actionable. Finally, we leverage topology for generative modeling. We propose a topology-guided diffusion framework that can be controlled using structural attributes like object count and connectivity.

Together, these contributions establish a unified approach to topology-informed, structure-preserving dense prediction models. By integrating topological reasoning with deep networks, this thesis advances models that are not only accurate, but also structurally consistent, interpretable, and controllable. The results from this thesis have been published in ECCV, NeurIPS, and ICLR.

Speaker: Saumya Gupta

Location: New Computer Science (NCS) 120


Zoom: https://stonybrook.zoom.us/j/93643318604?pwd=kv8DagpbayzizivU29UCYItnlzlYRM.1&jst=2