ABSTRACT: Recent progress in deep neural networks has revolutionized many computer vision tasks such as image classification, detection and segmentation. However, in addition to excelling in tasks that predict well-defined objective information, human-centered artificial intelligence systems should also be able to model subjective attributes, as defined by human perceptual behavior, that goes beyond the pure physical content of visual data. Example subjective tasks are the prediction of spatial or temporal regions that are interesting to humans (e.g., attract attention or are visually pleasing) and the recognition of subjective attributes (e.g., visually elicited sentiments). Better models for these tasks will improve the human-computer interaction experience in various applications. This thesis investigates several approaches to address the challenges in predicting those subjective attributes in visual data over a diverse set of tasks. I first present a novel framework for real-time automatic photo composition. The framework consists of a cost-effective data collection workflow, an efficient model training pipeline and a lightweight module to account for personalized preferences. Then I develop a novel and general algorithm to detect interesting segments in sequential data, which can be naturally applied to video summarization tasks. Furthermore, I propose methods that learn to represent sentiments elicited by images, in an unsupervised manner, using linguistic features extracted from large scale Web data. To conclude this thesis, I introduce a human-vision-inspired image classification algorithm that also predicts spatial visual attention even though no attention data was used for training it.
- Lav Varshney, PhD, MS, BS - Director of the AI Institute at Stony Brook University
- Meghan Reading Turchioe, PhD, MPH, RN, FAHA - Nurse Scientist at Columbia University
- Briana Shiri Last, PhD - Department of Psychology at Stony Brook University
This event is proudly sponsored by the Stony Brook University School of Nursing, the Stony Brook University Hospital Division of Nursing, and the Sigma Theta Tau Kappa Gamma Chapter.
As artificial intelligence transforms our world, what skills will remain uniquely human? How can we prepare for careers in an automated future?
Join Carnegie Mellon mathematics professor Po-Shen Loh for insights on navigating the AI revolution by embracing our humanity.
Dr. Loh brings a distinctive perspective shaped by his dual expertise: serving as national coach of the USA Mathematical Olympiad team (which has won four gold medals under his leadership) and developing innovative solutions for real-world challenges from pandemic response to educational technology.
Through his nationwide speaking tour that reached 250 audiences across 100 cities, he has refined a practical framework for thriving alongside AI.
In this presentation, Dr. Loh will explore how creative problem-solving, judgment, and communication become more valuable as automation grows -- and how students and professionals can build those strengths now.
The session includes real-world examples, guidance for education and careers, and a Q&A.
Speaker: Po-Shen Loh is a social entrepreneur and inventor, working across the spectrum of mathematics, education, and healthcare.
A math professor at Carnegie Mellon University, he also served a decade-long term as the national coach of the USA International Mathematical Olympiad (IMO) team, taking the team to gold on numerous occasions.
He has pioneered numerous innovations and has been featured in or co-created YouTube videos with more than 25 million views.
Location: Wang Center Theater
The series is offered by Stony Brook University's Institute for Creative Problem Solving in collaboration with the National Museum of Mathematics (MoMath) and Brookhaven National Laboratory.
The event is free but space is limited. Please register to reserve your space.
Abstract: Having been intensively studied over half a decade, computer vision has evolved as a broad research area and become mature in many applications. In this talk, we will summarize our work in computer vision in both core vision topics and application-oriented ones. In particular, for core vision problems, we will report studies on visual tracking, visual matching and visual detection; for applications, we will describe our work on medical image analysis, intelligent transportation, smart projector systems and preliminary work on material property prediction.
Bio: Haibin Ling received the BS and MS degrees from Peking University in 1997 and 2000, respectively, and the PhD degree from the University of Maryland, College Park, in 2006. From 2000 to 2001, he was an assistant researcher at Microsoft Research Asia. From 2006 to 2007, he worked as a postdoctoral scientist at the University of California Los Angeles. In 2007, he joined Siemens Corporate Research as a research scientist. From 2008 to 2019, he worked as a faculty member of Temple University. In fall 2019, he joined the Department of Computer Science of Stony Brook University where he is currently a SUNY Empire Innovation Professor. His research interests include computer vision, augmented reality, medical image analysis, and human computer interaction. He received the Best Student Paper Award at the ACM UIST in 2003, and the NSF CAREER Award in 2014. He serves as Associate Editor for several journals including IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Pattern Recognition (PR), and Computer Vision and Image Understanding (CVIU). He has served or will serve as Area Chair for CVPR 2014, 2016, 2019 and 2020.
The world of college is changing fast, and Artificial Intelligence (AI) is at the center of it. We are part of the Institute on AI, Pedagogy, and the Curriculum with AAC&U, and we need to hear from the people AI affects most: you!
This is an open discussion for all students to share their honest experiences, their top concerns, and their best ideas about AI in our academic environment. We'll be diving into these key questions:
- How can AI actually make learning better or easier? What opportunities do you see for using AI tools to enhance your assignments, research, or skills?
- What are your biggest worries about AI? Is it about cheating, being graded fairly, or preparing for the job market? How is AI impacting your workload or stress levels?
- What specific tools, workshops, or policies would help you use AI responsibly and successfully? (Think training, software, or clear rules.)
Time: 12:30pm-1:45pm
Location: West Campus - Location TBD
or
Date: Wednesday, December 3rd
Time: 10:30am-11:45am
Location: East Campus - HSC 2-154B
Please register in advance so we can confirm the room.
Note: Videos will not be shared publicly and comments will only be shared in aggregate.
Your voice matters. Come tell us how AI is affecting your studies, your stress, and your success!
- Dr. Rose Tirotta-Esposito (Assistant Provost; Director of CELT)
- Dr. Elizabeth Hewitt (Associate Professor in the Department of Technology and Society (DTS) in the College of Engineering and Applied Sciences)
- Chris Kretz (Associate Librarian and Head of Academic Engagement at SBU Libraries)
- Prof. Rajiv Lajmi (Assistant Professor in the School of Health Professions and Chair of Applied Health Informatics)
- Dr. Matthew Salzano (Assistant Professor in the Department of Communication in the School of Communication and Journalism)
Despite their promise, VLAs exhibit a critical limitation: they function primarily as trajectory learners rather than skill learners. Recent evaluations reveal that VLAs often fail when faced with even minor variations in object initialization or environmental conditions, suggesting they memorize specific trajectories rather than acquiring generalizable manipulation skills. Attempts to address this through 3D spatial representations have shown limited success, indicating that the missing component may be more fundamental than geometric understanding alone.
This work argues that World Models (WMs)---internal representations that predict future states given actions---constitute the missing piece for robust VLA systems. We present one completed contribution and two ongoing investigations.
We developed a dual-layer world model for human-robot interaction that anticipates both physical scene evolution and latent human preferences for assistive tasks. Building on these foundations, we present ongoing work probing VLA internal representations to verify implicit world model existence, and propose a WM-VLA integration approach operating in the native visual domain through embedding prediction and image decoding.
Together, these contributions and investigations establish a foundation for WM-VLA systems, pointing toward robust, generalizable robot policies.
Speaker: Jason Qin
Location: NCS 220
A lecture by-
Chris Wiggins
Columbia University and
Matthew L. Jones
Princeton University
The co-authors of the book How Data Happened will trace the dynamic relationships among data, truth, and power, exploring how data-empowered algorithms have come to shape our personal, professional, and political realities.
Location: 1008 Humanities