Abstract: Visual generation is a fundamental problem in computer vision and graphics, with applications ranging from 3D capture to content creation and image/video synthesis. Despite rapid progress in neural rendering and generative models, efficiency remains a key obstacle in practice: high-quality 3D reconstruction often depends on dense multi-view supervision; scalable 3D synthesis faces heavy optimization, training, and rendering costs; and modern image/video generators incur substantial computation as token grids grow with spatial resolution and temporal length.
This thesis targets efficient visual world modeling by improving sample efficiency in 3D reconstruction, representation efficiency in 3D generation, and computational efficiency in image/video synthesis. First, we improve sample efficiency for neural implicit surface reconstruction under sparse views by integrating multi-view stereo probability volumes as a geometric regularizer, enabling high-quality reconstruction from as few as three input images. Next, we introduce an explicit 3D representation for 3D generation, built from multi-view depth and RGB predictions with 3D Gaussian features, which enables the use of 2D generative priors while enforcing multi-view consistency via epipolar attention. We then address the computational bottleneck of image and video synthesis with importance-based token merging, using importance signals available during generation to preserve critical information while merging redundant tokens. Finally, we propose efficient mixed-resolution diffusion transformers via cross-resolution phase-aligned attention, aiming to improve attention stability under mixed token grids and support high-fidelity mixed-resolution generation.

Speaker: Haoyu Wu

Location: NCS120
The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 10 minutes) by multiple people. Students can register for 1 credit for CSE 656. Registered students must attend and present a minimum of 2 or 3 talks. Everyone else is welcome to attend. Fill in https://forms.gle/pCVXovgfMfQwGqG38 to subscribe to our mailing list for further announcement.

The Della Pietra Lecture Series is pleased to present this lecture by Scott Aaronson:

Abstract: I'll survey some areas where I think theoretical computer science, math, and statistics can potentially contribute to the urgent quest to align powerful AI with humane values. These areas include: the watermarking of AI outputs, mechanistic interpretability (including Paul Christiano's No-Coincidence Principle, and succinct digests of the training process to aid interpretability), and theoretical guarantees for out-of-distribution generalization.

Speaker: Scott Aaronson is Schlumberger Chair of Computer Science at the University of Texas at Austin, and founding director of its Quantum Information Center. He received his bachelor's from Cornell University and his PhD from UC Berkeley. Aaronson's research has focused mainly on the capabilities and limits of quantum computers. His first book, Quantum Computing Since Democritus, was published in 2013 by Cambridge University Press. He received the National Science Foundation's Alan T. Waterman Award, the United States PECASE Award, the Tomassoni-Chisesi Prize in Physics, and the ACM Prize in Computing, and is a Fellow of the ACM and the AAAS and a member of the National Academy of Sciences. He blogs at Shtetl-Optimized, https://www.scottaaronson.com/blog.

Location: Della Pietra Family Auditorium (SCGP 103)
Abstract: Large language models (LLMs) may exhibit unintended or undesirable behaviors. Recent works have concentrated on aligning LLMs to mitigate harmful outputs. Despite these efforts, some anomalies indicate that even a well-conducted alignment process can be easily circumvented, whether intentionally or accidentally. Does alignment fine-tuning yield have robust effects on models, or are its impacts merely superficial? In this work, we make the first exploration of this phenomenon from both theoretical and empirical perspectives. Empirically, we demonstrate the elasticity of post-alignment models, i.e., the tendency to revert to the behavior distribution formed during the pre-training phase upon further fine-tuning. Leveraging compression theory, we formally deduce that fine-tuning disproportionately undermines alignment relative to pre-training, potentially by orders of magnitude. We validate the presence of elasticity through experiments on models of varying types and scales. Specifically, we find that model performance declines rapidly before reverting to the pre-training distribution, after which the rate of decline drops significantly. Furthermore, we further reveal that elasticity positively correlates with the increased model size and the expansion of pre-training data. Our findings underscore the need to address the inherent elasticity of LLMs to mitigate their resistance to alignment.

Speaker: Huajian Zhang

Location: CS2311
What AI tools are available to help with the scholarly research process? Are they helpful? What do they do and is it worth the time and energy to try them out? Join librarian Christine Fena to explore and compare established and emerging AI research tools such as Elicit, Scite, Consensus, and Undermind. The online workshop will provide a starting point to understanding what these tools are, the basics of how they work, and how AI research assistants might bring changes to your search process in the future. All are welcome!



Register here via Zoom.
Talk by Zhenhua Liu to be followed by AI Institute updates


Abstract: Decision making with uncertainty has been studied in multiple communities extensively. Recently, online optimization has gained popularity partially because of its promising performance guarantees by incorporating predictions. In this talk, I will provide an overview of our work on algorithm designs for online optimization and its applications. Then, I will talk about our recent work in ACM Sigmetrics 2019 on choosing predictions and control algorithms simultaneously and dynamically. Finally, I will discuss some ongoing efforts and collaboration opportunities.

Bio: Zhenhua Liu is currently an assistant professor in the Department of Applied Mathematics and Statistics at Stony Brook University. He is also affiliated with the Department of Computer Science, the AI Institute and the Smart Energy Technology Cluster. He received his PhD degree in Computer Science from California Institute of Technology. His current research interests include cloud computing, online optimization and learning, smart grid, market design and distributed control. His research combines rigorous analysis and system design, and goes from theory, to prototype, and eventually to industry to make real impacts.
Abstract: The capacity to adapt machine learning models to various contexts, information, and objectives is particularly valuable. In this thesis, I focus on developing Class Conditional Guided Models. These are models that can be adaptively biased towards a class of interest via a conditional input. My primary focus lies in the efficiency of these models. They are constructed to require training only once, with the ability to quickly and conveniently adapt during testing time without necessitating fine-tuning or retraining.
Firstly, I propose RelationVAE, a novel generative model designed for few-shot scenarios, utilizing the prior knowledge of class similarity relationships. RelationVAE is designed to condition on the embeddings of the neighbor classes (i.e. classes with similarity relationships), to generate more reliable samples by making them more similar to the neighbor class. This enables adaptation of the generative model to the provided prior knowledge about class relationships.
As a second focus, I introduce scGAN, a shadow segmentation technique that enables adaptation to varying shadow distributions in different testing environments. scGAN is designed to condition on a sensitivity parameter, a scalar, to control the amount of the shadow detected. In the testing phase, the parameter is set to appropriate values, allowing the model to quickly adapt to specific test environments.
In my third contribution, I propose S-SEG, a methodology for fine-grained counting allowing adaptation to different granularities of fine-grained classes. In fine-grained problems, the distinction between classes is subtle and inconsistent across images, leading to variations in the granularity of the target class from one image to another. S-SEG is designed to be conditioned on an additional input, the sensitivity parameter, to control the granularities of the target class during inference.
My fourth contribution is a text-to-image synthesis method which allows controlling the number of the generated objects of a target class. I propose to generate an intermediate condition, the density map, which reflects the number of objects, together with their layout. This intermediate condition is used to effectively guide the generative model to generate objects with accurate counts.

Speaker: Vu Nguyen

Zoom: https://stonybrook.zoom.us/j/97114455337?pwd=Z4rB9dWcstlahUIs8PRrvQ9b2ZK2Df.1
Meeting ID: 971 1445 5337
Passcode: 272300
The Stony Brook Computing Society presents an exciting event featuring experts from Google (Danny Rosen - Technical Program Manager) and NVIDIA (Veer Mehta - Senior Solutions Architect), diving into the latest developments in generative AI. Learn how these industry leaders are shaping the future of technology and explore new ideas in a relaxed, engaging setting.

📍 Location: Frey 102
📅 Date: Monday, Nov 11
⏰ Time: 12 PM - 1:50 PM

Scan the QR code or register in the link.
Making sense of Twitter @ Bloomberg presented by Daniel Preotiuc-Pietro

ABSTRACT: The Bloomberg Terminal has provided ways for investors and journalists to sift through and understand the immense volume of tweets and discover financially-relevant content ever since the SEC approved the use of Twitter for company disclosures back in 2013.

In the first part of the talk, I will showcase how tweets impact financial markets and how Bloomberg is using Natural Language Processing methods to identify financially relevant tweets that move the markets. Our processing pipeline feeds directly to clients, journalists in the newsroom and powers several news analytic products offered by the company including trending companies and consumer sentiment for publicly traded equities.

However, understanding user pragmatic intent in individual tweets would allow us to gain deeper insights and enable new applications. I will present several recent research studies focused on understanding intent including identifying complaints and the roles with which vulgarity is used in social media and how these can help improve applications such as sentiment analysis and hate speech detection.

BIO: Daniel Preotiuc-Pietro is a Senior Research Engineer and Team Lead at Bloomberg LP, where he works on analyzing and building models for real-world large scale social media and news mining and information extraction. His research interests are focused on understanding the social and temporal aspects of text, especially from social media, with applications in domains such as Social Psychology, Law, Political Science and Journalism. Several of his research studies were featured in popular press including the Washington Post, BBC, New Scientist, Scientific American or FiveThirtyEight. He is a co-organizer of the Natural Legal Language Processing workshop series. Prior to joining Bloomberg LP, Daniel was a postdoctoral researcher at the University of Pennsylvania with the interdisciplinary World Well Being Project and obtained his PhD in Natural Language Processing and Machine Learning at the University of Sheffield, UK.
Topic: AI Seminar: Owen Rambow
Time: Mar 17, 2021 10:00 AM Eastern Time (US and Canada)
Join Zoom Meeting

https://stonybrook.zoom.us/j/93614644178?pwd=MzJtVDJYYmU5T1dtMzJiUFMxb0x4dz09
Meeting ID: 936 1464 4178.    Passcode: 965936






Natural Language Understanding and Semantic Parsing

(Partly joint work with former colleagues at Elemental Cognition)

Semantic parsing refers to the task of determining the propositional content of language: who did what to whom.  It is part of the larger task of natural language understanding (NLU).  I will start out by discussing what full NLU means, and argue that we are still far away, as a field, from solving full NLU, or even from knowing how to evaluate it.

In the second part of the talk, I will situate semantic parsing in the context of several other NLU subtasks.  Typically, the target representation of semantic parsing uses an ontology (such as PropBank or FrameNet).  Semantic parsing includes the subtasks of word sense disambiguation, argument detection, and argument role labeling.  I will discuss choices among possible target ontologies.  I will justify why we created a new ontology, Hector, based on FrameNet and the lexical resource NOAD, and explain some of its characteristics.

In the third part of the talk, I will present experiments we performed using transformer models.  We obtain best results using a two-phase model, in which we first choose the frame, and then, given the frame, choose the arguments.  We encode the problem for both tasks using indices in the sentence.  While we develop the parser for our new ontology Hector, this approach also beats the state of the art for FrameNet and PropBank parsing.Biography:  I am a professor in the Department of Linguistics at Stony Brook University with a joint appointment in IACS.

Until recently, I was a research scientist at Elemental Cognition. Elemental Cognition is working on deep natural language understanding.

I got my PhD with Aravind Joshi at the University of Pennsylvania in 1994. I have worked at CoGenTex, and at AT&T Labs -- Research, and for many years I was a research scientist at Columbia University in the Center for Computational Learning Systems.