Abstract: Visual generation is a fundamental problem in computer vision and graphics, with applications ranging from 3D capture to image and video synthesis. Despite rapid progress in neural rendering and generative models, efficiency remains a key bottleneck: high-quality 3D reconstruction often relies on dense multi-view supervision; high-fidelity 3D synthesis requires costly optimization, training, and rendering; and modern image and video generators require substantial computation as the number of tokens grows rapidly for high-resolution generation. This dissertation focuses on efficient visual generation by improving sample efficiency in 3D reconstruction, representation efficiency in 3D generation, and computational efficiency in image and video synthesis. First, we improve the sample efficiency of neural implicit surface reconstruction. We integrate multi-view stereo probability volumes as a geometric regularizer, enabling high-quality sparse-view reconstruction from as few as three input images. Next, we introduce an explicit 3D representation for 3D generation, built from multi-view depth and RGB images with 3D Gaussian features. This design allows the model to use 2D generative priors while enforcing multi-view consistency through epipolar attention. We then address the computational bottleneck in image and video synthesis with importance-based token merging. Our method uses importance signals available during generation to preserve critical information while merging redundant tokens. Finally, we enable efficient mixed-resolution diffusion transformers via phase-aligned attention. This approach stabilizes attention under mixed-resolution token grids and unlocks high-fidelity image and video generation at reduced cost. Taken together, these contributions reduce the data requirements, representational overhead, and computational demands, thereby providing a foundation for high-quality, scalable, and efficient visual generation.

Speaker: Haoyu Wu

Location: NCS 120
Abstract: Recent work in NLP uses debates between multiple LLMs to arrive at a more accurate conclusion. Earlier chain-of-thought prompting also shows improvements in accuracy when the model is asked to provide step-by-step reasoning in its response. Many publications since have developed strategies to improve the reasoning of model output with the goal of generating a more accurate result. However, even when asked to provide problem solving steps, the content of the reasoning provided by models is not well studied for all tasks and sometimes contains errors or conflicting statements even when the final result is correct. In fact, when evaluated across reasoning tasks, evidence shows that LLMs are not learning how to reason but are instead mimicking relevant solutions from their training sets.
By studying and evaluating the argumentation that LLMs provide, we can determine factors that may benefit or hinder the model's ability to give a complete, cohesive, and thorough answer. While there are signs that LLMs pattern match, finding where, when, and why this fails is valuable, as there may be ways to help the model imitate solutions that are more relevant to the task it is attempting to solve. Determining when pattern matching is not enough could show an area of improvement for future generations of LLMs. This research may separately aid in work on human-(AI)agent and inter-agent interaction. Specifically, frameworks could be used to determine when and why other models or humans are convinced by LLM-generated responses and which argument methods cause other models to change their response. Our current research in systematic versus heuristic cues shows that large language models sometimes present systematic or heuristic reasoning patterns based on prompting. Future research aims to explore other methods of classifying argumentation.

Speaker: Kiera Gross

Joining link: https://meet.google.com/xae-ywpv-udo
  • CEWIT's 6th annual hackathon sponsored by Major League Hacking, Hack@CEWIT2022, is taking place virtually on February 18-20, 2022. This year's theme is Hacking Into the Metaverse and will focus on NFT's, Blockchain, Crypto, and the Metaverse. To find out more about the event, mentoring, sponsoring, or to register, visit:

  • https://www.cewit.org/programs/events/hack.php

Synthetic Dreams and Ghost Machines

Dr. Steven Skiena will join the Cinema for a presentation exploring the rapidly evolving worlds of artificial intelligence and robotics.

Dr. Skiena is Distinguished Teaching Professor of Computer Science and Associate Director of the AI Innovation Institute at Stony Brook University. A leading researcher in data science and algorithms, he is the author of several influential books on AI and computation, including The Algorithm Design Manual and The Data Science Design Manual.

The presentation will be followed by a screening of Ex Machina, Alex Garland's provocative science-fiction thriller exploring the seductive and unsettling boundaries between human consciousness and artificial intelligence.

Prepare your Business for the AI-driven future with DocItUSA's Document Management Solutions.
In today's digital world, businesses need to leverage the benefits of the Digital Cloud Age by streamlining Document Organization, Storage, and Accessibility.
Michael Feingold, of Digital Onesource Consulting Solutions/DOCS Consulting, Inc. will show you how DocItUSA equips your company with the tools to efficiently capture, classify, and retrieve documents and enable seamless AI integration.
https://nysbdc.ecenterdirect.com/events/1019400
Abstract: AI has achieved remarkable advancements in image recognition and natural language processing. However, its applications in Earth and environmental sciences are still emerging. Unprecedented data from satellites, sensors, and in-situ measurements oIers new opportunities to improve physics-based models and forecasts of environmental systems with AI and to gain deeper insights into these phenomena. Extreme systems, such as weather and climate events, pose distinct challenges for AI, such as limited sampling of rare events, non-trivial data augmentation, errors-in-variables, and complexities of transfer learning across diverse tasks. In this talk, we will explore some of these challenges and showcase AI architectures designed to address them. We will use specific examples of forecasting dust storms, precipitation extremes, flash floods, and drought events in the Middle East. Finally, we will discuss a different AI approach for studying sinkhole formation in the Dead Sea.

Speaker: Prof. Yinon Rudich, Department of Earth and Planetary Sciences, Weizmann Institute, Israel


Join Zoom Meeting
ID: 98731258879
Passcode: cJjGQJqP

Visual Analytics and Machine Learning for Biomedical Imaging Diagnosis

 

Arie Kaufman

 

We present an integrated approach using visual analytics and machine learning (ML) to diagnose abnormalities in 3D radiological imaging and biological microscopes. The primary example will involve 3D virtual pancreatography (VP), a novel visualization-ML procedure and application for non-invasive diagnosis and classification of pancreatic lesions, the precursors of pancreatic cancer. Currently, non-invasive screening of patients is performed through visual inspection of 2D axis-aligned CT images, though the relevant features are often not clearly visible nor automatically detected. VP is an end-to-end visual diagnosis system that includes an ML-based automatic segmentation of the pancreatic gland and the lesions, a semi-automatic approach to extract the primary pancreatic duct, an ML-based automatic classification of lesions into four prominent types, and specialized 3D and 2D exploratory visualizations of the pancreas, lesions and surrounding anatomy. We combine volume rendering with pancreas- and lesion-centric visualizations and measurements for effective diagnosis. We designed VP through close collaboration and feedback from expert radiologists, and evaluated it on multiple real-world CT datasets with various pancreatic lesions and case studies examined by the expert radiologists. Other applications include virtual colonoscopy, COVID-19, pathology, brain neurites, etc.


Biography: Arie Kaufman is Distinguished Professor and formerChair of the Department of Computer Science at Stony Brook University, where he is also Director of the Center for Visual Computing (CVC), and Chief Scientist at the Center of Excellence in Wireless and Information Technology (CEWIT). 

He received his PhD in Computer Science at Ben-Gurion University of the Negev in 1977.   He is known for his work in visualization, graphics, virtual reality, user interfaces, multimedia, and their applications, especially in bio-medicine. He is especially well known for his work on the 3-dimensional virtual colonoscopy, a revolutionary low-risk technique for colon cancer screening, and for pioneering the use of Graphics Processing Units (GPUs) and GPU-clusters. In 2012, he presided over the development and opening of the Reality Deck, the largest virtual reality display in the world, at Stony Brook University.

Kaufman was the founding Editor in Chief of IEEE Transactions on Visualization and Computer Graphics (TVCG), co-founded the IEEE Visualization Conference and Volume Graphics series, and is currently the director of IEEE Computer Society Technical Committee on Visualization and Graphics. He is an IEEE Fellow, ACM Fellow, winner of many awards, including the IEEE Visualization Career Award, and member of the European Academy of Sciences.



Steven Skiena is inviting you to a scheduled Zoom meeting.

Topic: AI Seminar: Arie Kaufman
Time: Apr 21, 2021 10:00 AM Eastern Time (US and Canada)

Join Zoom Meeting
https://stonybrook.zoom.us/j/96017498640?pwd=SE0rdHB6ZVlCM2ZpY2RnRUxyVnR3Zz09

Hidden Biases. Ethical Issues in NLP, and What to Do about Them presented by Dirk Hovy of Bocconi University

ABSTRACT: Through language, we fundamentally express who we are as humans. This property makes text a fantastic resource for research into the complexity of the human mind, from social sciences to humanities. However, it is exactly that property that also creates some ethical problems. Texts reflect the authors' biases, which get magnified by statistical models. This has unintended consequences for our analysis: If our data is not reflective of the population as a whole, if we do not pay attention to the biases contained, we can easily draw the wrong conclusions, and create disadvantages for our users.

In this talk, I will discuss several types of biases that affect NLP models, their sources, and potential counter measures: (1) Bias stemming from data, i.e., selection bias (if our texts do not adequately reflect the population we want to study), label bias (if the labels we use are skewed) and semantic bias (the latent stereotypes encoded in embeddings); (2) Biases deriving from the models themselves, i.e., their tendency to amplify any imbalances that are present in the data; (3) Design bias, i.e., the biases arising from our (the researchers) decisions which topics to analyze, which data sets to use, and what to do with them. For each bias, I will provide examples and discuss the possible ramifications for a wide range of applications, and various ways to address and counteract these biases, ranging from simple labeling considerations to new types of models.

BIO: Dirk Hovey is an associate professor of Computer Science in the department of marketing at Bocconi University. He received his PhD from the University of Southern California in Los Angeles, where he worked as a research assistant at the Information Sciences Institute. 

He works in Natural Language Processing (NLP), a subfield of artificial intelligence. His research focuses on computational social science. His interests include integrating sociolinguistic knowledge into NLP models, using large-scale statistics to model the interaction between people's socio-demographic profile and their language use, and ethics for data science and algorithmic fairness.