Abstract: Reward hacking, where a reasoning model exploits loopholes in a reward function to achieve high rewards without solving the intended task, poses a significant threat. This behavior may be explicit, i.e. verbalized in the model's chain-of-thought (CoT), or implicit, where the CoT appears benign thus bypasses CoT monitors. To detect implicit reward hacking, we propose TRACE (Truncated Reasoning AUC Evaluation). Our key observation is that hacking occurs when exploiting the loophole is easier than solving the actual task. This means that the model is using less effort than required to achieve high reward. TRACE quantifies effort by measuring how early a model's reasoning becomes sufficient to obtain the reward. We progressively truncate a model's CoT at various lengths, force the model to answer, and estimate the expected reward at each cut-off. A hacking model, which takes a shortcut, will achieve a high expected reward with only a small fraction of its CoT, yielding a large area under the reward-vs-length curve. TRACE achieves over 65% gains over our strongest 72B CoT monitor in math reasoning, and over 30% gains over a 32B monitor in coding. We further show that TRACE can discover unknown loopholes during training. Overall, TRACE offers a scalable unsupervised approach for oversight where current monitoring methods prove ineffective.

Speaker: Dikshya

Location: Old Computer Science Building - CS2311
The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors.

Each seminar will consist of multiple short talks (around 10 minutes) by multiple people. Students can register for 1 credit for CSE 656. Registered students must attend and present a minimum of 2 or 3 talks. Everyone else is welcome to attend. Fill in https://forms.gle/pCVXovgfMfQwGqG38 to subscribe to our mailing list for further announcement.
University Libraries Presents

In a digital landscape flooded with hyper-realistic AI images and manipulated media, telling fact from fiction has never been harder. This workshop introduces the art of slow looking, teaching you how to pause, carefully analyze visual evidence, and critically evaluate the authenticity of what you see online.

Register here to join
Abstract: Sea ice is crucial to Earth's climate, Arctic communities, and ecosystems, yet climate change is driving significant losses, threatening polar stability. Quantifying the long-term impacts of a declining sea ice cover requires tools which improve climate-timescale prediction and bring new understanding of climate interactions. In this talk, I discuss how meeting this challenge requires a multi-disciplinary approach. Climate models, while essential, suffer from systematic biases due to missing or inaccurate physics, leading to uncertainty in future projections. I show how data assimilation (DA) offers a statistical framework for integrating satellite observations with climate models to quantify systematic sea ice model errors. Using convolutional neural networks (CNNs), we can learn these errors based on the model's atmospheric, oceanic, and sea ice conditions--what I term a state-dependent representation of the error. This approach enables real-time corrections to subsequent model simulations, which systematically reduces global sea ice biases. I highlight key successes and challenges in developing this hybrid ML+climate modeling framework, including transfer learning to enhance online generalization of ML models, and new methods for integrating Python-based ML frameworks with Fortran climate model code. Finally, I introduce GPSat, a scalable Gaussian process-based tool for reconstructing complete sea ice fields from sparse satellite altimetry data. Together, the DA+ML framework and GPSat offer future opportunities for improving targeted model physics errors for more robust climate simulation.


IACS Seminar Speaker: William Gregory, Princeton University

Location: IACS Seminar Room

Abstract: Human gaze behavior is a fundamental cue for understanding social intent, human-machine interaction, and cognitive processes. This dissertation addresses the challenges of gaze target estimation (GTE), also known as gaze following, by developing a holistic understanding of gaze across complex environments.
First, we improve GTE performance through Patch-level Distribution Prediction (PDP). Unlike traditional pixel-wise regression, PDP models gaze as a spatial distribution over patches, better accounting for annotation variance and regularizing the pixel-wise heatmap prediction through multi-scale modeling. Second, to mitigate the high cost of data labeling, we present GCDR, the first semi-supervised method for gaze following. By prompting large Visual Question Answering (VQA) models to generate initial Grad-CAM heatmaps and refining them via a diffusion model, GCDR achieves robust performance with minimal human annotation. Third, we expand the applicability of GTE to multi-camera environments. By introducing the Multi-View Gaze Target (MVGT) dataset, along with two novel frameworks for integrating information and predicting gaze targets across views, we explore a new direction that overcomes single-view limitations such as face occlusion and out-of-view targets. Finally, we propose OmniGF, a multi-person gaze following model built on Vision-Language Models (VLMs) that enriches gaze target localization with semantic and social reasoning. By leveraging the semantic capabilities of VLMs alongside structural innovations to ground the model with fine-grained cues for each individual, OmniGF achieves state-of-the-art performance across three gaze following tasks.
Collectively, by tackling the gaze following problem through the distinct yet complementary perspectives of probabilistic modeling, geometric reasoning, and multimodal learning, this dissertation builds a holistic understanding of human gaze, paving the way for more intuitive artificial intelligence systems in downstream applications.

Speaker: Qiaomu Miao

Location: NCS 220
Mind Brain Lecture: Constructing the World of Taste in Your Head You fork the morsel into your mouth and say yum...chocolate cake. The appreciation of your dessert's taste seems to follow directly, quickly and simply from the placement of the food on your tongue. The truth, however, is far more interesting and complex: your brain actually begins determining whether you will enjoy a bite of food even before the fork approaches your mouth and continues to work the problem well after. Information about your food's color, smell, texture and taste activates multiple parts of your brain, where that information collides with your pre-mouthful beliefs about how it should taste. The coming-together and shuffling of that information around the brain takes time, as networks of neurons work together to help you decide whether the morsel in your mouth is worth swallowing. Referring to work from psychology, biology and computational neuroscience, Professor Katz will de-mystify and reveal the beauty of these complexities of the neuroscience of taste. Donald Katz, Professor of Psychology, Departments of Neuroscience, Psychology, and the Volen National Center for Complex Systems, Brandeis University Free presentation intended for a general audience. Reception to follow. https://www.stonybrook.edu/commcms/mind/
Abstract: Large language models are prone to memorizing some of their training data. Memorized (and possibly sensitive) samples can then be extracted at generation time by adversarial or benign users. There is hope that model alignment---a standard training process that tunes a model to harmlessly follow user instructions---would mitigate the risk of extraction. However, we develop two novel attacks that undo a language model's alignment and recover thousands of training examples from popular proprietary aligned models such as OpenAI's ChatGPT. Our work highlights the limitations of existing safeguards to prevent training data leakage in production language models.

Speaker: Pegah Alipoormolabashi

Location: CS2311
Educational objectives:

1. Explain how AI represents a Cognitive Revolution in academic medicine, redefining thefundamental limits of human cognition and knowledge work.
2. Differentiate between superficial AI adoption (innovation theatre) and truetransformationthrough AI-native institutional design.
3. Analyze how AI fundamentally reshapes clinical, research, and educational work-from dataentry to verification, recall to recognition, and hypothesis generation to evaluation-andidentify implications for redesigning academic health systems.

Speaker: Jiajie Zhang, Ph.D.,Dean, Professor, and Glassell Family FoundationDistinguished Chair in Informatics Excellence,D. Bradley McWilliams School of BiomedicalInformatics, UTHealth Houston

Location: MART Building, Room: 7M-0602 (7th Floor)