Abstract: Reward hacking, where a reasoning model exploits loopholes in a reward function to achieve high rewards without solving the intended task, poses a significant threat. This behavior may be explicit, i.e. verbalized in the model's chain-of-thought (CoT), or implicit, where the CoT appears benign thus bypasses CoT monitors. To detect implicit reward hacking, we propose TRACE (Truncated Reasoning AUC Evaluation). Our key observation is that hacking occurs when exploiting the loophole is easier than solving the actual task. This means that the model is using less effort than required to achieve high reward. TRACE quantifies effort by measuring how early a model's reasoning becomes sufficient to obtain the reward. We progressively truncate a model's CoT at various lengths, force the model to answer, and estimate the expected reward at each cut-off. A hacking model, which takes a shortcut, will achieve a high expected reward with only a small fraction of its CoT, yielding a large area under the reward-vs-length curve. TRACE achieves over 65% gains over our strongest 72B CoT monitor in math reasoning, and over 30% gains over a 32B monitor in coding. We further show that TRACE can discover unknown loopholes during training. Overall, TRACE offers a scalable unsupervised approach for oversight where current monitoring methods prove ineffective.

Speaker: Dikshya

Location: Old Computer Science Building - CS2311

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes three short talks on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Tuesday, January 7, 2025, 12:00 pm -- CDS, Bldg. 725, Training Room

Speakers

Jianda Chen, EBNN - Improving the stability and accuracy of PDE-ML hybrid AGCMs

Boyang Li, CDS - Accelerating Materials Discovery using Machine Learning

Jaehye on Do, NPP Isotopes - Using LLMs for Isotopes Research and Production

Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1615289117?pwd=Hqkbj9itxWrFnkhZ8rQXHPInO2gxdF.1

Meeting ID: 161 528 9117
Passcode: 991382

The Collective Surgical Consciousness: Artificial Intelligence & the Future of Surgery Guest speaker Doctor Ozanan Meireles, the Director of the Surgical AI and Innovation Lab at Massachusetts General Hospital and a faculty member at Harvard Medical School, presents The Collective Surgical Consciousness: Artificial Intelligence & the Future of Surgery. Objectives: * Become familiar with the subfields of AI used in surgery * Understand the importance of a potential paradigm shift in surgical practice, training, and continue medical development * The importance of data acquisition, sharing and ownership, and development of machine learning algorithms
Abstract: Many foundation models for digital pathology have been released recently. Benchmarking available methods then becomes paramount to get a clearer view of the research landscape. For this reason, we introduce THUNDER, a tile-level benchmark for digital pathology foundation models, allowing for efficient comparison of many models on diverse datasets with a series of downstream tasks, studying their feature spaces and assessing the robustness and uncertainty of predictions informed by their embeddings. Such foundation models are often used as feature extractors and combined with Multiple Instance Learning (MIL) aggregators at downstream time. Such aggregation must be efficient and reliable. We will focus on two specific examples of this: (I) HistAug, a fast and efficient generative model for controllable augmentations in the latent space of foundation models to perform data augmentation for MIL, and (ii) CAR-MIL, a method based on counterfactual attention regularisation to improve the reliability of attention maps of MIL methods.

Short-bio: Pierre Marza is a Postdoctoral Researcher at CentraleSupelec in the Biomathematics team of the MICS lab, studying Computer Vision and Deep Learning for Medical Imaging, with a focus on Digital Pathology. Prior to this, he was a PhD student at INSA Lyon, in the LIRIS and CITI labs, advised by Christian Wolf, and co-advised by Laetita Matignon and Olivier Simonin. He studied Visual Navigation, Embodied AI, Spatial Reasoning, more specifically how to learn to represent 3D space, generalize to new environments and master diverse tasks from light supervision.

Location: NCS 220

Zoom: https://stonybrook.zoom.us/j/94798224254?pwd=CFraer25qnpORbJ14aAVHRwaSJOjJM.1

To truly understand human language, we must look at words in the context of the human generating the language. Factors such as demographics, personality, modes of communication, and emotional states have shown to play a crucial role in NLP models pre-LLMs era. Steps of mathematically defining the inclusion of human context in language modeling and more will be discussed with Nikita Soni, a PhD student at Stony Brook University co-advised by H. Andrew Schwartz and Niranjan Balasubramanian. She is the lead organizer of the workshop on human-centered large language modeling.

Please register for the STEM Speaker Series Zoom event here

Please RSVP for the STEM Speaker Series in-person event here
University Libraries Present: Analyzing quantitative data can feel overwhelming without the right tools. In this workshop, SBU Libraries' Data Literacies Lead, Ahmad Pratama will show you how to master the basics of exploratory data analysis for quantitative data using Python. This workshop covers several techniques to help you uncover patterns and insights in your datasets.

Online RSVP via link: https://stonybrook.zoom.us/meeting/register/vEPycmDrQoGjFqkmsYHgxw
CSE 656 Seminar in Computer Vision The overall purpose of this seminar is to bring together people with interests in Computer Vision theory and techniques and to examine current research issues. This course will be appropriate for people who already took a Computer Vision graduate course or already had research experience in Computer Vision. To enroll in this course, you must either: (1) be in the PhD program or (2) receive permission from the instructors. Each seminar will consist of multiple short talks (around 15 minutes) by multiple students. Students can register for 1 credit for CSE656. Registered students must attend and present a minimum of 2 talks. Everyone else is welcome to attend. Fill in https://forms.gle/q6UG9ygauLp2a8Po8 to subscribe to our mailing list for further announcement.
The Future Histories Studio welcomes Moontae Lee, LG AI Research.


Generative AI is transforming how we understand, create, and interact with information. Large Language Models (LLMS) comprehend contexts, answer non-trivial questions, and spark creative ideas. This talk introduces the evolution of these models, highlighting the most recent advancements in planning, reasoning, and evaluation. The talk also touches on the criticalconsiderations for both model developers and users, carefully addressing limitations of LLMs as well as ethical and societal implications. Finally, the talk provides ongoing directions in researchand production: from the rise of personalized AI agents to the future frontiers of AI.

Moontae Lee is the Director of the Superintelligence Lab at LG AI Research and an Assistant Professor of Information and Decision Sciences at the University of Illinois Chicago. His journey with Large Language Models began as a visiting scholar at Microsoft Research in 2019, continuously consulting the Deep Learning Group at Redmond until joining LG. He holds a PhD in Computer Science from Cornell, an MS from Stanford, and BS degrees in Computer Science, Mathematics, and Psychology from Sogang University. He has been an area chair for major AI conferences and earned recognition in Operations Research and Computational Social Science, including awards from INFORMS and Amazon.

His research interests include:
● Computational Creativity, Algorithmic Awareness
● Retrieval-Augmented Generation and Evaluation
● Code Generation, Reasoning, Planning
● Fine-grained Alignment from Human/AI Feedback in Generative AI
● Large Time-series Models, Diffusion/Consistency
● Machine Unlearning
● Ranking Monopoly, Voting Fairness
● AI Safety, Ethics, and Market Impacts

Join us in person @ Future Histories Studio Staller Center for the Arts, 4222