TITLE: Towards a Theory of Encode/Decoder Architectures by Andrej Risteski of CMU

ABSTRACT: A common choice of architecture in representation learning (i.e., learning a good embedding of the data) is an encoder/decoder architecture, which tries to map a part of the input into a good latent representation (via an encoder), and predict the remaining part of the input (via a decoder). Two common examples are universal machine translation: where one tries to learn to translate between any pair of a set of languages via a common latent language, given paired up corpora for only a part of the pairs; and contextual encoders -- where one tries to predict a part of the image, given the rest of the image.
 
We will give a framework for analyzing the sample complexity of such architectures -- i.e., how many pairs of languages do we need to have paired up corpora for? How many image prediction tasks do we have to solve to get a good representation?
AI can help you write, you hear. AI can save you time, leverage your skills, enhance your productivity. . . . But you also hear: AI output is not reliable, not adequate for advanced tasks/learning, not ethical to use -- you could get in deep trouble for using AI tools without adequate mastery and caution. Which way is it?
Come join this hands-on workshop where you will explore AI tools and their affordances. Engage in writing tasks to learn how to use AI tools effectively and responsibly.
Sign up for a seat now: https://docs.google.com/forms/d/e/1FAIpQLSd0iDTKkTYnkxFd4LkgqbtP97zQSS4FI_MiPVm7p6IY5SGwSg/viewform

Learn how to unlock the power of Image and visuals that will enhance your work by asking the experts questions in-person

No registration required - just stop by!

Location: Frank Melville Jr. Memorial Library Galleria (across from the Central Reading room)

You are cordially invited to attend the biweekly Brookhaven AI Mixer (BAM). BAM includes three short talks on AI research happening at BNL, followed by an open mixer over coffee and snacks for everyone to network and discuss all things AI. The first half hour will consist of presentations that will be available via ZOOM, and the second half hour will be for in person only networking.

Join us every other Tuesday at noon in CDSD's Training Room (building 725, 2nd floor) to learn about interesting AI methods and applications, engage with potential collaborators, prepare for pending FASST funding calls, and build a community of AI for Science at BNL.

Tuesday, January 7, 2025, 12:00 pm -- CDS, Bldg. 725, Training Room


Speakers

Sanket Jantre
Tao Zhang
Xi Yu


Join ZoomGov Meeting: https://bnl.zoomgov.com/j/1615289117?pwd=Hqkbj9itxWrFnkhZ8rQXHPInO2gxdF.1

Meeting ID: 161 528 9117
Passcode: 991382

CSE 600 Talk: Haibin Ling - Computer Vision Research and Applications


Abstract: Having been intensively studied over half a decade, computer vision has evolved as a broad research area and become mature in many applications. In this talk, we will summarize our work in computer vision in both core vision topics and application-oriented ones. In particular, for core vision problems, we will report studies on visual tracking, visual matching and visual detection; for applications, we will describe our work on medical image analysis, intelligent transportation, smart projector systems and preliminary work on material property prediction.

Bio: Haibin Ling received the BS and MS degrees from Peking University in 1997 and 2000, respectively, and the PhD degree from the University of Maryland, College Park, in 2006. From 2000 to 2001, he was an assistant researcher at Microsoft Research Asia. From 2006 to 2007, he worked as a postdoctoral scientist at the University of California Los Angeles. In 2007, he joined Siemens Corporate Research as a research scientist. From 2008 to 2019, he worked as a faculty member of Temple University. In fall 2019, he joined the Department of Computer Science of Stony Brook University where he is currently a SUNY Empire Innovation Professor. His research interests include computer vision, augmented reality, medical image analysis, and human computer interaction. He received the Best Student Paper Award at the ACM UIST in 2003, and the NSF CAREER Award in 2014. He serves as Associate Editor for several journals including IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Pattern Recognition (PR), and Computer Vision and Image Understanding (CVIU). He has served or will serve as Area Chair for CVPR 2014, 2016, 2019 and 2020.
Abstract: As computing and society become increasingly inseparable, we confront a fundamental design challenge: creating AI systems where human-machine interactions authentically embody our diverse values while thoughtfully evolving our social relationships. The recursive nature of these interactions--where human behavior shapes technology design and technological affordances influence human behavior--presents both profound risks and transformative opportunities as we reimagine our collective digital future. What interaction patterns emerge when algorithmic systems become active participants in societal decision-making? How can we design human-AI collaboration that ensures algorithmic systems align with diverse community values while serving the public interest? Through Public Interest AI, we explore a Pluralistic Design Language that creates interaction models for value-sensitive algorithmic ecosystems, strengthening AI-society alignment in both technology design and policy development. Through collaborative interaction with communities, we create systems that augment human capabilities while embedding ethical principles into the sociotechnical design of AI itself--ultimately redefining possibilities at the intersection of technology, policy, and society. This talk will examine the challenges of designing meaningful human-AI systems within social contexts through real-world applications that combine value-sensitive interaction design, human-inspired computing, and societal development to create technologies that advance our shared commitment to the public good.

Bio: Neil Gaikwad is an Assistant Professor of Data Science and Computer Science at UNC Chapel Hill. Additionally, he serves on the Faculty Advisory Council of the UNC Parr Center for Ethics and is a Fellow at the MIT Dalai Lama Center for Ethics and Transformative Values. Neil holds a Ph.D. in Society-Centered AI from MIT and is an alumnus of Carnegie Mellon University's School of Computer Science. Neil's scholarship, published in prominent AI and HCI conferences, has been recognized with several prestigious honors, including the Facebook Research Fellowship, UIST Best Paper Honorable Mention, MIT Engineering Fellowship, Human Rights & Technology Fellowship, Graduate Teaching Award, and the Karl Taylor Compton Prize, MIT's highest student honor. He has been recognized as a Rising Star by both Stanford University and the University of Chicago. Translating research into real-world impact, Neil is a dedicated educator and mentor who has taught over 500 students throughout his career. He has guided more than 30 students to publish influential papers on AI fairness, secure prestigious fellowships, and contribute to shaping AI policy through public interest research. Neil is also the founder of the AI Policy Global Initiative, which has successfully brought together academia, industry, government, and communities to address critical challenges in AI governance and develop collaborative approaches to responsible AI.

Location: Old Computer Science, room 1310
The Empirical Methods in Natural Language Processing (EMNLP) conference is a premier international academic conference in the field of artificial intelligence and natural language processing (NLP). Organized annually by the Association for Computational Linguistics (ACL) special interest group on linguistic data (SIGDAT), it focuses on research that uses empirical methods to solve language processing problems.

For more information, and registration, visit the official website.
A talk by Jerome Zhengrong Liang entitled, Machine Learning from Original Images to Texture Patterns: A Paradigm Shift from Non-Medical Application to Medical Diagnosis. Abstract: Artificial intelligence (AI) research for medical diagnosis started soon after human began to use computer, initially called artificial neural network (ANN) and now convolutional neural network (CNN). ANN has been mainly explored to classify the experts' handcrafted features from the original (or raw) images, while CNN has been mainly explored directly on the raw images for both tasks of extracting abstract features and classifying the features. Experimental evidences have been shown that CNN can be trained by a large number of the raw images with experts' scores (or labels) to match or even surpass the experts' performance for both non-medical and medical diagnosis applications. However, the performances of the CNN models as well as the experts on medical diagnosis dropped dramatically when the labels of the raw images were replaced by the corresponding medical pathological reports. Accumulated medical knowledge reveals that the lesion heterogeneity is a footprint of lesion evolution and ecology, and the heterogeneity is an indicator of lesion progress and response to medical intervention. The heterogeneity can be reflected by the image contrast distribution (or texture patterns) across the lesion volume. Image textures have been shown as an effective descriptor of the lesion heterogeneity for computer-aided diagnosis. Can we map the raw images into texture patterns (or images) and train CNN to learn from the texture images? This question is the central theme of this presentation with application to CT Colonography or virtual colonoscopy, a game from AlphaGo to PolypGo. Bio: Jerome Zhengrong Liang, PhD, IEEE Fellow Imaging Research and Informatics Laboratory Department of Radiology, Stony Brook University
Abstract: Large Language Model (LLM) agents have demonstrated remarkable generalization capabilities across multi-domain tasks. Existing agent tuning approaches typically employ supervised finetuning on entire expert trajectories. However, behavior-cloning of full trajectories can introduce expert bias and weaken generalization to states not covered by the expert data. Additionally, critical steps--such as planning, complex reasoning for intermediate subtasks, and strategic decision-making--are essential to success in agent tasks, so learning these steps is the key to improving LLM agents. For more effective and efficient agent tuning, we propose ATLAS that identifies the critical steps in expert trajectories and finetunes LLMs solely on these steps with reduced costs. By steering the training's focus to a few critical steps, our method mitigates the risk of overfitting entire trajectories and promotes generalization across different environments and tasks. In extensive experiments, an LLM finetuned on only 30% critical steps selected by ATLAS outperforms the LLM finetuned on all steps and recent open-source LLM agents. ATLAS maintains and improves base LLM skills as generalist agents interacting with diverse environments.

Speaker: Bijoy

Location: Old Computer Science Building - CS2311