CSE 600 Seminar Series | Fall 2025



Abstract:

We often talk about AI as if it begins with a dataset and ends with an application. But behind every model lie two sets of actors who are rarely acknowledged in technical documentation: the workers who train AI systems and the researchers who try to make sense of them. This talk brings both groups into view.
Dr. Ben Zhang will offer an on-the-ground examination of the prevailing values and invisible labor that underpin commercial AI production and data production. Drawing on ethnographic research inside AI data annotation centers in China, he introduces the concept of precision labor to unpack the labor dimension of constructing, managing, and performing technical accuracy. This concept highlights the hidden and excessive labor required to reconcile the ambiguity and uncertainty involved in AI training. A precision labor lens challenges the legitimacy and sustainability of the relentless pursuit of technical accuracy, raising new questions about its consequences and implications.
On the other end of the pipeline, as LLMs become embedded in society, social scientists like Dr. Jieshu Wang is scrutinizing their potential biases while employing them as research tools. She will present her recent work auditing LLM responses across different contexts, revealing that LLMs exhibit varying levels of environmental awareness and disproportionately reward institutional prestige in peer-review simulations. She also demonstrates how LLMs can serve as useful tools in social-science pipelines, e.g., extracting location information, inferring demographics, parsing citations, mapping social networks, and analyzing occupational data.
By placing these two worlds side by side - the labor of training AI and the scholarly efforts to study it - we show why responsible AI should go beyond the deployment phase - emphasizing fairness audits, and model explainability. It requires reimaging the values, labor regimes, and social science practices that shape AI systems from annotation to analysis.


Bios:

Dr. Jieshu Wang is an interdisciplinary researcher studying the human and social dimensions of artificial intelligence (AI) and how people can thrive in an AI-integrated future. She combines computational methods with qualitative insights to trace technology trends and understand their broader societal impact. She earned her Ph.D. in Human and Social Dimensions of Science and Technology from Arizona State University, after earlier degrees in Civil Engineering, Economics, and Science and Technology Studies. She has also worked as a patent examiner, an editor at a popular science magazine, and co-founded Synced (机器之心), an AI-focused media company in China. Her research looks both backward and forward. Backward-looking, she examines how AI are created, who creates them, and who is missing from the process. Forward-looking, she studies how AI is transforming the way we live, connect, invent, work, and adapt, as well as how AI might help address challenges such as climate change and workforce transitions.
Dr. Ben Zhang is an Assistant Professor in the Department of Technology. His research explores the production and sociotechnical impacts of AI systems in critical areas such as work, health, and sustainability. Drawing from his background in Human-Computer Interaction (HCI), Human-Centered AI, and Science and Technology Studies (STS), he employs a life-cycle-centered approach to holistically examine the promises and harms of these systems and to inform the design of responsible AI infrastructures across their development, deployment, and governance. Ben received his Ph.D. in Information Science from the University of Michigan. Ben's work has been supported by competitive awards and fellowships, including the University of Michigan Rackham Predoctoral Fellowship and the Weizenbaum Fellowship. His research has appeared in premier computing venues, including ACM CHI, ACM CSCW, and AAAI ICWSM.

Location: NCS 120


New York Scientific Data Summit (NYSDS) is a premier annual conference that brings together researchers and thought leaders from academia, national labs and industry to exchange ideas and foster collaboration focused on data-driven science and technology. Co-hosted by Brookhaven National Laboratory and the Institute for Advanced Computational Science (IACS) at Stony Brook University, NYSDS 2025 will take place on September 11-12, 2025, in the SUNY Global Center in New York City.

NYSDS 2025 will spotlight artificial intelligence (AI), machine learning (ML) and robotics - fields currently at a pivotal point with transformative impacts on science and technology. From accelerating computationally demanding simulations to discerning signals from noisy data, AI/ML has become an integral part of the scientific workflows. Despite many advances, challenges remain to ensure that AI/ML applications are reliable, explainable and trustworthy.

Robotics, a growing field that couples AI with physically actuated mechanical bodies, has seen increased interest in areas spanning science, technology and manufacturing. The need for real-time decision-making and control, along with the intricate morphology of robots, makes robotics an intriguing application of AI, advanced computing and optimization.


This NYSDS 2025 is open to the public. To be eligible to attend, all participants must register online by August 30, 2025. For questions or assistance with registering, please contact the Summit Coordinator.

Register here.

Abstract: Language offers a uniquely powerful lens for understanding the mind: one that can access latent psychological realities often missed by traditional measurement tools. However, as language models expand their ability to capture semantics through context length, expansion into deeper levels of semantics is less explored, especially with respect to understanding cognitive patterns of authors. This dissertation proposes that we can uncover deeper cognitive and affective patterns that reflect more accurate underlying mental states by analyzing language at higher levels of discourse semantics and by modeling latent states.


First, the dissertation focuses on uncovering cognitive styles or thinking patterns manifesting in language. We demonstrate that modeling language at deeper semantic levels such as discourse relations, can unveil latent psychological states and traits, including cognitive styles that influence both mental health and behavior. Introducing a novel blend of transfer and active learning, we efficiently curated a new set of linguistic data on cognitive styles like dissonance. This approach allows for more precise measurement when dealing with rare-classes and low-resource tasks. As a second contribution, effective validation methods are introduced to language-based assessments of the underlying cognitive styles. Controlled behavioral experiments and online studies show that cognitive styles detected through linguistic signals reliably predict real-world behaviors such as decision-making and engagement with extremist communities, both at the individual and community levels, sometimes months in advance

The research further moves beyond traditional measurement tools like questionnaires and expert judgments, which rely on Classical Test Theory, by establishing that language-based assessments more closely approximate true psychological states. The mechanisms by which these assessments outperform standard tools are explained, highlighting their predictive power for behaviors linked to underlying traits. Finally, a more sophisticated approach is explored by modeling psychological outcomes with Item Response Theory (IRT), an improvement over Classical Test Theory. Adaptive language-based assessments are introduced, showing that targeted, adaptive testing based on latent IRT scores can efficiently and accurately capture multiple psychological dimensions.

Taken together, these contributions argue for a shift towards language-based psychological assessments. By integrating deeper discourse-level semantics with measurement theory, this dissertation charts a path towards truer scores of mental states: ones that are more precise, and reflective of the complexity of human cognition and emotions.

Speaker: Vasudha Varadarajan

https://stonybrook.zoom.us/j/99180374682?pwd=w2zZTkQsfunrBZhHgEweR54NjKabZ2.1&jst=2

We invite faculty to deliver a 10-minute presentation during our afternoon session at the CELT Symposium on April 11, 2025. Showcase how you use emerging technology (i.e. AI, VR, etc.) to support diverse student populations and enhance learning experiences. Share your innovative strategies and inspire others!

CELT Symposium Theme: A New Era of Inclusivity and Innovation in Higher Education

https://t.e2ma.net/click/5w0gph/5wwlu4oe/9v63j6
Face Editing with Machine Learning presented by Zhixin Shu

ABSTRACT: The face is the most informative feature of humans and has been a long-standing research topic in Computer Vision and Graphics. Images of faces are also ubiquitous in photography and social media, and people have devoted significant resources to capturing and editing face images. Face editing can be broadly viewed as the encoding, manipulation and the decoding of some representations for face images. The challenges are that we want to manipulate an image in a controllable way and generate results that are both desirable and as realistic as possible. This thesis explores different Machine Learning-based face-editing approaches. I discuss the role of machine learning for achieving desirable edits by learning both the physical aspects as well as the statistical manifold of human faces. In my work for eye-editing, I discuss the importance of understanding multiple physical elements of a face image, such as shape, illumination, pose, etc. In a deep-learning-based approach, I introduce image formation domain knowledge to the construction and training of a neural network. This network provides transparent access to the disentangled representations of the aforementioned physical properties. With this network, we can achieve various face editing tasks in forms of representation manipulation. After that, I introduce Deforming Autoencoders, a network that learns to disentangle shape and appearance in an unsupervised manner. This disentanglement is beneficial for the learning of some other factors of variations, such as illumination and facial expression. In an extension of Deforming Autoencoders, we incorporate non-rigid structure-from-motion to learn a 3D morphable model for faces that only requires an image set for training. At last, I describe an image-to-image network for 3D face reconstruction, which also utilizes structure-from-motion in deep learning. With real face images in training, this network not only reconstructs 3D faces more accurately than prior art but also has better generalization ability in real-life testing cases.
Abstract: Modern language agents often need to solve tasks requiring long-horizon, multi-turn interactions, where they retrieve external information, adapt to observations, and answer interdependent queries. Yet, most LLM systems rely on full-context prompting, appending all past turns regardless of their relevance. This leads to un-bounded memory growth, increased computational costs, and degraded reasoning performance on out-of-distribution input lengths due to LLM forgetting the context. We introduce MEM1, an end-to-end reinforcement learning framework that enables agents to operate with constant context size when solving long multi-turn tasks. At each turn, MEM1 updates a compact shared internal state that jointly supports memory consolidation and reasoning. Leveraging reinforcement learning (RL) and rollout trajectory truncation, we train a MEM1 agent to develop internal states that integrate prior memory with new observations from the environment while strategically discarding irrelevant or redundant information. Experiments across three domains, including internal retrieval QA, open-domain web QA, and multi-turn web shopping, show that MEM1-7B improves performance by 3.5x while reducing memory usage by 3.7x compared to Qwen2.5-14B-Instruct on an augmented multi-hop QA dataset with 16 objectives in each task, and generalizes beyond the training horizon. Our results demonstrate the promise of reasoning-driven memory consolidation as a scalable alternative to existing solutions for training long-horizon task-solving agents that involve multiple interactions, where both efficiency and performance are optimized.

Speaker: Yiyang Feng

Location: CS2311
Abstract: Computational pathology has revolutionized cancer diagnosis and research through the analysis of digitized whole slide images (WSIs). However, their giga-pixel size creates two intertwined bottlenecks: computational inefficiency, as prohibitive GPU memory makes standard end-to-end (E2E) training infeasible, and label inefficiency, as expert annotation is tedious and expensive. This dissertation confronts both challenges through novel architectures, training paradigms, and self-supervised learning methods for efficient WSI analysis.
To improve computational efficiency, this dissertation first introduces a locally supervised learning paradigm that enables E2E training on entire WSIs by partitioning a network into gradient-isolated modules, circumventing the memory bottleneck of backpropagation. Second, it presents Prompt-MIL, a parameter-efficient fine-tuning framework training only a few prompts to guide large pre-trained models, reducing trainable parameters, memory, and training time. Third, this work proposes 2DMamba, the first intrinsic Mamba architecture that preserves the crucial 2D spatial structure of images, overcoming the spatial discrepancy in 1D models. Fourth, it presents Locally Bi-directional Mamba (LBMamba), whose hardware-aware local backward scan integrates bi-directional scanning into a single forward pass, improving the throughput-performance trade-off of Mamba models.
To improve label efficiency, this dissertation proposes a precise location based matching strategy for self-supervised dense contrastive learning, which allows a local patch in one augmented view to match multiple overlapping patches in another, producing more accurate correspondences and superior features for dense prediction tasks like segmentation and detection. Additionally, to better scale multi-channel cell imaging modalities, this dissertation introduces ChannelSFormer, a channel-agnostic vision transformer that disentangles spatial and channel-wise reasoning through divided attention and channel class token, enabling effective representation learning across variable channel configurations in both self-supervised and supervised settings.
In summary, this dissertation presents a holistic investigation into the efficiency bottlenecks in computational pathology. Through these combined contributions in model architecture, training paradigms, and self-supervised learning, this work establishes a more scalable, efficient, and powerful computational framework for analyzing giga-pixel pathology images.

Speaker: Jingwei Zhang

Location: NCS 220

Zoom: https://stonybrook.zoom.us/j/93175806292?pwd=xbtxnQyYGoThz5B1DyJxJxPF9lxiJE.1
Meeting ID: 931 7580 6292
Passcode: 314091
The North East AI Agents Day Organizing Committee invites you to '2026 AI Agents Day.'

The goal of this workshop is to offer a comprehensive overview of AI agents, bring ML, Systems, and HCI research communities together to share progress, discuss common problems and evaluation setups, and identify opportunities for collaboration. We aim to bring together attendees from diverse disciplines to foster interdisciplinary collaboration and discuss open research questions.

Location: Jane Street Offices, New York

Register here.