NLP Reading Group at Stony Brook | Spring 2026

Event Description

The Natural Language Processing Reading Group at Stony Brook University meets weekly to discuss recent research papers in NLP and related fields.
Join the Google Group here.

Title: J-Space and Mechanistic Interpretability

Abstract, J-Space: The rise of the term mechanistic interpretability has accompanied increasing interest in understanding neural models -- particularly language models. However, this jargon has also led to a fair amount of confusion. So, what does it mean to be mechanistic? We describe four uses of the term in interpretability research. The most narrow technical definition requires a claim of causality, while a broader technical definition allows for any exploration of a model's internals. However, the term also has a narrow cultural definition describing a cultural movement. To understand this semantic drift, we present a history of the NLP interpretability community and the formation of the separate, parallel mechanistic interpretability community. Finally, we discuss the broad cultural definition -- encompassing the entire field of interpretability -- and why the traditional NLP interpretability community has come to embrace it. We argue that the polysemy of mechanistic is the product of a critical divide within the interpretability community.

Abstract, Mechanistic Interpretability: If the mind is an ocean, we spend our lives floating at the surface. Beneath us, an enormous amount of processing takes place without our knowledge: our visual systems parsing the contours of a face, our motor circuits maintaining our posture. At any given moment, only a small fraction of this neural activity is accessible to us. Yet it is this privileged sliver of activity that we rely on to reason deliberately: to plan what ingredients to buy for a recipe, or to puzzle out why an engine won't start. Such thoughts can be articulated out loud, deliberately held in mind, and brought to bear on whatever task the moment demands. This distinction, between our accessible thoughts and our unconscious processing, is perhaps the most striking feature of human cognition. In this paper, we present evidence that an analogous functional distinction has emerged in modern AI models. Specifically, we observe that language models maintain a privileged set of internal representations, available for report, modulation, and flexible internal reasoning, atop a much larger volume of automatic processing. We identify these representations using a new interpretability technique, which surfaces the concepts a model is poised to verbalize at any point in its processing. Measuring and intervening on these representations provides us a window into a model's thought processes, uncovering internal reasoning and reactions that do not appear in its output.

Location: NCS 220

Zoom Link: https://stonybrook.zoom.us/j/92942833294?pwd=ysKMaQx7Hp0lkq5nOb3zJ9ivNZVvLv.1&jst=2

Date Start

Date End