How Language Makes us Smart (without Big Data) presented by Charles Yang

Abstract: Language provides the glue that combines simpler concepts into complex ones. To study how language guides conceptual development, we need precise accounts of how rules are learned from the child's linguistic experience, which is extremely limited in comparison to the amount of data available to current machine learning methods. In this talk, I discuss a mathematical model of inductive generalization, which enables language learning with very small amount of data. Such a view of learning has strong implications for the cross-cultural/linguistic variation of development. As a case study, I show that Hong Kong children learning Cantonese, which has a relatively simpler formal counting system, develop understanding of symbolic numbers a full year ahead of English-learning children in the United States, which is precisely predictable from the learning model. The new conception of learning adds another wrinkle to the eternal question of how language and thought are related to each other.

Bio: Charles Yang studied at the MIT AI lab and now teaches linguistics, computer science and psychology and directs the Program in Cognitive Science at the University of Pennsylvania. He is the author of several books: The Price of Linguistic Productivity (2016 MIT Press) won the Leonard Bloomfield Award from the Linguistic Society of America. His honors include a Guggenheim fellowship.
Abstract: Modern language agents often need to solve tasks requiring long-horizon, multi-turn interactions, where they retrieve external information, adapt to observations, and answer interdependent queries. Yet, most LLM systems rely on full-context prompting, appending all past turns regardless of their relevance. This leads to un-bounded memory growth, increased computational costs, and degraded reasoning performance on out-of-distribution input lengths due to LLM forgetting the context. We introduce MEM1, an end-to-end reinforcement learning framework that enables agents to operate with constant context size when solving long multi-turn tasks. At each turn, MEM1 updates a compact shared internal state that jointly supports memory consolidation and reasoning. Leveraging reinforcement learning (RL) and rollout trajectory truncation, we train a MEM1 agent to develop internal states that integrate prior memory with new observations from the environment while strategically discarding irrelevant or redundant information. Experiments across three domains, including internal retrieval QA, open-domain web QA, and multi-turn web shopping, show that MEM1-7B improves performance by 3.5x while reducing memory usage by 3.7x compared to Qwen2.5-14B-Instruct on an augmented multi-hop QA dataset with 16 objectives in each task, and generalizes beyond the training horizon. Our results demonstrate the promise of reasoning-driven memory consolidation as a scalable alternative to existing solutions for training long-horizon task-solving agents that involve multiple interactions, where both efficiency and performance are optimized.

Speaker: Yiyang Feng

Location: CS2311
LIN 627: Pragmatics Seminar
Th 3.30-6.20, Old Computer Science Building - CS2311, Zoom option
Instructor: Owen Rambow

Pragmatics is the study of how context (linguistic and non-linguistic) affects language use: the utterer (speaker or writer) chooses among many linguistic means licensed by the grammar of the language. Core explanatory notions in pragmatics are the common ground between utterer and addressee, and the use of speech acts to change the common ground.

This seminar will investigate the proposal of Grice, which has been called intentionalism. According to Grice, an utterer forms a communicative intention, performs a speech act to achieve the intention, and the speech act succeeds when the addressee recognizes the intention. The result is a change in common ground. With intentionalism, the notion of cognitive state becomes central to pragmatic theory. We will investigate Gricean intentionalism, including formalizations such as Stalnaker's, empirical evidence from cognitive science, and alternate pragmatic theories. In the second part of the course, we will then examine specific issues in pragmatics in light of the foundational theories we have discussed, including information structure, specific linguistic issues such as discourse particles, intonation, and politeness, and neuro-divergent communication. We will pay attention to empirical evidence throughout the course.

You can enroll for 0-3 credits, and further topics of interest may be included.

List of papers: https://protect.checkpoint.com/v2/r01/___https://docs.google.com/spreadsheets/d/1b7UC6leIHLxIa5yMWMAkUBZzT9Vg2SpS4XEfgH50ugY/edit?usp=sharing___.YzJ1OnN0b255YnJvb2s6YzpnOmE1Y2UxOTNiNTUyZWU4N2E0MzdkYzZmYWZlMTFmZWUxOjc6NGI2MjphZWY1NTUzNTdkMGJiN2U4MTI2MWMxYWQxZTU4OWNiY2NjYTlhNzk0NzlkZjU2ZDNmMzc5NDBkYjFhMjhhNTEwOnA6VDpG

Topic: AI Seminar: Stanley Bak
Time: Monday Nov 1, 2021 12:00 PM Eastern Time (US and Canada)

Join Zoom Meeting
https://stonybrook.zoom.us/j/91227496273?pwd=M3EyUDlzK3Vzd2pDOGpDU1ZjN0k1UT09

Abstract: The field of formal verification has traditionally looked at proving properties about finite state machines or software programs. The surge in deep learning has been accompanied by a surge of progress in trying to apply mathematical and algorithmic techniques to prove things about the function being computed by a neural network.

This talk formalizes the neural network verification problem and describes technical methods for neural network verification based on reachability analysis. Improvements to analysis efficiency will be given, as well as research directions for further exploration. We also include an objective comparison performed this last summer trying to evaluate the best existing verification methods in terms of speed and network size. The competition was performed on common hardware and involved the participation of twelve international teams (the tool authors) on a common set of benchmarks. 

Biography: Stanley Bak is an assistant professor in the Department of Computer Science at Stony Brook University investigating the verification of autonomy, cyber-physical systems, and neural networks. He strives to develop practical formal methods that are both scalable and useful, which demands developing new theory, programming efficient tools and building experimental systems.
Stanley Bak received a Bachelor's degree in Computer Science from Rensselaer Polytechnic Institute (RPI) in 2007 (summa cum laude), and a Master's degree in Computer Science from the University of Illinois at Urbana-Champaign (UIUC) in 2009. He completed his PhD from the Department of Computer Science at UIUC in 2013. He received the Founders Award of Excellence for his undergraduate research at RPI in 2004, the Debra and Ira Cohen Graduate Fellowship from UIUC twice, in 2008 and 2009, and was awarded the Science, Mathematics and Research for Transformation (SMART) Scholarship from 2009 to 2013. From 2013 to 2018, Stanley was a Research Computer Scientist at the US Air Force Research Lab (AFRL), both in the Information Directorate in Rome, NY, and in the Aerospace Systems Directorate in Dayton, OH. He helped run Safe Sky Analytics, a research consulting company investigating verification and autonomous systems, and performed teaching at Georgetown University before joining Stony Brook University as an assistant professor in Fall 2020.

The International Neuroethics Society (INS) Speaker Series on AI & Consciousness

AI has existed as a tool for a long time, performing simple tasks such as sorting documents, suggesting music, and so on. But with the development of new generations of AI, the perception of its value to society has been increasing, as it can bring potential and promising benefits in many areas of human life. AI is known to have errors or biases that result in strange or even dangerous responses, but what happens when in AI-human interaction, the latter have errors or biases? cultural errors or biases? And what could be the implications for human relationships?

Speaker Bio

Dr. Karen Herrera-Ferrá is an independent and global consultant on ethical, medical, psychological, legal, social, cultural, policy-making, human rights and political issues and concerns on the development and use of neuroscience, neurotechnology and AI. She is a former member of the Board of Directors of the International Neuroethics Society.

Register here

https://umaryland.zoom.us/meeting/register/tJMvfuqsqDspG9BKMLfUU49UbuUyP_IEvXRh


Abstract:
Deep neural network (DNN) is a powerful tool for solving image processing and computer vision tasks such as image and video reconstructions, object recognition, and scene understanding, etc. However, DNN have been used for only the digital domain in the imaging pipeline, such as the feature extractor and classifier models after an image is captured and digitized. In this research, we propose a new framework called deep sensing. The proposed framework also models the analog layer to the neural network model and jointly optimizes the parameters in optics and sensor designs of a camera, as well as reconstruction and classification models by the same training strategy.

Bio:
Hajime Nagahara is a professor at D3 Center, Osaka University, since 2024. He received Ph.D. degree in system engineering from Osaka University in 2001. He was a research associate of the Japan Society for the Promotion of Science from 2001 to 2003. He was an assistant professor at the Graduate School of Engineering Science, Osaka University, Japan from 2003 to 2010. He was an associate professor in Faculty of Information Science and Electrical Engineering at Kyushu University from 2010 to 2017. He was a professor at Institute for Datability Science, Osaka University, from 2017 to 2024. He was a visiting associate professor at CREA University of Picardie Jules Verns, France, in 2005. He was a visiting researcher at Columbia University in 2007-2008 and 2016-2017. Computational photography and computer vision are his research areas. He received IPSJ Nagao Special Researcher Award in 2012, ICCP2016 Best Paper Runners-up, and SSII Takagi Award in 2016. He is a program chair for ICCP2019, General Chair for upcoming ACCV2026 Osaka, Associate Editor for IEEE Transactions on Computational Imaging in 2019-2022, Associate Editor for IEEE Transactions of Pattern Recognisiton and Machine Intelligence, and Director of Information Processing Society of Japan in 2022-2024.

Zoom: https://stonybrook.zoom.us/j/4742917886?pwd=YiW3AL5lvPlEaWrd221YfLzYhlgsRK.1