Topic: AI Seminar: Owen Rambow
Time: Mar 17, 2021 10:00 AM Eastern Time (US and Canada)
Join Zoom Meeting

https://stonybrook.zoom.us/j/93614644178?pwd=MzJtVDJYYmU5T1dtMzJiUFMxb0x4dz09
Meeting ID: 936 1464 4178.    Passcode: 965936






Natural Language Understanding and Semantic Parsing

(Partly joint work with former colleagues at Elemental Cognition)

Semantic parsing refers to the task of determining the propositional content of language: who did what to whom.  It is part of the larger task of natural language understanding (NLU).  I will start out by discussing what full NLU means, and argue that we are still far away, as a field, from solving full NLU, or even from knowing how to evaluate it.

In the second part of the talk, I will situate semantic parsing in the context of several other NLU subtasks.  Typically, the target representation of semantic parsing uses an ontology (such as PropBank or FrameNet).  Semantic parsing includes the subtasks of word sense disambiguation, argument detection, and argument role labeling.  I will discuss choices among possible target ontologies.  I will justify why we created a new ontology, Hector, based on FrameNet and the lexical resource NOAD, and explain some of its characteristics.

In the third part of the talk, I will present experiments we performed using transformer models.  We obtain best results using a two-phase model, in which we first choose the frame, and then, given the frame, choose the arguments.  We encode the problem for both tasks using indices in the sentence.  While we develop the parser for our new ontology Hector, this approach also beats the state of the art for FrameNet and PropBank parsing.Biography:  I am a professor in the Department of Linguistics at Stony Brook University with a joint appointment in IACS.

Until recently, I was a research scientist at Elemental Cognition. Elemental Cognition is working on deep natural language understanding.

I got my PhD with Aravind Joshi at the University of Pennsylvania in 1994. I have worked at CoGenTex, and at AT&T Labs -- Research, and for many years I was a research scientist at Columbia University in the Center for Computational Learning Systems.

Nam Nguyen

4-5pm, Dec 17 2020

https://stonybrook.zoom.us/j/94214254415?pwd=K1VoQml4cFdlVW51VW41dWtid2tJdz09



The molecular mechanisms and functions in complex biological systems
currently remain elusive. Recent high-throughput techniques, such as
next-generation sequencing, have generated a wide variety of
multiomics datasets that enable the identification of biological
functions and mechanisms via multiple facets. However, integrating
these large-scale multiomics data and discovering functional insights
are, nevertheless, challenging tasks. To address these challenges,
machine learning has been broadly applied to analyze multiomics. In
particular, multiview learning is more effective than previous
integrative methods for learning data's heterogeneity and revealing
cross-talk patterns. Although it has been applied to various contexts,
such as computer vision and speech recognition, multiview learning has
not yet been widely applied to biological data--specifically,
multiomics data. Therefore, we have developed a framework called
multiview empirical risk minimization (MV-ERM) for unifying multiview
learning methods (Nguyen, et al., PLoS Computational Biology, 2020).
MV-ERM enables potential applications to understand multiomics
including genomics, transcriptomics, and epigenomics, in an aim to
discover the functional and mechanistic interpretations across omics.
Based on MV-ERM, we have developed the following methods:
ManiNetCluster, Varmole and ECMarker.



(1) ManiNetCluster (Nguyen, et al., BMC Genomics, 2019) is a manifold
learning method which simultaneously aligns and clusters gene networks
(e.g., co-expression) to systematically reveal the links of genomic
function between different phenotypes. Specifically, ManiNetCluster
employs manifold alignment to uncover and match local and non-linear
structures among networks, and identifies cross-network functional
links. We demonstrated that ManiNetCluster better aligns the
orthologous genes from their developmental expression profiles across
model organisms than state-of-the-art methods. This indicates the
potential non-linear interactions of evolutionarily conserved genes
across species in development. Furthermore, we applied ManiNetCluster
to time series transcriptome data measured in the green alga
Chlamydomonas reinhardtii to discover the genomic functions linking
various metabolic processes between the light and dark periods of a
diurnally cycling culture;



(2) Varmole (Nguyen, et al., Bioinformatics, 2020) is an interpretable
deep learning method that simultaneously reveals genomic functions and
mechanisms while predicting phenotype from genotype. In particular,
Varmole embeds multi-omic networks into a deep neural network
architecture and prioritizes variants, genes and regulatory linkages
via biological drop-connect without needing prior feature selections.
With an application to schizophonia, we demonstrate that Varmole
provides an effective alternative for recent statistical methods that
associate functional omic data (e.g. gene expression) with genotype
and phenotype and that link variants to individual genes in population
studies such as genome-wide association study;



(3) ECMarker (Jin*, Nguyen*, et al., Bioinformatics, 2020) is an
interpretable and scalable machine learning model that predicts gene
expression biomarkers for disease phenotypes and simultaneously
reveals underlying regulatory mechanisms. Particularly, ECMarker is
built on the integration of semi- and discriminative- restricted
Boltzmann machines, a neural network model for classification allowing
lateral connections at the input gene layer. With application to the
gene expression data of non-small cell lung cancer (NSCLC) patients,
we found that ECMarker not only achieved a relatively high accuracy
for predicting cancer stages but also identified the biomarker genes
and gene networks implying the regulatory mechanisms in lung cancer
development.



Finally, we propose a novel multiview learning method, Malignomics, to
predict phenotypes from heterogeneous multi-omic features. Malignomics
will first align multi-omic features by deep manifold alignment onto a
common latent space, better predicting nonlinear relationships across
omics. This deep alignment aims to preserve both global consistency
and local smoothness across omics and reveal higher-order nonlinear
interactions (i.e., manifolds) among cross-omic features. Second, it
uses these manifold structures to regularize the classifiers for
predicting phenotypes. This manifold-regularization allows
highlighting cross-omic feature manifolds and prioritizing the
features and interactions for the phenotypes. The prioritized
multi-omic features will further reveal underlying phenotypic
functions and mechanisms and thus enhance the biological
interpretation of Malignomics. We will apply Malignomics to
multi-omics data in neuropsychiatric disorders, and prioritize gene
regulatory networks linking risk variants, regulatory elements, and
genes for the disorders. We will also compare Malignomics with the
state-of-the-arts, and investigate how the manifold regulation will
potentially improve understanding of multi-omics functions and
predicting diseases.

The event will take place on Zoom and will feature two distinguished guest speakers: SBU alumnus, Velchamy Sankarlingam, president of Product and Engineering at Zoom, and Simeon Ananou, vice president for Information Technology and CIO at Stony Brook University. The discussion will be moderated by Haresh Gurnani, dean of the College of Business at Stony Brook University.

Exploring AI's Impact on Communication and Connection

Artificial Intelligence (AI) has rapidly evolved, becoming an integral part of various industries, including education and business. This event aims to delve into how AI is reshaping the way we learn and work, particularly in enhancing communication and fostering human connections. Velchamy Sankarlingam, an SBU alumnus and a key figure at Zoom, will share his insights on how AI-driven tools are revolutionizing virtual communication platforms, making interactions more seamless and effective.

Simeon Ananou, with his extensive experience in information technology, will provide a perspective on how AI is being integrated into educational institutions to improve learning outcomes and administrative efficiency. His role at Stony Brook University places him at the forefront of implementing innovative technologies that benefit both students and staff.

A Conversation Led by Expertise

Dean Haresh Gurnani, known for his leadership and expertise in business education, will guide the conversation, ensuring that the discussion remains focused on the practical implications of AI. He will explore how AI is not only boosting productivity but also enriching overall experiences in the workplace and educational settings. The event will include an interactive Q&A session, allowing attendees to engage directly with the speakers and gain deeper insights into the topics discussed.

As AI continues to develop, events like this are crucial for understanding its impact and potential. Stony Brook University's College of Business is committed to providing platforms for such important discussions, fostering an environment where innovation and education intersect.

This event is open to all. Please visit https://www.givecampus.com/schools/StonyBrookUniversity/events/artificial-intelligence-reshaping-learning-and-work to register.

Description:

As artificial intelligence and data science reshape the global information landscape, libraries are emerging as key players in both technological innovation and ethical stewardship. This international Zoom discussion brings together library professionals and educators from the U.S., Philippines, and Hong Kong to explore how institutions are integrating AI and data into their pedagogy and services.

Panelists will share concrete examples from their own libraries--ranging from data literacy initiatives to increasing discoverability. The conversation will also examine regional trends in librarianship, spotlighting how institutions in Asia are navigating the evolving role of data and AI.

Join us for a global conversation that highlights the transformative potential of libraries as hubs for innovation and critical inquiry in the age of AI.

Register for this free Zoom panel.

Panelists:

Ahmad Pratama is a Faculty Member and Associate Librarian at Stony Brook University Libraries, where he is working to build a comprehensive, campus-wide data literacy program within the Libraries. As the Data Literacies Lead, his work focuses on empowering students, faculty, and staff to critically and ethically engage with data and AI, including the development of a credit-bearing course in Critical Data & AI Literacies supported by an EDGE Fund Award from the Provost's Office. Previously, Dr. Pratama served as an Associate Professor of Information Technology, and his research and teaching explore the intersections of technology, policy, and society with a focus on data, AI, and innovation in higher education.

Dan Anthony Dorado is a full-time faculty member at the U.P. School of Library and Information Studies, where he teaches information technology, management and marketing, research methodology, and quantitative research. He was also the director of the Diliman Learning Resource Center under the Office of the Vice Chancellor for Student Affairs. Before that, he was an Information Specialist at the College of Engineering Library, in charge of the System and Network Administration and The Learning Commons. He completed his master's degree at the Technology Management Center in U.P. Diliman and is currently pursuing his PhD in Data Science. As a member of Sync.Bio.Optics laboratory and the Publics, Archives, and Data (PANDA) Lab, his research specialization covers Computational Methods, Open Education, Critical Data Studies, and Radical Statistics.

Ryun LEE is Associate University Librarian at The Chinese University of Hong Kong Library, leading Digital Initiatives and Library IT and Systems. He drives digital innovation through emerging technologies, particularly artificial intelligence to enhance services, streamline operations, and support CUHK's mission in research, education, and knowledge advancement. With a background in cataloging and digital repository development, Ryun leads projects in digitization, OCR, data visualization, text and network analysis, GIS, and digital scholarship. He actively promotes knowledge graph applications in Hong Kong studies and oversees efforts to digitize and preserve resources related to Hong Kong and Southern China. His recent work focuses on creating seamless digital experiences and developing data-driven infrastructure. He is currently exploring AI-driven approaches to digitization workflows and entity extraction, aiming to improve access, discovery, and long-term preservation of library materials.

Abstract: Language offers a uniquely powerful lens for understanding the mind: one that can access latent psychological realities often missed by traditional measurement tools. However, as language models expand their ability to capture semantics through context length, expansion into deeper levels of semantics is less explored, especially with respect to understanding cognitive patterns of authors. This dissertation proposes that we can uncover deeper cognitive and affective patterns that reflect more accurate underlying mental states by analyzing language at higher levels of discourse semantics and by modeling latent states.


First, the dissertation focuses on uncovering cognitive styles or thinking patterns manifesting in language. We demonstrate that modeling language at deeper semantic levels such as discourse relations, can unveil latent psychological states and traits, including cognitive styles that influence both mental health and behavior. Introducing a novel blend of transfer and active learning, we efficiently curated a new set of linguistic data on cognitive styles like dissonance. This approach allows for more precise measurement when dealing with rare-classes and low-resource tasks. As a second contribution, effective validation methods are introduced to language-based assessments of the underlying cognitive styles. Controlled behavioral experiments and online studies show that cognitive styles detected through linguistic signals reliably predict real-world behaviors such as decision-making and engagement with extremist communities, both at the individual and community levels, sometimes months in advance

The research further moves beyond traditional measurement tools like questionnaires and expert judgments, which rely on Classical Test Theory, by establishing that language-based assessments more closely approximate true psychological states. The mechanisms by which these assessments outperform standard tools are explained, highlighting their predictive power for behaviors linked to underlying traits. Finally, a more sophisticated approach is explored by modeling psychological outcomes with Item Response Theory (IRT), an improvement over Classical Test Theory. Adaptive language-based assessments are introduced, showing that targeted, adaptive testing based on latent IRT scores can efficiently and accurately capture multiple psychological dimensions.

Taken together, these contributions argue for a shift towards language-based psychological assessments. By integrating deeper discourse-level semantics with measurement theory, this dissertation charts a path towards truer scores of mental states: ones that are more precise, and reflective of the complexity of human cognition and emotions.

Speaker: Vasudha Varadarajan

https://stonybrook.zoom.us/j/99180374682?pwd=w2zZTkQsfunrBZhHgEweR54NjKabZ2.1&jst=2

The Provost's Spotlight Talks feature eminent visitors to the university as well as Stony Brook faculty members who have recently been recognized for outstanding contributions in their field.

Transmedia artist Stephanie Dinkins, Kusama endowed chair in art in the College of Arts and Sciences at Stony Brook University, brings her expertise in AI to the next Spotlight Talk with The Stories We Encode: AI, Love and the Future of Algorithmic Care on Tuesday, October 22, at 3:30 pm in the Charles B. Wang Center Theatre.

Working at the intersection of emerging technologies and social collaboration, Dinkins was named a 2023 TIME 100 Most Influential People in AI. She was recognized for her work with Not the Only One, an ongoing project in which she trained an AI on three generations of Black women to give it cultural roots, a deep history, and a perspective that existing systems do not offer.

The event is free and open to the public, and the discussion will be followed by a reception in the Wang Theatre lobby, hosted by the College of Arts and Sciences for new and promoted faculty.


About the Talk

AI's impact on society necessitates addressing longstanding human rights issues and prejudices. To ensure AI benefits humanity, we must confront institutional biases, rethink our relationship with other beings and emerging technologies, and reconcile ideals with actual power structures. This involves recognizing systemic inequalities, redefining human identity, and equitably distributing resources. AI, if developed and used ethically, offers an opportunity to reimagine a more equitable world for all inhabitants.

The Renaissance School Of Medicine Department of Scientific Affairs and its Single Cell Genomics facility are excited to host a special seminar and discussion on AI and single cell genomics analysis:

With the decreasing cost of sequencing, many biobanks and large research cohorts have moved to whole genome sequencing (WGS) and single-cell RNA-seq. However, making use of this deluge of data remains a challenge. I will discuss statistical and deep learning approaches that we are exploring to address the challenge of noncoding variant interpretation, including our work as part of the Alzheimer's disease sequencing project.

Speaker: David A. Knowles, PhD. Asst. Professor of Computer Science, Interdisciplinary Appointee in Systems Biology, Columbia University Core Faculty Member, New York Genome Center

Join us in person: Health Science Tower Level 3, Lecture Hall 5
Recently, large-scale language data combined with modern machine learning techniques have shown strong value as means for studying human psychology and behavior. For example, language alone has been shown predictive in mental health, personality, and health behaviors. However, many applications for such language-based assessments have readily available and important data beyond language (i.e. extra-linguistics), such as predicting the subjective well-being of a community using tweets, where one can take into account their age, education, and demographic attributes. Language may capture some characteristics while extra-linguistic variables captures others. We believe that effectively integrating linguistic and extra-linguistic data can yield benefits beyond either independently. In this thesis, we develop methods which effectively integrate extra-linguistic data with language data focused primarily on social scientific applications. The central challenge is dealing with the size and heterogeneity of, often sparse and noisy, language data versus the, often low-dimensional and non-sparse, extra-linguistic variables. First, we consider structured extra-linguistics, like socioeconomic (income and education rates) and demographics (age, gender, etc.), and propose two integration methods, named residualized controls (RC) and residualized factor adaptation (RFA), to be used in county-wise prediction tasks. Demonstrating techniques that integrate information at both the model-level and data-level, we found consistently strong improvement over naively combining features, for example, increasing county level well-being predictions by over 12%. Next, we consider unstructured extra-linguistic data. In the first part, we incorporate social network connections and language over time to propose a novel metric for quantifying the stickiness of words - their ability to spread across friendship connections in a social network over time (or in other words, stick in ones vocabulary after seeing friends use it). We obtain which language features are more probable to disseminate through friendship and show such a metric is useful for predicting who will be friends and what content will spread. In addition, we analyze language content over time by proposing a novel dynamic content-specific topic modeling technique that can help to identify different sub-domains of a thematic scope and can be used to track societal shifts in concerns or views over time.