Isabel Papadimitriou: What Can We Learn about Language from Exploring Multilingual Language Models?

Опубликовано: 29 Июнь 2026
на канале: SIGTYP
264
4

Isabel Papadimitriou: What Can We Learn about Language from Exploring Multilingual Language Models? #SIGTYP
Abstract: The development of successful language models has provided us with an exciting test bed: we have learners that can model language data they are given, and we can watch them do it. In this talk we will go over two sets of experiments that examine language representation and learning in language models, and discuss what we can learn from them. Firstly, we’ll look at subjecthood (the property of being the subject or the object of a sentence) in multilingual language models. By probing the embedding space, we show how a discrete feature like subjecthood can be encoded in a continuous space, affected but not fully determined by prototype effects, and also how these properties come into play with a feature being universally shared among many languages. Second, we approach questions of inductive learning biases and the abstract universals that underlie language by pretraining models on non-linguistic data and observing their language acquisition. Insofar as computational models of cognition act as hypothesis generators for inspiring and guiding our research into understanding language, language models are a very exciting tool to work with and understand.

For questions/discussion please visit our website: https://sigtyp.github.io/ws2022-sigty...