Training LLMs: Lessons from the Trenches

Опубликовано: 22 Июнь 2026
на канале: Toronto Machine Learning Society (TMLS)
362
11

Speaker: Bandish Shah: Engineering Manager, MosaicML/Databricks

Training large AI language models is a challenging task that requires a deep understanding of natural language processing, machine learning, and distributed computing. In this talk, we will go over lessons learned from training models with billions of parameters across hundreds of GPUs. We will discuss the challenges of handling massive amounts of data, designing effective model architectures, optimizing training procedures, and managing computational resources. This talk is suitable for ML researchers, practitioners, and anyone curious about the “sausage making” behind training large language models.