Developing a Data-Centric NLP Machine Learning Pipeline

Опубликовано: 20 Март 2026
на канале: Toronto Machine Learning Society (TMLS)
62
2

Speakers Bio:

Diego Castaneda, Data Scientist at Shopify
Diego is a Data Scientist with a background in computational astrophysics. He currently supports the needs of the messaging data team at Shopify.

Jennifer Bader, Content Strategist at Shopify
Jennifer Bader is a Content Designer. She works at Shopify, an e-commerce platform that makes it easier for people to start, run, and grow a business.

She joined Shopify to lead content strategy for Kit, Shopify’s AI-powered virtual employee. She now works on the Messaging team, which gives entrepreneurs tools to connect with their customers and collaborate with their team.

Jennifer has worked in UX content design and communications for the last 20 years. Before Shopify, she worked for OpenTable, IDEO, and a technology commercialization institute.

She’s passionate about the intersection of data science and content design. And using conversational interfaces to humanize e-commerce.

Abstract:
The number of components and level of sophistication in end-to-end ML pipelines can vary from problem to problem but there's one common element that is the key to make the whole system great and useful: your training data. The more time you spend developing the training dataset in your ML pipeline, the more positive results you'll get. In this talk, Diego and Jennifer will present the use case of a text classification pipeline they developed from scratch to integrate with one of our products. They'll show details of how they designed an appropriate classification taxonomy, a consistent annotated training dataset and how the end-to-end pipeline was pieced together to deploy a BERT-based model in a low latency real-time text classification system.