Smart Data Products From prototype to production

Опубликовано: 03 Август 2026
на канале: Toronto Machine Learning Society (TMLS)
154
1

💻 Abstract:
Smart Data Products From prototype to production. Notebooks are a great tool for Big Data. They have drastically changed the way scientists and engineers develop and share ideas. However, most world-class ML products cannot be easily engineered, tested and deployed just by modifying or combining notebooks. Taking a prototype to production with high quality typically involves proper software engineering and process. At Montevideo Labs we have many years of experience helping our clients to architect large systems capable of processing data at peta-byte scale. We will share our experience on how we productize ML starting from a prototype to production. Data Scientists should be truly free to use any tool and library available. On the other hand, engineers need artifacts that are modular, robust, readable, testable, reusable and performant. We'll outline strategies to bridge these two needs and aid the concepts with a live demo.

🔊🔊 Speaker bio
Senior Data Engineer of Montevideo Labs
Javier Buquet holds a degree in Computer Science from ORT University (Montevideo) and since 2015 has been working with Montevideo Labs as a Senior Data Engineer for large big data projects. He has helped top tech companies to architect their Spark applications, leading many successful projects from design and implementation to deployment. He is also an advocate of clean code as a central paradigm for development.

Founder and Chief Engineer of Montevideo Labs
Maximo Gurmendez holds a master’s degree in computer science/AI from Northeastern University, where he attended as a Fulbright Scholar. As Chief Engineer of Montevideo Labs he leads data science engineering projects for complex systems in large US companies. He is an expert in big data technologies and co-author of the popular book ‘Mastering Machine Learning on AWS.’ Additionally, Maximo is a computer science professor at the University of Montevideo and is the director of its data science for the business program.

If you enjoyed this talk, visit us at https://mlopsworld.com/ and come participate in our next gathering! 💼

Would you like to receive email summaries of these talks? Join our newsletter FREE here: http://bit.ly/MLOps_Summaries 📧

Timestamps:

0:00 Intro
0:11 Introduction of the host
2:26 Introduction of Maximo Gurmendez
3:15 Agenda
3:43 What is a Smart Data Product?
4:27 Issue #1 Assume assumptions on future data will remain the same
5:36 Issue #2 Let data scientists and engineers do their own thing independently
7:29 Issue #3 Metrics are outdated (or wrong)
9:28 Issue #4 Underestimate the importance of data dependencies
10:27 Issue #5 Not incorporate data science artifacts as part of the CI/CD
11:06 Issue #6 Not invest in local integration/debugging
11:47 Issue #7 Not having proper environmental/experimental hooks
12:34 Issue #8 Same users are subject to all A/B tests
13:02 Issue #9 Launch a product without proving the value through a prototype

14:00 Demo

29:16 Takeaways

❓ Q&A ❓

30:05 The audience was asking for a copy of the codes
32:08 Why do you decide to use a scala?
33:33 In training how do you measure the performance of the model?
34:54 What are your thoughts on future strategies to bring data scientists and engineers together?
36:15 Do you prefer to include your model and pre-processing together in a spark pipeline or maybe separate the two steps and why you have like that preference?
39:36 Do you create a prototype from Jupyter notebook?
40:49 Did you mean production operation metrics or model business metrics?

42:28 Closing remarks