Learn more about data pipelines: https://academy.starburst.io/explorin....
This video is part 4 of a 5-part series showing how data pipelines using modern data lakes are constructed using Starburst Galaxy and SQL. To do this, we will continue using the BlueBikes dataset: https://www.bluebikes.com/ to construct the Consume layer. In a modern data lake, also known as a data lakehouse, the Consume layer is the last of three data layers and stores a complete and finished copy of a dataset. The data inside has been fully transformed, cleansed, and normalized and is ready to be used. This is the layer that is usually used for queries, or as the basis for BI tools and dashboards. With Starburst Galaxy, the Consume layer can be constructed using SQL, and this step-by-step video tutorial shows you how to leverage this tool to perform data engineering tasks easily.
Starburst Galaxy makes creating data pipelines much easier than many platforms because it allows you to use SQL to perform many tasks that might normally require a general scripting language like Python. As you progress in this video training series, you can follow along with your own Starburst Galaxy account: https://www.starburst.io/platform/sta.... The aim is to provide a big data tutorial that lets you test out the software on real world datasets.
For more information on this topic check out the full course, Exploring data pipelines: https://academy.starburst.io/explorin.... It digs deeper into data pipelines at all levels, including a discussion of how Starburst Galaxy can help construct a pipeline using SQL. It is available on Starburst Academy: https://academy.starburst.io/.
The course is part of the Data foundations series: https://academy.starburst.io/page/dat.... The series provides a conceptual and contextual background to help you understand how Starburst fits within Big Data ecosystems, including Exploring data lakes: https://academy.starburst.io/explorin... and Exploring data lakehouses: https://academy.starburst.io/explorin.... Each course in the series is suitable for all audiences.
#datapipeline #etl #etlpipeline #datalake #datalakehouse #trino #starburst #dataengineering #datalake #sql #sqldatapipeline #data #datascience #dataconsumer #businessintelligence #datadashboards #datadrivendecisions