In this video, we delve into Lakeflow declarative pipelines, previously known as Delta Live Tables, introduced by Databricks. The latest features enhance the abilities of these pipelines and streamline data processing. We'll explore these enhancements, create simple SQL and PySpark pipelines, and work on building an end-to-end ETL pipeline, including data quality checks to ensure your data remains accurate and trustworthy.
Notebooks used are available at below path:
https://github.com/databeli/databrick...
Plan 1 on 1 with me at below URL:
https://topmate.io/narender_kumar_91/
What you will learn:
00:00 An overview of Lakeflow declarative pipelines and their new features.
03:29 How to create a basic SQL-based pipeline and its core functionalities.
10:02 The process of building a PySpark-based pipeline.
15:27 Building Custom PySpark Declarative Pipeline
18:17 Implement data quality checks (expectations) within declarative pipelines.
22:45 Scheduling and Orchesttration of Declarative Pipelines
If you enjoy this content, please like the video, leave a comment with your thoughts or questions, and subscribe to our channel for more insightful tutorials and discussions!
#databricks #etl #pyspark