Building Open Data Lakes: Debezium, Apache Kafka, Hudi, Spark, and Hive on AWS

Опубликовано: 05 Октябрь 2024
на канале: Gary Stafford
6,062
104

In this video demonstration, we will build a simple open data lake on AWS using a combination of open-source software, including Debezium for change data capture (CDC), Apache Kafka, Kafka Connect, Apache Hive, Apache Spark, and Apache Hudi and Hudi's DeltaStreamer.

All open-source files on GitHub: https://github.com/garystafford/kafka....

This video represents my own viewpoints and not of my employer, Amazon Web Services (AWS). All product names, logos, and brands are the property of their respective owners.

📣 Please subscribe to my YouTube channel for future videos.