833 подписчиков
126 видео
Apache Spark shuffle writers: SortShuffleWriter
Apache Airflow behavior for TriggerRule.ALL_DONE task
Apache Spark SQL and corrupted files management
Wildcard path and partition values in Apache Spark SQL
Apache Spark Structured Streaming - mapGroupsWithState with and without state removal
Apache Spark Structured Streaming, watermark and window processing
JIT compilation and Apache Spark SQL
Apache Spark 3 and Dynamic Partition Pruning
PySpark and serializers
Apache Flink and temporal table joinb
Delta Lake, commits and conflicts management
Lesson 03: Who is a data engineer?
Apache Airflow and DAG triggered with datetime.now()
Lesson 04: Data pipelines
Apache Spark 3.1.1 - state store metrics in Spark UI
Apache Spark Structured Streaming 3.3.0 - new features
Lesson 12: Data terms
Lesson 09: Data stores
Batch version of the transformWithState
Apache Kafka transactional producer with Apache Spark Structured Streaming
Master stream processing - Stream processing module
Master stream processing - Data ingestion online module
Connecting Apache Spark to Apache Kafka Schema Registry with ABRiS
State store role in stream-to-stream joins in Apache Spark Structured Streaming
Apache Spark Structured Streaming and file source
Apache Spark 3.0 and skew join optimization in the Adaptive Query Execution
Apache Spark 3.1 - global watermark correctness issue
Apache Airflow and changing start date for already running DAG
Lesson 10: Data processing - standalone
Dynamic Resource Allocation in Structured Streaming
Master stream processing - Data cleansing module
Lesson 07: Data architectures
Lesson 05: Batch processing
Lesson 01: Plan
Lesson 08: Data organization
Lesson 02: Data engineering and archeology
Apache Spark 3.0 and reuse subquery optimization in the Adaptive Query Execution
Apache Spark Structured Streaming 3.2.0 features - Apache Kafka
Apache Spark 3.1.1 - table API in Structured Streaming
Crypto-shredding - encryption and decryption for Java, Python and SQL
Apache Airflow data sensor example
Apache Spark on Kubernetes demo
Azure Data Factory controls flows mapped to Apache Airflow DAGs
Apache Spark Structured Streaming and broadcast join internals
Structured Streaming and Apache Kafka source - maxOffsetsPerTrigger impact on reprocessing, part 1
PySpark one-column DataFrame schema inference
Module 2: Data ingestion offline - plan
Scala - ExecutionContext impact on the program execution
Apache Spark Structured Streaming and broadcast variable
Writing data to multiple Apache Kafka topics from Apache Spark