Rundown of Flink's Checkpoints

Опубликовано: 21 Март 2026
на канале: Flink Forward
5,619
65

Apache Flink’s checkpoint-based fault tolerance mechanism is one of its defining features. Because of that design, Flink can easily scale to both very small and extremely large scenarios and provides support for many operational features like stateful upgrades with state evolution or roll-backs and time-travel. It’s been some time since the last time we explained how it works. At the same time over the years, it got a few improvements such as e.g. unaligned checkpoints or the latest effort to make it work intuitively with finite sources. The talk is aimed at people of all levels, but people new to the checkpointing mechanism are especially welcome as I will start with a very introduction of the concept and walk listeners through the newest additions. After the talk, I hope everyone will have the necessary understanding for running their pipeline reliably and leave the talk with helpful tips for doing so.