Netflix: A Tale of Two Flink Auto...

Netflix: A Tale of Two Flink Autoscalers

AI
The Daily Diff by Premchand Chidipoti
S1 · E21
Aug 25, 2026
09:16

Episode notes

Jordan and Riley compare the two autoscalers Netflix has run across its 30,000+ Apache Flink jobs — a 2019 Mantis-based scaler that reasoned over coarse cluster metrics and moved a single knob (total TaskManager count), versus the Apache Flink community autoscaler that reasons from inside the job using True Processing Rate to size each operator independently. They cover why Netflix is converging on the open-source option, the engineering gaps it had to close at scale (metric collection at high parallelism, preserving forward chaining, respecting sink limits), the safety checks in the "realizer," and a 58% (~$1.1M/yr) compute cut for one team. This is original commentary, not a reproduction of the original post.

Source: "A Tale of Two Flink Autoscalers" by Samuel Yeboah, Francesco Di Chiara, and Mingliang Liu, Netflix Tech Blog, Aug 21 2026 — https://netflixtechblog.com/a-tale-of-two-flink-autoscalers-e9f6a1b1492b

Keywords

Tech blog
Engineering blog
Software design
Software engineering