Netflix: A Tale of Two Flink Auto...

Netflix: A Tale of Two Flink Autoscalers

IA
The Daily Diff por Premchand Chidipoti
T1 · E21
25 ago 2026
09:16

Notas del episodio

Jordan and Riley compare the two autoscalers Netflix has run across its 30,000+ Apache Flink jobs — a 2019 Mantis-based scaler that reasoned over coarse cluster metrics and moved a single knob (total TaskManager count), versus the Apache Flink community autoscaler that reasons from inside the job using True Processing Rate to size each operator independently. They cover why Netflix is converging on the open-source option, the engineering gaps it had to close at scale (metric collection at high parallelism, preserving forward chaining, respecting sink limits), the safety checks in the "realizer," and a 58% (~$1.1M/yr) compute cut for one team. This is original commentary, not a reproduction of the original post.

Source: "A Tale of Two Flink Autoscalers" by Samuel Yeboah, Francesco Di Chiara, and Mingliang Liu, Netflix Tech Blog, Aug 21 2026 — https://netflixtechblog.com/a-tale-of-two-flink-autoscalers-e9f6a1b1492b

Palabras clave

Tech blog
Engineering blog
Software design
Software engineering