
A Tale of Two Flink Autoscalers
Netflix is migrating from a homegrown Flink autoscaler to the Apache Flink Autoscaler to better handle complex, stateful stream-processing jobs. The transition aims to improve resource efficiency and simplify their operational surface area.
Why it matters
Dynamic autoscaling reduces wasted cloud spend by adjusting resources to match actual traffic cycles. For example, one Netflix team saved approximately $1.1 million annually through these compute expenditure reductions.
The details
- The OSS autoscaler scales parallelism for individual operators instead of whole clusters.
- Netflix integrated the OSS library using Temporal workflows to isolate job failures.
- The system was optimized to support jobs with up to 3,000 Flink subtasks.
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.