ALEXKRI.NET
Resume Contact
← Back

From 40 seconds to under 10: rebuilding incident detection on OpenTelemetry, Apache Kafka, and Apache Flink on Kubernetes

The latency win is not the best part of Atlassian’s incident-detection rebuild. They went from a Node.js aggregator on ~90 VMs to a Flink job in 4 Kubernetes pods: detection from over 40s to under 10s, throughput from ~500M to over 1B events/day, monthly cost from ~$20,000 to ~$650. And they publish the awkward metric too, that automation caught 30.4% of major incidents over nine months. More of this, please.

Read the source ↗